Full text
ii DIREITOS DE AUTOR E CONDIÇÕES DE UTILIZAÇÃO DO TRABALHO POR TERCEIROS Este é um trabalho académico que pode ser utilizado por terceiros desde que respeitadas as regras e boas práticas internacionalmente aceites, no que concerne aos direitos de autor e direitos conexos. Assim, o presente trabalho pode ser utilizado nos termos previstos na licença abaixo indicada. Caso o utilizador necessite de permissão para poder fazer um uso do trabalho em condições não previstas no licenciamento indicado, deverá contactar o autor, através do RepositóriUM da Universidade do Minho. Licença concedida aos utilizadores deste trabalho Atribuição-NãoComercial CC BY-NC https://creativecommons.org/licenses/by-nc/4.0/
iii ACKNOWLEDGMENTS There are many people who have contributed to the accomplishment of this dissertation and supported me throughout the whole course of study. Firstly, I would like to express my sincere gratitude to my academic advisor, Professor Idalete Maria da Silva Dias, for supervising this thesis and continuous support in every step. This work would not be the same if not for her guidance, insightful and detailed comments, suggestions for improvement. I would also like to thank her for taking care of my accommodation issues and making my stay in Portugal as comfortable as possible. I would also like to show my deepest gratitude to the committee, including, Professor Álvaro Iriarte Sanromán, Professor Carlos Valcárcel Riveiro and Professor Idalete Maria da Silva Dias. My grateful thanks goes to the EMLex Coordinator, Professor Stefan Schierholz and all the EMLex Consortium for providing financial support and organizing the studying process in an interesting and efficient way. I will always carry positive memories about the EMLex. I take this opportunity to express my profound gratitude to my parents Mr. Vasyl Morys and Mrs. Svitlana Morys for their unceasing encouragement. To my son Dimitri and husband Gerardo for coming with me to Portugal and motivating me during the last semesters of the studies. Finally, I would like to say thanks to my friends and classmates of the EMLex program who have always helped me out to check the libraries of their home universities, advised new sources and shared their experiences. I am grateful to the God for bringing all of these people to my life.
iv STATEMENT OF INTEGRITY I hereby declare having conducted this academic work with integrity. I confirm that I have not used plagiarism or any form of undue use of information or falsification of results along the process leading to its elaboration. I further declare that I have fully acknowledged the Code of Ethical Conduct of the University of Minho.
v ABSTRACT Bilingual onomasiological online dictionaries: the preparation of the English-Spanish Summer Camp Dictionary Bilingual onomasiological dictionaries generate a considerable interest because of two main reasons. On the one hand, thematic dictionary research is a less explored field of lexicography. On the other hand, bilingual onomasiological dictionaries combine several dictionary types. Since the Internet has established itself as a platform for online dictionaries, it is important to study how onomasiological dictionaries are organized on an electronic medium. The current thesis aims at analyzing two monolingual thematic online dictionaries, namely BerufeNet and Computer Glossary . Search options and navigation structures of these dictionaries receive special attention. The key focus of the dissertation is to create a prototype of the English-Spanish Summer Camp Dictionary (ESSCD), which involves the detailed description of the dictionary’s sources, users and their needs, usage situations and functions. The methodology used to elaborate the ESSCD’s prototype involves six main steps: (1) create a corpus; (2) extract keywords; (3) extract collocations; (4) generate concordances; (5) create an XML database; (6) design a web page. The current thesis discusses the above-mentioned steps as well as presents the data categories of an ESSCD article. Keywords: electronic lexicography, lexicographical process, onomasiological dictionaries, specialized dictionaries
vi RESUMO Dicionários onomasiológicos bilíngues on-line: a preparação do Dicionário de Acampamento de Verão Inglês-Espanhol Dicionários onomasiológicos bilíngues geram um interesse considerável por duas razões principais. Por um lado, a pesquisa de dicionários temáticos é um campo pouco explorado da Lexicografia. Por outro, os dicionários onomasiológicos bilíngues combinam vários tipos de dicionários. Com a consolidação da Internet como uma plataforma para dicionários on-line, é importante observar como os dicionários onomasiológicos são organizados no meio eletrónico. A presente tese visa analisar dois dicionários online temáticos monolíngues, nomeadamente, o BerufeNet e o Computer Glossary . As opções de pesquisa e as estruturas de navegação desses dicionários recebem atenção especial. O foco principal da dissertação é criar um protótipo do Dicionário de Acampamento de Verão Inglês-Espanhol (ESSCD), que envolve a descrição detalhada das fontes do dicionário, utilizadores e suas necessidades, situações e funções de uso. A metodologia implementada na elaboração do protótipo do ESSCD envolve seis etapas principais: (1) criar um corpus; (2) extrair palavras-chave; (3) extrair colocações; (4) gerar concordâncias; (5) criar uma base de dados XML; (6) projetar uma página web . A presente tese analisa os passos acima mencionados, bem como apresenta as categorias de dados de um artigo ESSCD. Palavras-chave: dicionários especializados, dicionários onomasiológicos, lexicografia eletrónica, processo lexicográfico
vii TABLE OF CONTENTS DIREITOS DE AUTOR E CONDIÇÕES DE UTILIZAÇÃO DO TRABALHO POR TERCEIROS ........................ ii ACKNOWLEDGMENTS ...................................................................................................................... iii STATEMENT OF INTEGRITY ................................................................................................................ iv ABSTRACT .......................................................................................................................................... v RESUMO ............................................................................................................................................ vi Abbreviations ....................................................................................................................................... x List of Figures ..................................................................................................................................... xi List of Tables ...................................................................................................................................... xii CHAPTER 1 INTRODUCTION ...........................................................................................................13 1.1 Scope of the Dissertation ............................................................................................................. 13 1.2 Objectives .................................................................................................................................... 15 1.3 Significant Prior Research ............................................................................................................ 16 1.4 Terms and Definitions .................................................................................................................. 19 1.5 Structure of the Thesis ................................................................................................................. 21 CHAPTER 2 METHODOLOGY ...........................................................................................................22 2.1 The Dictionary Basis ....................................................................................................................22 2.1.1 Primary Sources .......................................................................................................................23 2.1.1.1 Corpora.................................................................................................................................23 2.1.1.1.1 Collecting the Comparable Summer Camp Corpus..............................................................24 2.1.1.1.2 Corpus of Contemporary American English .........................................................................28 2.1.1.1.3 iWeb Corpus ......................................................................................................................28 2.1.1.1.4 English Web Corpus 2015 ..................................................................................................28 2.1.1.2 Other Primary Sources and Lemma Selection ........................................................................29 2.1.1.2.1 International Exchange of North America Application Forms ................................................29 2.1.1.2.2 Indeed – a Job Search Engine as a Source for ESSCD’s Notes ............................................36 2.1.2 Secondary Sources ...................................................................................................................36 2.1.3 Preliminary Conclusions ...........................................................................................................36 2.2 Dictionary Subject Matter ............................................................................................................. 37 2.3 For whom is the ESSCD? User Profile ........................................................................................... 37 2.4 Usage Situations .........................................................................................................................39 2.5 User Needs .................................................................................................................................40
viii 2.6 Functions and the Genuine Purpose of the ESSCD ....................................................................... 41 2.7 Preliminary Conclusion ................................................................................................................42 CHAPTER 3 ESSCD AS A HYBRID LEXICOGRAPHIC GENRE .............................................................44 3.1 Dictionary Typologies .................................................................................................................. 44 3. 2 ESSCD as an Online Dictionary ................................................................................................... 47 3.2.1 Onine vs. Print Dictionaries .......................................................................................................47 3.2.2 User Research ..........................................................................................................................51 3.2.3 User Participation .....................................................................................................................54 3.3 ESSCD as a Specialized Dictionary ..............................................................................................56 3.4 ESSCD as an Onomasiological Dictionary ..................................................................................... 57 3.4.1 Analysis of Online Onomasiological Dictionaries ........................................................................58 3.4.1.1 BerufeNet – an Information Portal on Professions in Germany ................................................59 3.4.1.2 Computer Glossary ................................................................................................................64 3.5 ESSCD as a Bilingual Dictionary ..................................................................................................68 3.6 ESSCD as an Active vs. Passive Dictionary ...................................................................................69 3.7 Preliminary Conclusion ................................................................................................................70 CHAPTER 4 TRIADIC STRUCTURE OF THE ESSCD ..........................................................................72 4.1 The Lexicographical Database of the ESSCD ................................................................................ 72 4.1.1 ESSCD Header Database ..........................................................................................................72 4.1.2 ESSCD Text Database ..............................................................................................................75 4.2 The User Interface .......................................................................................................................80 4.2.1 Macrostructure ......................................................................................................................... 81 4.2.2 Microstructure .......................................................................................................................... 81 4.2.3 Outer Texts ...............................................................................................................................82 4.2.3 Search Options and Navigation structures. Access Structures of the ESSCD ...............................83 4.3 Preliminary Conclusion ................................................................................................................85 CHAPTER 5 DATA CATEGORIES OF THE ESSCD ARTICLES ..............................................................86 5.1 Comment on Form ......................................................................................................................86 5.1.1 Orthography .............................................................................................................................86 5.1.2 Pronunciation Data ...................................................................................................................87 5.1.3 Grammatical Data ....................................................................................................................87 5.2 Comment on Semantics ..............................................................................................................88
15 of test articles. The aspects of market research, staffing, calculations of expenses and benefits, lexicographer´s manual, financing, time plan with tasks and deadlines, online release preparation, archiving of dictionary data, dictionary version and its maintenance fall outside the scope of the thesis. . 1.2 Objectives The following are the principal objectives of the thesis: to provide an overall review of the access and cross-linking structures and data presentation within articles of thematic online dictionaries; to create a prototype of the ESSCD with a detailed description of the compilation process. Figure 1: Wiegand’s visualization for the computer-assisted lexicographical process (Wiegand, 1998, p. 234)
16 1.3 Significant Prior Research This section attempts to briefly review the previous research on thematic dictionaries. Onomasiological dictionaries have hitherto been addressed only by a small number of researchers (Kipfer, 1986; Reichmann, 1990; Hüllen 1994; McArthur, 1986, 1998; Jackson, 2002; Hartmann, 2005; Stark, 2011 etc.). One of the most recent works on onomasiological dictionaries was carried out by Martin Stark in 2011. The researcher argues that bilingual thematic dictionaries are rather at the beginning of their development. In his work, these dictionaries are defined as a combination “of three lexicographic traditions: the bilingual, the thematic and the pedagogical” (Stark, 2011, p. 379). Stark observes that in the past onomasiological dictionaries were often the product of one author and mentions the Longman Lexicon of Contemporary English (1981) by McArthur and the Oxford Learner’s Wordfinder Dictionary (1997) by Hugh Trappes Lomax as examples. Furthermore, the author discusses approaches to evaluating bilingual onomasiological dictionaries and provides recommendations for creating them. Although the survey on the usefulness of bilingual thematic printed dictionaries revealed that only 5% of learners of English and French as Foreign Languages use bilingual thematic dictionaries, 88% of informants find bilingual thematic dictionaries as useful as semasiological bilingual dictionaries. Interestingly, 65% of learners of English as a Foreign Language prefer bilingual onomasiological dictionaries to conventional bilingual dictionaries. The reason for such a result is the availability of notes and translation examples in L1. Onomasiological dictionaries have been addressed by Hartmann, who also conducted a survey and analyzed a number of onomasiological works published during the 20th century for about 20 languages. Hartmann has done tremendous work reviewing the International Encyclopedia of Lexicography and Lexicographica Series Maior , checking libraries and bookshops, involving more than fifty external experts from various countries to conduct an extensive survey on thematic dictionaries. The results of Hartmann’s survey are visualized in Figure 2 and Appendices 2, 3. Surprisingly, languages with a lower number of speakers, such as Danish and Swedish have a much higher average of onomasiological works per million of speakers than Spanish or English. As can be seen from Figure 2, Spanish has the highest number of onomasiological dictionaries.
17 Figure 2: Number of native speakers and onomasiological works per language. Based on Hartmann (2005). Visualized in Tableau Software In terms of English onomasiological reference works, Hartmann highlights that “[t]here is no single publication dedicated entirely to the discussion of onomasiological dictionaries […] the metalexicographic literature on this topic is often rather immature, episodic and superficial” (Hartmann, 2005, p. 8). Furthermore, Hartmann formulates desiderata for theoretical and practical lexicography. Particularly, the researcher calls for a more exhaustive bibliography of thematic dictionaries. Finally, Hartmann (2005, p. 15) concludes that “typological studies of onomasiological reference works other than the thesaurus genre are long overdue”. The chapter Abandoning the alphabet in Jackson's Lexicography: An Introduction discusses the drawbacks of an alphabetical listing. The researcher states that each word in A-Z arrangement is treated in isolation and “some words that belong together morphologically become separated” (Jackson, 2002, pp. 146-147). Another significant researcher Tom McArthur describes onomasiological dictionaries from two perspectives: as a metalexicographer and as a lexicographer. What does this mean exactly? McArthur compiled the first thematic learner’s dictionary, The Longman Lexicon of Contemporary English (1981),
18 and also provided broad perspectives, findings and results from the point of view of the actual dictionary-making process. Nevertheless, all the above-mentioned research projects were related to printed dictionaries. Most scholars disregard thematic online dictionaries. When it comes to specialized online dictionaries, Fuertes-Olivera and Tarp (2014) have analyzed sixteen sources including The Dictionary of Business and Management , The New Palgrave Dictionary of Economics Online , The Musikordbogen , The Glossary of Mortgage and Home Equity Terms , The Cambridge Business English Dictionary , TermFinder , The Business English-Spanish Glossary by A. D. Miles , The Glossary of FAO 4 Database and Information Systems (FAO Term Portal), Kicktionary , Interactive Terminology for Europe (IATE) , CercaTerm , The United Nations Multilingual Terminology Database (UNTERM) , The Multilingual Glossary , European Dictionary of Skills and Competences (Disco) , Genoma and EcoLexicon . However, only two of these, namely IATE and EcoLexicon , seem to have the onomasiological access structure. Three of the sixteen analyzed dictionaries allow the user to select subfield, discipline or thematic area ( FAO Term Portal , TermFinder , CercaTerm ). Some articles on onomasiological dictionaries were written by Sierra (2000, 2008) and Sierra & Hernández (2013). Notably, the author has elaborated and successfully tested the prototype of the Diccionario Electrónico para la Búsqueda Onomasiológica 5 with 33 terms of destructive phenomena. The dictionary is based on the following method: the user types keywords, and the system matches them by means of an inverted file in order to identify the possible term entries (Sierra, 2000). In addition, Sierra develops a method of automatic definition extraction for specialized onomasiological dictionaries (Sierra, 2008). Taking into consideration the previous research on thematic online dictionaries, we will attempt to discuss this dictionary type in this thesis. We have taken the Manual of Specialised Lexicography: The preparation of specialised dictionaries (Bergenholtz & Tarp (Eds.), 1995) and Theory and Practice of Specialised Online Dictionaries. Lexicography versus Terminography (Fuertes-Olivera & Tarp, 2014) as the main reference works for the thesis methodology. The theoretical frameworks of other lexicographers, Wiegand in particular, have been considered as well. 4 Food and Agriculture Organization. 5 Electronic Dictionary for Onomasiological Searching (translated by Sierra, 2000, p. 228).
19 1.4 Terms and Definitions In this thesis the terms onomasiological and thematic are used interchangeably. Candidates, participants and applicants (of a summer camp program) refer to the target users of the ESSCD. Text reception and text production are used as synonyms for the terms decoding and encoding correspondingly. Internet dictionary is used as synonym for the term online dictionary . Access structure “[d]ie Zugriffstruktur ist die Menge der Elemente und deren Ordnungsbeziehungen, auf die sich Benutzer während einer Zugriffshandlung stützen“ (Engelberg & Lemnitzer, 2009, p. 272). Collocations of a given word are statements of the habitual and customary places of that word (Firth, 1957, p. 181); Collocation is the cooccurrence of two or more words within a short space of each other in a text (Sinclair, 1991, p. 170). Comparable corpora are “texts which, though composed independently in the respective language communities, have the same communicative function” (Laffling, 1992); “collection of similar documents that are collected according to a set of criteria, e.g. the same proportions of texts of the same genre in the same domain from the same period (McEnery and Xiao, 2007) in more than one language or variety of languages (Sinclair, 1996) that contain overlapping information (Munteanu and Marcu, 2005; Hewavitharana and Vogel, 2008)” (Skadiņa et al., 2012, p.7). Corpus is “a collection of pieces of language that are selected and ordered according to explicit linguistic criteria in order to be used as a sample of the language” (Sinclair, 1996); “a collection of texts, of the written or spoken word, which is stored and processed on computer for the purpose of linguistic research” (Renouf, 1987, p. 1); “a collection of pieces of language text in electronic form, selected according to external criteria to represent, as far as possible, a language or language variety as a source of data for linguistic research” (Sinclair, 2004). Dictionary base is “die Menge der sprachlichen Dokumente und Wissensquellen, auf der die lexikographische Beschreibung des […] Wörterbuchgegenstandes basiert” (Engelberg & Lemnitzer, 2009, p. 272). Dictionary subject matter (Wörterbuchgegenstand) “ist die Sprache oder der Sprachausschnitt, die/den das […] Wörterbuch aufgrund der vorhandenen […] Wörterbuchbasis beschreibt.” (Engelberg & Lemnitzer, 2009, p. 272). Electronic dictionary “[t]he term electronic dictionary (or ED) can be used to refer to any reference material stored in electronic form that gives information about the spelling, meaning, or use of words. Thus a spell-checker in a word-processing program, a device that scans and translates printed words, a glossary for on-line teaching materials, or an electronic version of a respected hard-copy dictionary are
20 all EDs of a sort, characterised by the same system of storage and retrieval” (Nesi, 2000, p. 839); “[e]in e. W. [elektronisches Wörterbuch] ist ein Nachschlagewerk, das in digitalisierter Form auf einer CD-ROM, einer Diskette oder auf einem an das WWW angeschlossenen Server publiziert wird. Der Zugriff auf e. W.er ist nur mit Hilfe elektronischer Hilfsmittel möglich“ (Engelberg & Lemnitzer, 2009, p. 271). External subject classification is “a systematic arrangement of the subject field in question, delimiting this in relation to adjacent fields, with the purpose of identifying the material which is to form the empirical basis of the dictionary” (Bergenholtz & Tarp, 1995, p. 83). Internal subject classification “establishes an overview of the subject area in question, thus forming the basis of the systematic structuring of the dictionary” (Bergenholtz & Tarp, 1995, p. 84). Items are data-carrying entries from which the user can retrieve information relevant to fulfilling the genuine purpose of the specific dictionary. […] e.g. items giving the pronunciation, morphology, part of speech, paraphrase of meaning, translation equivalents, illustrative examples, etc. (Gouws, 2014, p. 161). Lemma selection „[m]it der Lemmaselektion werden diejenigen sprachlichen Einheiten identifiziert und ausgewählt, die Gegenstand der lexikographischen Beschreibungen sein sollen“ (Engelberg & Lemnitzer, 2009, p. 246). Lexicographical process „[e]in abgeschlossener lexikographischer Prozeß ist zu verstehen als die Menge derjenigen Tätigkeiten von Lexikographen, die ausgeführt wurden, damit ein bestimmtes Wörterbuch entsteht […] Kalkulierbarkeit, Zerlegbarkeit, Kontrollierbarkeit, Reglementierbarkeit, Lehrbarkeit und Prüfbarkeit sind zentrale Eigenschaften, die lexikographische Herstellungsprozesse und die gesamte lexikographische Praxis mit Herstellungsprozessen anderer Art teilen“ (Wiegand, 1989a, pp. 250-251). Macrostructure “[u]nter der MAKROSTRUKTUR eines Wörterbuches verstehen wir die geordnete Menge seiner Lemmata“ (Kunze & Lemnitzer, 2007, p. 79) 6 . Terminological classification is “a systematic listing of the LSP terminology of the subject field in question, for the purpose of ensuring that all LSP terms are captured” (Bergenholtz & Tarp, 1995, p. 84). 6 Emphasis in original.
21 1.5 Structure of the Thesis This dissertation is divided into five chapters. Following the introductory part, the thesis deals with the applied methodology in the second chapter: potential users and their information needs have been identified, usage situations and dictionary functions have been defined. Moreover, the second chapter describes material collection and data acquisition. The third chapter analyzes the ESSCD dictionary type and highlights the user research and user participation in online dictionaries. The discussion on online versus printed dictionaries gets special attention. Current online onomasiological dictionaries are considered in more detail. The fourth chapter focuses on the process of creation of an XML (Extensible Markup Language) database and the ESSCD structures. ESSCD’s articles and their data categories are presented in the fifth chapter. Finally, we critically analyze the limitations and suggest some insights into future work, followed by conclusions.
22 CHAPTER 2 METHODOLOGY I certainly do not know all lexicographic projects past and present; but one of those I know not a single one was finished in the time and for the money originally planned. Zgusta, 1971, p. 348 During the lexicographical process “different decisions must be taken, actions must be done and different methods must be used” (Schierholz, 2015, p. 328). This statement also applies to the ESSCD. A few questions arise in the process of dictionary creation: what will the base of the dictionary be; how will the lemma candidate list be selected; what data categories should the dictionary include; which examples should be taken into the dictionary; how should the dictionary database be created; how should the microstructure, macrostructure and access structure be organized. All in all, the lexicographical process of the ESSCD splits into the following six work steps: (1) corpus creation, (2) extraction of keywords, (3) extraction of collocations, (4) generation of concordances, (5) database creation, and (6) website design. A range of tools is used for each step (see Figure 3). Figure 3: Steps and working tools for the compilation of the ESSCD As can be seen from Figure 3, all the corpora used for the ESSCD will be searched with concordances . Concordances as a data extraction mode are commonly used in lexicographical practice. They extract phrases, where a certain headword occurs, and present it in the form of a list. To search corpora for the ESSCD we use the Keyword in Context ( KWIC ) concordance type. As has already been mentioned above, the ESSCD database is developed using Extensible Markup Language (hereafter XML ), which allows for the creation of the project specific tags. Extensible Stylesheet Language Transformations (henceforth XSLT ) and Cascading Style Sheets (hereafter CSS ) are likely to be employed for the dictionary web page design. 2.1 The Dictionary Basis Regarding the dictionary base for specialized dictionaries, Bergenholtz and Tarp (1995, p. 90) speak about empirical basis which consists of three basic types: (1) introspection , (2) existing literature , (3) texts. Introspection refers to the lexicographer’s own competencies. For the ESSCD, introspection
23 means first of all the lexicographer’s cultural knowledge of summer camps and linguistic competence in both source and target languages. Given that the ESSCD is a bilingual dictionary, introspection includes the lexicographer’s translation skills as well. Obviously, lexicographers cannot rely solely on their own skills. Hence, existing literature and texts are of paramount importance. For the preparation of the ESSCD, the following existing literature sources will be used: Linguee English-Spanish Dictionary for translation equivalents, web site Indeed for notes, grammar books of the English and Spanish languages for grammatical information. Texts for the ESSCD are corpora, which are described in the follow-up section of the thesis. 2.1.1 Primary Sources Wiegand (1998, p. 140) describes the primary, secondary and tertiary sources as follows: „[d]ie primären Quellen sind vor allem (aber nicht nur) Texte, welche aus natürlichen oder quasi-natürlichen Kommunikationssituationen stammen, oder größere zusammenhängende Ausschnitte aus solchen. […] Die Menge der primären Quellen, welche Texte oder größere zusammenhängende Ausschnitte sind, bildet das lexikographische Korpus […]. Zu den sekundären Quellen gehören alle Wörterbücher, die nach dem Instruktionsbuch entweder obligatorisch oder fakultativ konsultiert werden sollen, und zu den tertiären Quellen gehören alle sonstigen Sprachmaterialien, die benutzt werden wie z.B. linguistische Monographien und Grammatiken […]“. In the frame of the current project we use general language corpora (COCA, iWeb Corpus, English Web Corpus 2015) and a self-compiled comparable corpus as primary sources. Additional search engines, for example Indeed , and application forms support lemma selection and writing of the dictionary articles. Bilingual dictionaries, for example Linguee EnglishSpanish Dictionary , and grammar books belong to the secondary sources of the ESSCD. 2.1.1.1 Corpora Corpus evidence allows observing the behaviour of a word in its context, thus opening new perspectives for Lexicography. Several researchers have proposed definitions for the concept corpus (Atkins, Clear & Ostler, 1992; Sinclair, 1996; Pearson, 1998). In general, all the linguists agree that the corpus is a collection of authentic texts sampled in a systematic way which serves to represent a specific language. The corpus size is usually associated with its representativeness. The corpus is representative when “the lexical density does not alter when more texts are added” (Losey-León, 2015, p. 298). Sinclair (1996) defines three main characteristic of the corpus: quantity, quality and simplicity . The researcher determines a default value for each characteristic: large for quantity, authentic for quality and plain text for simplicity. In addition, Sinclair distinguishes between text corpus (or whole text corpus ) and a
24 samples corpus . The latter refers to specialized corpora, “which do not contribute to a description of the ordinary language, either because they contain a high proportion of unusual features, or their origins are not reliable as records of people behaving normally” (Sinclair, 1996). It goes without saying that corpora make it possible to identify the features of lexical units and to discover new collocations providing authentic examples of language in use. However, Geyken and Lemnitzer (2016, pp. 203-204) consider some significant limitations with respect to corpora as a dictionary base: “kein Korpus, egal welcher Größe kann eine lebende Sprache als Ganzes abbilden oder repräsentieren”. The editors of DWDS note that corpora often consist of newspaper texts solely. Only limited specialized corpora are likely to contain transcripts of spoken language. Moreover, the analysis of a huge amount of data can also lead to errors, because lexicographers tend to make subjective decisions (Geyken & Lemnitzer, 2016). 2.1.1.1.1 Collecting the Comparable Summer Camp Corpus This section deals with the compilation process of the English-Spanish Comparable Corpus and provides a comprehensive description of the steps followed in the corpus creation process. Creating a corpus for specialized dictionaries has its own requirements (Bergenholtz & Tarp, 1995, p. 94): 1. the corpus should cover all sub-fields of the subject matter in question; 2. the text types to be considered in the dictionary should be included in relation to their presumed relevance for the intended dictionary users and use situations. The English-Spanish Comparable Corpus is a web-based corpus. Firstly, the sources were selected: 20 web pages of Canadian and American summer camps, American Camp Association page, Canadian Camping Association and a Canadian placement agency pages for English and 18 Spanish and Mexican summer camp pages, camp searcher page and the Mexican Association of Camps for Spanish. It should be noted that the content of web pages for summer camps in Spain differ slightly from the classical organization of a summer camp. In Spain summer camps either focus on learning English, i. e. language camps such as Campamento Náutico Bilingüe, Irish Summer Camp or they are dedicated to a particular activity: Rockcamp dedicated to music, Campus Experience to football, Club Hipico Miracampos to horse riding. The main question which every lexicographer asks before compiling a corpus deals with the size of the corpus. Albeit the exact number of words to achieve a representative corpus has not been established yet, several suggestions have been made so far. For a specialized corpus it is generally agreed that “corpus size depends on the corpus aim and no minimum or
31 10. Emergency contact An internal subject classification functions as a drop-down menu of the external subject field classification , and is aimed “to ensure systematic representation of the subject field in the dictionary” (Bergenholtz & Tarp, 1995, p. 85). For the ESSCD it would look as follows: 1. Personal information 1.1. Basic info 1.2. Dates of availability 1.3. Permanent address 1.4. Mailing address 2. Education and background 2.1. Education 2.2. Dates of summer holiday 2.3. Background 3. Job preference 4. Work experience 4.1. Experience working with children 4.2. Employment and volunteer history 5. Skills 5.1. Camp Counselor 5.2. Support Staff 6. References 6.1. Reference 6.2. Applicant information 6.3. Rate the following 6.4. Overall 6.5. Reference signature 7. Documents 7.1. Health history form 7.2. Passport 7.3. Photo album 7.4. Photograph 7.5. Police background check
32 7.6. Proof of student status 7.7. Camp contract 8. Medical history 8.1. Medical history 8.2. Other background info 9. Visa information 10. Emergency contact 10.1. Contact address 10.2. Next of Kin's address One may wonder why dates of summer holidays is a separate subsection. In fact, the accuracy of the holiday duration is especially important, because if the dates the applicant has selected in the application form do not correspond to the dates on the proof of student status, two options are possible (1) the visa sponsor will not send documents required for the J-1 visa; (2) in case that the sponsor has sent the documents, there is a high risk of visa rejection. Terminological classification lists the Language for Special Purposes (LSP) terms related to the subject field of summer camps. Such listing ensures that the main vocabulary of the ESSCD is lemmatized. Research suggests that not only lemmata, but also non-lemmatic units ( nichtlemmatische Einheiten ) such as collocations, phrasal verbs, compounds and fixed expressions should be taken into the candidate list (Engelberg & Lemnitzer, 2009, p. 246). Dictionaries aimed at text production should include LSP terms and other non-common-language expressions according to Bergenholtz and Tarp (1995, p. 103). As can be seen from Table 3, the dictionary will lemmatize not only words but also phrases, i.e “a particular kind of lexeme which in itself consists of several lexemes or several grammatical words” (Bergenholtz & Tarp, 1995, p. 100).
33 5. Skills 5.1. Camp Counselor 5.2. Support Staff 5.1.1. Adventure 5.1.2. Aquatics and Waterfront 5.1.3. Arts and Crafts 5.1.4 Horse Riding 5.1.5. Music 5.1.6. Performing Arts 5.1.7. Special Skills 5.1.8. Sports 5.2.1. Administration 5.2.2. Housekeeping 5.2.3. Kitchen 5.2.4. Maintenance 5.2.5. Security camping diving arts and crafts stable management bagpipes ballet animal care aerobics office manager cleaning baking carpentry security guard climbing fishing camera work drumming cheer-leading camp fires archery office worker housekeeping cook electrical high ropes sailing candle making Guitar choreography Chess basketball laundry dish washer gardening hiking surfing ceramics Harp circus skills child care board diving food preparation lawn mowing low ropes swimming costume design Music dance Cookery boat driving general kitchen maintenance windsurfing fine arts Piano hip-hop digital editing cross country head chef plumbing canoeing glass blowing singing lighting Dressage cycling Waiter jet skiing jewellery making violin performing arts Ecology dodgeball kayaking knitting set design I.T. equestrian leather craft sound technician Leadership fencing loom weaving stage management Mechanic field hockey magic tap dance Nature fitness make up artist theater Newspaper go-carts mask making Nutrition golf
34 media-art print making gymnastics metalsmith Radio lacrosse model making recording studio land sports painting Robotics life guarding photography Rocketry martial arts sewing Science motocross stained glass trainable lifeguard mountain biking stationary making trainee lifeguard mountain boarding textiles trainee teacher open kayak Tye-Dye outdoor adventure videography paddle boarding wood working paint ball power boat driving quad bikes rack and field rib driving roller hockey roller skating rowing shooting
35 show jumping skate park ski boat driving soccer spinning swim coaching synchronized swimming tennis trapeze volleyball wakeboarding water skiing yoga instructor zip lines zumba instructor Table 3: Lemma candidate list for the field “Skills”
36 2.1.1.2.2 Indeed – a Job Search Engine as a Source for ESSCD’s Notes Indeed is a job search site, where instructor positions needed at camp are usually offered. The information on the website is used for the note sections of the ESSCD. In each entry of the Skills field there is a note General tasks of (name of the skill e.g. arts and crafts) instructor , which serves to prepare the users for the duties they will fulfill at the camp. Several job offers for arts and crafts instructors on the web page Indeed have been analyzed and essential duties and responsibilities have been summarized in a note. The descriptions of available work places on Indeed vary significantly: some provide a relatively detailed job specification and required qualifications, skills and experience, supplement functions as well as physical requirements whereas others just list the tasks and responsibilities. 2.1.2 Secondary Sources Our secondary sources include audio materials, the Linguee English-Spanish Dictionary and additional grammar books. The headwords are recorded by English native speakers. The Linguee English-Spanish Dictionary assists with translation. Linguee is the largest multilingual online dictionary in the world, available in more than 200 language pairs 20 . Linguee is especially attractive for translators offering a number of parallel texts in the selected language pair. Linguee gives access to more than a billion of human translations. The translations which were not verified by professionals are marked with a warning sign. 2.1.3 Preliminary Conclusion Corpora are the main source for the compilation of the ESSCD. It should be considered that crawling a corpus from the Internet usually involves a lot of noise, and words are frequently misspelled. In order to reduce the noise in the self-compiled corpus, external links were removed. Only the blog links were kept as a source of encyclopedic information. The English-Spanish Comparable Corpus tailored to the goals of our project, COCA, iWeb and the English Web Corpus 2015 form the text basis of the ESSCD, and act as the sources of collocations and corpus examples. The basis for grammatical information are grammar handbooks of each working language. Encyclopedic information is based on Indeed and the English-Spanish Comparable Corpus. Translation equivalents are taken from the Linguee EnglishSpanish Dictionary. 20 Linguee.com: The World’s Fastest Dictionary. Retrieved from https://www.linguee.com/press/EN/2015-02-09_PressRelease_EN.pdf
37 2.2 Dictionary Subject Matter Wiegand (1998, p. 302) defines the subject matter of a dictionary as follows: “[d]er Wörterbuchgegenstand eines bestimmten Wörterbuches ist die Menge der in diesem Wörterbuch lexikographisch bearbeiteten Eigenschaftsausprägungen von wenigstens einer, höchstens aber von endlichen vielen sprachlichen Ausdrücken, die zu einem bestimmten Wörterbuchgegenstandsbereich gehören.” The subject matter of the ESSCD is the summer camp application vocabulary. As has been mentioned before, summer camps are a part of the American and Canadian culture with a long tradition. Over 10 million children spend their summer in camps across the USA. Summer camp life is full of specific concepts, terms, activities etc. The ESSCD aims at describing them. 2.3 For whom is the ESSCD? User Profile In order to satisfy user’s needs, it is absolutely crucial to determine the profile of the consumer and the situations that trigger these needs: “[t]wo fundamental factors interact in the formation of [user] needs […]. The first of these factors is the characteristics of the concrete person experiencing the information needs, and the second is the social situation or context where these needs occur” (Fuertes-Olivera & Tarp, 2014, pp. 48-49). Prior to making a dictionary, several competences of the target users are to be taken into account: mother tongue, which is“[t]he pivot in any user profile is native language 21 ” (Bergenholtz & Tarp, 1995, p. 20), encyclopedic knowledge, native-language competence, foreignlanguage competence, native-language text production and reception, foreign-language text production and reception, translation from the foreign language into the native language and vice versa (Bergenholtz & Tarp, 1995, pp. 20-24). These aspects have been formulated in eight basic questions 22 suggested by the function theory 23 in order to build a profile of the target user (see Appendix 10). According to these questions, ESSCD’s users are learners of English as a Foreign Language with Spanish as their mother tongue. User needs are linked to their knowledge. The intended users should have the minimum level 24 of English: A2 for support staff positions and C1 for camp counselor positions. This also determines the content of the ESSCD, i.e. it includes vocabulary which falls into the A2-C1 frame. The intended users should have basic skills in translating between the languages. 21 Emphasis in original. 22 The list of 8 questions was further extended to include 11 questions in Fuertes-Olivera and Tarp, 2014 (see Appendix 11) and is open for additional questions. 23 The theory of lexicographic functions has been initiated and developed by researchers from the Center for Lexicography at the Aarhus School of Business since the 1990s. 24 According to the Common European Framework of Reference for Languages.
38 The participants are required to have at least a shallow level of the general cultural and encyclopedic knowledge. The intended users have to master the domain of summer camps at least at the beginner level: Stage of application Level of domain mastering Users sign up for a program, i.e. get informed about the summer camp project, available positions, application process and receive additional information via email or read more on the web page of the sponsor/recruiter25. beginner The candidate is hired and receives a set of educative videos about the summer camp life and his/her duties at the camp. basic The participant has lived the summer camp staff experience and wants to apply for a job for the next season. This type of applicant is called “returner”. semi-experts Users have worked in a summer camp continually for over three seasons and more; applicants with working camp experience who are employed as recruiters and want to work during their vacations in the camp. experts Table 4: Levels of mastering the summer camp domain Taking into account the encyclopedic knowledge of the users, the ESSCD is designed for laypeople and semi-experts. It is quite difficult to determine the mother tongue competence level of the intended users and the corresponding LSP in their mother and in a foreign language. In terms of consultation-relevant user characteristics (Fuertes-Olivera & Tarp, 2014, p. 50), the target users seem to lack experience in lexicographical consultations. 25 Recruiters are self-employed workers or agencies that are in charge of promoting the program, inform students about the program and follow and support participants with the hiring process. Figure 6: Levels of competence (Bergenholtz & Tarp, 1995, p. 21)
39 2.4 Usage Situations This section describes concrete lexicographic situations that motivate users to consult the ESSCD. The motivation for a dictionary consultation, as a rule, arises from lexical gaps, the search for translation equivalents or spelling (Müller-Spitzer, 2014). As claimed by the modern theory of lexicographic functions, two main groups of use situations can be distinguished: knowledge-orientated and communication-orientated 26 (Bergenholtz & Tarp, 2003, p. 174). In knowledge-oriented situations, users seek additional information to widen their knowledge about the subject field in question, for example, encyclopedic information. In this type of situation, the communication occurs between the lexicographer and the user (the user seeks knowledge and the lexicographer provides it), whereas during communication-oriented situations, two or more persons are engaged in written or oral text production/text reception. In the latter case, the lexicographer takes the role of an indirect mediator and helps to solve communication difficulties, which arise during the encoding or decoding of texts. Five forms of communication are presented in the figure below. Communication issues may arise during the phases of production, reception and translation styled in italics. Figure 7: Communication model (Bergenholtz & Tarp, 2003, p. 174) Keeping in mind the six types of communication-oriented use situation 27 suggested by Bergenholtz and Tarp (2003, p. 175), the ESSCD should be consulted under the following cicumstances: production of texts in English as a foreign language; reception of texts in English as a foreign language; translation of texts from English as a foreign language into Spanish as a native language. However, the above-presented use situations, e.g. reception, production are broad concepts. Consequently, it is important to define the context of the dictionary use, i.e. the extra-lexicographic one (Tarp, 2012, p. 114), which will be described in the next section. 26 The recent research of the Function Theory proposes four lexicographical situations: (1) communicative (e.g. text reception, production, translation etc.); (2) cognitive (users require information in order to save it in the memory as knowledge or to apply it for a particular task; (3) interpretive (users require information to interpret a special sign, symbol etc.); (4) operative (“a person needs information in order to perform an action of a physical, mental or linguistic nature” (Fuertes-Olivera & Tarp, 2014, p. 51). 27 Six basic types of communication-oriented user situations (Bergenholtz & Tarp, 2003, p. 175) are listed in Appendix 12.
40 2.5 User Needs Having defined the target users and the situations of use, we can proceed to the users’ needs. Dictionaries are designed to meet the users’ needs and to satisfy their preferences. For this reason, one of the leading tasks of the lexicographic team is to figure out the user necessities. User needs are connected to a specific user group and specific use situations (Bergenholtz & Tarp, 2003). FuertesOlivera and Tarp (2014, p. 48) point out that a distinction should be made between information needs and lexicographically relevant needs . The information needs “initially lead to a lexicographical consultation, but this consultation itself may give rise to a new type of need different from the former and related to the correct handling and use of a lexicographical tool, with the specific purpose of retrieving the information required to satisfy the original needs.” Nowadays there is a danger of information overload. We have too much data at our disposal and sometimes spend hours looking for a certain piece of information. The main reason for information overload in specialized dictionaries is the fact that the dictionaries regardless of their form try to satisfy many need types of varied user groups and propose all the possible solutions (Fuertes-Olivera & Tarp, 2014). As Tono (2010, p. 3) points out “[n]o user has specific needs unless they are related to a specific type of situation”. ESSCD is designed for certain users in particular non-lexicographic situations. Below is the description of the situations, when the user needs may arise and to which the ESSCD is linked. The user is willing to participate in a summer camp program, so he/she needs to fulfill the program requirements. One of the requirements is to fill in the application documents online. The application documents are available in English. Potential applicants may not understand the fields, so they would look for translations from English into Spanish. Furthermore, the intended user group may not be familiar with LSP terms in English, because they lack information about the summer camp subject field or some concepts existing only in the USA such as Next of Kin’s address . Another problem that the target group might encounter is the lack of essential vocabulary to come up with exhaustive answers. The mentioned problems have been personally observed among participants when working for a placement agency. The user needs have been identified through personal communication with applicants and through guidance them in the hiring process. Welker (2010) has reviewed 320 empirical studies on user research and concludes that the majority of studies involve limited number of participants. User research has been also criticized for the lack of usefulness and representatives. Tarp confirms the prospective of Sheatsley (1974) that the great number of surveys today is “a waste of time and money” (Tarp, 2009, p. 293). Apart from this, the
47 To sum up, dictionaries can be classified according to numerous criteria: from Kühn (1989) and Hausmann (1989) to Fuertes-Olivera and Tarp (2014). The great number of typologies have arisen due to the fact that a typology can be based on different criteria and parameters: “as the traditional image of the dictionary changes, new distinctions between dictionary types arise” (Nesi, 2000, p. 839). Some researchers even assume that “[d]ictionaries come in more varieties than can ever be classified in a simple taxonomy” (Béjoint, 1994, p. 37). 3.2 ESSCD as an Online Dictionary The number of Internet users worldwide is steadily increasing: from 1,024 billion in 2005 to 3,57 billion in 2017 as recent statistics prove (see Appendix 8). This tendency has an impact on Lexicography as well. We are living the transition from printed to online dictionaries and searching an online dictionary is considered as a “revolutionary experience” (Nesi, 2000, p. 839). Nowadays, the users opt for ever faster ways of information retrieval, and online dictionaries provide this opportunity. Electronic Lexicography offers new possibilities: “new kinds of evidence, new modes of description, new ways of organizing evidence, new possibilities for exploiting database structure and hypertext links, and the need for new theoretical foundations” (Hanks, 2012). In addition to the term “Electronic Lexicography”, Carr coins another term, Cyberlexicography , and defines it as “employing the Internet to compile or create a dictionary” (Carr, 1997, p. 209). 3.2.1 Onine vs. Print Dictionaries The best dictionary is probably the one rendering usable result in a short time. Bergenholtz and Gouws (2010, p. 114) The opinions of scholars regarding the medium of dictionary production are split. On the one hand, online dictionaries were strongly criticized as unreliable sources in the middle of the 1990s: “[w]er glaubt, mit einem Internet-Anschluß auch eine Vielfalt von qualitativ hochwertigen Nachschlagewerken erworben zu haben, die die Anschaffung von Wörterbüchern und Enzyklopädien in gedruckter Form oder auf CD-ROM überflüssig macht, wird sich zunächst enttäuscht sehen. Das Internet mit seiner dezentralen Organisationsform und seinem schnellen Wachstum ist bislang kein Ort der Verbindlichkeit und Verläßlichkeit“ (Storrer & Freese, 1996, p. 129). On the other hand, print dictionaries have obvious limitations, they “are likely to remain extremely conservative and command little or no serious investment and hence bring little or no serious innovation” (Hanks, 2012) and are “far from adequate as a medium for dictionaries” (Rundell, 2012, p. 16). In order to be an effective tool, both medium
48 forms should do the same: satisfy specific user needs within a short time and represent data as understandable as possible (Lew, 2012). The significant increase in online dictionary consultations has been observed: “electronic learners’ dictionaries already seem to be on the way to becoming a preferred alternative to the ‘fat’ dictionary in print” (Nesi, 1999, p. 65). This trend is growing thanks to innovative features of online dictionaries: portability, faster search, incorporation of multimedia elements (images, audio pronunciation, video clips etc.), mobile contents, immediate cross-references inside the dictionary and links to external sources, etc. New technologies speed up search queries and reduce the lookup time in online dictionaries as contrasted to leafing over the pages of their printed counterparts. These are the main reasons why many students favor the electronic medium (see the study of Taylor & Chan, 1994); the users appreciate “search speed and ease of use” (Dziemanko, 2012, p. 333). One of the latest developments that reduces the search time when reading a text in a foreign language is an electronic dictionary integrated into an application (e.g. Kindle). When the users come across an unfamiliar word, they can just click on it and a pop-up window with relevant lexicographical information will appear: “[i]t is already clear that the dictionary is moving from its current incarnation as autonomous ‘product’ to something more like a ‘service’, often embedded into other resources” (Rundell 2012, p. 29). Spelling and grammar checkers are seen as advanced lexicographical tools, because they detect mistakes and suggest corrections (Fuertes-Olivera & Tarp, 2014). On the other hand, distracting advertisements are present on the pages of some online dictionaries. Consequently, the consultation process might be slowed down. The overwhelming majority of online dictionaries provide audio pronunciation which is considered to be a benefit, since not all users are able to interpret International Phonetic Alphabet (IPA) transcriptions. One more advantage of online dictionaries is the free of charge consultation, as opposed to their paperbased counterparts, which are usually quite costly. Furthermore, an online dictionary can be frequently updated, whereas users normally wait several years for a new edition of a hardcopy dictionary to be published. The internet dictionary allows users to interact with each other, to elaborate dictionary articles on their own or to improve the existing material 33 in collaboration with other users. Online dictionaries can be accessed on any electronic device at any place and time, whereas their paper predecessors are usually bulky and difficult to carry around. The location of online dictionaries, i.e. URL, however, may change or disappear. Another option, offered by some online dictionaries, is the access to the most recent entries. Lew (2013) observes that paper dictionaries have non-immediate 33 See Section 3.2.3 User Participation.
49 crossreferences , while electronic dictionaries mostly provide immediate crossreferences, so the users can “jump” from one article to another with a mouse click, since articles in an online dictionary are linked with each other. Besides, the majority of online dictionaries incorporate error tolerant input , which helps the users to look up words with wrong spelling and gives them the possibility to review the search history. Although online dictionaries introduce a wider range of access routes than printed ones do, the content quality is not assured: “[t]he range and convenience of such search routes in EDs are, of course, no guarantee of the quality of the information content” (Nesi, 2000, p. 840). The size of a print dictionary leads to article density, therefore it is difficult for the user to retrieve information from condensed articles. In the print dictionaries era it was quite challenging to fit all the data on the paper pages. Therefore, many dictionaries were divided into volumes, which was not very practical for the user. Several volumes of a printed dictionary can be combined in one online dictionary: “one in which the electronic dictionary or dictionary site may encompass a single or a whole range of traditional dictionaries that can be adjusted in various ways to comply with the needs of particular user groups” (Trap-Jensen, 2010, p. 1133). In terms of space in online dictionaries, Lew (2013) distinguishes between storage, i.e. “capacity to hold the total content” and presentation space, i.e. “display of lexicographic information”. Similarly, Trap-Jensen (2010, p. 1133) speaks about the twosided nature of electronic dictionaries: (1) data that is saved in a database and (2) data represented on the screen. Regarding the presentation space, both dictionary forms (paper and electronic) have restrictions. Fuertes-Olivera and Tarp (2014) propose that the maximum amount of data that can be visualized in an online dictionary should allow the user to navigate without having to scroll down. Although online dictionaries have enough space to include images, tables or other additional data, users may experience information overload . This concept was introduced by the futurist Alvin Toffler in the book Future Shock (1970) 34 . Actually, users can experience information overload regardless of the medium. As mentioned earlier, the density in a print dictionary can also frustrate users. Reviewing hardcopy editions of the Oxford Advanced Learner's Dictionary of Current English (OALD4), Nesi (1999, pp. 55-56) states that “the more information the paper-based dictionary contains, the harder (and more time-consuming) it will become for learner users to find exactly what they need to know, without first having to negotiate a quantity of information that they do not need to know, or cannot process.” Now the question arises how to minimize the information overload and to identify “which particular elexicographic solutions work best (and for whom, and under what circumstances), so that future electronic dictionaries can be made more effective than their paper predecessors, and more effective 34 Future Shock. In: Wikipedia. Retrieved from https://en.wikipedia.org/wiki/Future_Shock
50 than the dictionaries available today” (Lew, 2012, p. 344). Fuertes-Olivera and Tarp (2014, p. 64) suggest possible solutions: “purely mono-functional; multi-functional allowing for mono-functional data access, mono-functional allowing for individualized data access or multi-functional allowing for individualized data access”. Dictionaries should give the users the opportunity “to select, filter and present the specific data needed by the user” (Fuertes-Oliverera & Tarp, 2014, p. 93). These lexicographers take Lexin , online dictionaries with Norwegian as a source language designed for immigrants in Norway, as an existing example for designing individual articles. Lexin allows for the inclusion or exclusion of certain data items in the article 35 , so that the problem of superfluous data may be solved, “electronic dictionaries […] do not have the organisational and spatial constraints of hardcopy dictionaries, and can retrieve and combine information according to the specifications of the user” (Nesi, 2000, p. 839). Figure 9: Pop-up window in Lexin To avoid information overload, the ESSCD will provide the precise data required by specific users in specific usage situations. Given that the ESSCD is an onomasiological dictionary, the online production form is significantly important, because print dictionaries organize entries mainly in a linear manner, which is “inadequate as a means of grouping and regrouping words according to their semantic and pragmatic similarities” (Nesi, 2000, p. 839). The common characteristic of all reference works is that they are not meant to be read linearly, but searched for selective information retrieval (Engelberg, Müller-Spitzer & Schmidt, 2016). Furthermore, the target users are young people, who are accustomed to access the Internet in case of any information need. Rundell (2012, p. 72) refers to the users in the age range between 17 and 24 as “digital natives”. From the lexicographer’s point of view, online 35 See Figure 9: Pop-up window in Lexin.
51 dictionaries are easier to edit, since a lexicographer does not need to review A-Z letters, but rather selects updates and additions. Thus, online dictionaries can be frequently updated. The conclusive argument in favor of creating an online dictionary for the ESSCD is the fact that the participants fill in the application documents online too. Therefore, it would be more convenient for them to access a new tab and start the dictionary search. As the first international study with 684 respondents on online dictionary use in 2010 reveals, the dictionary users are likely to consult dictionaries on laptops or computers rather than on the small screen devices (Koplenig & Müller-Spitzer, 2014). Having evaluated all the advantages and disadvantages, we have decided to compile the ESSCD in an online format. 3.2.2 User Research There is a wide range of studies on comparing electronic and print dictionaries in terms of their effectiveness and usefulness. Töpel (2014) summarizes important individual studies on the use of edictionaries and concludes that users tend to consult electronic dictionaries more often, and find the required information faster than in print dictionaries. The users are generally satisfied with electronic dictionaries. It is worth mentioning that more surveys were carried out with CD-ROM dictionaries and PEDs 36 than with online dictionaries. The results of important individual studies on electronic dictionaries from 1993 until 2016 are visualized in Figures 10 and 11. Figure 10 shows that more than half of the informants (57,78%) experience difficulties using electronic dictionaries, i.e. they are not familiar with the innovative functions of e-dictionaries and, consequently, need training. This tendency can be observed during the whole period of time (see Figure 11). 36 PEDs: Pocket Electronic Dictionaries; studies show that this type of electronic dictionaries is popular in Asia (see Boonmoh & Nesi, 2008; Chen, 2000; Boonmoh, 2012).
52 Figure 10: Results of important studies on electronic vs. printed dictionaries. Visualized in Tableau. Figure 11: Results of important studies on electronic vs. printed dictionaries over time. Visualized in Tableau It seems that users are not familiar with the novel features of e-dictionaries, unless they have been informed and trained on how to use innovative functions efficiently: “if a solution is unknown to the users, as is necessarily the case with any experimental feature we would like to test, their performance is likely to be negatively affected by the novelty of the feature. Depending on how steep a learning curve the new feature has, it may take more or less time and practice before users get more familiar with the innovation tested, and before the benefits, if any, get a chance to come to the surface” (Lew, 2011a, p. 11). The study by Taylor and Chan indicates that (1994) sometimes even English teachers have
53 questions regarding the use of electronic dictionaries. Even though the lexicographical data in online dictionaries implement innovative functions, the first two international studies on the use of online dictionaries prove that such novel features are not of greatest importance. The dictionary users still value conventional dictionary standards such as reliability of content, clarity 37 , updated content, speed and accessibility (Müller-Spitzer & Koplenig, 2014). Regardless of age, occupation, language version, the results remain the same (Müller-Spitzer, 2016, p. 321f). Similarly Trap-Jensen (2010, p. 1142) reports “[…] it may be disappointing that the users do not seem to take advantage of all these wonderful possibilities.” Reviewing the empirical surveys on innovative features of online dictionaries, Lew (2012, pp. 359-360) concludes that “available evidence invites optimism with respect to static pictures and audio recordings, but looks less optimistic when it comes to video and animation enhancements”. The focus of lexicographers on technical innovation has been criticized: “[..] it may be argued that the elements of customization implemented in electronic dictionaries so far result more from lexicographers’ ideas about how users should use e-dictionaries (to the point that it might be called a ‘lexicographer-oriented’ lexicography) rather than from the insights into the way dictionaries are actually used” (Verlinde & Peeters, 2012, p. 151). Nesi (2000) states that the recent dictionary research concentrates rather on technical novelty than “the source, quality and appropriacy of the definitions” or other relevant aspects. An additional observation is that the majority of conducted user research is dedicated to the intra-lexicographical consultation 38 phase i.e. when the users are selecting information, accessing, verifying and retrieving data (Fuertes-Olivera & Tarp, 2014). Summarizing the results on the survey “What makes a good online dictionary”, carried out in 2010 and involving 1074 participants, Müller-Spitzer and Koplenig (2014, p. 184) conclude that “being a reliable resource, and a clearly presented and understandable tool, which is kept as up to date as possible” are the key features of a good online dictionary. Unlike paper dictionaries, modern Internet dictionaries integrate their user into dictionary development, offering a broader choice of user participation. This have led to a new form of dictionary making – collaborative lexicography , that will be discussed in the next section. 37 Clarity is understood as “[t]he general structure of the website enables you to easily find the information you need” (Müller-Spitzer & Koplenig, 2014, p. 147). 38 Fuertes-Olivera and Tarp (2014, pp. 87-90) determine three phases that make up the lexicographical process from the user’s perspective: extralexicographical pre-consultation (the user experiences and becomes aware of the information need, as a result decides to start a dictionary consultation), intra-lexicographical consultation phase and extra-lexicographical post-consultation process (i.e. users apply retrieved information).
54 3.2.3 User Participation Lexicographers can use the Internet not only for corpus searching or dictionary consulting, but also to reach out to their potential users. It is becoming more and more common that the Internet community contributes to the creation of new dictionaries and the improvement of already existing lexicographic resources. User participation is possible thanks to the following two main factors: the lexicographical process in online dictionaries is continuous, and online dictionaries are open systems. This phenomenon is called bottom-up lexicography . The term goes back to Carr (1997, p. 214) and constitutes the opposite to the top-down process where dictionaries are made from editors, through publishers, to readers. Lew (2014) proposes a new term for the modern user “ prosumer ” as a blend of the words “producer” and “consumer” for the users contribute to the dictionaries and profit from their content at the same time. Such user collaboration is significant for specialized lexicography in particular, including dialects, youth language or endangered languages. This section presents a review of interaction possibilities between the users and lexicographers. Editors benefit largely from collaborative lexicography. They can gather first-hand information on user needs and preferences. The users can provide direct feedback in relation to the dictionary as a whole or to the single articles contained in it. In fact, the Oxford English Dictionary has a long tradition of collaborative editing. Its user participation practice reaches back into the 19th century when the Philological Society of London started working on the New English Dictionary on Historical Principles . The citizens of Great Britain, its colonies and North America were asked to collect and submit the examples for common word usage while reading: “[d]iese freiwilligen Helfer wurden gebeten, Bücher zu lesen, Belege auf Zetteln in einem bestimmten Format zu notieren und diese Zettel bei der Philological Society einzureichen” (Thier, 2014, p. 63). At the very beginning, the collection was unsystematic; further the public received the list of references that had already been processed and was asked to explore other text genres. Nowadays, users are encouraged to work on a specific domain e.g. current American texts, historical texts, scientific literature (see Their, 2014, pp. 63-64). Béjoint (1979) also encourages the involvement of users in lemma selection. The user contributions have increased the number of headwords in Collins, Longman and Cambridge dictionaries (Nesi, 2000). Abel and Meyer (2016) determine three main types of user participation: direct, indirect and accompanying user participation (direkte, indirekte und begleitende Nutzerbeteiligung). Direct user participation refers to the possibility of the user creating, modifying and erasing dictionary articles. In the direct user participation, users can contribute to (Abel & Meyer, 2016, pp. 253-263):
55 open-collaborative dictionaries (e. g. Wiktionary). Usually collaborative dictionaries have their predefined article structure, for instance, Wiktionary provides templates of Wiki Markup. In open-collaborative dictionaries, users are allowed to create, edit or delete articles. collaborative-institutional dictionaries . These dictionaries belong to a publisher e.g. MerriamWebster Open Dictionary, Macmillan Open Dictionary. Unlike the previous category, the users can not directly edit or delete dictionary articles. Many user contributions in these dictionaries do not receive lexicographical attention. semi-collaborative dictionaries (e.g. OpenThesaurus, LEO). Lexicographers carefully check user contributions before including them into the dictionary. The ESSCD, like any other reference work, strives to achieve that the users find what they are looking for. Hence, the ESSCD will implement indirect user participation 39 , aimed at gathering feedback on existing or missing dictionary content, dictionary usage, or on the dictionary as a whole (Abel & Meyer, 2016, p. 263f). If users have questions about the functions or the content of the ESSCD, they are welcome to send an e-mail to the editorial board. Secondly, the user can fill in a blank form, where the request will be specified through the question and answer method. The processing of log files as implicit feedback is definitely too expensive for this project. Moreover, log files certainly do not output exact results as they do not distinguish between search engine access and users’ access (see Verlinde & Binon, 2010). The Internet dictionaries allow communication between users. The third type of user participation, namely accompanying user participation refers not only to the interaction between users and lexicographers (e.g. blogs, newsletters, language games), but also to the interaction between the users themselves. Fora, discussions within the user community, comments in social media are some examples of the interaction between users themselves. As a rule, the users ask for help with translation equivalents in bilingual dictionary fora. Interestingly, Duden offers linguistic consulting ( Sprachberatungsdienste ) on the phone, which is also considered a complementary participation (Abel & Meyer, 2016). The ESSCD will allow users to create their own user account, where they can mark their favorite words or multiword expressions with a star. This can provide data on the most read articles, motivating the lexicographers to improve the preferred articles. Another option will be to add lemmata to the “To Learn List” or marking them as “Already Learnt”. On the one hand, user participation might speed up the lexicographical process: publishing houses could save money, users become more familiar with the dictionary structure. On the other hand, 39 See Section 6.2 Discussion on Future Directions.
56 dictionaries take a risk of including erroneous user contributions. As stated by Carr (1997, p. 214), “[t]he Internet is content-neutral: misinformation becomes concurrent with information”. The result will depend on the quality of work users have done. Not everybody has a linguistic competence, consequently spelling and grammar mistakes or other errors often occur. Regarding the quality of the user contribution, Abel and Meyer (2016) point out two quality issues: spam and vandalism on the one hand, and wrong, outdated or too complicated descriptions on the other. To hinder vandalism, Wiktionary, for example, provides detailed instructions to avoid copyright violation and encourages the users to report the cases of plagiarism. It also explains user rights and obligations. Besides, Wiktionary allows permanent access to the users who have written at least 200 articles. Hanks (2012) notes that the English Wiktionary contains outdated definitions, and suggests professional revision and corpora evidence in Wiktionary articles: “[i]n the English Wiktionary, the etymologies are taken from or based on those in older dictionaries; as are definitions, which are extremely old-fashioned and derivative” (Hanks, 2012, pp. 77-78). It turned out that these “stilted and archaic in wording” definitions were taken from the Webster’s Revised Unabridged Dictionary (1913), which is available under public domain. Hanks also highlights the advantages of Wiktionary e.g. hypertexts to Wikimedia. Although Wiktionary is relatively new (the English language Wiktionary has its roots in December 2002), the number of user contributions to Wiktionary have increased rapidly 40 . Over 5 million users have written more than 30 million articles. Now the Wiktionary is available in 174 languages 41 . 3.3 ESSCD as a Specialized Dictionary Hartmann and James define specialized dictionaries as “[…] reference works devoted to a relatively restricted set of phenomena” and set them in opposition to general dictionaries: “[i]n contrast to the general dictionary which is aimed at covering the whole vocabulary for the ‘general’ user, special (or ‘segmental’) dictionaries concentrate either on more restricted information, such as idioms or names, or on the language of a particular subject field, such as the jargon of the drug scene or the technical terms of mechanical engineering” (DoL, 1998, p. 129). Fuertes-Olivera and Tarp (2014) made some observations 42 regarding this definition. In particular, the researchers note that Hartmann and James do not clearly distinguish between the terms specialized dictionary and LSP dictionary . Additional concepts 40 A detailed study on the motivation of collaborative editing can be found in Abel and Meyer (2016). 41 Wiktionary. In Wikipedia. Retrieved from https://en.wikipedia.org/wiki/Wiktionary 42 More on the discussion related to the definition of specialized dictionary see Fuertes-Olivera, P. A. & Tarp, S. (2014). Theory and Practice of Specialised Online Dictionaries: Lexicography versus Terminography . Berlin: de Gruyter: 4-8.
63 Figure 14: The results of the search query Friseur The search may be also conducted according to the subjects, which are assigned to one or more fields of study. This option helps the users in the step-by-step search for a specific field of study. In particular, dictionary articles on courses of study inform users about the enrollment requirements, course contents, financial aspect, length of the study and list universities that offer a chosen field of study and employment possibilities after graduation. The MINT-Search 54 option outputs professions and subjects, for which mathematics knowledge is of great importance. The last search option, Suche nach reglementierten Berufen , is performed according to regulated professions such as medical, legal, teaching professions, occupations in the public service. This means that these types of jobs depend on governmental acknowledgment for professional qualifications. BerufeNet is being updated continuously. As a rule, the updates are based on requests of federal employment agencies and on official sources, for example, training regulations. The feedback from employers, employees and other public institutions also supports development of BerufeNet (Janser, 2018). If questions concerning the functions or content of BerufeNet arise, the users are encouraged to send an e-mail or call the editors directly. The current version number can be found on every page in the footer. There is also a possibility of adding dictionary articles to a watch list ( Merkliste ), which appears as a separate window on the left-side of the screen and allows users to access already marked articles immediately. Moreover, BerufeNet incorporates a wide range of internal links to job trends, data 54 MINT stands for Mathematik, Informatik, Naturwissenschaften and Technik (mathematics, informatics, natural sciences and technology).
64 sources, legal regulations (as every single state in Germany can have its own regulations) etc. and external links to related web pages, documents, brochures etc. The dictionary is primarily designed for the workers of employment agencies: “[t]he purpose of this database is two-fold: it is used by vocational counselors and job placement officers at local employment agencies for career guidance and job placement, but it also serves the general public as a free database for career orientation” (Janser, 2018, p. 18). Looking at the content coverage of BerufeNet, it seems to be especially useful for high school graduates. BerufeNet provides detailed information about possible subjects to study, entry requirements and career paths. University or vocational training graduates can also benefit from BerufeNet by becoming familiar with an exhaustive list of potential work places and requirements. This information system provides young people with some idea of what to expect from their future jobs and what the working environment looks like. Furthermore, people with work experience are likely to make use of the dictionary as they might be interested in further education, job market, alternative professions. The BerufeNet can be compared to the American Dictionary of Occupational Titles (DOT) by the US Department of Labour. One of the significant drawbacks of the DOT 55 is that it is not kept up to date, even though its web page states that it was revised in April 2011. Like all paper dictionaries, DOT provides a Table of Contents, which is considered to be outdated. When users want to access articles thematically, they have to go through a long search path to access them. The structure of articles in DOT is not clear. The dictionary interface is not userfriendly, since all the contents of the dictionary articles are located mainly on the left-side of the web page To sum up, BerufeNet is an innovative dictionary with a rich content. It provides detailed information on the current job requirements, prerequisites, tasks to fulfil in a particular occupational field, working equipment, the working conditions, as well as numerous professional alternatives. The dictionary offers a wide range of lookup options and a user-friendly interface. Although the dictionary contains a large amount of data, this data is clearly structured. 3.4.1.2 Computer Glossary Computer Glossary is “the IT encyclopedia and research library in TechTarget's network” 56 . It consists of two parts “Browse Definitions by Topic” and “Quick Study Resources”. The current description focuses on the first part of the dictionary. 55 Also called Occupational Information Network ( O*NET). 56 Business Partners. In Computer Glossary. Retrieved from https://whatis.techtarget.com/about/partners
65 The dictionary provides users with the definitions of more than 10 000 terms and around 1 000 fast references, quizzes and cheat sheets. Computer Glossary is daily updated. The terms in the glossary can be accessed via the following “topics”: Application Development Business Software Computer Science Consumer Tech Data Center IT Management Networking Security Storage and Data Management Each topic is subdivided into various sections. Figure 15: Computer Glossary start page. Subfields highlighted in blue. The target group of the Computer Glossary is information technology and business professionals. Various abbreviations such as AppDev instead of Application Development, Data Mgmt instead of Data Management prove that the dictionary is aimed at experts. The genuine purpose of Computer Glossary is to assist IT specialists and TechTarget customers with comprehension issues, which are usually technology companies such as Dell, Intel, Cisco, Microsoft. In order to provide an excellent Business-toBusiness service (also called B2B), both parties should understand each other. Since IT and technology
66 professionals use highly specialized vocabulary, the dictionary definitions should explain technical terms and business concepts clearly and in a compact way. The outer texts of Computer Glossary are situated on the right-hand side and on the bottom of the web site. They include: Word of the Day. Users can subscribe to the Word of the Day and see the Word of the Day archive. Quote of the Day Meet the Editor Recently Published Definitions Newest and Updated terms Quiz Yourself Trending Topics Buzzword Alert. This feature shows a word which is fashionable at the moment of lookup. View Related Content. At the end of the dictionary articles there is a suggestion of further reading. Dig Deeper on e.g. telecommunication networking Related Terms Users are encouraged to participate in dictionary editing in many ways. Firstly, in the Meet the Editor window, the users are asked to share their feedback, to suggest a new term or an update. The photo of the editor and the title “Meet the Editor” convey an impression that lexicographers are closer to the users, making it easier to start collaboration. Moreover, even the editor may initiate a discussion. Below the dictionary article there is usually a question from the editor and an invitation to join the discussion, followed by a list of related terms and a list of sources which can be read by users for additional information. Another option of user contribution is the forum Ask Your Peers a Question . On the bottom of the page the users are encouraged to comment dictionary articles with imperatives “Start the conversation”, “Share your comment”. Users can receive notifications when other members comment the same article. The dictionary incorporates internal and external links to TechTarget's other IT-specific Web sites. Computer Glossary does not provide a user manual, images or the incorporation of sound files. The advertisement of books or services provided by TechTarget appears on the dictionary page. However, these advertisements are not as bright and annoying as the regular Web advertisement. Like BerufeNet, Computer Glossary provides the user with the search path followed to obtain a specific
67 result: from the start page to a particular article. Both dictionaries, BerufeNet and Computer Glossary, are easy to navigate; the users can always maintain the overview of their search actions. However, BerufeNet incorporates a broader variety of visual means such as images, font type and size that, in our opinion, facilitate information retrieval and help to avoid the “lost in hyperspace” effect (Kemmer, 2000, p. 16). It is also worth mentioning that the design and content of the analyzed dictionaries is adaptive to mobile device consultation. Computer Glossary offers several search modi. Users can browse definitions by index. Compared to BerufeNet, where alphabetical search results are presented in a two column table (professional title and professional group correspondingly) and the headwords are situated on several pages, Computer Glossary splits the results of one letter into “subgroups” (e.g. A – ACC, ACC – ACT, ACT – ADV etc.) on the left-hand side as a scroll down menu. Phrases in Computer Glossary are presented as separate lemmata. The user can also search terms typing them in the search field. However, the automatic term completion 57 from an index is not available as is the case of BerufeNet. Notwithstanding, Computer Glossary incorporates a search technique with a combination of two or more words and allows the results to be sorted by relevance or date. None of the examined dictionaries implements spellingtolerant search. Each dictionary article in the Computer Glossary consists of lemma, definition and the date when the article was created or updated for the last time. Apart from defintions, many dictionary articles include encyclopedic information (e.g. common features, benefits, comparison with similar terms). Computer Glossary also provides some country-specific information, for example: “A-Law is the type of PCM used in most of the world. The other type, mu-Law, is used in the United States and Japan.” The dictionary indicates who posted a certain article and who contributed to it. The following observations can be made regarding the lemmata. Firstly, editors seem to be inconsistent with abbreviations which are a part of lemma. Sometimes the acronym is placed before the corresponding full form in parenthesis, for instance AUI ( attachment unit interface ), sometimes the other way around: the full form precedes the acronym in parenthesis e.g. Access Network Query Protocol (ANQP). Furthermore, a headword is sometimes composed of two synonyms, for example, cable TV or CATV (community antenna television), even though the CATV already exists as a separate article. The internal link from cable TV to CATV is implemented in the definition. The alternative terms for the designation of the same concept are separated with “or” as in the example cable TV or CATV and parenthesis e.g. cloud telephony 57 Also called type-ahead search , search-as-you-type , incremental search , inline search , or instant search (Lew, 2013, p. 24).
68 (cloud calling). In order to achieve consistency in all articles, editors should follow standardization principles. 3.5 ESSCD as a Bilingual Dictionary Bilingual dictionaries have a long tradition. In fact, a bilingual dictionary is older than a monolingual one. For example, the first bilingual dictionary for the language pair English-Latin appeared before 1450, English-French in 1570, whereas the first monolingual English dictionary was published in 1604 (DoL, 1998). A bilingual dictionary is defined in contrast to a monolingual one as follows: “[a] type of DICTIONARY which relates the vocabularies of two languages together by means of translation EQUIVALENTS, in contrast to the MONOLINGUAL DICTIONARY, in which explanations are provided in one language” 58 (DoL, 1998, p. 14). A bilingual dictionary is sometimes called a translation dictionary, however, the second term rather refers to the conventional type whose main function is to provide translation equivalents (Stark, 2011). Regarding the access to translation equivalents, bidirectional and monodirectional dictionaries can be distinguished. The translation equivalents in bidirectional dictionaries can be reached from L2 to L1 and vice versa (DoL, 1998). The ESSCD is designed as a monodirectional dictionary. The user can look up words from English to Spanish. The fundamental reason for this is that the applicants receive the application documents in English. Bidirectionality may be suitable once the dictionary is further developed and goes outside of the frame of the application documents, i.e. has a high number of lemmata. On the other hand, Bergenholtz and Tarp (1995) suggest the compilation of bilingual specialized dictionaries as monodirectional, because some concepts do not always find their equivalents in other languages. Compared to monolingual dictionaries, the bilingual dictionaries have two absolute advantages. Firstly, they provide translation equivalents. Secondly, they focus on the particular language pair, thus considering the special features of these languages, e.g. false friends. On the other hand, the syntactic and morphological characteristics are not explicit enough. Definitions and examples are not as comprehensive as in the monolingual dictionaries (Lemnitzer & Engelberg, 2009). Bogaards (1996, p. 300) reports that monolingual dictionaries are more useful for reception than for production: “If you do not already know the L2 item you want to investigate, how will you find it in an MLD?” Although it is generally accepted that language learners would benefit more from monolingual dictionaries, they usually opt for bilingual ones. One of the broadest studies in learners dictionary use 58 Emphasis in original.
69 (Atkins & Knowles, 1990) with more than 1000 participants from seven countries has proven that the language learners mostly (75%) use bilingual dictionaries. However, this does not imply higher efficiency. The survey actually demonstrated that the use of monolingual dictionaries for task fulfillment leads to better results. In contrast, Lew’s survey (2016) revealed that active bilingual learners dictionaries significantly improve the quality of writing (for about 33%). The surveys of Laufer and Melamed (1994) have proven that efficient dictionary use strongly depends on the user competence and dictionary search abilities. Competent users scored better with monolingual dictionaries whereas unskilled users with bilingual ones. Stark (2011) considers three main advantages of the bilingual thematic dictionaries over the monolingual ones. First of all, it is much easier for the users to follow information in their mother tongue. In addition, bilingual dictionaries can resolve language-specific difficulties such as false friends. Last but not least, it is sometimes difficult to provide satisfactory definitions in the same language. Béjoint and Moulin (1987) point out that the bilingual dictionaries are suitable for a brief consultation, whereas the monolingual ones represent the lexical system of a foreign language. As mentioned before, the language learners, especially beginners, avoid consulting monolingual dictionaries and find them too difficult, albeit the foreign language teachers always encourage their students to use the latter. Moreover, think-aloud protocols from Wingate’s survey (1999) suggest that the participants using the monolingual printed dictionaries sometimes forget which word they were looking for. Lexicographers differentiate between active and passive bilingual dictionaries that will be introduced in the next section. 3.6 ESSCD as an Active vs. Passive Dictionary The idea to classify dictionaries as active and passive goes back to Shcherba (1940). Engelberg and Lemnitzer (2009, p. 126) summarize the difference 59 between active and passive dictionaries as follows: “[d]as aktive Wörterbuch erfordert eine umfangreiche Mikrostruktur, das passive eine umfangreiche Makrostruktur“. In the lexicographic reality we can hardly find examples which would stand as prototypes for active or passive dictionary implementation model. Hence, it is more adequate to use terms active or passive dictionary functions (Mugdan, 1992). As Frankenberg-Garcia (2015) highlights, learner’s dictionaries do not separate examples which support comprehension and those which facilitate production. The survey on usefulness of bilingual thematic dictionaries (Stark, 2011) shows that 45% of English as a foreign language learners use this dictionary type for writing. The ESSCD is primarily designed as a 59 For more on the differences between active vs. passive dictionaries see Engelberg and Lemnitzer (2009, pp. 125-133).
70 production dictionary. As explained above, the ESSCD combines both types, as its main purposes are to assist the learners with translation (decoding) and text production in the English language (encoding). The passive and active functions of the ESSCD have been described in Section 2.6 Functions and the Genuine Purpose of the ESSCD. It is commonly agreed that the text production process is more challenging for language learners than the comprehension one (Rundell, 1999). Lew (2013) stresses that very few dictionaries are designed as production resources and explains the reason for this. Firstly, this dictionary type should have richer content and, as a consequence, requires more lexicographical effort. Secondly, paper-based dictionaries do not have enough space for “redundant” content. Engelberg and Lemnitzer (2009, pp. 120-121) distinguish four reasons for dictionary consultation for text production: Lexeme usage e.g. valency, collocations, inflexion, connotations. The best option to solve this problem is a learner’s dictionary. Lexeme finding (passive vocabulary, language learners do not remember words). In this case thesauri and dictionaries of synonyms are useful. Alternate lexemes. The user is looking for alternative expressions for a particular word or phrase. Vocabulary. The user desires to master vocabulary of a particular concept for example, feelings, rail travel. General language dictionaries can assist with language production in several ways providing grammar, collocations and examples in dictionary articles (Frankenberg-Garcia, 2015). However, as several studies have proven, many users are not aware that dictionaries include information on syntactic patterns (Béjoint, 1981). This statement even applies to English university students (Herbst & Stein, 1987). Several studies reveal that the users may encounter problems using the dictionary for writing because of the lack of essential dictionary search skills (Tomaszczyk, 1979; Nesi & Meara, 1994; Atkins & Varantola, 1997). Furthermore, it is often the case that users do not search the whole entry and concentrate only on the first sense (Nuccorini, 1994) or misinterpret the data at all (Nesi & Meara, 1994). Problems may be encountered due to native language interference. 3. 7 Preliminary Conclusion The ESSCD combines several lexicographic dictionary types. First of all, it is a specialized dictionary, since it is devoted to covering a restricted range of information, namely the language of application forms, as mentioned above. The ESSCD is a learner’s dictionary i.e. applicants are learners of English
71 as a Foreign Language. The ESSCD fulfils functions of an active and passive translation dictionary, since it helps with text encoding and decoding. Last but not least, the ESSCD is an onomasiological dictionary, because it presents words according to concepts (e.g. personal information, skills, education etc.).
72 CHAPTER 4 TRIADIC STRUCTURE OF THE ESSCD Based on the example of Accounting Dictionaries , the Function Theory of Lexicography addresses the triadic structure of bilingual specialized dictionaries that consists of (1) a lexicographical database; (2) a user interface where one or more dictionaries are integrated; (3) a search engine that mediates between the database and the user interface (Fuertes-Olivera & Tarp, 2014, p. 195). The current chapter discusses the three structures of the ESSCD. 4.1 The Lexicographical Database of the ESSCD We should distinguish between the dictionary database and the presentation of dictionary data. The database is a storage system. Kunze and Lemnitzer (2007, p. 12) define lexical databases as „digitale lexikalische Ressourcen, die in einer Form abgespeichert sind, dass die einzelnen Datensätze konsistent im Hinblick auf eine formale Beschreibung ihrer Struktur sind”. A single dataset may correspond to a dictionary article or its part. Lexical databases are especially beneficial for lexicographers because some lexicographical data such as example sentences can be related to different articles. In this section, we provide the description of the logical structure of the ESSCD and introduce all the components on the basis of concrete examples. Markup languages have been developed for the purpose of annotating documents and texts in electronic form. As mentioned earlier, the Extensible Markup Language (XML) was used to annotate the ESSCD. The Text Encoding Initiative (TEI) proposes guidelines for annotating different text types in electronic form based on the XML syntax rules. The TEI Guidelines include a chapter on the representation of dictionaries and lexicographic resources. One of the major advantages of XML is the fact that it is a descriptive markup language, i.e. it describes the role a specific piece of information plays in a text or document, for example: the author of a text can be marked up with the descriptive tag or element <author>; a date that appears in a text can be annotated with the tag <date>; a place name can be marked using the tag <place>. The extensible nature of XML allows for dictionary-specific tags to be created. Documents based on the TEI are composed of two main sections: TEI header and TEI text . These sections of the ESSCD will be discussed in detail in the following subchapters. 4.1.1 ESSCD Header Database The lexicographical elements are grouped hierarchically. The TEI Header consists of four parts: File Description <fileDesc> groups the bibliographic description of the ESSCD.
79 142. </cit> 143. <cit xml:lang="es"> 144. <quote>explicar la seguridad de la navegación</quote> 145. </cit> 146. </unit> 147. <unit> 148. <cit xml:lang="en"> 149. <quote>to provide assistance to campers using the equipment</quote> 150. </cit> 151. <cit xml:lang="es"> 152. <quote>asistir a los campistas que utilizan el equipo</quote> 153. </cit> 154. </unit> 155. </note> 156. </entry> In order to ensure that all articles are well-formed and follow the standardization principles, a Document Type Definition (DTD) was created. The DTD allows us to define the elements used to annotate metadata and the dictionary content: (1) the structure and content of the dictionary; (2) the order in which these elements must appear; (3) if the elements are mandatory or optional; (4) which attributes are used with the respective elements. The DTD contains the rules that must be followed during the annotation process. As can be seen from the example below, the content model defines the content of each entry as follows: <!ELEMENT entry (form, gramGrp*, domain, sense*, wordClass*, collocGrp*, corpusExamples*, note*, image*)>. The root element <entry> is composed of the following information units: <form> is composed of orthography and pronunciation; <gramGrp> includes elements part-ofspeech, subcategorization, gender and tense; <domain> specifies which (sub)field the headword belongs; <sense> contains the translation equivalent(s) in Spanish with grammatical information on the translation equivalent; <collocGrp> contains collocations according to their patterns, e.g. <pattern id="noun"> for whitewater kayaking ; <corpusExamples> includes examples that explain the usage of collocations; <note> consists of encyclopedic information; <image> indicates the title and the copyright of the image used to represent the concept. Once all the elements are determined, we can move to the declaration of the attributes. Attributes are declared for an element in a sequence/list form. Each attribute declaration contains the element, the attribute associated with the element, the type of value admitted by the attribute and its status as optional or mandatory. The attribute declaration <!ATTLIST cit xml:lang (en|es) #REQUIRED> specifies that the element <cit> admits the mandatory (#REQUIRED) attribute xml:lang and its value is either "en" or "es" (English or Spanish). The DTD of the ESSCD articles is an external DTD:
80 1. <?xml version="1.0" encoding="UTF-8"?> 2. <!ELEMENT entry (form, gramGrp*, domain, sense*, wordClass*, collocGrp*, corpusExamples*, note*, ima ge*)> 3. <!ELEMENT id (#PCDATA) > 4. <!ELEMENT form (orth,pron)> 5. <!ELEMENT orth (#PCDATA)> 6. <!ELEMENT pron (#PCDATA |pc|ptr)*> 7. <!ELEMENT pc (#PCDATA)> 8. <!ELEMENT ptr EMPTY> 9. <!ELEMENT gramGrp (pos, subc*, gen*, tns*)> 10. <!ELEMENT pos (#PCDATA)> 11. <!ELEMENT subc (#PCDATA)> 12. <!ELEMENT gen (#PCDATA)> 13. <!ELEMENT tns (#PCDATA) > 14. <!ELEMENT domain (usg+) > 15. <!ELEMENT usg (#PCDATA)> 16. <!ELEMENT sense (cit)> 17. <!ELEMENT cit (gramGrp|quote |synset| bibl|usg)*> 18. <!ELEMENT quote (#PCDATA | phrase)*> 19. <!ELEMENT phrase ( #PCDATA)> 20. <!ELEMENT synset (quote+) > 21. <!ELEMENT bibl ( author| title| date)*> 22. <!ELEMENT author (#PCDATA)> 23. <!ELEMENT title (#PCDATA)> 24. <!ELEMENT date (#PCDATA)> 25. <!ELEMENT wordClass (gramGrp,sense,collocGrp,corpusExamples)> 26. <!ELEMENT collocGrp (pattern*, colloc*)> 27. <!ELEMENT pattern (colloc+)> 28. <!ELEMENT colloc (cit+)> 29. <!ELEMENT corpusExamples (example+)> 30. <!ELEMENT example (cit+)> 31. <!ELEMENT note (unit+)> 32. <!ELEMENT unit (cit| source)*> 33. <!ELEMENT source (#PCDATA |ref)*> 34. <!ELEMENT ref (#PCDATA)> 35. <!ELEMENT image (imageDesc, ptr)> 36. <!ELEMENT imageDesc (imageTitle, copyright)> 37. <!ELEMENT imageTitle (#PCDATA)> 38. <!ELEMENT copyright (#PCDATA)> 39. <!ATTLIST entry id CDATA #REQUIRED > 40. <!ATTLIST form type (lemma| headword) #REQUIRED> 41. <!ATTLIST pron notation CDATA #REQUIRED> 42. <!ATTLIST ptr type (audioFile|image) #IMPLIED> 43. <!ATTLIST ptr target CDATA #IMPLIED > 44. <!ATTLIST ptr title CDATA #IMPLIED > 45. <!ATTLIST usg dom CDATA #IMPLIED > 46. <!ATTLIST usg subdom CDATA #IMPLIED > 47. <!ATTLIST cit type (translation |translation_lemma) #IMPLIED> 48. <!ATTLIST cit xml:lang (en |es) #REQUIRED > 49. <!ATTLIST pattern id (noun | verb | adjective|adverb|preposition) #REQUIRED> 50. <!ATTLIST colloc n CDATA #IMPLIED > 51. <!ATTLIST example n CDATA #IMPLIED > 52. <!ATTLIST phrase rend (italics) #REQUIRED > 53. <!ATTLIST note type CDATA #IMPLIED > 54. <!ATTLIST tns type (past_simple|past_participle) #REQUIRED> 4.2 The User Interface When a query result appears on the screen, “the data ceases to be a part of the database and becomes a dictionary” (Nielsen & Almind, 2011, p. 143). One of the most important tasks of the lexicographer is to decide where particular data should be located. In terms of data presentation, the
81 lexicographer needs to bear in mind the reference skills of the potential users: “[i]deally, an online dictionary interface will combine simplicity (for those who cannot be bothered) with sophistication (for those who can). A reasonable way to achieve this is to offer a simple default interface with an optional advanced alternative” (Lew, 2013a, p. 29). Taking into consideration the general theory of lexicography, the section User Interface will describe the macrostructure, microstructure and outer texts of the ESSCD, despite the fact that there is no clear separation between microstructure, macrostructure and outer texts in electronic dictionaries (Engelberg & Lemnitzer, 2009, p. 166). In contrast to HTML, XML does not specify how the content should be displayed, i.e. XML being a descriptive markup language does not specify how text should be formatted or displayed, instead it describes the role the units of information play in the text. In the current project the styling language XSLT (eXtensible Stylesheet Language) can be used to format the text. 4.2.1 Macrostructure Normally, dictionaries imply semasiological order for the macrostructure, however such organization does not reflect conceptually grouped vocabulary. Therefore, the ESSCD will offer a different thematic macrostructure that better reflects the semantic arrangement of the vocabulary. According to the application documents of users, the ESSCD’s macrostructure is composed of 10 conceptual fields, which are subdivided into other conceptual subfields. The macrostructure of the ESSCD will include around 1000 articles. In the frame of this dissertation, seven dictionary articles, namely kayaking , camping , child care , arts and crafts , cook , guitar and dance from the field “Skills” were elaborated. 4.2.2 Microstructure Microstructure refers to the hierarchical internal structure of a dictionary article. The microstructure of online dictionaries clearly differs from the microstructure of print ones, i.e. the user can select whether to display or hide specific article items (Engelberg & Lemnitzer, 2009). Engelberg and Lemnitzer differentiate between the real and virtual microstructure. The latter appears directly on the screen. The former includes all microstructural items that are available for a dictionary article. Data in the ESSCD has to be presented in a maximally comprehensible form, therefore the main task is to find a way to present dictionary articles so that the target audience can retrieve information as fast as possible. The study of Tono (2010) revealed that if the microstructure differs from the users’ expectations, they became confused and perform a slow search. Lew’s study (2004) with Polish learners of English as a Foreign Language has proved that a microstructure, which is too rich, might cause information overload. When the user is overwhelmed with information, this leads to information stress (Bergenholtz
82 & Bothma, 2011, p. 55). Hence, in the ESSCD we use menus to enable users to select the item providing the data that they need. The headword level contains information on spelling, pronunciation, grammatical characteristics of the headword and its translation. The collocations level includes all the collocations grouped by patterns (noun, verb, adjective etc. as collocates). The example level illustrates collocation entities within corpus examples. The last level corresponding to the notes section provides encyclopedic information related to the headword. Furthermore, the ESSCD should allow the users to choose the interface language, either English or Spanish. The ESSCD incorporates typographical structural indicators, i.e. it styles lemmata in bold, collocations within examples in italics, makes use of font types and colors. 4.2.3 Outer Texts In terms of printed dictionaries, “[d]ictionary users are known to allocate little time to the study of these prefatory matters” (Busane, 1990, p. 28). Similarly, the first two international studies on online dictionary use prove that the detailed information on searching and navigating is not essential for most users. Only 4,4% of informants reported that the presence of a clear introduction is important (MullerSpitzer & Koplenig, 2014). The components beyond the dictionary articles are known as outer texts in printed dictionaries. To distinguish commonly accepted outer texts from their print counterparts, Klosa and Gouws (2015, p. 148) suggest the term outer features for online dictionaries since “ the elements presented in the outer domain of online dictionaries do not all belong to the broad category of texts”. With regard to the ESSCD’s genuine purpose and its functions, the dictionary will provide the following outer features: User guidelines will be found under the link “About the project” that provides information on the dictionary content and search functions. The ESSCD should incorporate hyperlinks from the user guidelines to the corresponding entries. The ESSCD will also include question mark symbols in places where questions may arise. Grammar will be incorporated as mouse-over effect. If users do not understand what is for example a mass noun , they will place the mouse over the grammar information and a brief explanation will appear. Questions and Answers User profile, where users can adjust the user interface according to their needs, add dictionary articles to the list or play some language games, i. e. the ESSCD fulfills edutainment functions. Special attention has been paid to the outer text “Questions and Answers” which consists of:
83 Types of camps . For the users to prepare themselves for a stay in a summer camp, it is essential to know in what type of camp they are going to work. The majority of camps are traditional ones. However, many applicants are placed in other camp types such as religious camps, girl guide camps, wilderness camps, underprivileged camps (also called camps for disadvantaged children), special needs camps (also called camps for children with physical or mental disabilities), agency summer camps, day camps etc. (for the brief camp description see Appendix 14). List of the documents for the J-1 visa* . This section contains an asterisk because the politics of the US department of State may change. The obligatory documents for all the nationalities to obtain this visa type are DS-2019 Form and SEVIS 61 Receipt. The former document is issued by a visa sponsor that verifies the participant’s eligibility to take part in cultural exchange programs, summer camp programs among them. Paperwork at camp . Once the applicants arrive at the camp, they need to obtain a Social Security Number and fill in further forms, e.g. W-4 which is the tax form. By the way, applicants can claim their taxes back, as long as they are not US citizens. What should I pack? The applicants should have their passport with visa and the DS-2019 Form, otherwise they will not enter the USA. In terms of suitcase packing, we recommend applicants take the following: copies of all documents, pairs of underwear and socks, warm sweater, bathing suits, Flip Flops, t-shirts, trainers or walking shoes, jeans, rain coat, toothbrush and toothpaste, deodorant. These recommendations are very useful for young people who go to the camp for the first time. Pre-departure orientation should contain presentations or videos about the potential cultural differences and camp life in detail. The Questions and Answers section will also provide answers to questions which users have asked in a forum. 4.2.3 Search Options and Navigation structures. Access Structures of the ESSCD The search engine helps users to retrieve the data from the database. The ESSCD should provide four search options. The macrostructure is an elementary access structure in the dictionary (Kunze & Lemnitzer, 2007), so users can navigate the dictionary using mainly a conceptually-oriented access 61 SEVIS stands for Student Exchange Visitor Information System, where the US Department of State controls the process of all J-1 visa applicants.
84 structure (see Figure 16). In addition, the ESSCD should offer a semasiological A-Z search. The indexbased search will provide clickable letter sections. Users need to click on the initial letter and then further navigate the target letter section. The third search option enables users to type their query directly in the search field. The ESSCD search engine should include the feature corresponding to automatic term completion from an index. This function becomes active after the user has typed in a minimum of four characters. Fuzzy-spelling search will be implemented as the Did you mean? technique, i. e. the search results will show suggestions for a query that was misspelled. The last search option of the ESSCD should incorporate internal and external cross-referencing. Thanks to the former feature any word within an entry may be looked up in the dictionary immediately. Such crossreferences such as synonyms, antonyms or superordinate terms are considered to be helpful for the users. Internal references will not only refer to other entries, but also lead to dictionary components outside the dictionary articles. For example, if a user would like to know what a noun phrase is, he/she can access the corresponding grammatical information with just one click. Moreover, each article of the ESSCD includes the field(s) and subfield(s) to which the lemma belongs to. By clicking on the respective (sub)field, the user has access to a list of all the lemmata belonging to the same (sub)field. Figure 16: Draft of the ESSCD start page
85 4.3 Preliminary Conclusion The ESSCD database has been annotated with the Markup Language XML developed to represent texts and documents in electronic format. The Text Encoding Initiative (TEI) Guidelines follow the syntactic rules of XML and propose tags for the electronic encoding of a wide variety of text types including dictionaries. Chapter 9 of the Guidelines proposes tags for encoding lexicographic information. However, the TEI Guidelines for lexicographic resources focus primarily on the encoding of print dictionaries. This means that the TEI Guidelines present some lacunae when it comes to encoding born digital lexicographic materials and products. Therefore, during the database creation process, new lexicographical elements were included into “the relational hierarchy” (Nielsen & Almind, 2011, p. 143). During the encoding process some elements were modified or added. In order to save consultation time, the ESSCD’s microstructure should be organized clearly and provide various search options.
86 CHAPTER 5 DATA CATEGORIES OF THE ESSCD ARTICLES When it comes to the user needs, Bergenholtz and Tarp (1995, p. 24) indicate the following required types of information to support foreign-language text production in specialized dictionaries : (1) orthography, gender, pronunciation, irregularity, collocations, usage information; (2) standard label (DIN, ISO etc.), field label or brief explanation. Taking into consideration the translation from foreign language , the following information types will be needed (Bergenholtz & Tarp, 1995, p. 24): (1) on the foreign language: word class, gender, pronunciation, collocations, irregularity; (2) on the native language: orthography, gender, pronunciation, irregularity, collocations, usage information. The ESSCD follows the suggestions of Bergenholtz and Tarp (1995) 62 but also takes into account the ESSCD user profile and their specific needs. Since the ESSCD is a monodirectional dictionary, it is unlikely that the Spanish native speaker will need information on Spanish pronunciation. This section describes the entry items of the ESSCD and its organization. The main task is to select which information categories to be included. According to Wiegand, Feinauer and Gouws (2013, p. 343), a dictionary article consists of two structural components, called comments, namely a comment on form and a comment on semantics . 5.1 Comment on Form The comment section on the form of the ESSCD articles includes orthography, pronunciation and grammatical data on the word class and inflection. 5.1.1 Orthography The spelling variants of the lemma, such as go-cart and go-kart ; tie-dye and tye dye, tennis racket and tennis racquet etc., are included into the ESSCD: <form type="lemma"> <orth>go-cart</orth> or <orth>go-kart</orth> </form> The spelling form which is proposed by the IENA application appears first followed by the spelling variant. Spelling variants sometimes belong to different varieties of English language. An example is counselor for American English and counsellor for British English. The ESSCD includes American forms 62 More recent research on the Function Theory of Lexicography (Fuertes-Olivera & Tarp, 2014, p. 101) proposes the following data to be included into the dictionary article of online specialized dictionaries: collocations/word combinations; example sentences; contextual definitions and background text.
87 as lemmata and British ones as spelling variants. The British spelling variants have corresponding indications: <form type="lemma"> <orth>couselor</orth> <form> <usg type="geo">UK</usg> <orth>counsellor</orth> </form> <pron notation="ipa"> <pc>/</pc>ˈkaun(t)-s(ə-)lər <pc>/</pc> <ptr type="audioFile" target="counselor" title="listen to pronunciation"/></pron> </form> 5.1.2 Pronunciation Data The ESSCD includes both phonetic transcription according to IPA and the audio representation. Lew (2013) determines two principal benefits of keeping the graphic transcription in online dictionaries. The first one is clarity i.e. the language learners are unlikely to perceive all the phonological details at once as they do it from the phonetic system of the mother tongue. Another significant advantage is the indexical function e.g. IPA transcriptions are based on commonly accepted criteria/rules. As can be seen from the element <form> above, the phonetic transcription in the ESSCD is placed between slashes. The audio pronunciation of the ESSCD should be recorded by native English speakers. The pronunciation as a data category is included into the ESSCD because it serves for text production purposes during the interview with recruiters and camp directors. 5.1.3 Grammatical Data Grammatical data of the ESSCD contains, as a rule, the word class and verb inflection. Since many nouns and verbs in the English language have the same orthographic representation, the word class also designates the meaning, thus contributing to text reception. Most of the dictionaries encode concomitant metalinguistic information such as abbreviations, codes or symbols. This causes comprehension problems among the users. The aim of the ESSCD is to avoid contracted forms as the ones above-mentioned and to make the dictionary use as transparent as possible. Therefore, the ESSCD offers worded grammar, indicating the part-of-speech in its full form e.g. “noun” instead of “n”. Recent studies with the participation of 117 Dutch and 606 Polish learners of English (Bogaards & van der Kloot, 2002; Dziemianko, 2012) also reveal that the majority of informants prefer worded grammar patterns.
88 Some lemmata contain subcategorization information. For example, wakeboarding is a mass noun, next of kin – a noun phrase, arts and crafts – a plural nominal word combination: <gramGrp> <pos>noun</pos> <subc>mass</subc> </gramGrp> The majority of the ESSCD lemmata are nouns 63 as the dictionary is organized conceptually. However, if the lemma sign for a verb and for a noun is the same, the ESSCD provides explicit word forms for verb, for example tie-dye : <gramGrp> <pos>verb</pos> <tns type="present_simple">tie-dyes</tns> <tns type="present_participle">tie-dyeing</tns> <tns type="past_simple">tie-dyed</tns> <tns type="past_participle">tie-dyed</tns> </gramGrp> Bergenholtz and Tarp (1995, p. 51) recommend that “specialised dictionaries intended for translation from the foreign language into the native language should provide, in addition to relevant collocations, a minimum of grammatical information on the native language” 64 . The ESSCD includes the word class and the gender of Spanish equivalents. 5.2 Comment on Semantics The comment section on semantics of the ESSCD articles consists of translation equivalents, collocations, examples and notes, which will be discussed in the following sections. The mentioned data categories help users to decode the meaning of the lemma and indicate its usage. 5.2.1 Translation Equivalents The selection of equivalents for lemmata with cross-cultural differences poses a serious challenge for lexicographers. Concepts in one language do not always have their counterparts in another. Bergenholtz and Tarp (1995) propose two solutions for the lack of target language equivalents: paraphrase and partial equivalence. Regarding the latter option, the difference between the lemma and partial equivalent must be explained. The ESSCD also contains headwords which have zero correspondence in 63 L’ Homme (2003) claims that specialized dictionaries usually contain nouns that define activities or processes and leave out the corresponding verbs. 64 Emphasis in original.
95 table once again demonstrates the difference between American and British English, for example, room and board vs. full board and illustrates the presence of this variety in the self-compiled corpus. 5.2.2 Collocations You shall know a word by the company it keeps. Firth, 1957, p. 11 There is no commonly-agreed definition for the term “collocation”. In this thesis we use it to refer also to multi-word keywords and terms that were extracted from the self-compiled corpus. Regarding the collocation structure, Hausmann distinguishes between the base and the collocate: “Die Basis ist ein Wort, das ohne Kontext definiert, gelernt und übersetzt werden kann. Der Kollokator ist ein Wort, das beim Formulieren in Abhängigkeit von der Basis gewählt wird und das folglich nicht ohne die Basis definiert, gelernt und übersetzt werden kann (Hausmann, 2007, p. 218). Bilingual dictionaries are criticized for including redundant information, which users do not need, as well as for the lack of collocations (Atkins, 1996). Mastering collocations is one of the greatest challenges in second language learning. From the point of view of dictionary research on collocations, various studies regarding collocation search and users’ capability of applying the information in production tasks were carried out (e.g. Laufer & Waldman, 2011; Nesselhauf, 2005). In particular, their results revealed that more than 50% of the errors are caused by the influence of the native language, i.e are interlingual. Therefore, one of the main purposes of the ESSCD is to provide an exhaustive collocation list and correct collocation equivalents. The main selection criteria for collocations was their relevance to the subject of summer camps, thus child care credit , parental child care are not relevant terms in summer camp domain. The corpora which we have searched, enrich each other. For example, for the lemma kayaking , the English Web 2015 proposes go kayaking . On the other hand, the COCA corpus suggests types of kayaking depending on its venue e. g. sea kayaking , ocean kayaking , river kayaking , whitewater kayaking . General language corpora offer a wide range of collocates, therefore we had to select which ones are relevant for the ESSCD. 5.2.2.1 Collocations from the Self-compiled Corpus Various statistical approaches and association measures determine collocations within the corpus. Corpus collocation extraction goes back to the 1980s (Sinclair, 1991). However, at that time the corpus size 10-50 million tokens was too small to retrieve useful examples (Geyken & Lemnitzer, 2016, pp. 221-222). How large should a corpus be to enable the extraction of representative collocations? 85
96 million words were used for the Oxford Learner’s Dictionary of Academic English 79 . Another problem that we still have to tackle today is insufficient accuracy. The above-mentioned term “collocation” is set in a wide framework and the extraction tools frequently obtain “weak” word combinations such as a new shirt (Geyken & Lemnitzer, 2016, p. 222). Therefore, it is helpful to use association measures. Keywords and terms extracted from the self-compiled corpus with higher mutual information 80 are likely to better represent the summer camp domain, e.g. activity specialists (MI 8,7) vs. intake specialist (MI 14), age group (MI 5,5) vs. cabin group (MI 5,7). Age group is a common collocation, while cabin group relates directly to summer camp life. Here are some interesting collocations extracted from the English subcorpus: sleeping bag , intersecting chords , tennis court , canoe trip , day session , camp season , staff member , family camp , camp spirit , camp experience , leadership skills , camp staff , senior staff , former camper , personal growth , camp placement , etc. The last term refers to the hiring process, i.e. if the candidate is hired, it means that he/she is placed in a camp. An illustrative word combination is hacky sacks that refers to a game played with footbag, which is filled with sand, rice or other materials. An interesting example taken from the output of the candidate term extraction process is t-shirts tanks . Even though the expression itself does not exist, it leads us to two synonyms sleeveless shirt and tank top , “named after tank suits, onepiece bathing suits of the 1920s worn in tanks or swimming pools” 81 . An additional incorrect compound is shield cage. Used separately, hockey shield or hockey cage refer to a helmet for hockey. In nearly every camp, campers go canoeing and kayaking. For their personal belongings not to get wet, they use dry bags , a type of container that protects things from getting wet. An interesting multi-word expression found in the English subcorpus is staff-to-camper ratio that refers to the number of counselors divided by the number of campers, e.g. 1:2. A camp trunk or foot locker are primarily used to store the belongings and “become a dresser, stepstool, seat and table for [...] camper” 82 . Another term found in the self-compiled corpus, in loco parent is a legal term meaning “in place of the parent”: “[w]hile state laws vary, camp professionals generally serve in loco parentis (in place of the parent)” (see American Camp Association). An additional term that is worth mentioning is seasonal staff . Only few staff members work in camp for the whole year. The majority of camp staff is seasonal, i.e. works during the camp season. One more identified keyword is a polar bear swim / polar bear dip that refer to “early 79 Oxford Learner’s Dictionary of Academic English . Retrieved from https://www.oxforddictionaries.com/oldae 80 Mutual information is computed in Cygwin terminal using the script ngrams.py . 81 Tank top . In The Free Dictionary . Retrieved from http://www.thefreedictionary.com/tank+top 82 Camp Pathfinder. See in English-Spanish Comparable Corpus Sources.
97 morning swims, a camp activity where you jump in the lake in the early morning before breakfast” 83 ; “a refreshing optional dip before each breakfast” 84 . Lists of keywords and terms extracted from the English subcorpus were transformed into a network (see Figure 17) using Gephi software 85 . Once collocations were extracted, the data was formatted in one CSV file for the importation to Gephi’s “data laboratory” (see Tables 7, 8). Gephi constructs a network composed of edges and nodes. Nodes are collocation bases and edges are connections between the base and its collocate. At the beginning of the process, the information in the graph appears in random order. Therefore, the Modularity Algorithm implemented in Gephi was applied. It attributes nodes to clusters of highly interconnected nodes. In Figure 17, few collocation pairs are “isolated” such as dining hall , sleeping bag , while the majority belong to clusters, composed of pairs in which one collocation element is connected to other word(s). Part-of-speech (PoS) category was included into the database because the network can be further elaborated with adjectives and verb collocations. Other categories, e.g. association or statistical measures like mutual information or log-likelihood can be employed. ID Label PoS Frequency camp_NN Camp NN 15 326 summer_NN Summer NN 4 008 staff_NN Staff NN 3 818 Table 7: Collocation data for import to Gephi as nodes table 86 Source Target Type Frequency summer_NN camp_NN Direct 1 270 camp_NN staff_NN Direct 457 camp_NN employment_NN Direct 260 Table 8: Collocation data for import to Gephi as edges table The next network, SkillNet , illustrates collocates of the base skill . The graph is based on the results of the self-compiled corpus and includes 34 nodes (see Figures 18 and 19). Each node has its value, which corresponds to the frequency of the word in the corpus. The edges’ value equals the frequency with which the base appears with its collocate. It was possible to include more nodes, however the graph would look overloaded. The color scheme was applied to distinguish between word classes: light 83 NYQUEST Camp Canada. See in English-Spanish Comparable Corpus Sources. 84 Camp Pathfinder. See in English-Spanish Comparable Corpus Sources. 85 Brett (2017) uses Gephi in the paper Collocate networks in the language of crime journalism . 86 See Appendices 15, 16 for the full database of collocation networks.
98 green for adjectives ( outdoor skills ), light blue for verbs ( to master skills , to demonstrate skills ) and red for nouns ( skill acquisition , communication skills ). As can be observed from Figures 18 and 20, the thicker the edge is, the more frequent the words occur together, for example, skill development , leadership skills , social skills . The Fruchterman Reingold layout algorithm and the Expansion layout algorithm have been used to make the appearance of SkillNet user-friendly for retrieval. Some properties of the mentioned layouts, namely the speed and scale factor have been manipulated before running the algorithms. The edges were ranked by weight. The nodes were ranked by degree and value. We consider the use of such visualizations particularly important in terms of the current study. They not only inform about the typical collocations of summer camps but also demonstrate the strength between the base and its collocate as well as emphasize the connection between collocations.
99 color of the node frequency of the base Figure 17: Bigrams extracted from the English part of the self-compiled corpus with frequency 35 and above*. * Networks in Figures 17, 18, 19 and 20 were created using Gephi by the author of the thesis.
100 Figure 18: SkillNet Figure 19: Expansion of the central part of the SkillNet Figure 20: SkillNet. Edges visualization.
101 5.2.3 Example Sentences 5.2.3.1 Corpora vs. Lexicographer-made Examples It is commonly assumed that large corpora are the most representative sources for dictionary examples because they include the language of native speakers and illustrate real life communication, i.e. are authentic. On the one hand, lexicographer-made examples are criticized as isolated and overloaded with information (Fox, 1987). On the other hand, lexicographers have linguistic intuition and are also native speakers of a particular language. Furthermore, the “artificial” examples might be preferred in learner’s dictionaries (Laufer, 1992). Some studies have proven that neither language teachers nor native speakers were able to distinguish between “natural and made-up examples” (Rundell & Maingay, 1990). Similarly, Laufer (1992) examined whether corpus examples contribute to better comprehension by learners compared to example sentences written by lexicographers. The experiment has shown that lexicographers’ examples are remarkably more helpful to understand new words and slightly more useful for production purposes. The advocate of corpus examples is Sinclair, who also describes the role of corpora for the compilation of the Collins Cobuild Dictionary (1987). Other lexicographers may state that the lexicographers’ examples are focused on the usage point (sense item) that should be illustrated and authentic examples only confuse the user (Engelberg & Lemnitzer, 2009, p. 236). Rundell (2010) suggests linking corpora to online dictionaries 87 . On the one hand, corpora allow users to search authentic examples. On the other hand, it could lead to information overload as the task of the lexicographer is to process the corpus data, evaluate it and select the best examples. Moreover, the target group of the ESSCD is not skilled at corpus searching. As Bothma (2011) and Tarp (2011) claim, data should be repackaged and adopted to the specific information needs. 5.2.3.2 Corpus Examples in the ESSCD ESSCD’s examples aim to assist users with text production and collocation use. We have followed four main principles in selecting corpus examples. First of all, they need to illustrate the use of collocations in a sentence. Secondly, if the example is too long, the part of the example which does not contain the collocation is omitted in a way that the sense of the sentence is not lost. We believe that too long examples increase information costs 88 . As a rule, examples contain bibliographic information with author, source and date. However, not all concordances in corpora are fully annotated with resource details. Moreover, web-based corpora include only links to the web pages which often cannot be accessed as 87 Base lexicale du français offers access to such an external database and gives the user the possibility to select the corpus. For more about the integration of online dictionaries and corpora, see Heid, Prinsloo and Bothma (2012). 88 Nielsen and Fuertes-Olivera (2013) distinguish between search-related information costs (the efforts which the user makes to look up the needed information) and comprehension-related information costs (the efforts the user makes to understand the data found).
102 the web site was removed. In these cases, the ESSCD does not include bibliographic information. Last but not least, examples should preferably contain encyclopedic information. For instance, for the article guitar: Then, as you tune the guitar up for the first time, give the strings a good tug to help stretch them in. Until one learns to tune their guitar manually they can always use the principle of beats to tune their stringed instrument. One day he picked up an old classical guitar he'd bought at a garage sale, tuned it up, and played a few chords. The examples give instructions on how to prepare a guitar before playing and explain the differences between guitar types. Ideally examples describe camp life: Campers will participate in arts and crafts, sports, organized games, field trips and much more. By default 5 examples will be shown, users have an option to click on the plus sign, thereby expanding the listed examples. 5.2.4 The Notes Section The notes section usually contains the general tasks and duties of a particular working position, for example, those of a dance instructor. Moreover, notes sometimes include idioms and the equipment related to a certain activity such as camping supplies . The note tab might offer action verbs , e.g. for the lemma cook , it lists stir , simmer , chop , etc. Notes can be also designed for error prevention clarifying the use of easily-confused words. Encyclopedic information will be given in the notes section as well. The Did you know that? feature will offer interesting etymological facts of some entries, for example, lacrosse goes back to Canadian French la crosse , that means the crooked stick 89 . The term go-carts underwent a semantic change from the baby walker of the 17th century to nowadays a small racing car 90 . 5.3 Preliminary Conclusions The dictionary is an utility tool, hence the main aim of the ESSCD is to ensure that the user obtains information in a straightforward way: “[th]e perfect dictionary is one in which you can find the thing you are looking for preferably in the very first place you look” (Haas, 1962, p. 48). The ESSCD contains data that respond to the needs of the users: a word’s pronunciation, its spelling, basic grammatical 89 See Lacrosse . In Merriam-Webster Dictionary . Retrieved from https://www.merriam-webster.com/dictionary/lacrosse 90 See Go-cart . In Online Etymology Dictionary . Retrieved from https://www.etymonline.com/word/go-cart
103 information, syntactic behavior in a sentence and collocational restrictions. In sum, the ESSCD offers accurate and complete lexicographic data and allows the users to retrieve the exact required information that they need 91 . A good dictionary shall cover the most frequent collocations and illustrative examples with the collocations (Laufer & Waldman, 2011). Examples are selected carefully and contribute to a better understanding of the meaning of the word and illustrate collocational behavior within the sentence. 91 Most lexicographers welcome the possibility of showing exactly the relevant information categories in a particular lookup situation, no less and no more, tailored to the specific needs and skills of the user (Trap-Jensen, 2010, p. 1142).
104 CHAPTER 6 GENERAL CONCLUSION 6.1 Limitations The current project has certainly several limitations and some aspects can be improved. One of them is the corpus size. Due to the time constraints, we limited the information extraction process to 38 camp websites out of the more than 3 600 available camps. During the extraction of the webpage URLs, external links, for example, to foundations, Facebook or Twitter camp pages appeared. This led to some noise in the corpus, fact that reduced the number of representative examples. Another aspect that can be developed is the insertion of labels for regional variety of American and Canadian English. The dictionary compilation requires an immense input of effort and time. It goes without saying that the compilation of the ESSCD should involve experts with subject-field knowledge, skilled and experienced translators, experts with high level of proficiency in the source (English) and target (Spanish) languages, web page programmers and designers. Given that the current lexicographic project is the result of the work developed as part of the Master’s thesis, subject to a limited time span, it was not possible to compile a dictionary manual. Since the ESSCD is not a finished reference work, we have not conducted any research on dictionary use and could not conclude whether the users’ needs are satisfied. It is implied that the indended users of the ESSCD should assess and evaluate the dictionary. 6.2 Discussion on Future Directions Lemmata of dictionary articles that were written in the frame of the current thesis did not pose a serious comprehension problem, hence no definitions were included and the translation equivalents were sufficient for the understanding of the meaning of the headwords. However, the users are likely to lack assistance to understand particular terms, for example tetherball , tye-dye , lacrosse , doughboy 92 . Therefore, such headwords should contain the explanation of meaning in a definition . We suppose that this data category should be included into the article structure. One of the most important future tasks will be to implement novel technologies which will allow users to optimize consultation time. Several lexicographers have already suggested several options. In 1996, Atkins (1996, pp. 13-14) foresees that bilingual dictionaries of the future might operate two modes of information . The equivalence mode assists the user in the performance of “specific tasks such as translation, comprehension or self-expression” and contrast mode “offers ways of contrasting the meaning and syntactic behaviour of chosen words across languages”. Since every user has individual 92 Also called stick biscuit .
111 https://www.researchgate.net/publication/280064849_Old_Needs_New_Solutions_Compara ble_Corpora_for_Language_Professionals Berti, B., & Pinnavaia, L. (2012). Towards a Corpus-driven Italian-English Dictionary of Collocations. In R. Facchinetti (Ed.), English Dictionaries as Cultural Mines (pp. 201-222). Newcastle-UponTyne: Cambridge Scholars Publishing. Berti, B., & Pinnavaia, L. (2014). Creating a Bilingual Italian-English Dictionary of Collocations. In: A. Abel, C. Vettori, & N. Ralli (Eds.), Proceedings of the 16th EURALEX International Congress (pp. 515-524). Bolzano. Retrieved from http://euralex.org/wpcontent/themes/euralex/proceedings/Euralex%202014/euralex_2014_038_p_515.pdf Bogaards, P. (1996). Dictionaries for Learners of English. International Journal of Lexicography , 9 (4), 277-320. Bogaards, P., & van der Kloot, W. (2002). Verb Constructions in Learners’ Dictionaries. In A. Braasch & C. Povlsen (Eds.), Proceedings of the 10th Euralex International Congress , (pp. 747–757) . Copenhagen: CST. Bogaards, P., & van der Kloot, W. (2002). Verb Constructions in Learners’ Dictionaries. In A. Braasch & C. Povlsen (Eds.), Proceedings of the 10th Euralex International Congress , (pp. 747–757) . Copenhagen: CST. Boonmoh, A. (2012). E-dictionary use under the spotlight. Student’s use of pocket electronic dictionaries for writing. Lexikos , 22, 43-68. Retrieved from http://146.232.21.2/pub/article/viewFile/997/514 Boonmoh, A., & Nesi, H. (2008). A survey of dictionary use by Thai university staff and students, with special reference to pocket electronic dictionaries. Horizontes de Lingüística Aplicada , 6 (2), 79-90. Bothma, T.J.D. (2011). Filtering and adapting data and information in an online environment in response to user needs. In P. A. Fuertes-Olivera & H. Bergenholtz (Eds .), e-Lexicography: The Internet, Digital Initiatives and Lexicography (pp. 71-102). London, New York: Continuum. Bothma, T.J.D. (2011): Filtering and adapting data and information in an online environment in response to user needs. In: Fuertes-Olivera P.A., Bergenholtz H. E-Lexicography: The Internet, Digital Initiatives and Lexicography . London/New Delhi/New York/Sydney: Bloomsbury Academic, 71-102. Brett, D. (2017). Collocate networks in the language of crime journalism. Studii de lingvistica , 7, 125144. Retrieved from http://studiidelingvistica.uoradea.ro/docs/7-2017/pdf_uri/Brett.pdf.
112 Busane, M. (1990). Lexicography in Central Africa: The user perspective, with special reference to Zaïre. In R. R. K. Hartmann (Ed.), Lexicography in Africa. Progress reports from the Dictionary Research Center Workshop at Exeter (pp. 19-35). Exeter, England: University of Exeter Press. Carr, M. (1997). Internet Dictionaries and Lexicography. International Journal of Lexicography , 10 (3), 209-230. Chen, Y. (2010). Dictionary use in EFL learning. A contrastive study of pocket electronic dictionaries and paper dictionaries. International Journal of Lexicography , 23 (3), 275-306. De Schryver, G. M. (2003). Lexicographers’ Dreams in the Electronic-Dictionary Age. International Journal of Lexicography , 16 (2), 143-199. Dengler, K., Matthes, B., & Paulus W. (2014). Berufliche Tasks auf dem deutschen Arbeitsmarkt. Eine alternative Messung auf Basis einer Expertendatenbank. FDZ-Methodenreport. Methodische Aspekte zu Arbeitsmarktdaten , 12. Retrieved from http://doku.iab.de/fdz/reporte/2014/MR_12-14.pdf Dictionaries. In Text Encoding Initiative (TEI) Guidelines. Retrieved from http://www.teic.org/release/doc/tei-p5-doc/en/html/DI.html Dziemianko, A. (2012). Noun and Verb Codes in English Monolingual Dictionaries for Foreign Learners: A Study of Usefulness in the Polish Context . Poznań: Adam Mickiewicz University Press. Dziemianko, A. (2012a). On the use(fulness) of paper and electronic dictionaries. In S. Granger & M. Paquot (Eds.), Electronic lexicography (pp. 320-341). Oxford: Oxford University Press. Engelberg, S., & Lemnitzer, L. (2009). Lexikographie und Wörterbuchbenutzung (4th ed.). Tübingen: Stauffenburg. Engelberg, S., & Storrer, A. (2016). Typologie von Internetwörterbüchern und -portalen. In A. Klosa & C. Müller-Spitzer (Eds.), Internetlexikografie. Ein Kompendium (pp. 31-63). Berlin, Boston: de Gruyter. Engelberg, S., Müller-Spitzer C., & Schmidt, T. (2016). Vernetzungsund Zugriffsstrukturen. In A. Klosa & C. Müller-Spitzer (Eds.), Internetlexikografie. Ein Kompendium (pp. 153-195). Berlin, Boston: de Gruyter. Firth, J. R. (1957). Papers in Linguistics 1934-1951 . Oxford: Oxford University Press. Fox, G. (1987). The case for examples. In J. Sinclair (Ed.), Looking Up. An account of the COBUlLD project in lexical computing , (pp. 137-19). London, Glasgow: Collins. Frankenberg-Garcia, A. (2015). Dictionaries and encoding examples to support language production. International Journal of Lexicography , 24 (4), 490-512.
113 Fuertes-Olivera, P. A. (2013). E-lexicography: The continuing challenge of applying new technology to dictionary making. In H. Jackson (Ed.), Bloomsbury Companion to Lexicography (pp. 323-340). London: Bloomsbury. Fuertes-Olivera, P. A., & Tarp, S. (2014). Theory and Practice of Specialised Online Dictionaries: Lexicography versus Terminography. Berlin, Boston: de Gruyter. Geyken, A., & Lemnitzer L. (2016). Automatische Gewinnung von lexikographischen Angaben. In A. Klosa & C. Müller-Spitzer (Eds.), Internetlexikografie. Ein Kompendium (pp. 197-247). Berlin, Boston: de Gruyter. Gouws, R. (2014). Article Structures: Moving from Printed to e-Dictionaries. Lexikos , 24, 155-177. Retrieved from http://lexikos.journals.ac.za/pub/article/view/1256 Haas, M. R. (1962). What belongs in a bilingual dictionary? F. W. Householder & S. Saporta (Eds.), Problems in Lexicography , (pp. 45-50) . Bloomington: Indiana University. Hanks, P. (2012). Corpus evidence and electronic lexicography. In S. Granger & M. Paquot (Eds.), Electronic lexicography (pp. 57-82). Oxford: Oxford University Press. Retrieved from http://www.patrickhanks.com/uploads/5/1/4/9/5149363/hanks_2012f.pdf Hartmann, R. R. K. (2005). Onomasiological Dictionaries in 20th-Century Europe. Lexicographica. International Annual for Lexicography , 21, 6-19. Hartmann, R. R. K. (2005a). Pure or hybrid? The development of mixed genres. Facta Universitatis, 3, 193-208. Retrieved from http://facta.junis.ni.ac.rs/lal/lal2005/lal2005-06.pdf Hausmann, F. J. (1984). Wortschatzlernen ist Kollokationslernen. Zum Lehren und Lernen französischer Wortverbindungen. Praxis des neusprachlichen Unterrichts , 31, 385-406. Hausmann, F. J. (1989). Wörterbuchtypologie. In F. J. Hausmann, O. Reichmann, H. E. Wiegand, L. Zgusta (Eds.), Dictionaries: An International Encyclopedia of Lexicography (vol. 1, pp. 968-981). Berlin, New York: de Gruyter. Hausmann, F. J. (2007). Die Kollokationen im Rahmen der Phraseologie – Systematische und historische Darstellung. Zeitschrift Für Anglistik und Amerikanistik , 55, 217-234. Retrieved from http://www.zaa.uni-tuebingen.de/wp-content/uploads/05-Hausmann_ffinal.pdf Heid, U., Prinsloo, D. J., & Bothma, T. J. D. (2012). Dictionary and corpus data in a common portal: state of the art and requirements for the future. Lexicographica , 28, 269-291. Heid, U., Prinsloo, D. J., & Bothma, T. J. D. (2012). Dictionary and corpus data in a common portal: state of the art and requirements for the future. Lexicographica , 28, 269-291.
114 Herbst, T., & Stein G. (1987). Dictionary-Using Skills: a plea for a new orientation in language teaching. In A. P. Cowie (Ed.), The Dictionary and the Language Learner (pp. 115-127). Tübingen: Niemeyer. Hildenbrandt, V., & Klosa, A. (Eds.). (2016). Lexikographische Prozesse bei Internetwörterbüchern. Opal – Online publizierte Arbeiten zur Linguistik, 1. Retrieved from https://ids-pub.bszbw.de/frontdoor/deliver/index/docId/4730/file/Hildenbrandt_Klosa_Lexikographische_Prozes se_bei_Internetw%C3%B6rterb%C3%BCchern_2016.pdf Hüllen, W. (1994). The world in a list of words . Tübingen: Niemeyer. Hüllen, W. (1999). English Dictionaries 800–1700: The Topical Tradition . Oxford: Clarendon. Jackson, H. (2002). Lexicography: An Introduction. London: Routledge. Janser, M. (2018). The greening of jobs in Germany. First evidence from a text mining based index and employment register data. IAB-Discussion Paper , 14, 1-76. Retrieved from http://doku.iab.de/discussionpapers/2018/dp1418.pdf Kemmer, K. (2010). Onlinewörterbücher in der Wörterbuchkritik. Ein Evaluationsraster mit 39 Beurteilungskriterien. In OPAL – Online publizierte Arbeiten zur Linguistik, 2, 1-33. Kenning, M. (2010). What are parallel and comparable corpora and how can we use them? In M. McCarthy & A. O'Keeffe (Eds.), The Routledge Handbook of Corpus Linguistics , (pp. 487-501). Abingdon: Routledge. Kilgarriff, A. (2013). Using corpora as data source for dictionaries. In H. Jackson (Ed.), The Bloomsbury Companion to Lexicography (pp. 77-96). London: Bloomsbury. Kipfer, B. A. (1986). Investigating an onomasiological approach to dictionary material. Dictionaries: Journal of the Dictionary Society of North America, 8, 55-64. Kirkpatrick, B. (1989). User’s Guides in Dictionaries. In F. J. Hausmann, O. Reichmann, H. E. Wiegand, L. Zgusta (Eds.), Dictionaries: An International Encyclopedia of Lexicography (vol. 1, pp. 754761). Berlin, New York: de Gruyter. Klosa, A. (2013). The lexicographical process (with special focus on online dictionaries). In R. H. Gouws, U. Heid, W. Schweickard & H. E. Wiegand (Eds.). Dictionaries. An International Encyclopedia of Lexicography. Supplementary Volume. Recent developments with focus on electronic and computational lexicography (pp. 517-524). Berlin, Boston: de Gruyter. Klosa, A., & Müller-Spitzer, C. (Eds.). (2016). Internetlexikografie. Ein Kompendium. Unter Mitarbeit von Martin Loder. Berlin, Boston: de Gruyter.
115 Klosa, A., & Rufus, G. (2015). Outer features in e-dictionaries. Lexicographica , 31, 142-172. Retrieved from https://core.ac.uk/download/pdf/83653976.pdf Koplenig, A., & Müller-Spitzer, C. (2014). General issues of online dictionary use. In C. Müller-Spitzer (Ed.). Using Online Dictionaries (pp. 127-141). Berlin, Boston: de Gruyter. Kühn, P. (1989). Typologie der Wörterbücher nach Benutzungsmöglichkeiten. In F. J. Hausmann, O. Reichmann, H. E. Wiegand, L. Zgusta (Eds.), Dictionaries: An International Encyclopedia of Lexicography (vol. 1, pp. 111-127). Berlin, New York: de Gruyter. Kunze, C., & Lemnitzer, L. (2007). Computerlexikographie: Eine Einführung. Gunter Narr Verlag: Tübingen. L’ Homme, M. (2003). Capturing the lexical structure in special subject fields with verbs and verbal derivatives. A model for spacilized lexicography. In International Journal of Lexicography 16 (4), 403-421. Laffling, J. (1992). On constructing a transfer dictionary for man and machine. Target , 4(1), 17–31. Landau S. (2001). Dictionaries. The Art and Craft of Lexicography. Cambridge: Cambridge University Press. Laufer B., & Melamed L. (1994). Monolingual, bilingual and ‘bilingualised’ dictionaries: which are more effective, for what and for whom? In W. Martin, W. Meijs, M. Moerland, E. ten Pas, P. van Sterkenburg, & P. Vossen (Eds.), Proceedings of the 6th Euralex International Congress (565576). Amsterdam: Euralex. Laufer, B. & Kimmel, M. (1997). Bilingualised dictionaries: how learners really use them. System , 25 (3), 361-369. Retrieved from https://www.academia.edu/10246535/Bilingualised_dictionaries_how_learners_really_use_t hem Laufer, B. (1992). Corpus-based versus lexicographer examples in comprehension and production of new words. In H. Tommola, K. Varantola, T. Salmi-Tolonen, & J. Schopp (Eds.), Euralex ’92 Proceedings (pp. 71-76). Retrieved from http://www.euralex.org/elx_proceedings/Euralex1992_1/013_Batia%20Laufer%20-Corpusbased%20versus%20lexicographer%20examples%20in%20comprehension%20and%20productio n%20of%20n.pdf Laufer, B., & Waldman, T. (2011). Verb-noun Collocations in Second Language Writing: A Corpus Analysis of Learners’ English. Language Learning , 61, 647-672.
116 Lehr, A. (1996). Zur neuen Lexicographica-Rubric “Electronic dictionaries”. Lexicographica, 12, 310317. Lew, R. (2004). Which dictionary for whom? Receptive use of bilingual, monolingual and semi-bilingual dictionaries by Polish learners of English . Poznań: Motivex. Lew, R. (2011). Online Dictionaries of English. In P. A. Fuertes-Olivera & H. Bergenholtz (Eds .), eLexicography: The Internet, Digital Initiatives and Lexicography (pp. 230-250). London, New York: Continuum. Lew, R. (2011a). User studies: Opportunities and limitations. In A. Akasu & S. Uchida (Eds.), Lexicography: Theoretical and practical perspectives. Papers submitted to the seventh ASIALEX Biennial International Conference (pp.7-16). Kyoto: The Asian Association for Lexicography. Lew, R. (2012). How can we make electronic dictionaries more effective? In S. Granger & M. Paquot (Eds.), Electronic lexicography (pp. 343-361). Oxford: Oxford University Press. Lew, R. (2013). Space restrictions in paper and electronic dictionaries and their implications for the design of production dictionaries. Retrieved from https://repozytorium.amu.edu.pl/bitstream/10593/799/1/Lew_space_restrictions_in_paper _and_electronic_dictionaries.pdf Lew, R. (2013a). Online dictionary skills. In K. Iztok, J. Kallas, P. Gantar, S. Krek, M. Langemets & M. Tuulik (Eds.), Electronic lexicography in the 21st century: Thinking outside the paper. Proceedings of the eLex 2013 conference (pp. 16-31). Tallinn, Estonia: Trojina, Institute for Applied Slovene Studies. Retrieved from http://eki.ee/elex2013/proceedings/eLex2013_02_Lew.pdf Lew, R. (2014). User-generated content (UGC) in English online dictionaries. Opal – Online publizierte Arbeiten zur Linguistik , 4, 8-26. Retrieved from https://pub.idsmannheim.de//laufend/opal/pdf/opal2014-4.pdf Lew, R. (2016). Can a dictionary help you write better? A user study of an active bilingual dictionary for Polish learners of English. International Journal of Lexicography , 29 (3), 353–366. Losey-León, M. A. (2015). Corpus-based Contrastive Analysis of Keywords and Collocations across Sister Specialized Subcorpora in the Maritime Transport Field . Procedia - Social and Behavioral Sciences. 198, 526-534. Maingay, S., and Rundell, M. (1990). What makes a good dictionary example? Paper presented at the 24th LATEFL Conference . Dublin.
117 Maingay, S., and Rundell, M. (1990). What makes a good dictionary example? Paper presented at the 24th LATEFL Conference . Dublin. McArthur, T. (1986). Worlds of Reference . Lexicography, Learning and Language from the Clay Tablet to the Computer. Cambridge: Cambridge University Press. McArthur, T. (1998). Living Words: Language, Lexicography and the Knowledge Revolution . Exeter: Exeter University Press. Meliss, M., & Sánchez Hernández, P. (2015) Theoretical and methodological foundations of the DICONALE project: a conceptual dictionary of German and Spanish. In Planning non-existent dictionaries (pp. 163-179). Lisboa: Centro de Linguística da Universidade de Lisboa. Meyer, P. (2014). Meta-computerlexikografische Bemerkungen zu Vernetzung in XML-basierten Onlinewörterbüchern – am Beispiel von ELEXIKO. Opal – Online publizierte Arbeiten zur Linguistik , 2, 9-21. Retrieved from https://pub.idsmannheim.de/laufend/opal/pdf/opal20142.pdf Mugdan, J. (1992). Zur Typologie zweisprachiger Wörterbücher. In G. Meder, & A. Därner (Eds.), Worte, Wörter, Wörterbücher: Lexikographische Beiträge zum Essener Linguistischen Kolloquium (pp. 25-48). Tübingen: Niemeyer. Müller-Spitzer, C. (2003). Ordnende Betrachtungen zu elektronischen Wörterbüchern und lexikographischen Prozessen. Lexicographica , 19, 140-168. Müller-Spitzer, C. (2007). Der lexikographische Prozess. Konzeption für die Modellierung der Datenbasis. Tübingen: Narr. Müller-Spitzer, C. (2014). Empirical data on contexts of dictionary use. In C. Müller-Spitzer, (Ed.), Using Online Dictionaries (pp.85-126). Berlin, Boston: de Gruyter. Müller-Spitzer, C. (2016). Nutzerbeteiligung. In A. Klosa & C. Müller-Spitzer (Eds.), Internetlexikografie. Ein Kompendium (pp. 291-342). Berlin, Boston: de Gruyter. Müller-Spitzer, C., & Koplenig, A. (2014). Online dictionaries: expectations and demands. In C. MüllerSpitzer (Ed.). Using Online Dictionaries (pp. 143-188). Berlin, Boston: de Gruyter. Nesi, H. (1999). The specification of dictionary reference skills in higher education. In R. R. K. Hartmann (Ed.), Dictionaries in language learning. Recommendations, national reports and thematic reports from the Thematic Network Project in the Area of Languages, sub-project 9: dictionaries (pp. 53-67). Berlin: Freie Universität Berlin. Retrieved from https://www.researchgate.net/publication/272167806_The_Specification_of_Dictionary_Refer ence_Skills_in_Higher_Education
118 Nesi, H. (2000). Electronic Dictionaries in Second Language Vocabulary Comprehension and Acquisition: the State of the Art. In: U. Heid, S. Evert, E. Lehmann & C. Rohrer (Eds.), Proceedings of the 9th Euralex International Congress (pp. 839-847). Stuttgart: Institut für Maschinelle Sprachverarbeitung, Universität Stuttgart. Retrieved from http://www.euralex.org/elx_proceedings/Euralex2000/099_Hilary%20NESI_Electronic%20Dic tionaries%20in%20Second%20Language%20Vocabulary%20Comprehension%20and%20Acquisiti on_the%20State%20of%20the%20Art.pdf Nesi, H., & Meara P. (1994). Patterns of misrepresentation in the productive use of EFL dictionary definitions. System , 22 (1), 1-15. Nesselhauf, N. (2005). Collocations in a learner corpus . Amsterdam: John Benjamins. Nielsen, S., & Almind, R. (2011). In P. A. Fuertes-Olivera & H. Bergenholtz (Eds .), e-Lexicography: The Internet, Digital Initiatives and Lexicography (pp. 141-167). London, New York: Continuum. Nielsen, S., & Fuertes-Olivera, P. (2013). Development in Lexicography: From Polyfunctional to Monofunctional Accounting Dictionaries. Lexikos 23 , 323-347. Nuccorini, S. (1994). On Dictionary Misuse. In W. Martin, W. Meijs, M. Moerland, E. ten Pas, P. van Sterkenburg, & P. Vossen (Eds.), Proceedings of the 6th Euralex International Congress (586597). Amsterdam: Euralex. Osselton, N.E. (2000). Murray and his European counterparts. In L. Mugglestone (Ed), Lexicography and the OED: Pioneers in the Untrodden Forest (pp.59-76). Oxford: Oxford University Press. Pearson J. (1998). Terms in Context . Amsterdam, Philadelphia: John Benjamins. Reichmann, O. (1990). Das onomasiologische Wörterbuch: Ein Überblick. In F. J. Hausmann, O. Reichmann, H. E. Wiegand & L. Zgusta (Eds.), Dictionaries: An International Encyclopedia of Lexicography (vol. 2, pp. 1057-1067). Berlin, New York: de Gruyter. Renouf, A. (1987). Corpus Development. In Sinclair, J. (Ed.), Looking up. An account of the Cobuild Project in lexical computing (pp. 1-40). London: Collins. Rundell, M. (1999). Dictionary use in production. International Journal of Lexicography , 12 (1), 35–53. Rundell, M. (2002). Introduction. In Macmillan English Dictionary for Advanced Learners . Oxford: Macmillan Education. Rundell, M. (2010). What future of the learner’s dictionaries? In I. Kernerman, & P. Bogaards (Eds.), English Learner's Dictionaries at the DSNA 2009 , (pp. 169-175). Jerusalem: Kdictionaries. Rundell, M. (2012). ‘It works in practice but will it work in theory?’ The uneasy relationship between lexicography and matters theoretical. In R. Vatvedt Fjeld & J. M. Torjusen (Eds.). Proceedings of
119 the 15th EURALEX International Congress , (pp. 47-92). Oslo: UiO. Retrieved from http://www. euralex.org/elx_proceedings/Euralex2012/pp47-92%20Rundell.pdf Sánchez Hernández, P. (2013). Zur Konzipierung eines deutsch-spanischen kombiniert onomasiologisch-semasiologisch ausgerichteten Verbwörterbuchs mit online-Zugriff— ausgewählte Aspekte. Aussiger Beiträge , 7, 135–155. Schierholz, S. J. (2015). Methods in Lexicography and Dictionary Research. Lexicos , 25, 323-352. Sharoff, S., Rapp, R., Zweigenbaum, P., & Fung, P. (Eds.). (2013). Building and using comparable corpora . Heidelberg: Springer. Shcherba, L.V. (1940). Towards a general theory of lexicography. International Journal of Lexicography, 8 (4), 1995, 315-350. Sierra, G. & Hernández, L. (2013). Automatic construction of the knowledge base of an onomasiological dictionary. Retrieved from https://www.researchgate.net/publication/283773324_Automatic_construction_of_the_knowl edge_base_of_an_onomasiological_dictionary Sierra, G. & McNaught J. (2000) Extracting Semantic Clusters from MRDs for an Onomasiological Search Dictionary. International Journal of Lexicography , 13 (4), 264-286. Sierra, G. (2000). The onomasiological dictionary: a gap in lexicography. In U.Heid et al. (Eds.). Proceedings of the 10th Euralex International Congress , (pp. 223-235). Stuttgart: Universität Stuttgart. Retrieved from https://www.researchgate.net/publication/265099932_The_onomasiological_dictionary_a_g ap_in_lexicography Sierra, G. (2008). Natural Language Searching in Onomasiological Dictionaries. Retrieved from https://www.researchgate.net/publication/238529713_Natural_Language_Searching_in_Ono masiological_Dictionaries. Sinclair, J. (1991). Corpus, Concordance, Collocation . Oxford: Oxford University Press. Sinclair, J. (1996). Preliminary Recommendations on Corpus Typology. EAGLES Document EAG-TCWGCTYP/P. Retrieved from http://www.ilc.cnr.it/EAGLES96/corpustyp/corpustyp.html Sinclair, J. (2004). Intuition and annotation - the discussion continues. In K. Aijmer, & B. Altenberg (Eds.). Advances in corpus linguistics. Papers from the 23rd International Conference on English Language Research on Computerized Corpora (ICAME 23). (pp. 38-59). Amsterdam, New York: Rodopi.
120 Skadina, I., Vasiļjevs, A., Skadiņš, R., Gaizauskas, R., Tufiş, D., & Gornostay, T. (2010). Analysis and Evaluation of Comparable Corpora for Under Resourced Areas of Machine Translation. Proceedings of the 3rd Workshop on Building and Using Comparable Corpora (pp. 6-14). Retrieved from https://www.researchgate.net/publication/228517099_Analysis_and_evaluation_of_compara ble_corpora_for_under_resourced_areas_of_machine_translation Stark, M. (2011). Bilingual Thematic Dictionaries . Berlin, Boston: de Gruyter. Storrer, A., & Freese, K. (1996). Wörterbücher im Internet. Deutsche Sprache, 2, 97-153. Tarp, S. (2008 ). Lexicography in the Borderland between Knowledge and Non-Knowledge. General Lexicographical Theory with Particular Focus on Learner’s Lexicography . Tübingen: Niemeyer. Tarp, S. (2009). Reflections on lexicographic user research. Lexicos , 20, 450-465. Tarp, S. (2011). Lexicographical and Other e-Tools for Consultation Purposes: Towards the Individualization of Needs Satisfaction. In P. A. Fuertes-Olivera, & H. Bergenholtz (Eds.). eLexicography: The Internet, Digital Initiatives and Lexicography , (pp. 54-70). London, New York: Continuum. Tarp, S. (2011). Lexicographical and Other e-Tools for Consultation Purposes: Towards the Individualization of Needs Satisfaction. In P. A. Fuertes-Olivera, & H. Bergenholtz (Eds.). eLexicography: The Internet, Digital Initiatives and Lexicography , (pp. 54-70). London, New York: Continuum. Tarp, S. (2012). Theoretical challenges in the transition from the lexicographical p-works to e-tools. In S. Granger & M. Paquot (Eds.), Electronic Lexicography (pp. 107-118). Oxford: Oxford University Press. Taylor, A., & Chan, A. (1994). Pocket electronic dictionaries and their use. In W. Martin, W. Meijs, M. Moerland, E. ten Pas, P. van Sterkenburg, & P. Vossen (Eds.) Proceedings of the 6th Euralex International Congress (598-605). Amsterdam: Euralex. Text Encoding Initiative (TEI) Guidelines. Retrieved from https://tei-c.org/ Thier, K. (2014). Das Oxford English Dictionary und seine Nutzer. Opal – Online publizierte Arbeiten zur Linguistik, 4, 63-70. Retrieved from https://pub.idsmannheim.de//laufend/opal/pdf/opal2014-4.pdf Tomaszczyk, J. (1979). Dictionaries: users and uses. Glottodidactica , 12, 103-119. Tono, Y. (2010). A critical review of the theory of lexicographical functions. Lexicon , 40, 1-26.
127 Appendix 7: Two-step technical-(meta)lexicographic electronic-dictionary typology (Lehr 1996: 315, redrawn and translated here). (de Schryver 2003, p. 148) Appendix 8: Number of Internet users worldwide from 2005 to 2017 3 3 Number of internet users worldwide from 2005 to 2017 (in millions). In Statista –The portal for statistics. Retrieved from https://www.statista.com/statistics/273018/number-of-internet-users-worldwide/
128 Appendix 9: Parameters for the description and evaluation of online dictionaries (Kemmer, 2010, p. 30) Appendix 10: Questions to ask in order to identify the users’ characteristics (Bergenholtz & Tarp, 2003, p. 173) 1. Which language is their mother tongue? 2. At what level do they master their mother tongue? 3. At what level do they master a foreign language? 4. How are their experience in translating between the languages in question? 5. What is the level of their general cultural and encyclopaedic knowledge? 6. At what level do they master the special subject field in question? 7. At what level do they master the corresponding LSP in their mother tongue?
129 8. At what level do they master the corresponding LSP in the foreign language? Appendix 11: Questions to ask in order to draw up lexicographically relevant user characteristics (Fuertes-Olivera & Tarp, 2014, pp. 49-50) - Function-relevant user characteristics: Which language is the user’s mother tongue or the first language? What is the user’s proficiency level in the mother tongue? With which method is the user learning the mother tongue or first language? What is the user’s proficiency level in a second, third, etc., language? With which method is the user learning a second, third, etc., language? What is the user’s general cultural and encyclopaedic level? What is the user’s experience in translation between a specific set of languages? What is the user’s proficiency level in a specific specialized language? What is the user’s experience in translation between a specific set of specialized languages? Etc., etc… - Consultation-relevant user characteristics: What is the user’s experience of lexicographical consultations? Is the user blind, deaf, or suffers from any other handicap which may limit the use of specific types of lexicographical tools? Does the user have electricity and electric light? Does the user possess a device with access to the Internet? Does the user know how to distinguish between right and left? Etc., etc…. Appendix 12: Six basic types of communication-oriented user situations (Bergenholtz & Tarp, 2003, p. 175) 1. Production of texts in the mother tongue (or first language) 2. Reception of texts in the mother tongue (or first language) 3. Production of texts in a foreign language (or second, third language etc.) 4. Reception of text in a foreign language (or second, third language etc.)
130 5. Translation of texts from the mother tongue (or first language) into a foreign language (or second, third language etc.) 6. Translation of texts from a foreign language (or second, third language etc.) into the mother tongue (or first language). Appendix 13: Data Types in the ESSCD database 4 Data Type Rationale Lemma All dictionaries contain lemmata. Language code to lemma Indicates language; Assists users in identifying variants of English language Pronunciation Indicates the correct pronunciation of the headword Domain and subdomain Allows user to identify to what field belongs the headword and to access all the lemmata of a particular field or subfield. Grammatical data addressed to lemma Assist users in text reception and production; offers inflections, countability Translation equivalent Spanish equivalents assist in text reception Language code for equivalent Indicates the language Grammatical data addressed to equivalent Assist user in translation Collocations Assist users in production Language code of collocation Indicates the language Translation of collocations Assist users in translation and reception Language code to translation of collocation Indicates the language Examples Full sentences showing the use of collocations. Assist in production Source Contain bibliographical information of example sentences Translation of examples Full sentences showing equivalents in use. Assist in translation and reception Notes Assist in all user types of communicative and cognitive situations Translation of notes Assist in translation and reception Images Illustrate notes or lemmata Cross-references Hyperlinks to internal and external texts 4 Based in analogy to the Table 9.2 Data Types in the Lexicographical Accounting Database (Fuertes-Olivera & Tarp, 2014, pp. 199-200).
131 Appendix 14: Types of summer camps 5 Active Faith-Based Camps are operated by Christian or Jewish organizations Daily prayer, worship, and religious study are typically not major camp activities, however, may be part of daily camp life. Applicants don't have to be religious to work in these camps just need to be open-minded. Special Needs Camps serve children and adults with disabilities. Some camps focus on providing adaptive facilities for physical limitations, some serve campers with developmental disabilities/behavioral challenges and others integrate special needs campers into traditional camp settings. Applicants don't need previous training in working with special needs, just the willingness to learn. Underprivileged Camps are run by non-profit organizations, these traditional-style camps provide lower cost or free programming for underprivileged/disadvantaged youth who may never have been in wilderness before. Wilderness camps: applicants will be sleeping and living in platform tents out in the wilderness. No electricity in the cabin: the kitchen and office will have electricity for example to charge a camera and phone. The rest of the camp gives you the chance to go back to basics and forget about the stresses of everyday life whilst in nature. Traditional Camp: are most similar to what we have seen in the movies. These camps blend a variety of sports, wilderness, creative arts and specialist activities into their daily programs. 5 Source: Application form of Camp Canada. Retrieved from https://www.campcanada.co.uk/
132 Appendix 15: Collocation data for importation to Gephi as nodes table id label PoS freq camp_NN camp NN 15326 summer_NN summer NN 4008 counsellor_NN counsellor NN 1606 staf f_NN staf f NN 3818 experience_NN experience NN 2819 skill_NN skill NN 481 development_NN development NN 1653 program_NN program NN 1588 employment_NN employment NN 722 member_NN member NN 638 interview_NN interview NN 614 part icipant_NN part icipant NN 1174 canoe_NN canoe NN 426 trip_NN trip NN 692 problem_NN problem NN 1046 solving_NN solving NN 386 youth_NN youth NN 1027 director_NN director NN 420 community_NN community NN 1739 dining_NN dining NN 101 hall_NN hall NN 113 season_NN season NN 265 environment_NN environment NN 661 life_NN life NN 1340 leadership_NN leadership NN 864 placement_NN placement NN 207 day_NN day NN 2379 team_NN team NN 777 session_NN session NN 379 training_NN training NN 653 applicat ion_NN applicat ion NN 442 process_NN process NN 723 sleeping_NN sleeping NN 105 bag_NN bag NN 181 age_NN age NN 709 group_NN group NN 2296 cabin_NN cabin NN 547 camping_NN camping NN 557 health_NN health NN 569 care_NN care NN 877 child_NN child NN 569 insurance_NN insurance NN 234 tripping_NN tripping NN 165 head_NN head NN 228
133 Appendix 16: Collocation data for importation to Gephi as edges table source target type freq summer_NN camp_NN direct 1270 camp_NN counsellor_NN direct 993 camp_NN staff_NN direct 457 camp_NN experience_NN direct 438 skill_NN development_NN direct 334 camp_NN program_NN direct 355 camp_NN employment_NN direct 260 staff_NN member_NN direct 254 interview_NN participant_NN direct 166 camp_NN participant_NN direct 142 employment_NN experience_NN direct 142 canoe_NN trip_NN direct 136 problem_NN solving_NN direct 134 youth_NN development_NN direct 127 camp_NN director_NN direct 98 camp_NN community_NN direct 93 dining_NN hall_NN direct 81 camp_NN season_NN direct 73 camp_NN environment_NN direct 71 camp_NN life_NN direct 59 leadership_NN skill_NN direct 58 camp_NN placement_NN direct 55 day_NN camp_NN direct 53 staff_NN team_NN direct 48 camp_NN session_NN direct 47 staff_NN training_NN direct 46 leadership_NN development_NN direct 46 application_NN process_NN direct 45 sleeping_NN bag_NN direct 44 age_NN group_NN direct 42 cabin_NN group_NN direct 41 camping_NN experience_NN direct 39 health_NN care_NN direct 59 child_NN care_NN direct 59 health_NN insurance_NN direct 37 summer_NN staff_NN direct 36 child_NN development_NN direct 36 canoe_NN tripping_NN direct 35 head_NN staff_NN direct 35