scieee AI-readable full text Open interactive document viewer

Multi-word patterns in the corpus of Information and Communication Technology. Terminological bundles in the genre – ‘Textbooks in ICT’

Marek Weber

Full text

39Terminologija | 2012 | 19 Multi-word patterns in the corpus of Information and Communication Technology. Terminological bundles in the genre – ‘Textbooks in ICT’ MAREK WEBER BialystokUniversityofTechnology KEYWORDS: corpus linguistics, Information and Communication Technology, lexical bundles, functional classification of lexical bundles, terminological bundles, lexical scrutiny DATA FOR THE STUDY, THE CORPUS OF INFORMATION AND COMMUNICATION TECHNOLOGY Data for the study consists of an electronic corpus of the broad domain of Information and Communication Technology (ICT) an umbrella term that includes all technologies developed for the purposes of processing and the transfer of data. The texts were selected to represent a crosssection of the field of ICT. The texts were divided into seven categories which can be regarded as genres within the domain of ICT: Genre 1: Professional articles in ICT; Genre 2: Academic articles in ICT; Genre 3: Technical documentation of software applications in ICT; Genre 4: Technical documentation of ICT hardware; Genre 5: Textbooks in ICT; Genre 6: Technical documentation of programming languages and programming environments; Genre 7: Technical documentation of network technologies. The corpus contains a collection of 973 texts totaling almost 24 million words. The choice of data for a corpus is one of the most important concerns for the researcher as the corpus should be representative of the analyzed domain of use. As Douglas Biber suggests the quantity of texts is crucial for research in which the scholar concentrates on the text as the primary unit of scrutiny. A sufficient number of texts should be included in each genre to account for variation between categories and authors (Biber, 2006). 40 Marek Weber Multi-wordpatternsinthecorpusofInformation  andCommunicationTechnology.Terminological  bundlesinthegenre–‘TextbooksinICT’ table 1. Composition of the Information and Communication technology corpus genre Number of texts Number of tokens (running words) Professional articles 391 225551 Academic articles 95 1272001 Technical documentation of software applications in ict 92 2472839 Technical documentation of ICT hardware 100 2471772 Textbooks 99 15315237 Technical documentation of programming languages and programming environments 90 1282175 Technical documentation of network technologies 106 725418 TOTAL 973 23764993 LEXICAL BUNDLE: THE TERM Recurrent sequences of words which appear together more frequently than expected by chance are an important part of linguistic output. The term lexicalbundle was first used by Douglas Biber, Stig Johansson, Geoffrey Leech, Susan Conrad and Edward Finegan in the now classic Longman grammarofspokenandwrittenEnglish to refer to multi-word combinations identified on the basis of the sole criterion of frequency (Biber et al. 1999) while Mike Scott calls them wordclusters in the manual of the software application WordSmithTools version 4 (1996) and WordSmithTools version 5 (2010). The only consideration in identifying lexical bundles is their frequency. “Clusters are words which are found repeatedly together in each others’ company, in sequence” (Scott 2010: 337). Scott points out that word clusters may involve semantic prosody i.e. the tendency for certain lexemes to co-occur with certain other lexemes (e.g. the tendency for cause to come with negative effects such as accident,trouble, etc.) (Scott 2010: 337). Biber, Johansson, Leech, Conrad and Finegan define lexicalbundles as recurrent expressions regardless of their structural properties, i.e. they are sequences of word forms that frequently co-occur in discourse and point out that shorter bundles are often incorporated into more than one longer lexical bundle (Biber et al. 1999). They also observe that in most cases lexical bundles do not form structural units and most of them are 41Terminologija | 2012 | 19 not expressions that language users would recognize as idioms or other fixed lexical expressions. Biber, Johansson, Leech, Conrad and Finegan establish the following cut off thresholds for multi-word sequences to qualify as lexical bundles: they must occur in at least 10 times per million words and at least in five texts (Biber et al. 1999). Biber notes that a surprising result of a frequency driven approach is that lexical bundles have two unexpected characteristics (Biber 2006: 134). Firstly, most of them are not idiomatic in meaning but the meanings are usually transparent from individual word-forms. Ken Hyland also points to the fact that most bundles are semantically transparent and “formally regular, providing the building blocks of coherent discourse” (Hyland 2008: 6). Secondly, most bundles are not complete grammatical structures. Biber, Johansson, Leech, Conrad and Finegan observed that approximately 15 % of the lexical bundles in conversation formed complete structural units while only approximately 5 % of the lexical bundles in academic prose could be considered as complete phrases or clauses (Biber et al. 1999: 1000). A valuable finding into the syntactic nature of lexical bundles by Biber is that they are lexical units that frequently cut across grammatical structures, for example they can bridge two phrases or clauses in such a way that the last words of a bundle are the beginning parts of a second syntactic structure (Biber 2006: 135). Hyland also emphasizes that lexical bundles being identified solely on the basis of their frequency usually span structural units (Hyland 2008: 6). A range of corpus studies have been devoted to the analysis of recurrent strings of uninterrupted word-forms and they demonstrate how important they are in various types of discourse as well as they show considerable variation of lexical bundles in different genres and registers (e.g. Biber 2006; Biber, Conrad, Cortes 2004; Hyland 2008; Scott, Tribble 2006; Gozdz-Roszkowski 2011). METHODOLOGY USED IN THE STUDY The decision was made to concentrate on the analysis of 4-word bundles because on the one hand the frequencies of 4-word bundles are substantially higher than 5-word sequences and on the other hand they frequently contain 3-word strings and enable to identify more apparent patters and structures than 3-word expressions. Lists of 4-word bundles in each of the 42 Marek Weber Multi-wordpatternsinthecorpusofInformation  andCommunicationTechnology.Terminological  bundlesinthegenre–‘TextbooksinICT’ seven ICT genres were generated using the WordSmithTools v. 5.0 – the software package for searching patterns in corpora. The following two cut-off points were set in this study: a minimum frequency of occurrence of 20 times per million words and another criterion that a bundle should occur in at least five texts. The process of applying the cut-off criteria required appropriate calculations in each of the seven files containing the list of bundles in a given genre. The calculations are described in the following steps: 1. Creating an additional column named “Frequency per million” in the spreadsheet obtained from the WordSmithTools software. 2. Inserting the running words value into a free Excel spreadsheet cell within a given category. 3. Filling in the first cell of the “Frequency per million” column with the following formula: 4. Copying the formula to other cells by clicking on the right bottom corner of the filled-in cell and dragging in downwards to other cells. 5. Highlighting the “Number of texts” column. 6. Choosing the “Data” menu and selecting the field “Sort &Filter” in the sorting tool. 7. Choosing the descending sorting order. 8. Deleting all rows of the table, for which the values of the “Number of texts” field is lower than 5. 9. Highlighting the “Frequency per million” column. 10. Choosing the “Data” menu and selecting the field “Sort & Filter” in the sorting tool. 11. Choosing the descending sorting order. 12. Deleting all rows of the table, for which the values in the “Frequency per million” field are lower than 20. 1886 different bundles were found in the corpus after applying the above cut-off criteria. The total number of bundles identified in the entire data amounted to 197 553. 43Terminologija | 2012 | 19 table 2. Distribution of lexical bundles in ICt genres genre The total number of bundles after applying cut-off criteria Number of different bundles after applying cut-off criteria % of running words in bundles Professional articles 1057 131 1,8 Academic articles 5548 120 1,7 Software applications 44669 403 7,2 Hardware 55296 672 8,9 Textbooks 72483 119 1,9 Programming languages and programming environments 10984 180 3,4 Network technologies 7516 261 4,1 Total 197553 1886 The percentage of running words in bundles was obtained by multiplying the number of total cases for a given genre by 4 (4-word bundles are analyzed) and dividing by the number of running words in a given genre and then multiplying by 100 % according to the formula: In our view two parameters: the range of different bundles and the percentage of running words in bundles can be regarded as indicators of the degree to which a given genre is formulaic and repetitive in comparison to other genres. In other words the degree of formulaicity and repetitiveness of a genre can be measured by the range of different bundles employed in a genre and the percentage of running words in bundles. The ICT genres can be divided into three groups according to the criterion of formulaicity and repetitiveness. Genre 3 – ‘technical documentation of software applications in ICT’ and genre 4 – ‘technical documentation of ICT hardware’ are marked by high degrees of formulaicity and repetitiveness as they display comparable patterns by employing the biggest scope of different bundles (403 and 672 respectively) and the highest percentage of running words in bundles (7,2 % and 8,9 % respectively). 44 Marek Weber Multi-wordpatternsinthecorpusofInformation  andCommunicationTechnology.Terminological  bundlesinthegenre–‘TextbooksinICT’ Figure 1: Genres with high degrees of formulaicity and repetitiveness In contrast, genre 1 – ‘professional articles in ICT’, genre 2 – ‘academic articles in ICT’ and genre 5 – ‘textbooks in ICT’ are characterized by relatively low degrees of formulaicity and repetitiveness as they employ the lowest scope of different bundles (131, 120 and 119 respectively) and the lowest percentage of running words in bundles (1,8 %, 1,7 % and 1,9 % respectively). The remaining two genres: genre 6 – ‘technical documentation of programming languages and programming environments’ and genre 7 – ‘technical documentation of network technologies’ form the third group with the values of both parameters in the middle of the range (numbers of different bundles: 180 and 261 respectively; percentage of running words in bundles 3,4 % and 4,1 % respectively). However, it is worth bearing in mind that as Stanislaw Gozdz-Roszkowski rightly points out direct comparisons can only be safely made between genres with similar word counts. In order to account for that consideration the parameter percentageofrunningwordsinbundles is computed for each genre in such a way as to reflect the word counts of each of the subcorpora representing respective genres (Gozdz-Roszkowski 2011: 111). FUNCTIONAL CLASSIFICATION OF BUNDLES A framework for the functional analysis of the bundles obtained in this corpus was established from Biber’s (Biber 2006; Biber et al. 2004), Hyland’s (2008) and Gozdz-Roszkowski’s (2011) taxonomies. Biber’s (2006) classification resulted from the scrutiny of a broad corpus of spoken and written registers which covered among others such types of discourse as: casual conversations, class sessions tapes, classroom teaching, office hours, study groups, on-campus service encounters, textbooks, course packs, institutional texts (e.g. university catalogs, brochures). 45Terminologija | 2012 | 19 The corpus in Hyland’s (2008) study (size 3,5 million words) comprises research articles, PhD dissertations and MA/MSc theses from four disciplines: electrical engineering and microbiology from the applied and pure sciences, and business studies and applied linguistics from the social sciences. Gozdz-Roszkowski (2011) analyzed a corpus of American Law containing over 5,5 million words and representing seven genres within American legal culture and education: academic articles, briefs, contracts, legislation, opinions, professional articles and textbooks. Drawing upon the abovementioned taxonomies lexical bundles in ICT were functionally divided into three broad categories with respect to their meanings in the texts: research-centered, text-centered and participant-centered. Bundles grouped in the first category help writers to organize their activities and experiences in the domain of ICT. Textcentered bundles are employed to indicate the organization of the text and its meaning. Finally participant-centered bundles are used to signal different attitudes or assessments and they are focused on the writer or reader of the text. Figure 2: Functional taxonomy of lexical bundles Each of the three major functional categories of lexical bundles was further subdivided into a number of subcategories. Research-centered bundles include the following subcategories: Quantity bundles (e.g. oneofthebiggest;alargenumberof;oneofthe following;thenumberofelements); Time reference bundles (e.g. atthetimeof;endoftheyear;thenext fewweeks;inthecomingweeks); Place / direction reference bundles (e.g. intheUnitedStates;ofthe mainwindow;bottomofthewindow;inthestatusbar); 46 Marek Weber Multi-wordpatternsinthecorpusofInformation  andCommunicationTechnology.Terminological  bundlesinthegenre–‘TextbooksinICT’ Procedure bundles – used to describe diverse functions pertaining to ICT (e.g. theuseofthe;theuseofa); Topic indicator bundles – pertaining to the area of research (e.g. proceedingsoftheIEEE;programminglanguagesandsystems;thedropdown menu;thefollowingconfigurationoptions); Multi-functional reference bundles – bundles which can be used as time / place / text reference (e.g. theendofthe;thestartofthe;theleft ofthe;atthebeginningof); Description bundles – specify characteristics of the following noun (e.g. thecomplexityofthe;thesurfaceofthe;thescopeofthe;theheightofthe). Text-centered bundles include the following subcategories: Elaboration bundles – further elaborate on the analyzed topic and clarify it (e.g. atthesametime,aswellasthe,thismeansthatthe,inthis casethe); Transition bundles – provide additive or contrastive connections between portions of texts (e.g. ontheotherhand,asopposedtothe,inadditiontothe,incontrasttothe); Framing attributes bundles – “situate arguments by specifying limiting conditions” (Hyland 2008: 14) for making claims or arguments (e.g. intermsofthe,inthecontextof,inthepresenceof,thecontentsofthe); Conditions bundles – express conditions (e.g. ifyouwantto,ifyoudo not,ifyouhavea,ifyouhavean); Results bundles – indicate logical links between elements in terms of cause and result relationships (e.g. sothatyoucan,asaresultof,asa resultthe,asafunctionof); Structure markers bundles – are used to point to other parts of the text (e.g. asshowninfigure,asdiscussedinsection,isshowninexample, laterinthischapter). Participant-centered bundles include the following subcategories: Engagement bundles – address readers directly; “actively address readers as participants in the unfolding discourse” (Hyland 2008: 18) (e.g. do oneofthe,selectthetypeof,isrecommendedthatyou,selectoneofthe); Stance bundles – convey emotions, attitudes, value judgments and assessments; “provide a frame for the interpretation of the following proposition” (Biber, 2006: 139) (e.g. itisnecessaryto,isnoguaranteethat,itis importantto,itisnotpossible,itispossibleto); 47Terminologija | 2012 | 19 Modality bundles – express that something is probable, permissible or necessary (e.g. mustbefulfilleda,canbeusedto,ascanbeseen,inwhich youcan,maynotbedisplayed,maybereproducedor,mustacceptanyinterference,interferencethatmaycause,thatmaycauseundesired); Prediction bundles – express the writer’s prediction of some future action (e.g. isexpectedtobe,youwillbeprompted,willbedisplayedin,this willopenthe,willbeaskedto). Table 3 shows the numbers of bundles belonging to the three major functional categories in each genre. table 3: Distribution of lexical bundles across functional categories Researchcentered Textcentered Participantcentered Others Professional articles 45 23 13 45 Academic articles 48 35 928 Software applications 123 55 122 83 Hardware 207 51 138 187 Textbooks 32 35 18 20 Programming languages and programming environments 39 39 26 62 Network technologies 91 24 11 91 Figure 3: Classification of research-centered bundles