scieee AI-readable full text Open interactive document viewer

A corpus-based study on near-synonymy: The concept PLEASANT SMELLING in 19th- and 20th-century American English

Pettersson-Traba, Daniela

Abstract

This dissertation examines the distributional patterns of the five adjectival near-synonyms fragrant, perfumed, scented, sweet-scented, and sweet-smelling, which designate the concept PLEASANT SMELLING in American English by paying special attention to their diachronic development, as represented in the Corpus of Historical American English (1810-2009). The distribution of the selected adjectives is analyzed across a wide range of semantic, morphosyntactic, and stylistic contexts which function as a proxy for semantic similarity. Three separate analyses are conducted, focusing on different aspects of the semantic structure of the near-synonyms. Results indicate that the set is undergoing processes of convergence and substitution, possibly as a result of extralinguistic factors. Therefore, the analyses shed light on the diachronic development of lexical near-synonyms, a dimension that has up to now been relatively disregarded in the specialized literature.

Full text

  7(6('('28725$0(172 #%14275$#5'&567&;10 0'#45;010;/;6*'%10%'26 2.'#5#065/'..+0)+06* #0&6*%'0674;#/'4+%#0 '0).+5*  'DQLHOD%HDWUL]3HWWHUVVRQ7UDED  '5%1.#&'&17614#/'061+06'40#%+10#. 241)4#/#&'&17614#/'061'0'567&15+0).'5'5#8#0<#&15.+0)gı56+%# .+6'4#674#'%7.674# 6$17,$*2'(&203267(/$  DECLARACIÓN DO AUTOR/A DA TESE D./Dna. Daniela Beatriz Pettersson-Traba Título da tese: A corpus-based study on near-synonymy: The concept PLEASANT SMELLING in 19thand 20th-century American English Presento a miña tese, seguindo o procedemento axeitado ao Regulamento, e declaro que: 1) A tese abarca os resultados da elaboración do meu traballo. 2) De ser o caso, na tese faise referencia ás colaboracións que tivo este traballo. 3) Confirmo que a tese non incorre en ningún tipo de plaxio doutros autores nin de traballos presentados por min para a obtención doutros títulos. E comprométome a presentar o Compromiso Documental de Supervisión no caso de que o orixinal non estea na Escola. En Santiago de Compostela, ϳ de džĂŶĞŝƌŽ de 202ϭ. Sinatura electrónica AUTORIZACIÓN DO DIRECTOR/TITOR DA TESE D./Dna. María José López Couso En condición de: Titor/a e director/a Título da tese: A corpus-based study on near-synonymy: The concept PLEASANT SMELLING in 19thand 20th-century American English INFORMA: Que a presente tese, correspóndese co traballo realizado por D/Dna Daniela Beatriz Pettersson-Traba, baixo a miña dirección/titorización, e a utorizo a súa presentación , considerando que reúne os r equisitos esixidos no R egulamento de Estudos de Doutoramento da USC, e que como director/titor desta non incorre nas causas de abstención establecidas na Lei 40/2015. En Santiago de Compostela, ϳ de džĂŶĞŝƌŽ de 202ϭ Sinatura electrónica ABSTRACT The last couple of decades have witnessed a renewed interest in the semantic phenomenon of near-synonymy (e.g. Gries 2001; 2003; Taylor 2003; Divjak 2010; Liu 2010; 2013). In particular, recent distributional corpus-based approaches and techniques used for semantic analysis, such as Behavioral Profiles (e.g. Divjak & Gries 2006; 2008) and correspondence analysis (e.g. Desagulier 2014; Krawczak 2018), have successfully uncovered subtle distinctions in meaning between near-synonyms by analyzing, among other factors, their collocational and stylistic preferences. Some semantic domains have received particular attention, for instance, those of SIZE and AMOUNT (cf. Biber, Conrad & Reppen 1998: Section 2.6 on big,large, and great; Taylor 2003 on high and tall; and Gries & Otani 2010 on big, large, and great and little,small, and tiny), while other near-synonym sets have been relatively underresearched. Moreover, most studies so far have dealt with the semantic structure of sets of near-synonyms from a synchronic perspective (e.g. Divjak & Gries 2006; 2008; Liu 2010; 2013), whereas their diachronic evolution has generally been neglected, with only a handful of investigations adopting a historical approach (e.g. Kaunisto 2001; Primahadi-Wijaya-R & Rajeg 2014). Against this backdrop, the aim of the present dissertation is to examine five adjectival near-synonyms in the history of American English from the understudied semantic domain of SMELL, namely fragrant,perfumed,scented,sweet-scented, and sweet-smelling, which designate the concept PLEASANT SMELLING. To this end, instances of these five adjectives are retrieved from a large historical corpus of this variety of English, to wit, the Corpus of Historical American English (Davies 2010–), which covers the timespan 1810–2009. The distribution of the adjectives is analyzed over time across a wide range of contexts and functions, including semantic, morphosyntactic, stylistic, and collocational variables, since distributional patterns of this type have been shown to serve as a proxy for semantic (dis)similarity (e.g. Divjak & Gries 2006; 2008; Gries & Otani 2010). The data is submitted to various univariate and multivariate statistical techniques in order to uncover fine-grained DANIELA PETTERSSON-TRABA ii (dis)similarities among the members of this near-synonym set, as well as possible changes in their prototypical structures, from both an onomasiological and a semasiological perspective. Three separate analyses are conducted, each focusing on a different aspect of the internal semantic structure of the near-synonym set. The first one (Chapter 5) examines the distribution of the five adjectives across senses and semantic categories of the nouns that they modify to discover their prototypical uses in both semasiological and onomasiological terms. The second analysis (Chapter 6) delves deeper into the competition between the three most common of these near-synonyms (fragrant,perfumed, and scented) by expanding the range of contexts considered to cover also non-semantic and stylistic factors (e.g. syntactic function of the adjectives, degree, text-type). Finally, Chapter 7 accounts for the idiosyncratic collocational behavior of these three adjectives, hence zooming in on their specific collocational preferences by using techniques which are specifically geared towards this issue. The results demonstrate that the near-synonym set under analysis is undergoing a process of semantic convergence, whereby the adjectives are progressively used in more similar semantic contexts, becoming more and more frequent over time to designate artificial smells as opposed to natural ones. This change is here claimed to be motivated by extralinguistic factors, to wit, the social and technological transformations experienced by American society after the First and Second Industrial Revolutions, which have led to an ever-increasing need to refer to artificially scented products rather than to naturally fragrant plants and flowers. Moreover, this process of convergence is accompanied by one of substitution, in which the initially most dominant adjective of the set (fragrant) is gradually being replaced by another one (scented). Therefore, the findings obtained shed valuable light on the diachronic development of lexical near-synonyms, a dimension that has up to now been relatively disregarded in the specialized literature and that is yet to receive the attention it certainly deserves. Keywords: (near-)synonymy, historical linguistics, semantic change, distributional corpusbased approach, collocation, PLEASANT SMELLING, American English ACKNOWLEDGEMENTS I would like to express my most sincere and profound gratitude to all the people who, in one way or another, have played a part in the successful completion of this PhD dissertation. Without their wholehearted support and commitment during these years, this would undoubtedly not have been possible. First of all, I am greatly indebted to my supervisor Dr. María José López Couso. From the time when I started working under her supervision six years ago to write my BA thesis on this same topic, she has always encouraged me to follow my own research interests and to continue on the dark path of semantics. Her invaluable advice and guidance ever since are the main reasons why I decided to continue in the academic world, hence aiding me in discovering my vocation. Her help has transcended academic matters, as she has also been a source of moral support on a more personal level, always ready to lend a helping hand with any problem that might arise. I feel immensely privileged to have learnt —and still be learning— from such an exceptional professional. I would also like to acknowledge the useful feedback received from Dr. Belén Méndez Naya, Dr. Paloma Núñez Pertejo, and Dr. Susana Doval Suárez, who served as examiners of a pilot study of this work, my MA thesis. Thanks to their insightful suggestions, I have been able to continue polishing this piece of work. Additionally, I would like to show my genuine gratitude to Dr. Kathryn Allan, Dr. Dirk Speelman, and Dr. Augusto Soares da Silva for their expertise and wise counsel received during my research stays at University College London, KU Leuven, and the Catholic University of Portugal. The time they all generously devoted to discussing my research with me has certainly enhanced its quality. Special thanks are due to Dr. Dirk Speelman for everything he has taught me about statistics, both concerning the use of the program Rand concerning the multiple techniques and methods for semantic analysis. His continued patience for bearing with me, a statistics rookie, every week for three months in Leuven is gratefully acknowledged. I would also like to extend my thanks to all those scholars I have had the pleasure to meet and learn from at conferences, seminars, and research meetings LIST OF FIGURES Figure 1. Main types of semantic change distinguished in the historical-philological tradition .....................................................................................................................13 Figure 2. Development of corpus-based research on lexical synonymy over the last 30 years ..........................................................................................................................63 Figure 3. Proportions of the five near-synonyms across periods ...........................................152 Figure 4. Proportions of the concept PLEASANT SMELLING in each sense across periods.......156 Figure 5. Proportions of the concept PLEASANT SMELLING in each semantic category across periods ..........................................................................................................161 Figure 6. Average frequency per million words in COHA of the semantic domains INDUSTRY AND TECHNOLOGY,ADVERTISING AND MEDIA,andCOSMETICS AND HYGIENE in the period 1810–2009...........................................................................169 Figure 7. Overall sense distribution (in proportions) per near-synonym ...............................171 Figure 8. Sense distribution (in proportions) per near-synonym across periods....................173 Figure 9. Overall semantic category distribution (in proportions) per near-synonym...........175 Figure 10. Semantic category distribution (in proportions) per near-synonym across periods ....................................................................................................................177 Figure 11. Continuum of senses of the five adjectives...........................................................181 Figure 12. Overall proportions of the five near-synonyms in each sense ..............................183 Figure 13. Proportions of the five near-synonyms in each sense across periods ...................184 Figure 14. Overall proportions of the five near-synonyms in each semantic category..........186 Figure 15. Proportions of the five near-synonyms in each semantic category across periods ....................................................................................................................187 Figure 16. Variable importance of predictors according to random forest analysis...............207 Figure 17. Effect of Concreteness rating on the probabilities of the adjectives fragrant,perfumed, and scented.............................................................................212 Figure 18. Effect of Degree on the probabilities of the adjectives fragrant,perfumed, DANIELA PETTERSSON-TRABA xii and scented...............................................................................................................213 Figure 19. Effect of Countability on the probabilities of the adjectives fragrant,perfumed, and scented..............................................................................................................214 Figure 20. Effect of Syntactic function on the probabilities of the adjectives fragrant, perfumed, and scented.............................................................................................215 Figure 21. Effect of Text-type on the probabilities of the adjectives fragrant,perfumed, and scented across periods......................................................................................216 Figure 22. Effect of Sense on the probabilities of the adjectives fragrant,perfumed, and scented across periods.............................................................................................217 Figure 23. Effect of Semantic category on the probabilities of the adjectives fragrant, perfumed, and scented across periods.....................................................................219 Figure 24. Collocational preferences of fragrant....................................................................225 Figure 25. Collocational preferences of perfumed..................................................................226 Figure 26. Collocational preferences of scented.....................................................................228 Figure 27. The use of fragrant,perfumed, and scented in the artificial, indeterminate, and natural senses over time..........................................................................................233 Figure 28. Dendrogram of the collocational preferences of fragrant,perfumed, and scented throughout the period 1810–2009..............................................................250 Figure 29. Two-dimensional MDS map of the collocational preferences of fragrant, perfumed, and scented throughout the period 1810–2009......................................252 Figure 30. Significant correlations between MDS dimensions and semantic (sub-)categories in USAS .......................................................................................257 Figure 31. Collocational network of fragrant,perfumed, and scented in P1 (1810–1859)....260 Figure 32. Collocational network of fragrant,perfumed, and scented in P2 (1860–1909)....262 Figure 33. Collocational network of fragrant,perfumed, and scented in P3 (1910–1959)....264 Figure 34. Collocational network of fragrant,perfumed, and scented in P4 (1960–2009)....266 Figure C.1. Three-dimensional MDS map of the collocational preferences of fragrant, perfumed, and scented throughout the period 1810–2009.....................................358 LIST OF TABLES Table 1. Classifications of synonymy by Lyons and Cruse .....................................................56 Table 2. Sense division of the five adjectives according to nine dictionaries and thesauri......97 Table 3. Number of words in COHA across decade and genre..............................................104 Table 4. Number of instances retrieved per query..................................................................107 Table 5. Major semantic classes in USAS...............................................................................116 Table 6. Major semantic classes in the HTOED.....................................................................117 Table 7. Classification of nouns into semantic categories .....................................................127 Table 8. Variables in the first dataset and their levels ...........................................................145 Table 9. Absolute and relative frequencies of the five near-synonyms..................................151 Table 10. Frequency clines of the five adjectives in the period 1810–2009 ..........................153 Table 11. Distribution of the concept PLEASANT SMELLING across senses.............................157 Table 12. Sense distribution of the semantic categories BODY AND PEOPLE,OBJECT, SENSATION,SPACE, and SUBSTANCE AND MATERIAL ..............................................160 Table 13. Distribution of the concept PLEASANT SMELLING across semantic categories........162 Table 14. Distribution of animate and inanimate referents per Semantic category ...............198 Table 15. Absolute and relative frequencies of fragrant,perfumed, and scented..................201 Table 16. Hypothetical scenario of reordering of levels in random forest analysis...............202 Table 17. Model summary of multinomial logistic regression Model A (Sense) ..................203 Table 18. Model summary of multinomial logistic regression Model B (Semantic category).................................................................................................................205 Table 19. Model summary of multinomial logistic regression with only semantic predictors (Sense)...................................................................................................205 Table 20. Model summary of multinomial logistic regression with only semantic predictors (Semantic category)...............................................................................205 Table 21. Model A coefficients (Sense).................................................................................209 Table 22. Model B coefficients (Semantic category).............................................................210 DANIELA PETTERSSON-TRABA xiv Table 23. Model summary of the binary mixed-effects logistic regression (fragrant)...........222 Table 24. Model summary of the binary mixed-effects logistic regression (perfumed).........222 Table 25. Model summary of the binary mixed-effects logistic regression (scented)............222 Table 26. Structure of the dataset............................................................................................241 Table 27. Similarity matrix based on cosine similarity scores for the data ............................245 Table 28. Distance matrix for the data....................................................................................246 Table 29. Settings for the construction of collocational networks..........................................249 Table 30. AU p-values for the five-cluster solution in Figure 28...........................................251 Table 31. Significant correlations between the dimensions in Figure 29 and semantic (sub-)categories in USAS........................................................................................254 Table 32. Examples of types of noun collocates in the semantic (sub)-categories in USAS.......................................................................................................................255 Table 33. Number and percentage of consistent, initiating, terminating, and transient collocates of fragrant,perfumed, and scented........................................................268 Table 34. Consistent, initiating, terminating, and transient noun collocates of fragrant........269 Table 35. Consistent, initiating, terminating, and transient noun collocates of perfumed......272 Table 36. Consistent, initiating, terminating, and transient noun collocates of scented.........273 Table A.1. Chi-square residuals of the frequency of the five adjectives in the period 1810–2009 ...........................................................................................................317 Table A.2. Chi-square residuals of the concept PLEASANT SMELLING across senses in the period 1810–2009 ................................................................................................317 Table A.3. Diachronic development of the semantic category SPACE according to sense......317 Table A.4. Chi-square residuals of the concept PLEASANT SMELLING across semantic categories in the period 1810–2009 .....................................................................318 Table A.5. Average frequency per million words in COHA of the semantic domains INDUSTRY AND TECHNOLOGY,ADVERTISING AND MEDIA, and COSMETICS AND HYGIENE in the period 1810–2009........................................................................318 Table A.6. Sense distribution of fragrant...............................................................................319 Table A.7. Sense distribution of perfumed .............................................................................319 Table A.8. Sense distribution of scented.................................................................................320 Table A.9. Sense distribution of sweet-scented ......................................................................320 Table A.10. Sense distribution of sweet-smelling...................................................................321 List of tables xv Table A.11. Semantic category distribution of fragrant ........................................................321 Table A.12. Semantic category distribution of perfumed.......................................................323 Table A.13. Semantic category distribution of scented..........................................................324 Table A.14. Semantic category distribution of sweet-scented................................................325 Table A.15. Semantic category distribution of sweet-smelling..............................................326 Table A.16. Overall absolute and relative frequencies of the five near-synonyms in each sense ...................................................................................................................327 Table A.17. Absolute and relative frequencies of the five near-synonyms in each sense in P1 (1810–1859)..................................................................................................327 Table A.18. Absolute and relative frequencies of the five near-synonyms in each sense in P2 (1860–1909)..................................................................................................328 Table A.19. Absolute and relative frequencies of the five near-synonyms in each sense in P3 (1910–1959)..................................................................................................328 Table A.20. Absolute and relative frequencies of the five near-synonyms in each sense in P4 (1960–2009)..................................................................................................329 Table A.21. Overall absolute and relative frequencies of the five near-synonyms in each semantic category..............................................................................................330 Table A.22. Absolute and relative frequencies of the five near-synonyms in each semantic category in P1 (1810–1859)................................................................331 Table A.23. Absolute and relative frequencies of the five near-synonyms in each semantic category in P2 (1860–1909)................................................................332 Table A.24. Absolute and relative frequencies of the five near-synonyms in each semantic category in P3 (1910–1959)................................................................333 Table A.25. Absolute and relative frequencies of the five near-synonyms in each semantic category in P4 (1960–2009)................................................................334 Table B.1. Distribution of the near-synonyms according to Animacy...................................335 Table B.2. Mean and median Concreteness rating.................................................................335 Table B.3. Distribution of the near-synonyms according to the Countability........................336 Table B.4. Distribution of the near-synonyms according to Syntactic function ....................336 Table B.5. Distribution of the near-synonyms according to Degree......................................337 Table B.6. Distribution of the near-synonyms according to Text-type..................................337 Table B.7. Period contrasts in fiction.....................................................................................338 DANIELA PETTERSSON-TRABA xvi Table B.8. Period contrasts in non-fiction ..............................................................................339 Table B.9. Period contrasts in periodicals...............................................................................340 Table B.10. Period contrasts in the artificial sense.................................................................341 Table B.11. Period contrasts in the figurative sense...............................................................342 Table B.12. Period contrasts in the indeterminate value.........................................................343 Table B.13. Period contrasts in the natural sense ...................................................................344 Table B.14. Period contrasts in the semantic category ABSTRACT..........................................345 Table B.15. Period contrasts in the semantic category BODY AND PEOPLE.............................346 Table B.16. Period contrasts in the semantic category CLEANING ..........................................347 Table B.17. Period contrasts in the semantic category COSMETICS ........................................348 Table B.18. Period contrasts in the semantic category EARTH,ATMOSPHERE,AND WEATHER .............................................................................................................349 Table B.19. Period contrasts in the semantic category FOOD AND DRINK...............................350 Table B.20. Period contrasts in the semantic category OBJECT...............................................351 Table B.21. Period contrasts in the semantic category PLANTS AND FLOWERS.......................352 Table B.22. Period contrasts in the semantic category SUBSTANCE AND MATERIAL...............353 Table B.23. Period contrasts in the semantic category SENSATION.........................................354 Table B.24. Period contrasts in the semantic category SPACE ................................................355 Table B.25. Period contrasts in the semantic category TEXTILE AND CLOTHING.....................356 Table C.1. Goodness-of-fit statistics comparison of three SVS analyses...............................357 Table C.2. Most frequent noun collocates of sweet-scented in the four periods ....................357 Table C.3. Most frequent noun collocates of sweet-smelling in the four periods...................358 Table C.4. PMI score of the noun collocates of fragrant in the four periods distinguished...359 Table C.5. PMI score of the noun collocates of perfumed in the four periods distinguished.360 Table C.6. PMI score of the noun collocates of scented in the four periods distinguished....361 ABBREVIATIONS AND CONVENTIONS ABBREVIATIONS Corpora, dictionaries, and thesauri AHDOE = American Heritage Dictionary of the English Language BNC =British National Corpus CD = Cambridge Dictionaries COCA =Corpus of Contemporary American English CoD =Collins Dictionary COHA =Corpus of Historical American English HTOED =Historical Thesaurus of the Oxford English Dictionary LDOCE = Longman Dictionary of Contemporary English MD =MacMillan Dictionary MW =Merriam-Webster NHDAE =Newbury House Dictionary of American English OED =Oxford English Dictionary USAS =UCREL Semantic Analysis System Periodization in the history of English OE = Old English ME = Middle English eModE = Early Modern English lModE = Late Modern English PDE = Present-day English Semantic categories ABS =ABSTRACT B&P=BODY &PEOPLE DANIELA PETTERSSON-TRABA xviii CL=CLEANING COS =COSMETICS EAW =EARTH,ATMOSPHERE,AND EARTH F&D=FOOD &DRINK OBJ =OBJECTS P&F=PLANTS &FLOWERS S&M=SUBSTANCES &MATERIAL SEN =SENSATIONS SPACE =SPACES T&C=TEXTILE MATERIAL &CLOTHING USAS Semantic tags (only those referred to in the text) B=THE BODY AND THE INDIVIDUAL B1=ANATOMY AND PHYSIOLOGY B4=CLEANING AND PERSONAL CARE B5=CLOTHES AND PERSONAL BELONGINGS F1=FOOD F2=DRINKS F4=FARMING AND HORTICULTURE H=ARCHITECTURE,BUILDINGS,HOUSES AND THE HOME H5=FURNITURE AND HOUSEHOLD FITTINGS L3=PLANTS M3=MOVEMENT/TRANSPORTATION:LAND M7=PLACES O1=SUBSTANCES AND MATERIALS GENERALLY O1.1 = SUBSTANCES AND MATERIALS GENERALLY:SOLID O2=OBJECTS GENERALLY Q1=COMMUNICATION Q1.2 = PAPER DOCUMENTS AND WRITING Q4=THE MEDIA S2=PEOPLE W=THE WORLD AND OUR ENVIRONMENT Abbreviations and conventions xix W3=GEOGRAPHICAL TERMS W4=WEATHER X3=SENSORY Other AmE = American English AU = Approximately Unbiased p-values BrE = British English HAC = Hierarchical Agglomerative Cluster HCFA = Hierarchical Configural Frequency Analysis MDS = Multidimensional Scaling NFs = Normalized frequencies PMI = Pointwise Mutual Information POS = Part of speech SVS = Semantic Vector Space CONVENTIONS Italic type is employed for example words and phrases in the running text, as well as for the names of dictionaries and corpora, both in full and in abbreviated form. SMALL CAPITALS are employed for semantic concepts or domains and for semantic categories, both those included in my own semantic classification and those provided by the HTOED and USAS. Additionally, SMALL CAPITALS are used for conceptual metaphors and metonymies. Bold type is employed for the purpose of highlighting words in the examples provided. Dictionary citations appear in the format (1538. OED, s.v. perfumed adj. 1), indicating year — if applicable—, name of the dictionary, headword, part of speech, and sense number. Single inverted commas (‘’) are used for meanings and for search queries made in the corpus. The first word of variable names is capitalized (e.g. Semantic category), whereas variable levels are not capitalized (e.g. animate). The font Lucida console is used when referring to packages and functions in R. DANIELA PETTERSSON-TRABA 6 Firth and Sinclair in the 1950s and 1960s. This approach and its emphasis on collocational behavior lie at the core of the methodological and theoretical assumptions of the present dissertation, as well as of much recent work in lexical semantics. Finally, Section 2.4 is concerned with cognitive semantics, a maximalist approach to the study of language, in general, and of meaning, in particular. Two of the most important contributions of this school figure prominently in this section, to wit, prototype theory and usage-based theories of semantic change, which turn out to be relevant for the interpretation of the findings of the analyses carried out in Chapters 5–7. Chapter 3 zooms in on the semantic relation of synonymy. First, the most relevant classifications of synonyms into types and subtypes are discussed in Section 3.1. The main conclusion drawn from these classifications is that a distinction between absolute synonymy and near-synonymy is the most adequate for an empirical investigation into the semantic phenomenon of synonymy. Then, Section 3.2 moves on to an in-depth review of synchronic usage-based corpus studies on specific pairs or groups of near-synonyms that have been published over the last 30 years. Two separate waves of research are first identified and explained in Section 3.2.1, which suggest that considerable developments have recently taken place in this area of research. Then, in Section 3.2.2, the focus turns to studies on nearsynonymous adjectives, since the items belonging to the synonym set analyzed in this dissertation are precisely adjectives. Finally, Section 3.3 provides an overview of diachronic research on lexical synonymy by first considering the few usage-based corpus studies on the historical development of particular pairs or groups of near-synonyms, and then shifting the attention to general processes of semantic change relevant for the diachronic study of synonymy. Chapter 4 is concerned with issues related to the data employed in the analyses provided in subsequent chapters. The chapter opens with a description in Section 4.1 of the synonym set object of study; it outlines the main motivations for its selection and then reviews the information on these lexical items in various present-day and historical dictionaries and thesauri. Next, Section 4.2 discusses the corpus used as a source of data, as well as the data retrieval process. Lastly, Section 4.3 deals with the data annotation process. Considerable space is devoted here to the explanation of the language-internal and language-external variables used for the codification of the data. 1. Introduction 7 Chapter 5 moves on to the first corpus-based study presented in the dissertation, which constitutes a descriptive univariate analysis of the data. The chapter is divided into four sections, each focusing on a different aspect of the synonym set. Section 5.1 provides the overall frequency of the five near-synonyms, as well as their frequency developments over the time span 1810–2009. Then, Section 5.2 turns to the concept PLEASANT SMELLING as a whole, that is, as designated by the five adjectives, and its distribution across senses and semantic categories of modified nouns. Next, Section 5.3 examines the behavior of each adjective separately, thus offering a semasiological perspective of the data. Finally, Section 5.4 deals with the competition between the near-synonyms in different semantic contexts, hence adopting an onomasiological orientation. Chapter 6 delves deeper into the onomasiological structure of the synonym set. This is done by measuring the competition between the near-synonyms across a large number of intralinguistic and extralinguistic factors. Here, the perspective adopted is a multivariate one and hence a series of more sophisticated statistical techniques are employed. These statistical methods are explained in Section 6.1, together with the dataset used for this particular analysis. Section 6.2, in turn, discusses the results of the study by first focusing on the effects of the predictors included (Section 6.2.1) and then offering a preliminary enquiry into the collocational behavior of the adjectives (Section 6.2.2), where their individual noun collocates are taken into consideration. The chapter closes with a discussion of the main findings in Section 6.3, highlighting the importance of the idiosyncratic collocational preferences of the adjectives under scrutiny, a topic which is taken up in Chapter 7. Chapter 7 opens with a description of the dataset and the methods used in the collocational analyses (Section 7.1). Section 7.2 deals with the results of these analyses and their discussion. A bird’s-eye view of the collocational behavior of the near-synonymous adjectives is provided in Section 7.2.1 by having a look at their general co-occurring patterns before moving to a more qualitative examination of the most prominent individual noun collocates in Section 7.2.2. The dissertation concludes with Chapter 8, a summary of the most relevant generalizations that can be drawn from the studies conducted in the preceding chapters and the identification of some areas that can be further improved and taken up in future research on the topic. 2 APPROACHES TO AND NOTIONS IN LEXICAL SEMANTICS Lexical semantics has been a more or less prominent discipline within linguistics since the beginning of the 19th century when it emerged as an established field of study. From a historical perspective, the research conducted in this field over approximately the last 200 years can be classified into five broad theoretical movements or traditions, to wit, historical-philological semantics, structuralist semantics, generative semantics, neostructuralist semantics, and cognitive semantics. The perspectives that these schools have adopted to the study of semantics are varied and there is yet no agreement on how best to describe lexical meaning. Some scholars (e.g. Geeraerts 2010: 277) suggest that the historical development of this field can be characterized on the basis of a series of oppositions that serve to describe the main interests of each of these five movements. First, the schools’ approaches to the study of word meaning vary depending on whether meaning is conceptualized as a purely linguistic entity, separate from general human cognition, or whether the latter, in the form of encyclopedic knowledge, should also be contemplated in semantic description. Second, the five movements diverge in terms of the prominence they assign to language use (i.e. parole and performance in the parlance of Saussure 1916 and Chomsky 1965, respectively). A third dimension along which the traditions differ is that of semasiology (i.e. from word form to meaning) vs. onomasiology (i.e. from meaning to word form), with some schools focusing primarily on the study of the former and others on the latter. Lastly, the emphasis given to either synchronic or diachronic research varies from one school to another: while, for instance, historical-philological semantics is mainly interested in semantic change, embracing a full-fledged diachronic orientation, structuralist semantics generally opt for a synchronic approach. These five schools should not be considered in isolation but rather in relation to one another, as the transition from one movement to the next can be characterized either in terms of continuation or of reaction along the four dimensions just described. For example, whereas the historical-philological tradition provides many ideas that are later revisited and further DANIELA PETTERSSON-TRABA 10 developed within the framework of cognitive semantics, many of its main assumptions are flatly rejected by structuralist semantics. Therefore, to understand the concepts and principles present in one tradition, it is often necessary to have some knowledge of the key developments in the preceding ones. In what follows, a brief overview of some of the main ideas of each of the five movements is presented. However, as the approach adopted in this dissertation follows the neostructuralist distributional corpus-based methodology, while at the same time embracing the usage-based onomasiological stance of cognitive semantics, only those aspects relevant for the development of the assumptions which lie at the core of both traditions are here discussed in depth. Therefore, the aim of this chapter is not to give an exhaustive account of the research conducted within the different schools, but rather to comment on those issues and concepts that have contributed significantly in one way or another to the principles that characterize the approach adopted here.1Due to the diverging foci of the five movements, some of them play a more crucial role than others in the evolution of the theoretical foundations which are of central importance throughout the analyses in Chapters 5–7 and are, consequently, described in greater detail. For instance, in view of the diachronic nature and primarily onomasiological orientation of the analyses conducted in this dissertation, many of the central issues of both historical-philological and structuralist semantics are highly relevant to understand subsequent developments within neostructuralist and cognitive semantics. On the other hand, the contributions of generative semantics are much less pertinent for our purposes. This chapter is structured into four main sections, each of them focusing on the major developments of the five movements separately, with the exception of structuralist and generative semantics, which are treated together, given the great importance attributed to formalization in both movements. As such, the focus of Section 2.1 is on historical-philological semantics, while structuralist and generative semantics are discussed in Section 2.2, where more emphasis is placed on structuralist approaches than on generative ones. In Section 2.3 the neostructuralist methods are described, paying special attention to the principal assumptions of the distributional corpus-based approach. Section 2.4, in turn, offers an overview of the main theoretical concerns of cognitive semantics, focusing particularly on the concepts of prototypicality, salience, and entrenchment, as well as on its major contributions to diachronic 1The information included in Sections 2.1 and 2.2 of the present chapter is mainly drawn from Geeraerts (2010). For a detailed historical overview of the field of lexical semantics, the reader is referred to this monograph. 2. Approaches to and notions in lexical semantics 11 semantics. Finally, Section 2.5 provides a summary of the whole chapter, as well as some conclusions that can be drawn from it. The four oppositions mentioned earlier (i.e. semantic vs. encyclopedic knowledge, use vs. system, semasiology vs. onomasiology, and synchrony vs. diachrony) serve as a guiding theme in the remainder of this chapter in the sense that the five traditions are described in terms of their positions with respect to each of these dimensions. 2.1 HISTORICAL-PHILOLOGICAL SEMANTICS Historical-philological semantics, sometimes also known as traditional diachronic semantics or prestructuralist semantics, dominated the field for approximately 100 years between 1830 and 1930. As mentioned in the preceding paragraphs, this school takes a diachronic perspective towards semantics and thus displays a great interest in meaning change, especially at the level of individual words. As the evolution of isolated words and their senses constitute the principal object of study, the main concern within this tradition is on semasiological structure in the lexicon. It is in order here to explain the difference of perspective between the study of semasiology and that of onomasiology. Whereas the former takes as its starting point particular words in order to study their senses and the concepts that they designate, the latter starts from specific concepts and explores the various expressions which can be used to represent them (Baldinger 1980: 278; Grondelaers, Speelman & Geeraerts 2007: 988–989; Kay & Allan 2015: 2–3, 185, 188). It is worth noting here that these two different views are commonly associated with the semantic phenomena of polysemy and synonymy, respectively (Levshina 2011: 11): whereas polysemy refers to the existence of various related word meanings or senses of one lexical item (i.e. semasiology), synonymy involves the existence of various lexical items denoting the same concept (i.e. onomasiology). One of the major contributions of historical-philological semantics to contemporary diachronic approaches concerns the various classifications of semantic change (cf., for instance, Carnoy 1927; Stern 1931), which are in fact still in use today in refined and further developed forms.2In order to analyze the relations holding between different senses of particular words and the changes that they undergo over time, a great deal of the work within this movement 2Although various other classifications of semantic change within historical-philological semantics have been proposed, Carnoy’s and Stern’s are more elaborate and provide a more fine-grained depiction of the major types of changes. Another influential classification for contemporary research in semantics is that of Ullman (1957; 1962), which despite falling outside the period of the historical-philological tradition, is worth mentioning here. DANIELA PETTERSSON-TRABA 12 revolve around the identification of regular mechanisms of semantic change. Without going into too much detail about the numerous classifications and subclassifications postulated by particular scholars, a brief overview of the main types of changes is provided in what follows (for a visual summary of the different types, see Figure 1 below). A first distinction is typically made between onomasiological and semasiological mechanisms of meaning change. The former involve changes that lead to the association of a newly coined expression to already existing ones in the lexicon. These changes, typically also known as lexicogenetic mechanisms, include morphological processes such as derivation, composition, clipping, and blending, as well as lexical borrowing from other languages (not to be confused with semantic borrowing, see below). However, since the focus within this tradition is primarily on semasiology, most of the types distinguished in these classifications fall within the semasiological group of changes, i.e. those that bring about new senses of already existing words or expressions. Changes of this kind can, in turn, be further divided into connotational and denotational changes, and within the latter scholars typically differentiate between non-analogical and analogical mechanisms. Analogical processes are those in which the new sense of a given word is established in relation to the semantics of a pre-existing expression, either in terms of similarity or in terms of opposition, as in the case of semantic borrowing and differentiation, respectively. Contrariwise, this is not the case in non-analogical changes, such as specialization and generalization, among others (cf. also Geeraerts 1997: 93–102).3 3Since a detailed description of the semantic changes relevant for the study of synonymy, and more specifically to the analyses conducted in the present dissertation, is provided in Section 3.3, the analogical changes included in Figure 1 (e.g. specialization, generalization, and differentiation) are not further discussed at this point. 2. Approaches to and notions in lexical semantics 13 Figure 1. Main types of semantic change distinguished in the historical-philological tradition Concerning the conceptualization of meaning, that is, the opposition between semantic vs. encyclopedic knowledge, the perspective adopted in traditional diachronic semantics is a psychological one, in which language is seen as being closely connected to our general cognitive abilities, which together help us understand the world we live in. Bréal (1897), who brought the mental position of linguistic meaning to the fore, is one of the most important scholars within this movement. Three of Bréal’s main points should be mentioned here to shed further light on this perspective of meaning. First, Bréal contemplates linguistic meanings as thoughts or ideas, that is, as psychological entities, which permit language users to categorize their experiences of the world. It is in this sense that language is directly connected to, and thus not independent from, the rest of human cognition. Second, it follows from this conception of language in general, and of meaning in particular, that semantic changes are also cognitively motivated and thus result from mental processes. As such, the general mechanisms of semantic change identified above (cf. Figure 1) can also be understood as a result of the way in which human cognition works, for instance, by means of metaphorical and metonymical thought processes. Third, Bréal argues that the motivations behind semantic change are ultimately functional in nature because changes occur in order to adjust language to the communicative needs and experiences of its users, which are themselves ever evolving. This psychological conceptualization of meaning, which Bréal emphasized and explored already at the end of the DANIELA PETTERSSON-TRABA 14 19th century, still figures prominently within cognitive semantics more than a century later, as we will see in Section 2.4. A further relevant feature of historical-philological semantics is the standpoint it takes on the importance of language usage and of specific contexts of use in the conception of meaning, especially when it comes to semantic change. In this sense, Paul (1920: 75) argues that we need to differentiate between what he calls usual and occasional meanings. Usual meaning refers to the conventional senses of individual word forms, whereas occasional meaning is applied to the specific readings of the established senses in particular contexts of use. To further illustrate this idea, consider the following two made-up examples featuring the expression red jacket: (6) I really like the red jacket that you are wearing today. (7) Can you please give this book to the red jacket sitting in the corner over there? In (6) we find an instance referring to one of the various usual meanings of jacket listed in the OED, namely ‘[a]n outer garment for the upper body […]’ (OED, s.v. jacket n. I. 1a), whereas (7) displays an occasional meaning of the same word, in which it is used to designate the person sitting in the corner over there who is wearing a red jacket in that specific speech situation. In this case, the occasional meaning emerges from a metonymical pattern which can be characterized as PIECE OF CLOTHING FOR PERSON (e.g. Waag 1908; Nyrop 1913; Paul 1920; Esnault 1925). However, even though this particular reading of jacket can be inferred in this context, it has not been conventionalized as an established sense of the word in question. This example serves to demonstrate two fundamental aspects of Paul’s conception of meaning. On the one hand, it reveals the significance that he ascribes to language use so as to understand the meaning of particular instances of words since the occasional meaning of a word may differ substantially from its usual meaning. In addition, context also plays a fundamental role, given that the realization of “the occasional meaning amounts to selecting the appropriate reading from among the multiple established senses of a [polysemous] word” (Geeraerts 2010: 15). Thus, for instance, the English word hot is bound to be interpreted differently if it serves as a modifier of nouns referring to liquids such as water or coffee than if it serves as a modifier of nouns denoting people.On the other hand, the distinction usual vs. occasional meaning shows that Paul views meaning as extremely dynamic and flexible due to the concrete interpretations that lexical items acquire in specific contexts of use. 2. Approaches to and notions in lexical semantics 15 According to Paul, language use is also essential to understand semantic changes taking place at the level of individual words, that is, semasiological changes. By making use of the distinction he establishes between the two types of meaning, he goes on to explain the way in which occasional meanings may develop into usual ones over time. Paul argues that this happens by means of repetition, in which case an occasional meaning of a word, by being used more and more frequently, comes to establish itself as another conventional sense of that word alongside others. In other words, the occasional meaning becomes independent from the usual meaning it derives from and can, consequently, be interpreted autonomously or without context. In this sense, we may say that it has become decontextualized. Moreover, Paul suggests that the opposite process, that is, the modulations of usual into occasional meanings, is directly linked to the classifications of semantic change introduced above (e.g. metonymy, metaphor, and specialization), since these mechanisms are precisely those through which an occasional meaning arises from a usual one. A case in point are examples (6) and (7), in which a metonymical extension resulted in a new occasional meaning derived from an established one. If contextually dependent readings such as this occur frequently enough in actual speech, they may become conventionalized and over time establish themselves in the language system. In fact, Paul’s view of semantic change serves as a clear historical precedent of the usage-based model of change of cognitive semantics, which is explained in Section 2.4.2. As we have seen, both Bréal’s (1897) and Paul’s (1920) conceptions of meaning as a cognitive and contextual phenomenon clearly emphasize the causes leading to semantic changes and can thus be said to go a step beyond the focus dominating the historicalphilological movement until the 1860s, namely that of identifying and classifying semantic changes without paying much attention to the explanatory forces underlying such developments. Bréal’s and Paul’s psychological models of meaning prevailed in the historicalphilological school from the 1860s onwards, although different views of these models coexisted. Scholars such as Waag (1908), Erdmann (1910), Nyrop (1913), and Esnault (1925) postulated their own views and opinions on the matter. A case in point, and perhaps a crucial one to understand later developments in the history of cognitive semantics, is Erdmann’s claim about the widespread presence of vagueness, and not just polysemy, in the lexicon (1910: 5, quoted in Geeraerts 2010: 22): DANIELA PETTERSSON-TRABA 22 either Stuhl or Sessel without any of the two prevailing over the other. The organization of fields, according to Gipper, is consequently based on central or focal areas, with the most clearcut examples of a given field, and peripheral ones, where the borderline between related fields is not as obvious. Another important point made by Gipper is that the naming process is not based on random choice or personal whim, but rather displays a structured system. In the case of Stuhl and Sessel, for instance, the principal function of the seat to be named plays a major role, with ‘comfort’ being associated with Sessel and ‘practicality’ with Stuhl. Gipper’s conclusion about the fuzziness of field structure, just as in the case of Erdmann’s conception of semantic vagueness, is in agreement with the prototypical organization of categories which is later postulated within the theoretical framework of cognitive semantics (cf. Section 2.4.1). Componential analysis, which evolved somewhat later than Trier’s (1931; 1934) original field theory and can be regarded as a further development of it, focuses primarily on the formalization of the existing relations between the interconnected items in a semantic field, with particular emphasis on the oppositions between neighboring words (e.g. Coseriu 1964). As such, studies within this approach are devoted to the ways in which these relations and contrasts should be given formal expression. Alternative descriptive theories have been put forward to account for conceptual or semantic substance (cf., for instance, Nida 1951; Goodenough 1956; Pottier 1964; 1965), but essentially they all proceed by dividing the semantics of lexical items into smaller units called semantic components. This is done with the intention of trying to establish a fixed number of characteristics or primitives that could be used to distinguish between related lexical items. An example would be to decompose the words man and woman into the components [+HUMAN +MALE +ADULT] and [+HUMAN -MALE +ADULT], respectively, and to establish that the difference between them lies in their position regarding the component MALE (Kay & Allan 2015: 33–34). Given that the formalization of semantic content lies beyond the scope of this dissertation, componential analysis is not further described here. The third branch of the structuralist movement especially relevant for our purposes is relational semantics, which is primarily associated with the work of Lyons (1963; 1977), although in-depth evaluative reviews of this approach are also provided in Cruse (1986) and Murphy (2003). Relational semantics, which builds on Trier’s (1931; 1934) field theory and componential analysis, is devoted to the description of the structural connections that hold between neighboring words within semantic fields. According to Lyons, these relations are 2. Approaches to and notions in lexical semantics 23 more intricate than those of mere opposition present in compositional analysis and the work of scholars such as Coseriu (1962; 1964). Consequently, a more all-encompassing apparatus is needed to adequately explain the structure of language, particularly one in which a clear separation is made between the autonomous linguistic level and the extralinguistic, encyclopedic one. Lyons therefore postulates a broader model of meaning relations in which the structure of language is comprised by a complex network of multiple types of semantic links that connect word senses to one another. As such, it brings to the fore the systematic study of a wider range of lexical —or perhaps more accurately termed senseor semantic— relations than just antonymy (i.e. semantic oppositeness), which are deemed equally relevant, including hyponymy, meronymy, and, more crucially for this dissertation, synonymy. To this purpose, much of the work within relational semantics is devoted to the identification and description of these four semantic relations, among others, as well as the different subtypes established within them: for instance, gradable as opposed to non-gradable antonymy or total as opposed to partial synonymy. For our goals, it is here enough to provide a brief definition of the various semantic relations distinguished without a detailed consideration of each of them. The different types of synonymy established principally in the works of Lyons (1963; 1968), Cruse (1986), and Murphy (2003) are further discussed in Chapter 3, which is devoted to synonymy in general and to the existing distributional corpus-based studies on the internal semantic structure of specific pairs and groups of synonyms in particular. Starting with antonymy, this semantic relation is best described as an umbrella term for various and slightly diverging kinds of oppositions of meaning, such as those linking together terms which are (i) placed on opposite ends of a gradable cline (e.g. good vs. bad;young vs. old) or (ii) mutually and logically exclusive in the sense that they contradict one another rather than exist on a continuous cline (e.g. mortal vs. immortal;single vs. married). Second, hyponymy is typically defined in terms of semantic inclusion which couples a more general with a more specific term—or rather concrete senses of the terms in question— belonging to the same semantic field. As an example consider the terms tulip and flower. In this case, the former term (i.e. tulip), typically called hyponym or subordinate, is included in the definition or description of the latter (i.e. flower), which is called hypernym or superordinate. Other lexical items which are also hyponyms of flower (e.g. gardenia and forget-me-not)are known as co-hyponyms of tulip, since they are situated at the same level within the hierarchy. DANIELA PETTERSSON-TRABA 24 Contrasting with hyponymy, we find the relation of meronymy, a part-to-whole or wholeto-part relation such as that existing between the terms eye and face (i.e. a component part of a material whole) or between captain and crew (i.e. member meronymy). Although these relations can be conceptualized as representing inclusion, they are to be distinguished from hyponymy since the types of inclusive relations that hold in the two cases differ. To further elucidate this distinction, consider again the example of flower introduced above. As already mentioned, this term holds a relation of hyponymy with tulip, as the latter is a particular type of flower.In contrast, the relation between flower and petal is a meronymic one, as a petal is an integral element of a flower, thus designating a specific part of it. Finally, synonymy is sometimes thought of, in layman’s terms, as a relation of identity of meaning between two or more words or word senses. However, a somewhat less strict definition of synonymy is typically given in the specialized literature since, as will become clear in the course of Chapter 3, synonyms do not often have identical meaning and can thus not always substitute one another. If two words or word senses can always be used interchangeably, we have a case of total or absolute synonymy (as opposed to cognitive or near-synonymy). Moreover, a distinction is made in the specialized literature depending on whether the terms are semantically equivalent in all their senses or only in one or some of them. Near-synonyms, on the other hand, even though equivalent or even identical on the denotational dimension of meaning, often differ in other aspects of meaning such as connotation, style, or collocation, which prevents them from being used interchangeably in all contexts.7 The strong emphasis of componential analysis on formalization is also a prominent feature of generative semantics, which came after structuralism, that is, from around the 1960s onwards. In fact, componential analyses were frequent also within this movement, for instance, in the work of Katz & Fodor (1963; see also Katz 1972). However, despite the common ground with structuralism, the conception of meaning changed yet again in generativism, with meaning being once more conceived of as having a psychological nature. In other words, a mental orientation towards the study of meaning is now brought to the fore again. Two approaches to semantics can be distinguished within generativism, namely interpretative semantics and generative semantics. In interpretative semantics, semantics has a secondary role, since emphasis lies mainly on syntax. Issues of semantics considered important within this branch are in fact related to syntax, for instance, argument structure or quantification, which aid in 7Further details about the different types of synonymy and dimensions of meaning are provided in Section 3.1. 2. Approaches to and notions in lexical semantics 25 syntactic analysis. On the contrary, lexical semantics is completely disregarded. This interpretation of semantics is the one held by the majority of scholars in this tradition (Geeraerts 2006a: 143–144). Contrariwise, in the second branch of generative semantics, the goal is primarily to establish a semantically-based syntax.8This approach was favored by scholars such as Lakoff (e.g. 1970; 1971), Fillmore (e.g. 1968), and Langacker, who later became important figures in cognitive linguistics. One of their core assumptions concerns the impossibility of drawing a clear line between semantic and encyclopedic knowledge, a distinction which is rejected in cognitive semantics, as we will see in Section 2.4. The major contribution of structuralist semantics has been to provide a theoretical framework to understand the relations between neighboring words in the language system, that is, onomasiology, which although not entirely lacking in historical-philological semantics, had not received sufficient attention before. However, several problems exist regarding how onomasiological structure is perceived and studied by structuralists. First, the emphasis put on semantic relations leads to an underestimation of the importance of semasiological processes. For instance, whereas Lyons clearly highlights that the semantic relations identified and established within the relational approach hold between senses rather than between words, almost no attention is paid to sense discrimination of particular words. Moreover, the strict division between a supposedly independent semantic level and an encyclopedic one emphasized in structuralist semantics cannot be maintained. Gipper’s (1959) study on Stuhl and Sessel discussed above clearly demonstrates the need for extralinguistic criteria to adequately describe the differences between related concepts, since the factor ‘functionality’ (i.e. practicality vs. comfort), which is an aspect that relates to the actual referent of the concept INDIVIDUAL SEAT WITH A BACK, is crucial in order to explain the alternation between these two words. This is just one instance in which the encyclopedic level permeates what are supposed to be purely semantic descriptions.9Finally, and more importantly for our purposes, the excessive focus on the 8The heated debates and differences of opinion between the two strands of generativism, i.e. generative semantics vs. interpretative semantics, concerning the demarcation between grammar and semantics is often referred to as the “linguistics wars,” which dominated part of the linguistic scene during the 1960s (cf., for instance, Harris 1993; Geeraerts & Cuyckens 2007: 12). 9 In a recent study carried out within the framework of cognitive semantics, Levshina (2015: 373–384) shows that further factors related to the encyclopedic level of knowledge apart from ‘functionality’, such as ‘material’ (e.g. fabric, leather, and wood), ‘texture’ (e.g. soft vs. not soft), and ‘seat height’ (i.e. high, low, or adjustable), to mention a few, also play a major role in the choice between Sessel and Stuhl. DANIELA PETTERSSON-TRABA 26 language system and the general lack of concern that this onomasiological orientation towards meaning shows for actual usage and language context is highly problematic. This creates major complications when it comes to explaining the existence of alternative forms such as synonyms, since it leads to descriptions which are too restrictive, as they merely focus on the identification and classification of types of semantic relations. However, for an in-depth analysis of the relations of particular synonym sets, it is certainly not sufficient to state that such relations exist between the members of the set, but an analysis of the multiple factors determining the choice between such expressions is also required. In other words, even more interesting than knowing about synonymous relations and the impact they have on the structure of language is knowing in which contexts we use one word instead of another and the reasons why we do so. This change of perspective, however, involves usage considerations and required a shift from a systemic onomasiological orientation, such as that postulated by structuralists, to a usage-based onomasiological orientation, which emerges in later traditions of semantics (cf., for instance, Grondelaers, Speelman & Geeraerts 2007). The lack of examination of the motivations underlying lexical variation is one of the main causes for the abandonment of the structuralist view of onomasiology by current scholars in the field. With regard to the contributions of generativism, it is enough for our purposes to mention that since the major force within this framework was interpretative semantics, the study of meaning was relegated to a very minor topic, secondary to syntax. Nonetheless, one enduring and significant impact of generative semantics on the development of the field of lexical semantics is the postulation of several assumptions which lie at the core of contemporary theories, first put forward by scholars such as Lakoff, Fillmore, and Langacker, who rejected the division between an autonomous semantic level of knowledge and an encyclopedic, extralinguistic one. 2.3 NEOSTRUCTURALIST SEMANTICS The fourth major school to the study of meaning, typically known as neostructuralist semantics, comprises a series of approaches which are possibly of a more heterogenous nature than those in the traditions discussed so far. This is because, as a matter of fact, these approaches do not adhere to any particular theoretical framework, but are instead grouped under the label of neostructuralism due to their affinity with the structuralist approaches and conceptions of meaning introduced in Section 2.2. As the prefix neosuggests, they all can, in one way or 2. Approaches to and notions in lexical semantics 27 another, be contemplated as a renewed or at least modified form of the structuralist methodologies. At the same time, several of the neostructuralist approaches to semantics build on the generativist idea mentioned at the end of Section 2.2, according to which the need to incorporate the psychological nature of meaning into semantic description is again revived. As such, two broad groups are typically distinguished among the neostructuralist approaches, to wit, those with a compositional orientation and those with a relational orientation. A brief overview of the former (i.e. compositional approaches) is provided in Section 2.3.1, which examines the problems posed by such approaches for diachronic studies due to their failure to account adequately for semantic change. In turn, relational approaches are the focus of Section 2.3.2. Within the relational orientation, the distributional corpus-based analysis is discussed in more length, since it provides a methodological basis for the analyses in the practical part of the present dissertation (cf. Chapters 5–7). 2.3.1 The compositional orientation Among the approaches taking a compositional orientation we find those that continue with the formalization of meaning set out in structuralism and generativism (cf. Section 2.2), but which explicitly do so by adopting a cognitive orientation to semantics in the sense of admitting that encyclopedic knowledge and contextual flexibility of meaning somehow need to be borne in mind. This, however, does not mean that scholars advocating such approaches abandon the restrictive tendencies of earlier types of compositional analyses, but rather that they try to find ways to accommodate this reduction while at the same time accounting for the fuzziness of meaning and its cognitive reality. Four well-known theories adopting this stance are Wierbicka’s natural semantic metalanguage (e.g. 1972; 1996; 1997; cf. also Goddard & Wierzbicka 2002), Jackendoff’s conceptual semantics (e.g. 1990; 1996; 2006), Bierwisch’s two-level semantics (1983a; 1983b; cf. also Bierwisch & Lang 1989; Lang 1993), and Pustojevsky’s generative lexicon (1995). Despite openly acknowledging the existence of different levels of knowledge and the importance of accounting for this multi-level reality, all four models arrive at the same conclusion, namely that the fuzziness, vagueness, and flexibility of meaning are due to the extralinguistic or encyclopedic level. In sum, what they imply is that whereas the semantic level is tidy and neat, the encyclopedic and contextual one is vague and fuzzy. Such a conclusion is what permits these models to keep semantic descriptions and definitions at a manageable size, by focusing only on those aspects of meaning of a word or DANIELA PETTERSSON-TRABA 28 concept which are invariable —as opposed to those aspects which fluctuate— in the case of Wierzbicka, or only on those features of meaning which are supposed to be stored in the mental lexicon —as opposed to those features that can be contextually derived— in the case of Bierwisch. As a consequence, we can clearly see a parallelism with the problems mentioned above in relation to structuralist approaches (cf. Section 2.2), since the same limitations are applicable also in these cases. First, the research interest again lies not on the explanation of the inherent variability of concepts and categories of meaning, which is the most problematic to account for, but instead on the more stable characteristics of meaning. Second, these compositional approaches face the same difficulties in arriving at sufficiently general and satisfactory definitions that are able to encompass all members of a given concept or category while at the same time adequately distinguishing neighboring concepts and categories from one another. Another point to highlight and which is possibly more important for our purposes is the disadvantages such approaches pose when employing a diachronic perspective to the study of meaning, especially when semantic changes taking place over time are accounted for. This is so because drawing such a clear division between the two levels of knowledge (i.e. semantic and encyclopedic), and limiting the study to what is invariant or stored in the lexicon, consequently disregarding contextually derived meanings in semantic descriptions, is not feasible when dealing with diachronic data, as what is considered an invariable feature of meaning today may in fact become a variable one in the future, or vice versa. This idea of a close connection —rather than a strict separation— between variable and invariable or between stored and derived aspects of meaning, which would allow us to adequately explain semantic changes, was already present in the prestructuralist approach discussed in Section 2.1, especially in Paul’s (1920) notions of usual and occasional meanings (cf. examples (6) and (7) above). This conception of meaning in which the two levels of knowledge are brought together lies at the heart of cognitive semantics, which finds a much more suitable way of dealing with and explaining semantic changes by drawing on empirical data rather than introspection, as we will see in Section 2.4. However, before turning to this full-fledged psychological theory of meaning, we will have a look at how the relational approach postulated by Lyons (1963; 1968; 1977) within structuralist semantics (cf. Section 2.2) was further refined by neostructuralist scholars, especially within the method known as distributional corpus-based analysis. 2. Approaches to and notions in lexical semantics 29 2.3.2 The relational orientation The second type of approach distinguished within neostructuralist semantics is said to develop the systematic description of sense relations first initiated by Lyons (1963; 1968; 1977). As in the case of the approaches building on the compositional method discussed in Section 2.3.1, we can distinguish two different branches of these more up-to-date relational traditions. On the one hand, we find two approaches with a strong focus on the formalization and the codification of semantic knowledge into a formal language, namely WordNet (cf., for instance, Miller 1990; Fellbaum 1998; 2005) and Mel’þuk’s meaning-text theory (e.g. 1988; 1996; 1998). Due to their emphasis on formalization, these two approaches indeed share some common ground with the four neostructuralist traditions of componential analysis referred to above (i.e. natural semantic metalanguage, conceptual semantics, two-level semantics, and generative lexicon), but at the same time they go a step further by introducing computational devices to formalize meaning. WordNet is an online lexical database in which words are structured on the basis of the semantic relations they maintain with other words, including the traditional relations mentioned in Section 2.3, to wit, antonymy, hyponymy, meronymy, and synonymy.10 As such, a search in this database of a particular lexical item yields a whole range of examples of other words with which it is connected in its different senses through any of the aforementioned relations. The meaning-text theory, put into practice in the Explanatory Combinatorial Dictionary (Mel’þuk, Clas & Arbatchewsky-Jumarie 1984), also links words which hold any of the four paradigmatic sense relations, but it includes additional types of relations such as grammatical and morphological ones. Moreover, syntagmatic relations of co-occurrence are also contemplated. In this way, the meaning-text theory moves further away from Lyons’ original relational approach, whereas WordNet adopts a more classical perspective. As already mentioned, given that formalization is not an important issue in this dissertation, these two approaches are not discussed any further. The second type of approach distinguished within the neostructuralist relational orientation consists in a full-blown usage-based distributional methodology which, instead of emphasizing the formalization of meaning, pays particular attention to the recurrent patterns of words. This is done by closely analyzing how lexical items are used in actual language and by examining different types of co-occurrence information, thus embracing a syntagmatic perspective. Although syntagmatic relations were not totally absent from the structuralist and the 10 For more details about this database, see https://wordnet.princeton.edu/. DANIELA PETTERSSON-TRABA 30 generativist views of meaning, comprehensive research adopting a full-fledged syntagmatic orientation only arrived with the advent of neostructuralist semantics in the form of the distributional approach advocated by scholars such as Levin (1993) and Sinclair (1966; 1987), among others.11 However, we can distinguish various ways of implementing the distributionalist view in actual semantic research. The first form, which is of less relevance for our purposes, is embraced by Levin in her 1993 study of English verbs, in which she arrives at a classification of verbs depending on their syntactic patterns. In particular, she analyzes the patterns of alternations in which verbs can appear, demonstrating that these patterns are recurrent for a whole range of individual items. As we can see, although Levin’s approach indeed assumes a close link between syntax and lexis, this type of distributional research mainly takes a syntactic perspective. The second form of distributionalism is that set out by Sinclair (1966; 1987; 1991a; 2004) and later described and applied by a wide number of contemporary scholars working in the field of lexical semantics, including the pivotal works by Stubbs (1993; 2002a), Biber et al. (1998), and Partington (1998). Even though Sinclair is considered to be the pioneer of this approach, its main principles are based on ideas already present in earlier work, including that by Wittgenstein (1953), who asserts that meaning is equal to actual usage, and, more importantly, by Firth (e.g. 1935; 1957a). Firth was one of the first scholars to suggest that distributional patterns and, in particular, collocation, which are dealt with in detail below, are essential for the analysis of lexical semantics. Many of Firth’s statements on the importance of contextual relations still enjoy wide currency in contemporary lexical semantic research. In two of his most cited works, he claims that “[…] the complete meaning of a word is always contextual […]” (1935: 37) and that language users “shall know a word by the company it keeps” (1957a: 11). In other words, the semantic and functional features of lexical items closely correlate with their distributional patterns. However, the methodology invoked by Firth in such claims did not start gaining much importance until the 1960s. Sinclair (1966) extends Firth’s ideas by stating that the principal duty of researchers devoted to lexico-semantic analysis is to describe the collocational tendencies of lexical items by means of the examination of large collections of natural language data, i.e. corpora, which were starting to be compiled at that time. In fact, the 11 &IIRULQVWDQFHWKHGLVFXVVLRQRIWKHZRUNFRQGXFWHGE\3RU]LJDQG+DUULVLQ6HFWLRQDQG0HO¶þXN¶V meaning-text theory in the present section. 2. Approaches to and notions in lexical semantics 31 most important breakthroughs of the distributional method spring from its implementation in corpus analysis, which Sinclair vigorously advocates. The tradition originating in Firth’s and Sinclair’s work is often said to rest on three major tenets, each of which is discussed individually below in further depth: (i) a radical usage-based orientation, (ii) a paramount importance of collocation, and (iii) a strong link with statistical analysis (cf. Geeraerts 2010: 167). The first of these three principles is clearly reflected in both of Firth’s claims cited above. The emphasis on a usage-based perspective in order to reach an adequate description of meaning remains a fundamental doctrine. The usage-based orientation was first highlighted by Firth in his explicit rejection of the introspective methods defended by both structuralists and generativists. In fact, he completely dismisses the analyses of invented sentences which are presented in isolation and states that it is unrealistic to examine meaning without access to the whole context in which these sentences occur if one wants to achieve serious and significant results. Despite his valuable critiques of the limitations of introspective data, Firth’s own analyses are not based on much actual text, mainly because of the lack of resources to conduct such systematic research of usage patterns that corpus data would have allowed (cf. Stubbs 1993: 8–9). In this sense, much of his research suffers from the same limitations as that of historical-philological semantics (cf. Section 2.1), which prevents it from becoming a fully developed usage-based approach when put into practice. Consequently, even though Firth’s conception of the study of meaning clearly reflects, in theory, the principles which later came to lie at the core of the distributional corpus-based theory, it is not until the arrival of Sinclair’s work that the actual reappraisal of the Saussurean and Chomskyan dichotomies, that is, langue vs. parole and competence vs. performance, takes place. Firth, Sinclair, and other scholars such as Halliday instead opt for a view of language in which langue/competence and parole/performance are closely connected and truly interdependent, as they affect one another. Here, the weather and climate comparison provided by Halliday (1991: 33–34; 1992: 66–67) serves as a clear depiction of the interrelation of system and use —or instance, in Halliday’s terminology (1991: 66): [T]here is language as a system, an abstract potential, and there are spoken and written texts, which are instances of language in use. But the “system” and the “instance” are not two distinct phenomena. There is only one phenomenon here, the phenomenon of language […]. If I may use again the analogy drawn from the weather: the instance- DANIELA PETTERSSON-TRABA 38 earlier allow the quantification of the probability of association between collocates. PMI is a measure of collocational strength that compares the likelihood of two words co-occurring visà-vis the likelihood of the individual items (Church & Hanks 1990: 23): a score of zero implies that the words do not typically co-occur, whereas a score of three or higher indicates that they are strongly associated. There exists a variety of association measures to compute collocational strength, including t-score, log-likelihood, and ȴP (e.g. Gries 2013), among others, all of them tools to move away from introspective and impressionistic judgements about what constitutes a collocation. Additional statistical techniques have been used within the distributional corpusbaserd approach not only to measure collocational strength, but also to model the semantic dis(similarities) between related lexical items on the basis of their collocational behavior and other distributional patterns, including stylistic and syntactic ones. These include regression analyses (Speelman 2014; Levshina 2015: Chapters 7, 12–13), both linear and logistic models, various clustering techniques, such as hierarchical agglomerative cluster or HAC analysis (e.g. Gries 2009: 316–319) and conditional inference trees (e.g. Levshina 2015: 291–300), as well as semantic vector space (SVS) modeling (e.g. Sahlgren 2006; Heylen et al. 2008; Peirsman, Heylen & Geeraerts 2008; Hilpert & Correia Saavedra 2017). Some of these statistical techniques are used in the semantic analyses provided in Chapters 6 and 7, where they are explained in depth. The distributional corpus-based approach has been successfully applied not only to lexical semantics, but also to research in fields such as language teaching and learning (cf., for instance, Nattinger & DeCarrico 1992; Nesselhauf 2005; Ellis, Simpson-Vlach & Maynard 2008; Paquot & Granger 2012; Gablasova, Brezina & McEnery 2017) and computational linguistics (cf., e.g., Kilgarriff 2006; Kilgarriff & Rychlý 2010; Schulz & Aziz 2016), among others, especially over the last two decades. Moreover, different linguistic phenomena have been analyzed following this approach, particularly polysemy and synonymy. For instance, by examining the distributional patterns of polysemic words, it is possible to distinguish their different senses (e.g. Church et al. 1991; Biber 1993). More importantly for our purposes, an ever-increasing number of studies have demonstrated the effectiveness of the distributional corpus-based approach to uncover fine-grained semantic and usage distinctions between sets of synonyms. A detailed overview of such research is provided in Sections 3.2 and 3.3, which deal with synonymy in depth. Undoubtedly, one of the major contributions of this radically distributionalist orientation to meaning has been the incorporation of collocation first into 2. Approaches to and notions in lexical semantics 39 semantic description and later into linguistic research in general. In fact, most software employed nowadays in corpus linguistics includes tools for the statistical analysis of collocational behavior, as is the case of WordSmith Tools (Scott 2020). Thus, due to its wide range of applications, the neostructuralist distributional approach has undoubtedly meant a significant improvement from a methodological viewpoint over structuralist theories. However, as argued by Geeraerts (2010: 177–178), the approach is ultimately a method rather than a theory, as it is not always clear how the distributional patterns uncovered relate to theoretical issues in lexical semantics. Recently, distributional methods have been adopted for the study of meaning under the theoretical framework of cognitive semantics, which embeds the study of meaning in the study of human cognition in general. As we will see in the next section (cf. Section 2.4), this theory converges with the original neostructuralist distributional method in many ways, in particular due to their common reliance on actual usage-based data and their quantitative orientation. Given the rapid increase of studies in cognitive semantics with a distributional stance over the last two decades, the importance and utility of this methodology cannot be underrated. 2.4 COGNITIVE SEMANTICS The fifth and last tradition to be discussed and probably the approach that enjoys the widest currency in contemporary semantic research is that of cognitive semantics, a subfield within the framework of cognitive linguistics. As mentioned earlier (cf. Section 2.2), cognitive linguistics emerged in the 1980s with the work of Lakoff, Fillmore & Langacker, who are considered to be the founding fathers of this school (cf. Geeraerts & Cuyckens 2007: 3). Contrary to previous traditions, particularly generativism, this framework assumes that meaning is of chief importance, given that language is, as made clear in the editorial statement of the number one issue of the journal Cognitive Linguistics published in 1990, “an instrument for organizing, processing and conveying information” (quoted in Geeraerts 2006b: 3). Thus, semantic description is brought to the fore in linguistic research. As the name of the approach suggests, and as has been mentioned at several points throughout this chapter, meaning is conceptualized in this tradition as a psychological phenomenon whose study should therefore be integrated into the study of cognition in general. With regard to the oppositions or dichotomies that have characterized the development of the field of semantics and that have served as the guiding theme of the present chapter, to wit, semantic vs. encyclopedic DANIELA PETTERSSON-TRABA 40 knowledge, use vs. system, semasiology vs. onomasiology, and synchrony vs. diachrony, cognitive semantics takes a maximalist approach. The maximalist perspective implies that some of these oppositions are completely abolished (namely, semantic vs. encyclopedic knowledge and use vs. system), whereas others are maintained (namely, semasiology vs. onomasiology and synchrony vs. diachrony) but none is given priority over the other. On the one hand, the elimination of the first two oppositions is crucial in order to go beyond the restrictive types of semantic description put forward by approaches in which a clear-cut separation is established between a pure semantic level and an encyclopedic one, and between language structure and language use. On the other, a reevaluation of the latter two oppositions is important to provide a more adequate description of meaning in which the interconnection of semasiological and onomasiological processes, and of the synchronic and diachronic orientations, are systematically examined, thus leading to a more comprehensive picture of meaning and of meaning relations. This maximalist stance towards semantics brings about four central principles that lie at the core of the conception of meaning in cognitive semantics, which according to Geeraerts (2006b: 3–6), can be summarized as follows: (i) Linguistic meaning implies perspectivization. (ii) Linguistic meaning is characterized by its dynamicity and flexibility. (iii) Linguistic meaning is encyclopedic and therefore not independent from the rest of human cognition. (iv) Linguistic meaning is experiential in nature and heavily dependent on actual use. First, given that the principal function of language is to categorize the world and our experiences in it, meaning indubitably involves perspectivization. This is so because, as a vehicle for categorization, language does not merely reproduce reality in an objective way, but rather reflects it from the subjective standpoint of language users. It is in this sense that language is said to “[…] impose a structure on the world” (Geeraerts & Cuyckens 2007: 5) or to “construe the world” (Geeraerts 2006b: 4). The way in which meaning imposes such a structure depends on the experiences, requirements, and desires of language users, which vary from one society to another, from one culture to another, and from one period to another. This leads to the second belief about meaning, which is directly related to the first one, namely its dynamicity and 2. Approaches to and notions in lexical semantics 41 flexibility. If language —and meaning— is a way of modeling the world, and since the world we live in is ever changing, language and meaning must change as well. As new realities and experiences arise, new categories and concepts need to be created or existing ones need to be modulated to make room for such adjustments. Therefore, language cannot be an inflexible structure, but must instead be flexible and dynamic, characterized by great and constant adaptability. Connected to the first principle, the third tenet implies that meaning necessarily involves encyclopedic knowledge and, therefore, a purely semantic level of knowledge does not exist. This links up with the idea that meaning is a cognitive phenomenon and, as such, it obviously holds a close relation with the rest of human cognitive abilities. Consequently, language in general and meaning in particular do not constitute autonomous entities that have to be analyzed independently of other systems, as is believed in structuralism (cf. Section 2.2). Rather, the study of meaning needs to be embedded into the study of cognition. Finally, because meaning is built upon the experiences of the world of language users, it is undoubtedly experiential in nature, and thus language use is key to understand the language system. The linguistic experiences of speakers do not directly involve abstract syntactic patterns, but instead concrete instantiations of those structures, that is, actual words which are recurrently used in certain patterns. This principle directly echoes the ideas postulated in the neostructuralist distributional corpus-based approach introduced in Section 2.3.2. These four central tenets constitute the foundation of the major models and theories within cognitive semantics, such as the conceptual theories of metaphor and metonymy, idealized cognitive models, frame theory, prototype theory, and the various usage-based approaches to meaning change. In what follows, only the latter two contributions are discussed, as they are the most relevant for the subsequent analyses in Chapters 5–7.16 Section 2.4.1 explains the 16 The first three do not figure prominently in the present dissertation and are therefore only discussed briefly here. The conceptual theories of metaphor and metonymy focus on the interconnection between the different senses of words which are linked by means of the mechanisms of semantic extension of metaphor and metonymy, understood within this theory as cognitive rather than purely linguistic processes. A metaphor is a cognitive process whereby one domain (i.e. target domain) is understood in terms of a different domain (i.e. source domain) due to some resemblance perceived between them. Therefore, a mapping emerges from the source domain to the target domain following the pattern TARGET IS SOURCE, for instance, IDEAS ARE FOOD, as in I just can’t swallow that claim (Lakoff & Johnson 2003: 43). Metonymy, in turn, is a cognitive process which connects related concepts within the same domain. In this case, we speak about vehicle and target, the former denoting the concept which is used to refer to the latter along the pattern VEHICLE FOR TARGET, for instance, AUTHOR FOR WORK,as inI’m reading Shakespeare (Kövecses 2010: 171–172; cf. also Lakoff 1987; 1993 for metaphor and Panther & Radden 1999 for metonymy). Idealized cognitive models refer to the set of assumptions and beliefs of the world, on which linguistic and other DANIELA PETTERSSON-TRABA 42 concepts of prototypicality,gradience,entrenchment, and salience, while Section 2.4.2 deals with usage-based theories of semantic change, which constitute a continuation or reappraisal of the diachronic research set out in the historical-philological tradition to semantics (cf. Section 2.1). 2.4.1 Prototype theory: Prototypicality, gradience, entrenchment, and salience As became evident in the introductory paragraphs to this section, in cognitive semantics meaning is essentially a means by which we categorize the world. As such, research concerning the internal structure of categories is of fundamental importance in this framework. One case in point is prototype theory, which since its origin in psycholinguistics in the 1970s in the work of Rosch (cf., for instance, 1973a; 1973b) has become extremely influential also in the fields of linguistics and semantics, where it has been applied to both semasiological and onomasiological structures and phenomena. The aim of this section is not to provide a detailed account of the theory, but merely to introduce some of the basic notions associated with it, to wit, prototypicality, gradience, salience, and entrenchment, concepts which will aid in the explanation of some of the changes undergone by the synonyms object of study in subsequent chapters.17 However, before explaining these notions, a few words seem in order concerning some of the fundamental assumptions of prototype theory. Prototype theory emerged and gained importance in the 1970s due to the limitations of the models based on the classical orientation to category structure postulated by Aristotle, which had dominated the field of semantics in the structuralist and the generativist compositional approaches, but also in Trier’s (1931; 1934) original field theory (cf. Section 2.2). Such approaches assumed that categories —and the members belonging to them— can be defined by a set of necessary and sufficient conditions, thus reducing category membership of a particular object or experience to an ‘all vs. nothing’ or ‘either vs. or phenomenon’. However, as we have cognitive processes depend (Geeraerts 2010: 224; cf. also Lakoff 1987). This theory, therefore, emphasizes the strong connection between linguistic and encyclopedic knowledge. Frame semantics shares many of its main assumptions with idealized cognitive models, but goes beyond it by arguing that our interpretation of the world is filtered not only through cognitive models, but also through the different ways in which we give linguistic expression to these models (Fillmore 1977; 1985; Fillmore & Atkins 1992). Here, the notions of perspectivization and construal play a crucial role (cf., for instance, Verhagen 2007). 17 For comprehensive accounts of prototype theory, see, for instance, Lakoff (1987), Lehmann (1988), and Taylor (1995). For an overview, see Geeraerts (2006a; 2010: 183–203). 2. Approaches to and notions in lexical semantics 43 seen when discussing Erdmann’s (1910) example of GERMAN (cf. Section 2.1) and Gipper’s (1959) study of Stuhl vs. Sessel (cf. Section 2.2), this is hardly ever the case in the real world, where it is frequently difficult to assign a particular object or experience, and consequently the lexical item(s) denoting it, either to category Xor to category Y. Problems of this kind explain the shift from an interest in determining the essential features of members of categories to the establishment of the best or most typical examples of particular categories. This shift in focus, though, implied in conclusions such as those drawn by Erdmann and Gipper, among others, did not receive due attention until the emergence of prototype theory (Taylor 1995: Chapters 2 & 3; Geeraerts 2006b; 2010: 183–184; Kay & Allan 2015: 37–39). The most essential notion in such a conception of category structure is that of prototypicality, which basically refers to the degree of representativeness of the members of categories, including both the actuals entities in the world as well as the lexical items used to refer to them. Whereas some members are more central and therefore more prototypical, others are more peripheral and therefore less prototypical. In this view, categories are said to exhibit the following four prototypicality effects that can help us explain how categories are organized (cf., for instance, Geeraerts, Grondelaers & Bakema 1994: 44–50; Geeraerts 1997: 11; 2006b: 146–147; 2010: 187–189; Grondelaers, Speelman & Geeraerts 2007: 989–990): (i) Non-equality or gradience: There is no clear-cut delimitation between neighboring linguistic categories and concepts, but instead a cline with a set of central examples and a set of less central ones, as we saw in the case of Stuhl and Sessel in Section 2.2. Thus, rather than about absolute membership, we talk about degrees or gradience of membership within categories. (ii) Family resemblance: Category membership depends on a series of features, but most members do not exhibit all of these features. On the contrary, most exhibit a subset of these features, and therefore share some of them with other members, but rarely coincide in all. For instance, if one category is defined by features X,Y, and Z, then one member may exhibit features Xand Y, another Yand Z, and yet another Xand Z. Here, all members are somewhat related, but none of them are entirely identical as they do not exhibit the same set of features. DANIELA PETTERSSON-TRABA 44 (iii) Non-discreteness: The limits of linguistic categories and concepts are only clear-cut in the center, where the more prototypical members are located, not in the periphery, where membership is often uncertain. (iv) Absence of necessary and sufficient conditions: No set of necessary and sufficient characteristics holds to define all members of a given category or concept, while at the same time adequately distinguishing that category or concept from neighboring ones. Related to the notion of prototypicality just described are those of entrenchment and salience. Entrenchment refers to the process whereby a linguistic unit, such as a new word or a new sense of an already existing word, becomes more cognitively established in the grammar or lexicon due to its more frequent occurrence in actual language use. In the words of Langacker (1987: 59), [there exists a] continuous scale of entrenchment in cognitive organization. Every use of a structure has a positive impact on its degree of entrenchment, whereas extended periods of disuse have a negative impact. With repeated use, a novel structure becomes progressively entrenched, to the point of becoming a unit; moreover, units are variably entrenched depending on the frequency of their occurrence. Langacker thus equates entrenchment with frequency of use of a lexical item. However, nowadays the concept does not typically refer to the overall frequency of a word, but rather to its frequency with respect to a particular sense or function, especially when alternative choices exist to express that same sense or function (e.g. Schmid 2007: 119). Consequently, entrenchment is particularly relevant when dealing with onomasiological processes such as the relation of synonymy, since members within a synonym set commonly exhibit different degrees of entrenchment, with one or some being used more frequently than others to refer to the concept they denote. The degree of entrenchment of a lexical item in comparison with its alternatives, for which we can also use the term onomasiological salience,18 can be calculated as follows, according to Geeraerts (2010: 202; cf. also Geeraerts 1997: 44–45): “the ratio 18 Onomasiological salience is not to be confused with semasiological salience, which is defined by Grondelaers & Geeraerts (2003: 71) as “the degree of prototypicality of the referent with regard to the semasiological structure of the category.” (cf. Chapter 5). 2. Approaches to and notions in lexical semantics 45 between (a) the frequency with which the members of a category are named with an item that is a unique name for that category, (b) and the total frequency with which the category occurs in a corpus.” Given that categorization processes are experientially based, we can easily observe how features such as prototypicality and entrenchment may vary across both space and time depending on the different experiences of language users. For instance, varying degrees of exposure to objects or situations lead to differences in prototypicality and entrenchment. A case in point is found in the examples mentioned by Levshina (2011: 15) regarding the nouns kangaroo and kiwi on websites from Australia and New Zealand. Whereas kangaroo is 8.5 times more common on the former sites, kiwi is 14 times more common on the latter. An example of the influence of time on changes in prototypicality is that of the concept or category SHIP, as its prototypical members nowadays probably differ substantially from those in the Middle Ages, even though the core meaning is still ‘vehicle for transportation across masses of water’ (Kay & Allan 2015: 39). This example demonstrates how sociocultural changes and technological advances lead to differences in degrees of prototypicality. Such changes and advances often bring about modifications of an existing category, the disappearance of some concepts associated with a given category, or the incorporation of new categories (Blank 1999: 71–73). Given the impact of experience on category structure, research within cognitive semantics has attempted to account for semantic variation and change by means of usage-based corpus approaches that allow the quantification of features such as the prototypicality and entrenchment of specific concepts and categories, as well as the items employed to refer to them (e.g. Geeraerts 1988; Geeraerts, Grondelaers & Bakema 1994; Soares da Silva 2015). This issue, among others related to semantic variation and change, are discussed in the next section. 2.4.2 Usage-based theories of semantic change As mentioned earlier, both synchronic and diachronic perspectives to the study of meaning abound in cognitive semantics. Consequently, with the increasing popularity of this framework, the interest in semantic change is currently back on the agenda after having been somewhat disregarded in most structuralist and generativist approaches, especially during the 1950s and 1960s. This renewed interest is mainly due to the appeal of issues regarding the relation between cognition and linguistics, the flexibility and dynamicity of meaning as represented in actual usage, and the relation of meaning with society and culture, which, as Kay & Allan (2015: 73) DANIELA PETTERSSON-TRABA 46 claim, “offer convincing ways of approaching the ‘messiness’ and unpredictability of lexical semantic change.” Another likely reason for the resurgence of diachronic approaches to semantics is the existence of material such as electronic corpora which comprise a great amount of actual language use from previous historical periods, thus easing the process of studying diachronic developments in language. Lexico-semantic changes are often considered to be more difficult to categorize than other types of linguistic change in the sense that each word follows idiosyncratic developmental patterns. The reasons for this are twofold. First, semantic changes are possibly caused by a more extensive range of factors than other types of change, including both intraand extralinguistic determinants (Durkin 2009: 222–223; cf. also Kay & Allan 2015: 72–74). Secondly, extralinguistic factors do in fact play a major role in many meaning changes, which adds more complexity to the task of describing the evolution of particular lexical items. As such, nonlinguistic history, including technological and cultural developments, is often resorted to in order to help explain the motivation underlying these linguistic changes. According to Durkin (2009: 222–223): Semantic changes are notoriously difficult to classify or systematize […]. [A]lthough some semantic changes occur in clusters, with a change in one word triggering another, we do not find anything comparable to a regular sound change […]. Additionally, semantic change is much more closely connected with change in the external, nonlinguistic world, especially with developments in the spheres of culture and technology. This does not mean, however, that recurrent patterns or mechanisms of semantic change cannot be identified. As we have already seen, comprehensive classifications of semantic change were already proposed in the historical-philological tradition (cf. Section 2.1), and many of these classifications still enjoy wide currency in contemporary semantic research. In fact, diachronic analyses conducted within cognitive semantics echo many of the principles of historicalphilological semantics, as the conception of meaning in both frameworks is very similar (i.e. close relation between cognition and meaning, no separation between semantic and encyclopedic knowledge, flexibility and dynamicity of meaning, and importance of context and usage). Therefore, it can safely be maintained that the diachronic branch of cognitive semantics is both a continuation and an extension of the research set out in the prestructuralist theory of semantics (Geeraerts 2010: 229–230). 2. Approaches to and notions in lexical semantics 47 Language use figures prominently in explanations of semantic change in cognitive semantics. Building on Paul’s (1920) distinction between usual and occasional meanings mentioned in Section 2.1, semantic scholars belonging to this tradition highlight that changes in meaning, for instance, the emergence of new senses and nuances of meaning, occur in actual language use. Therefore, they also distinguish between decontextualized and contextualized meanings. In fact, one explanation that has been postulated for the development of new meanings involves the concept of invited inference (cf. König & Traugott 1988; Traugott & Dasher 2002). This model argues that many utterances show a certain degree of ambiguity since, besides the more established and central sense of a linguistic expression, additional ones may also arise in a particular context of use. The theory of invited inference builds on Grice’s (1989 [1975]) notion of conversational implicature, which refers to the meanings that go beyond the conventional meanings of words, but that are implied by speakers when using those words in a specific structure, either consciously or unconsciously. Conversational implicatures therefore arise when participants in a communicative exchange violate or flout the so-called cooperative principle by not adhering to one or more of its maxims (i.e. the maxims of quantity, quality, relation, and/or manner).19 With time and repetition, such inferred meanings —or conversational implicatures in Grice’s terminology— may become conventionalized and established as a new sense beside the original one. This is the case of temporal markers such as since and after, among others, which have developed causal meanings over time (Traugott & Dasher 2002: 16; Hopper & Traugott 2003: 80–81). Such inferred meanings, by being implied repeatedly, need to be stored at some point in the mental lexicon, thus leading to a higher degree of entrenchment and prototypicality. Consequently, there emerges a clear convergence between usage-based approaches to meaning change and prototype theory described above: while conventionalized senses correspond with central and prototypical senses, contextualized and less conventionalized senses correspond with peripheral and less entrenched ones. However, as we have already seen, features such as prototypicality, salience, and entrenchment are not constant, but rather vary across time and space. However, not all semantic changes imply the appearance of new senses but may instead bring about alterations to the central and prototypical meaning so that its range of application becomes either enlarged or more restricted. Consider the example of Dutch legging ‘leggings’ discussed in Geeraerts (1997: 33–47; 2010: 234). When the form entered the language at the 19 For an overview of inferential-based approaches, see, for instance, López-Couso (2017). DANIELA PETTERSSON-TRABA 54 to be regarded as synonymous, while they are allowed to vary in other aspects of meaning. To provide an example, woman and lady are considered synonyms as they both denote an ‘adult human female’. Nevertheless, the two words differ as regards expressive meaning: whereas woman is fairly neutral, connoting a vast set of properties, lady is expressively more restricted, since it is closely associated with features such as class, status, elegance, and manners, and thus carries a positive connotation (Leith 1983: 70–72). However, as shown by this example, all non-denotational dimensions of meaning are crucial when an onomasiological orientation is adopted, since as Grondelaers, Speelman & Geeraerts (2007: 994) point out, “[…] the very definition of nonreferential [i.e. non-denotational] meaning involves the concept of onomasiological alternatives”. It is precisely along these other dimensions of meaning that the great majority of synonyms differ and, consequently, they must be accounted for in a semantic analysis of synonymy. Different classifications of synonymy have been put forward in the literature and, as a result, a variety of labels have been used to refer to the same types of synonyms, in some cases even by the same author. Here, we focus exclusively on the different types of synonymy established in the most important and well-known works on the topic, namely those by Lyons (1968; 1981a; 1981b; 1995) and Cruse (1986; 2000; 2002), but also that of Murphy (2003). Lyons’ (1968) early classification is probably the most comprehensive one as it distinguishes between two dimensions: total synonymy, that is, those words which are interchangeable in all contexts of use, and complete synonymy, i.e. those words which are equivalent both on the denotative and connotative dimensions of meaning. By crossing these two dimensions, four types of synonyms can be established, namely (i) complete and total synonyms, (ii) complete and non-total synonyms, (iii) incomplete and total synonyms, and (iv) incomplete and non-total synonyms. For example, the first type, complete and total, includes those synonyms which share both their denotational and expressive meanings and are interchangeable in all contexts. This type refers to what is usually termed absolute synonymy.22 However, although this classification is the most exhaustive one, it can also be said to provide unnecessary divisions. This is so because type (iii), incomplete and total synonymy, would refer to words or word senses which are interchangeable in all contexts of use without being denotationally and expressively equivalent. Such a relation is quite unlikely —if not 22 Absolute synonymy is sometimes also known as perfect synonymy (e.g. Stern 1931: 225), pure synonymy (e.g. Ullmann 1957: 110), exact synonymy (e.g. Ziff 1964: 172), or real synonymy (e.g. Lyons 1963: 74). 3 Synonymy 55 impossible— to occur in language.23 Lyons’ (1995) latest classification differs considerably from his previous ones (i.e. 1968; 1981a; 1981b), as he establishes a different division into absolute synonymy, cognitive synonymy, near-synonymy, and partial synonymy. Partial synonymy roughly corresponds with type (ii) in his early classification (i.e. complete and nontotal). The reason for drawing a distinction between absolute and partial synonyms, according to Lyons, is that whereas absolute synonyms are most certainly non-existent in language, partial synonyms, though infrequent, do exist. This is because, in contrast to absolute synonyms, partial synonyms are equivalent both as concerns denotational and expressive meaning, but need not be interchangeable in all contexts. Cruse (1986), in turn, proposes a three-way split into absolute synonymy, cognitive synonymy, and plesionymy, which partially overlap with Lyons’ (1995) typology.24 In subsequent classifications by Cruse (2000; 2002), the same three types of synonyms are present, although the labels for the latter two vary. For instance, Cruse uses the labels cognitive synonymy and plesionymy in his 1986 monograph, while in later work (2000) he refers to the same types of synonymy as propositional synonymy and near-synonymy, respectively. Table 1 below summarizes the various classifications of synonymy by the two authors and the terminology discussed so far: 23 Some subsequent formulations of Lyons’ classification overlap substantially with the original one while at the same time differing in certain respects, both regarding the labels used and the types identified (for more details, cf. 1981a; 1981b). 24 Near-synonymy, rather than plesionymy (e.g. Hirst 1995; Storjohann 2009), is the label adopted in the present dissertation as it is the most frequent term employed in contemporary research on synonymy. DANIELA PETTERSSON-TRABA 56 Table 1. Classifications of synonymy by Lyons and Cruse Author Definition Lyons (1968) Lyons (1995) Cruse (1986) Cruse (2000; 2002) Identical on all dimensions of meaning and interchangeable in all contexts of use Complete & total Absolute synonymy Absolute synonymy Absolute synonymy Identical on all dimensions of meaning but not interchangeable in all contexts of use Complete & non-total Partial synonymy -- Not identical on all dimensions of meaning but interchangeable in all contexts of use Incomplete & total --- Identical on the denotative dimension of meaning, but not on other dimensions, and not interchangeable in all contexts of use Incomplete & non-total Cognitive synonymy Cognitive synonymy Propositional synonymy Similar but not identical on the denotational dimension of meaning and not interchangeable in all contexts of use Nearsynonymy Plesionymy Nearsynonymy In addition to Lyons’ and Cruse’s classifications of synonymy, Murphy (2003) postulates yet another typology, in which she distinguishes between synonyms in terms of denotative meaning only and synonyms which share more than their denotation. As regards denotative meaning, she considers two dimensions, namely (i) how many senses two or more synonymous words share and (ii) how similar the shared senses are. Full synonyms are those which share all their senses and these senses are all identical (e.g. groundhog and woodchuck). Sense synonyms, in turn, are those which share one or more senses, but not all. As in the case of full synonyms, the shared senses are also identical. One example would be couch and sofa, which share the sense ‘[a] long, stuffed seat with a back and ends or end, used for reclining’ (OED s.v. sofa n. 2), but not the meaning ‘[a] couch upon which a patient reclines when undergoing psychoanalysis or psychiatric treatment’(OED, s.v. couch n. 3b), which only the former term can designate. Full synonyms and sense synonyms are included within the general category of cognitive synonyms, or logical synonyms in Murphy’s terminology. Finally, lexical items that are similar, but not identical in one or more of their senses are called near-synonyms or plesionyms, as in the case of fog(gy) and mist(y), which differ as regards the density of the weather phenomenon they denote (see below). Concerning non-denotative dimensions (e.g. register, dialect, and connotation), synonyms may also differ. According to Murphy (2003: 151), such aspects of meaning can either enhance or diminish their degree of similarity. For 3 Synonymy 57 instance, the sense synonyms drunk and inebriated, despite sharing the sense ‘[a]ffected by alcohol to the extent of losing control of one's faculties or behavior’ (Lexico, s.v. drunk adj. 1), differ in register, the former being more informal than the latter. Notwithstanding the existence of these and other classifications of synonymy, what is relevant here is that three types are recurrently mentioned in most classifications, though with different labels. These are absolute synonymy, cognitive synonymy, and near-synonymy. The first of these three types is thought to constitute one of the endpoints of the synonym continuum, namely that of absolute identity of meaning. Exactly what counts as absolute identity of meaning depends largely on the types of meaning that are considered to be relevant. Nevertheless, most scholars agree that absolute synonymy refers to those words or word senses which are identical on all four dimensions of meaning mentioned above (i.e. denotational meaning, expressive meaning, stylistic meaning, and collocational meaning). Ullmann (1957: 109–110), for instance, claims that two criteria need to be met for absolute synonymy to occur. On the one hand, absolute synonyms must be able to substitute for each other in all contexts of use and, on the other, they must not differ in neither denotational nor non-denotational aspects of meaning. Cruse (1986; 2000) goes one step further by stating that these two criteria are not sufficient, and adds yet a third one, to wit, that of contextual normality. This is a particularly useful criterion for a contextual approach such as the one adopted in the present dissertation. According to such a perspective, in addition to being interchangeable, absolute synonyms must also be equinormal in all contexts of use. Along these lines, Cruse (2000: 157) argues that […] absolute synonyms can be defined as items which are equinormal in all contexts: that is to say, for two lexical items X and Y, if they are to be recognized as absolute synonyms, in any context in which X is fully normal, Y is, too; in any context in which X is slightly odd, Y is also slightly odd, and in any context in which X is totally anomalous, the same is true of Y. This definition evidences that very severe requirements are imposed on words or word senses in order for them to comply with the criteria for absolute synonymy and demonstrates why this type of synonymy is very rare —if not inexistent— in language. Following this definition, if a context exists in which two synonyms are not equally normal, that is, in which one is more natural or idiomatic than the other, absolute synonymy is ruled out. Cruse (2000: 268–269) tests some plausible candidates, for instance, begin and commence,almost and nearly, and DANIELA PETTERSSON-TRABA 58 scandalous and outrageous, and draws the conclusion that for all these synonymous pairs contexts in which one of the members is more normal than the other can be observed. In fact, it is commonly claimed that there is no reason why a semantic relation of this type should exist and be maintained in language, as it goes against the economy principle. As Taylor (2003: 264) states, “[absolute] synonymy would have to be regarded as an extravagant luxury, even dysfunctional, in that limited symbolic resources get squandered on the designation of one and the same semantic unit.” In this statement, Taylor indirectly refers to the well-known nosynonymy rule or isomorphic state of languages, i.e. one form, one meaning (e.g. Bolinger 1977: ix–x, 9; Wierzbicka 1988: 13–14; Croft 2000: 176; Nuyts & Byloo 2015: 62–63), which makes languages work against absolute synonymy if it at some point comes to exist through processes such as borrowing. This is an issue that has clear diachronic implications and is therefore further discussed in Section 3.3 below. Concerning cognitive synonymy, this is a relation which holds between words or word senses that are identical on the denotational dimension —that is, which have the same conceptual content— and therefore mutually entail one another. In other words, cognitive synonyms, if substituted for one another in a particular context of use, generate clauses or sentences with the same truth-values. Nevertheless, such pairs or groups of synonyms do not meet the criteria for absolute synonymy because they differ in non-denotational traits, for instance, connotation (e.g. firm and stubborn), register (e.g. drunk and pissed), or style (e.g. kick the bucket,die, and pass away), or the language variety in which they occur, such as British English (BrE) autumn and AmE fall (Cruse 1986: 278–280; Desagulier 2014: 153). In many cases, cognitive synonyms differ in more than one of the non-denotational types of meaning simultaneously (e.g. both expressive and stylistic meaning). Given that cognitive synonyms differ in non-denotational traits, they are typically used in different speech situations, text-types, or fields of discourse, and using one lexical item in a context in which its synonym is more typical or frequent will most probably result in an odd sentence. Areas of vocabulary which are related to human experiences that stir strong emotions and opinions among people, for instance, death, religion, and sex in Western cultures and societies, tend to display a wide array of cognitive synonyms (Quine 1951: 28–31; Cruse 1986: 88, 270–285; 2000: 158–159).25 Finally, near-synonymy, which is undoubtedly the most frequent type of synonymy in language, refers to those words or word senses that differ slightly in conceptual content and are 25 Cf. Cruse (1986: 285) for a comprehensive list of cognitive synonyms in these semantic domains. 3 Synonymy 59 thus not denotationally identical. As such, if substituted for one another, the generated sentences yield somewhat different truth-conditions. However, near-synonyms are still sufficiently semantically similar to be interchanged in many contexts of use. The denotational traits in which near-synonyms vary must, according to Cruse (2000: 159), “[…] be either minor or backgrounded”. Differences between near-synonyms include, but are not limited to, aspects such as the following (examples taken from Cruse 2000: 157–160; Murphy 2003: 147; Divjak 2010: 3–4; Desagulier 2014: 153): (i) Contiguity on a gradable continuum or degree, e.g. fog(gy) and mist(y): Both fog and mist refer to ‘[a] cloud of water substance present in the atmosphere at or close to the surface of the earth’ (OED, s.v. fog n. 2, I 1a; mist. n. 1, I 1a). However, fog is denser and thicker than mist, the latter being lighter and therefore not limiting visibility to the same extent as the former. (ii) Nuances involving specialization, e.g. laugh and giggle: the verb laugh denotes any type of spontaneous movement of the lower part of the face, accompanied by an explosive vocal sound, to display emotions such as joy or amusement, arising in a wide variety of situations (OED, s.v. laugh v. 1a) . In turn, the verb giggle is more specialized, as it refers to a particular way of laughing,namely a light and foolish type which occurs in a more restricted set of circumstances, including amusement but also uneasiness and embarrassment (OED, s.v. giggle v.1, a).26 (iii) Prototypicality, e.g. brave and courageous: While brave is used to describe a fearless person in physical terms, courageousemphasizes the moral or intellectual dimension. (iv) Viewpoint or connotation, e.g. slim/slender and skinny: The three terms are all used to refer to people who are thin, but whereas slim and slender have positive connotations, being associated with grace and elegance, skinny has a clearly negative connotation, being linked to unattractiveness and negative esthetic qualities in general. (v) Aspectual variation, e.g. calm and placid: The former adjective is used to denote a determinate state, which is ephemeral, while the latter refers to a more inherent feature of a person or entity. 26 Some authors would argue that giggle is a hyponym of laugh, as giggle is included in the definition of laugh. This example demonstrates that the dividing line between the two semantic relations (i.e. synonymy and hyponymy) is not always clear-cut. DANIELA PETTERSSON-TRABA 60 In contrast to absolute synonyms, near-synonyms, which are particularly frequent in language, cannot be considered uneconomical and unnecessary, especially since language users are more likely to devise differences between such related words than to abandon one or more of the terms. In fact, a great majority of scholars claim that near-synonyms add substantially to the lexical expressivity of language, as they enable speakers to convey different nuances of meaning in a very precise manner (cf., for instance, Edmonds & Hirst 2002: 107–108; Murphy 2003: 165–166). However, given that near-synonyms can vary on any dimension of meaning and since that variation is frequently multidimensional, even native speakers sometimes find it difficult to master the fine-grained differences that exist between them. Although the distinction between cognitive synonyms and near-synonyms appears to be clear in theory, as the former entail each other while the latter do not, the boundary between the two types becomes much more problematic in practice. This is so because it is not often obvious whether two or more words differ solely in their non-denotational semantic traits or whether, on the contrary, slight differences in denotation also exist between them. Due to the difficulty in distinguishing between cognitive synonymy and near-synonymy when specific pairs or sets of synonyms are concerned, some scholars have argued for a two-fold, instead of a three-fold, division into synonym types, namely absolute synonymy vs. non-absolute synonymy or nearsynonymy, thus dismissing the distinction between cognitive synonyms and near-synonyms, as this opposition is often troublesome. This is the standpoint adopted by Edmonds & Hirst (2002: 116–117) and Desagulier (2014: 153), among others, and the one followed in the present dissertation. A great amount of research has been conducted on specific near-synonymous pairs or groups in the last few decades. Given that near-synonyms differ in one or more aspects of meaning, it is certainly not sufficient to state that a relation of partial equivalence exists between the members of a synonym set, but it is also necessary to examine the multiple factors determining the choice between them in order to reach an in-depth and comprehensive understanding of their semantic relation. Such an approach is the only one which allows us to know for certain in which contexts and the reasons why we use one word instead of another one. Disentangling the differences between near-synonyms, which are often multivariate and related to semantic, stylistic, and/or morphosyntactic dimensions, among others, has been and continues to be the main aim of many studies with a usage-based onomasiological orientation. The remainder of this chapter is concerned with a review of such studies. Section 3.2. discusses 3 Synonymy 61 synchronic research on lexical synonymy, while Section 3.3 considers studies with a diachronic perspective. 3.2 USAGE-BASED CORPUS STUDIES ON SPECIFIC SYNONYMS Although synonymy in general has figured prominently in theories of lexical semantics, dating back to the structuralist school (cf. Section 2.2), the internal semantic structure of specific pairs or larger sets of lexical near-synonyms had until quite recently not received much attention, especially if compared to constructional near-synonyms and other semantic phenomena such as metaphor and polysemy, which have been the subject of an ever-increasing body of literature (Edmonds & Hirst 2002: 106; Taylor 2003: 264; Geeraerts 2010: 264; Liu 2010: 57). Consequently, several scholars have lately stressed the need for more research in this field, which has led to the emergence over the last 30 years of studies following diverse types of methodologies within the distributional approach originally set out by Firth and Sinclair in the middle of the 20th century (cf. Section 2.3.2). The aim of this section is to offer an overview of the existing synchronic corpus-based research on specific synonym sets that has been carried out since the end of the 1980s and the evolution of this line of research on methodological grounds. A detailed account of all the existing studies is not provided here because not all the methods proposed in the literature are equally relevant for our purposes. This is so because the approach adopted in such studies differs widely depending on the specific groups of synonyms analyzed, the POS the synonyms belong to, and the factors that significantly influence the choice between them.27 This is to be expected since, as argued by Liu (2010: 61), […] the use of diverse procedures makes perfect sense in the study of internal structures of near-synonym sets because the way near synonyms differ often varies from synonym to synonym and from set to set. […] Thus a researcher has to decide, based on a close scrutiny of the features of the synonyms being examined, what micro-procedures to employ.28 27 See Glynn (2010: Tables 1 and 2) for a list and a methodological summary of corpus-based studies in cognitive semantics dealing with lexical and constructional onomasiology and semasiology, including research on synonymy and polysemy, among other semantic relations. 28 For a similar argument, cf. also Hanks (1996: 92; 96). DANIELA PETTERSSON-TRABA 62 Therefore, the next section provides some generalizations that can be extracted from previous studies, primarily focusing on those dealing with adjectival near-synonyms, as the factors they consider have more in common with the analyses conducted in this dissertation (cf. Chapters 5–7). 3.2.1 Two waves of research The research which is reviewed in this section can be classified into different groups depending on three parameters related to the scope of the analyses and to methodological issues: (i) Number of synonyms examined: pairs vs. larger sets (i.e. three or more synonyms). (ii) Amount and type of factors conditioning the choice between synonyms, ranging from collocational and/or stylistic patterns only to a wider array of determinants of variation which include semantic, morphosyntactic, and/or extralinguistic factors. (iii) Statistical techniques and measures used: from mere frequencies and percentages, to association measures of collocational strength (e.g. PMI, t-score, log likelihood), multivariate tests (e.g. Hierarchical Configural Frequency Analysis (HCFA), logistic regression analysis, multiple correspondence analysis), and clustering techniques (e.g. HAC). Along these parameters, corpus-based synonymy research can be divided, with some exceptions, into two main waves. The first one includes mainly those studies published approximately before the mid-2000s, most of which focus on pairs of near-synonyms by making use of raw frequencies, percentages, or association measures in order to examine their collocational and stylistic patterns. By contrast, in the second wave, which can be said to take off with the works of Divjak (2006) and Divjak & Gries (2006; 2008), the focus shifts to larger groups of near-synonyms. Here, in addition to accounting for collocational and stylistic behavior, further factors from different linguistic levels are also considered, including semantics and morphosyntax. A direct corollary of this multidimensional perspective is the use of more advanced and sophisticated techniques, which allow for the inclusion of a larger number of variables. The development of corpus-based research on lexical synonymy over the last 30 years can thus be summarized as shown in Figure 2: 3 Synonymy 63 Pairs Sets Collocation & style Wider range of factors Frequencies/percentages Association measures Multivariate & clustering techniques Figure 2. Development of corpus-based research on lexical synonymy over the last 30 years As has already been mentioned, most of the early research in this field considered only pairs of near-synonyms rather than larger groups (cf., for instance, Persson 1989; Church et al. 1991; Kennedy 1991; Gries 2001; 2003; Arppe 2002; Kjellmer 2003; Taylor 2003). One case in point is Biber, Conrad & Reppen (1998: Section 4.3), who analyzed the grammatical preferences of the verbs begin and start. Although begin and start can be used in exactly the same syntactic structures, both intransitive and transitive, and even with the same type of objects when used transitively (i.e. noun phrases, to-clauses, and -ing clauses), closer inspection of their distributional patterns in such contexts reveals that the way in which they are typically used in actual discourse differs considerably. While the frequency of start almost doubles that of begin in intransitive structures in fictional texts, begin is three time more common than start with to-clauses as objects. In fact, in fictional texts more than half (60%) of the total examples of begin are transitive uses followed by to-clauses. Contrariwise, start occurs in this particular structure in only 17% of the cases, which demonstrates that it is preferred in other patterns, be it intransitive structures or transitive ones with other types of direct objects. Therefore Biber, Conrad & Reppen (1998) show that by examining the valency of the two verbs, remarkable differences in use emerge. Their results point to a division of labor between the two nearsynonymous verbs: despite being interchangeable in most cases, each verb is clearly the preferred option in different contexts of use. Additionally, language users are not able to predict the individual preferred grammatical associations of the two verbs, which they generally consider to be completely interchangeable. It is therefore only by means of detailed analyses of usage data such as that in Biber, Conrad & Reppen, made possible by the availability of large corpora, that these preferred structures can be discovered, while native speaker intuitions are often unreliable. Despite the valuable findings of studies on near-synonymous pairs such as the one conducted by Biber, Conrad & Reppen (1998), it is clear from the information presented in dictionaries and thesauri that most lexical items have more than one synonym and thus synonyms typically come in larger groups (e.g. Divjak 2006; Gries 2010). For instance, the verb DANIELA PETTERSSON-TRABA 70 Contemporary American English or COCA (Davies 2008–), which comprises five different text-types showing different degrees of formality, to wit, spoken, fiction, newspaper, magazine, and academic writing, Liu (2010) establishes a cline of formality for the near-synonymous adjectives main,major,chief,principal, and primary, with main being the most informal and primary the most formal. While most corpus-based studies on near-synonyms during the 1990s and early 2000s exclusively examined collocational and/or stylistic preferences, a majority of studies from the second wave analyze a wider range of factors and therefore go far beyond previous corpusbased research on synonymy. Again, Divjak (2006) and Divjak & Gries (2006; 2008) represent a clear break with previous analyses, this time in the sense that they examine a large number of factors.33 In fact, Divjak & Gries proposed a new approach to the study of semantic phenomena, mainly synonymy and polysemy, which they called the Behavioral Profile (BP) approach, following Hanks (1996: 79), the first author to use this label to refer to the typical patterns and contexts in which a given word is used. In brief, the BP approach consists in analyzing sets of near-synonymous words or senses of a polysemous word by considering many different types of co-occurrence information (e.g. morphological, syntactic, semantic, and stylistic) in order to determine their conventional uses. This is done by means of different types of statistical techniques, such as correlations, HAC, and logistic regression (Gries & Divjak 2009; Gries 2010).34 In one of the earliest BP studies, namely Divjak & Gries (2006), referred to earlier on in this section, they analyze the distributional patterns of nine Russian verbs of trying, including 87 different variables —or as they call them, ID-tags, following Atkins (1987)— pertaining to the domains of semantics and morphosyntax. For instance, among the morphosyntactic contextual clues, they consider verb related characteristics, such as aspect, mood, and tense, as well as clause and subject related characteristics, such as sentence type and clause or subject structure. Semantic co-occurrence information, on the other hand, refers to features concerning the nominative subject paradigms, that is, oppositions such as animate vs. inanimate or concrete vs. abstract. Moreover, they analyze the adjacent verb collocates of the nine near-synonyms by grouping them into semantic categories such as PHYSICAL ACTION,PHYSICAL PERCEPTION, 33 Even though some earlier studies (e.g. Hanks 1996; Biber, Conrad & Reppen 1998: Sections 4.2 and 4.3; Arppe 2002), also include some syntactic and morphological co-occurrence features, the amount of factors considered in these works cannot be compared to those by Divjak (2006) and Divjak & Gries (2006; 2008), who include 47 and 87 variables, respectively. 34 These two papers constitute in-depth summaries of the BP approach. The reader is referred to them for further details on the steps and procedures used, as well as for examples of variables to be included in such analyses. 3 Synonymy 71 SPEECH, and INTELLECTUAL ACTIVITY, among others, thus paying special attention to the complementation patterns and argument structures of the verbs under analysis. As mentioned above, the results point to the existence of three clearly delineated clusters or groups within this synonym set, each of them containing three of the verbs. This study proved the usefulness of the BP approach for the delineation of the internal semantic structure of near-synonyms and can be said to have paved the way for further BP studies along the same lines on other POS: adjectives (e.g. Gries & Otani 2010; Liu 2010; Yatandu Uba 2015), nouns (e.g. Liu 2013b), and adverbs (e.g. Liu & Espino 2012). The BP approach is not the only existing corpus-based distributional method in cognitive semantics that has been used to analyze a wider range of contextual factors that determine the choice between near-synonyms. Other methods can be found which share the same main goals, but differ in the specific procedures employed along three stages distinguished by Levshina (2011: 24–29). These stages are data collection, data exploration, and confirmatory testing. Within the first stage (i.e. data collection), we find methods in which the data is collected manually, as in most BP studies, as opposed to automatically, as in SVS modeling (e.g. Heylen et al. 2008). SVS can be used to model the behavior of polysemous or near-synonymous words by analyzing their collocational patterns with the help of association measures. In the second stage (i.e. data exploration), the data is explored. Depending on the method adopted, different statistical techniques can be used to visualize the results. For instance, BP studies tend to use dimensionality-reduction techniques, such as HAC, Multidimensional Scaling or MDS for short (Wickelmaier 2003; Levshina 2015: Chapter 17), and (multiple) correspondence analysis (e.g. Desagulier 2014; Krawczak 2018). Finally, in the third stage (i.e. confirmatory testing) statistical tests are applied to determine whether the results can be generalized to other data. Logistic regression analysis is most commonly employed to see which variables and variable interactions have a statistically significant effect on the choice between near-synonyms (e.g. Arppe 2008; Krawczak 2014; 2018), but other tests can also be used, for instance, HCFA (e.g. Liu 2010; Liu & Espino 2012; Liu 2013b). Many BP studies do not include this last stage, mainly due to the extremely large number of variables to be analyzed, and are therefore of a more exploratory nature (but cf., for instance, Divjak 2010 for an exception). Some of the techniques mentioned here are further explained in Chapters 6 and 7, where they are used to examine the near-synonymous adjectives from the olfactory domain fragrant,perfumed, scented,sweet-scented, and sweet-smelling. DANIELA PETTERSSON-TRABA 72 3.2.2 Near-synonyms and POS: Studies on adjectives The distributional corpus-based research reviewed in the previous section reveals that synonym sets belonging to different POS have been analyzed over the last few decades, including verbs (Schmid 1993; Hanks 1996; 2013; Biber, Conrad & Reppen 1998: Section 4.3; Arppe 2002; 2008; Divjak 2006; 2010; Divjak & Gries 2006; 2008; Arppe & Järvikivi 2007; Chung 2011), nouns (Schmid 1993; Liu 2013b), adjectives (Persson 1989; Church et al. 1991; Biber, Conrad & Reppen 1998: Sections 2.6 and 4.2; Partington 1998: Section 2.4; Gries 2001; 2003; Taylor 2003; Gries & Otani 2010; Liu 2010; Krawczak 2014; 2018; Yatandu Uba 2015), adverbs (Partington 1998: Section 3.6; Liu & Espino 2012; Desagulier 2014), and prepositions (Kennedy 1991). As such, a wide variety of different types of factors from different linguistic levels (e.g. semantic, syntactic, and morphological) have been examined in order to determine their effect on the choice between particular near-synonyms. However, not all factors are equally relevant for all synonym studies, since the way synonyms differ from one another depends largely on their POS and conceptual nature. Thus, most studies on near-synonymous verbs, besides including collocational patterns and semantic features in the analysis, also consider a great amount of morphosyntactic contextual clues, for instance, tense, aspect, and mood, or sentence and clause type in which the verbs occur. In fact, in many of these cases the morphosyntactic variables have proven to be of utmost importance to distinguish between nearsynonymous verbs (cf., for instance, Divjak 2010: 183–193). On the other hand, studies on adjectives and adverbs generally consider fewer morphosyntactic factors, whereas semantic features seem to play a more crucial role in determining the choice between near-synonyms belonging to these POS, for instance, semantic preference regarding the elements they modify (e.g. Gries 2001; Liu 2010; Liu & Espino 2012). In what follows, studies on near-synonymous adjectives are dealt with in more depth, concentrating on the types of factors that are typically included in these analyses, as such determinants are the most relevant for our purposes in the present dissertation. A great majority of the existing distributional corpus-based research on pairs or sets of near-synonymous adjectives consider their collocational and/or stylistic behavior only. More specifically, as has already been mentioned in Section 3.2.1, such studies tend to focus on the R1 collocates of the adjectives in order to concentrate on the nouns they modify. This is so because, as has been demonstrated in previous research on adjectives in general, analyzing the nouns that adjectives modify is one of the best ways to reveal the nature of the semantic content 3 Synonymy 73 of the adjectives, and modified nouns are therefore considered to be more informative collocates (cf. Geeraerts 1986: 282–284 on the Dutch adjective vers ‘fresh’ in Section 2.3.2). However, other studies go beyond mere individual collocates and consider other features of the modified nouns, including primarily semantic aspects, but also morphological ones. For example, as mentioned earlier, Gries (2001; 2003) includes information concerning the dichotomies concrete vs. abstract and specific/basic level vs. general/superordinate level of the modified nouns of -ic and -ical adjectival pairs, but he also contemplates that of animacy (i.e. animate vs. inanimate) for those adjectives in which such a division is relevant, as in the case of optic vs. optical. In his study on chief,main,major,primary, and principal, Liu (2010) discovers a division of semantic labor between the five adjectives by classifying their noun collocates into six semantic categories:35 (i) ABSTRACT:change,problem, and reason. (ii) CONCRETE:dish,entrance, and street. (iii) DUAL (can be either concrete or abstract): character,component, and source. (iv) INSTITUTION:city,corporation, and school. (v) POSITION TITLE:deputy,investigator, and officer. (vi) NON-POSITION TITLE:author,owner, and sponsor. By doing so, Liu concludes that whereas dual and abstract nouns are often modified by all five near-synonymous adjectives, the other four semantic categories tend to be dominated by one or two of them: for instance, while concrete nouns prefer main, position title nouns are dominated by chief. Furthermore, as all five adjectives seem to converge in meaning when they modify dual and abstract nouns, Liu investigates whether other fine-grained differences exist between them when used with nouns belonging to these two specific semantic categories. To this end, he explores two morphological features of the modified nouns or noun phrases in general, to wit, number (singular vs. plural) and definiteness (indefinite vs. definite). This is relevant for the adjectival synonyms object of study in order to establish the degree of importance that they convey: one main issue is more important than two main issues and the main issue is more important than a main issue. Thus, Liu identifies a cline of importance, with primary,chief, and 35 For further information about the nouns belonging to each category and the decisions regarding specific nouns, see Liu (2010: 84). DANIELA PETTERSSON-TRABA 74 main being used to refer to entities that are considered to be more important by speakers, followed by major, and lastly principal. By far the most comprehensive work on synonymous adjectives is the BP study conducted by Gries & Otani (2010), in which they examine two sets of near-synonymous from the semantic domain of SIZE, namely big,great, and large, on the one hand, and little,small, and tiny, on the other. They analyze a great amount of factors from different linguistic levels, including morphological, syntactic, and semantic ones. Among morphological factors, they consider the aspect, tense, voice, and transitivity marking of the finite verb of the clause in which the adjectives occur. As regards syntactic factors, these are similar to those considered in BP studies on near-synonymous verbs, for instance, the type (i.e. main vs. dependent) and function (e.g. direct object, noun phrase postmodifier) of the clauses where the adjectives are located. This syntactic information is automatically retrieved, as Gries & Otani use the British component of the International Corpus of English, which is a parsed corpus. Another syntactic feature that they examine, which is also present in other studies on adjectival near-synonyms (e.g. Biber, Conrad & Reppen 1998: Section 2.6; Gries 2001; 2003; Liu 2010), is the syntactic function of the adjectives, that is whether they are used in attributive, predicative, or adverbial function. Finally, Gries & Otani analyze several semantic features of the noun collocates, such as countability (count vs. non-count) and a wide range of semantic categories including, among others, CONCRETE,ABSTRACT,HUMAN,ORGANIZATION/INSTITUTION, and QUANTITY. Another semantic feature they consider is how SIZE is modified, i.e. literally, metaphorically, quantitatively, or evaluatively. The results for each of the variables are then aggregated by means of HAC and visualized with the use of dendrograms. Findings point to a clustering solution according to different parameters, including both sameness of meaning, so that tiny and smallest are grouped together, and oppositeness of meaning, with big and little forming one cluster, and large and small another. Although Gries & Otani’s (2010) study includes a wide range of different types of factors, the effects of each of the levels of the variables are not discussed. This is so because it is practically impossible to statistically test for all the variables included in their analysis without an extremely large amount of data for each of the six adjectives at issue. This is, in fact, a recurrent problem in many BP studies, not just the one carried out by Gries & Otani, as the third step of analyses of this type, i.e. testing for significance, is unviable (cf. Section 3.2.1 above). Therefore, even though such studies provide valuable results concerning the general behavior of the synonyms examined, specific details 3 Synonymy 75 about exactly how they differ from one another are seldom offered. However, it must be borne in mind that the goal of many BP studies is not that of giving an in-depth description of each of the variables explaining the choice between near-synonyms, but rather that of presenting a bird’s-eye view of their overall contextual behavior. Finally, some studies on near-synonymous adjectives, besides considering intralinguistic factors pertaining to their co-occurrence preferences, also examine the effects of extralinguistic variables. This is the case of, for instance, Krawczak (2014; 2018), who analyzes adjectives from the semantic domain of SHAME. Her 2014 paper examines three lexical items, namely ashamed,embarrassed, and humiliated, in both AmE and BrE in order to verify Wierzbicka’s (1992; 1999) claims as to their semasiological and onomasiological structures. To this end, Krawczak includes three different semantic features, to wit, cause of emotion (e.g. bodily causes, insecurity, social failure), type of emotion (i.e. internal vs. external), and temporal scope of the emotion (present, past, and general). Moreover, she includes the factor ‘dialect’ to ascertain whether there exist extralinguistic usage differences. By means of two multivariate methods, namely correspondence analysis and multinomial logistic regression, she establishes three clearly delineated usage-profiles for the three adjectives: ashamed is connected to internal and atemporal causes related to insecurities of social status, emotional or bodily problems, and social failures; embarrassed, in turn, is associated with interactive factors, for instance, selfesteem difficulties or deviation from social conventions regarding politeness; finally, humiliated is linked to damages to one’s social status due to external motivations. With regard to the extralinguistic variable examined (i.e. dialect), no significant differences are found, as the lexical items display basically the same behavior in the two reference varieties of English. Krawczak’s (2018) study constitutes a continuation of the author’s previous research. She adopts a cross-linguistic and cross-cultural perspective to the same social emotion, namely SHAME, but restricts the analysis to two of the adjectives: ashamed and embarrassed, and their counterparts in French (i.e. honteux and embarrassé) and Polish (i.e. zazenowany and zawstydzony). The same factors are examined, with the exception of dialect, which is replaced by country and language, and a variable coding for whether an audience is present or absent from the speech situation in which the social emotion emerges. The findings point to the existence of a cline of communities from Poland through France to the UK and the US that is linked to their respective cultures and, more particularly, to whether the societies are more collectivist, as in the case of Poland, or more individualistic, as in the case of the UK and the DANIELA PETTERSSON-TRABA 76 US. France, in turn, occupies an intermediate position. Krawczak thus finds a significant effect of the extralinguistic factors on the choice between the two near-synonymous terms in the different languages. Consequently, although the concept EMBARRASSED in general is more interactive than that of ASHAMED, being closely linked to the presence of an audience, this is not the case in Polish, where the line between the two emotions is not as clear-cut. Krawczak’s (2018) findings may indeed be relevant for the present dissertation since, in addition to intralinguistic factors of a primarily contextual nature, extralinguistic ones may need to be considered in order to reach a full understanding of onomasiological structure.36 As we have just seen, some studies resort to sociolinguistic and cultural aspects to explain onomasiological variation. However, many of the works that have been reviewed throughout this section, despite offering valuable descriptive information about differences between nearsynonymous expressions, do not connect the patterns uncovered to a specific theoretical framework, nor do they discuss the implications of their findings. As mentioned in Section 2.3.2, Geeraerts (2010: 177–178) states that the distributional corpus-based approach is ultimately a method rather than a theory, given that it is not always evident how the distributional patterns relate to theoretical issues in lexical semantics. Similarly, Gries (2010: 324–325) points to three areas in which this line of research could be further improved, one of which has to do precisely with the lack of theoretical background in many of these studies.37 Therefore, such distributional methods have recently tended to adopt the framework of cognitive semantics and make use of concepts and principles within this theory, such as prototypicality and entrenchment, to explain the findings obtained. One rather early corpus-based study on near-synonymy which provides a cognitive explanation for its findings is Taylor (2003) on the adjectives high and tall. Both terms are positive polarity items indicating the extent of an entity on the vertical dimension. However, while high dates back to the Old English (OE) period, having been inherited from Germanic, the spatial sense of tall is rather recent, approximately from the mid-16th century according to 36 As mentioned in Section 2.4.2, many other studies on usage-based onomasiological variation with a sociolinguistic orientation have been conducted in the domain of cognitive semantics over the last few decades, which demonstrate the usefulness and importance of extralinguistic variables related to cultural and social differences (cf., for instance, Speelman, Grondelaers & Geeraerts 2003; Levshina 2011; Soares da Silva 2013; 2014). 37 The other two areas mentioned by Gries (2010: 324–325) concern (i) the number and range of synonyms object of study and (ii) issues related to data and methodology, more specifically to the type of co-occurrence information of the synonyms taken into consideration. 3 Synonymy 77 the OED (OED, s.v. tall n. II, 6). The adoption of tall to denote an entity of great vertical extent did not imply a decrease in the usage range of high, which was retained for all types of entities with which it was used before. The adjective tall, on the other hand, offered a new way of conceptualizing positive polarity verticality that is specific to the human body (e.g. tall man, tall girl). In order to explain this distribution, Taylor makes use of the so-called vantage theory (e.g. MacLaury 1987; 1995; 1997), originally applied to the semantic domain of COLOR, in which a distinction is made between dominant and recessive terms. A dominant term within a pair is that which can be applied to a wide range of entities or situations, as it emphasizes the similarities between them. On the other hand, a recessive term emphasizes their differences, and as such can only be applied to a narrow range of entities or situations (Taylor 2003: 280). By using this distinction, Taylor concludes that high is the dominant term in the pair, while tall is the recessive one, as it is limited to denote a specific type of verticality, i.e. that of the human body. Taylor (2003) thus demonstrates how findings from distributional corpus-based studies on near-synonymy can be embedded within a theoretical framework on language and human cognition in general, and lexical semantics, in particular. Even though Taylor (2003) briefly considers the diachronic development of tall and high in his discussion of the findings, the rest of the studies reviewed until now adopt a fully synchronic orientation towards lexical onomasiological variation. The reason for this is the scarcity of diachronic studies on lexical near-synonymy from the perspective of usage-based onomasiology with a cognitive stance. The next section deals with the scant diachronic corpusbased analyses on pairs or sets of lexical near-synonyms, as well as with some generalizations about synonymy and diachrony that have been proposed in the specialized literature, including the types of semantic change that near-synonymous words or expressions have been said to undergo over time. 3.3 SYNONYMY AND DIACHRONY As mentioned in Section 2.4.2, both synchronic and diachronic approaches to the study of meaning abound within cognitive semantics. However, while a considerable amount of diachronic work within this tradition has recently been devoted to constructional synonymy and other morphosyntactic changes (e.g. Shank, Plevoets & Cuyckens 2014; Hilpert 2016; De Smet et al. 2018; Breban & De Smet 2019; D’hoedt, De Smet & Cuyckens 2019), we find only a DANIELA PETTERSSON-TRABA 78 handful of diachronic studies on lexical synonymy from a distributional corpus-based perspective. In what follows, a brief summary of this research is provided. Kaunisto (2001) investigates the evolution of twelve pairs of -ic and -ical adjectives from the second half of the 16th century to the first half of the 19th century by drawing on data from the prose section of Chadwyck-Healey Literature Online (1996–2000): angelic(al), authentic(al), comic(al), domestic(al), fantastic(al), heroic(al), magic(al),majestic(al), philosophic(al), poetic(al), tragic(al), and tyrannic(al).First, he analyzes the productivity of each of the two suffixes from the 1300s to the present-day by counting the number of first attestations of -ic/-ical adjectives in the OED. By doing so, Kaunisto is able to determine that, whereas until the 17th century both suffixes are approximately equally productive, from 1750 and until about 1900 -ic is used considerably more often to coin new adjectives, and thus seems to be the dominant suffix throughout this period, especially in scientific domains such as chemistry (e.g. acidic,cyanic). Nevertheless, in the 20th century the productivity of both suffixes becomes similar again. As regards the specific diachronic patterns of pairs of -ic/-ical near-synonyms, Kaunisto identifies five evolutionary trends: (i) First, in some pairs, which are all characterized by exhibiting the semantic feature NOBILITY (e.g. angelic(al), heroic(al), majestic(al)), there is a move from -ical to - ic mainly in the 17th century. Moreover, in previous periods, when adjectives with both suffixes are common, there appears to be no change in meaning or function between them, as they are used with the same senses and in the same syntactic functions and even seem to collocate with the same types of nouns. (ii) Second, pairs such as comic(al), fantastic(al), and magic(al) come to be more often used with -ic in the 18th century. (iii) Third, in the case of tyrannic(al), the number of occurrences with -ic increase as in the preceding two cases, but the -ical form remains the most frequent. (iv) Fourth, philosophic becomes more common at the expense of philosophical in the second half of the 18th century, when the two near-synonyms seem to undergo a process of differentiation: the former has a more popular meaning (e.g. philosophic look), while the latter occurs mainly with nouns referring to science, such as research and theory. Nevertheless, the two terms then converge again in meaning and philosophic became the dominant adjective in both uses. 3 Synonymy 79 (v) Finally, tragic(al) undergoes a change from -ical to -ic, but much later than the rest of the terms discussed, after the first half of the 19th century. Thus, in almost all cases we witness a shift from -ical to -ic, although at different points in time. In some -ic/-ical pairs semantic differentiation takes place in a given period. However, Kaunisto concludes that, in the end, the main result of these changes seems to be that of discarding one of the options rather than maintaining both forms with different meanings and/or functions. Primahadi-Vijaya-R. & Rajeg (2014) examine the nominal collocational profiles of two near-synonymous adjectives from the domain of TEMPERATURE, namely hot and warm, in AmE throughout the last one and a half centuries (i.e. 1860s–2000s) by drawing on data from COHA (Davies 2010–). To this purpose, they extract the top 100 R1 noun collocates of the two adjectives, excluding those that appear less than five times in each decade in the corpus. The data is visualized by means of motion charts (Hilpert 2011; 2013: 66–74), which is a method to display a series of graphs in order to plot the diachronic changes of a given phenomenon. Primahadi-Vijaya-R. & Rajeg show, for each decade in COHA, the absolute co-occurrence frequencies of the different R1 noun collocates and identify several diachronic trends that have contributed to the differential use of these two near-synonyms in present-day AmE. Whereas in the 1860s some nouns, for example, blood,water, and weather, are commonly modified by both hot and warm in the literal sense ‘of high temperature’, other nouns are already clearly associated with one of the two adjectives. For instance, warm typically collocates with nouns such as smile,heart,affection, and friend, thus pointing to its frequent use in the metaphorical extension of ‘friendliness’, although this sense seems to decrease in frequency over time, as shown by the decline of warm with some of these nouns (e.g. heart and welcome). On the other hand, hot seems to be typically used in the 1860s in another metaphorical sense, viz. that of ‘excitement’ and ‘intensity’, as in hot haste or hot pursuit. With the passing of time, the collocational profile of hot changes somewhat. First, from the 1920s onwards, it undergoes a lexicalization process with the noun dog (i.e. hotdog), which is today a compound noun rather than a combination of an adjective and a head noun. This is also the case of hotspot, in which the original meaning of hot is no longer transparent. Second, from the 2000s onwards, hot becomes strongly associated to nouns denoting people (e.g. girl,guy,andwoman) in order to refer to ‘sexually attractive’ individuals, thus indicating a rise in this particular sense of the adjective. Primahadi-Vijaya-R. & Rajeg (2014) thus uncover some interesting differentiating DANIELA PETTERSSON-TRABA 86 whereas the Anglo-Saxon counterparts are more colloquial or neutral and thus used in a wider range of stylistic contexts (Jackson 1988: 66–67; Kay & Allan 2015: 12). Although the competition theory is widespread and has proved successful to understand many changes in language, not only at the lexical level but also at the syntactic, morphological, and phonological ones (cf., for instance, Mondorf 2010; Berg 2014), some scholars consider it an oversimplification, as it fails to account for situations in which synonyms become more similar over time, rather than more dissimilar, a process that has been labeled attraction (De Smet et al. 2018: 203). De Smet et al. (2018) provide evidence of several pairs of nearsynonymous constructions that have come to share more semantic space across time. In one of their case studies, they show that the verb begin has come to be frequently complemented by - ing-clauses at the expense of to-infinitive clauses. In the 1840s begin was usually followed by to-infinitive clauses, with both agentive and non-agentive subjects, though particularly with the latter. Progressively, -ing-clauses became more frequent, at first only with agentive subjects, and later on also with non-agentive ones. To-infinitive complements simultaneously experienced a significant downward tendency in frequency, but did not retract from contexts with agentive subjects, as would be expected in a process of differentiation. On the contrary, the decrease in frequency of to-infinitives took place in non-agentive contexts, which led to the semantic profiles of the two constructions becoming more similar. Thus, while an instance of ongoing replacement of to-infinitive clauses by -ing-clauses, the development described here is also a clear case of attraction. The example provided by De Smet et al. (2018) suggests that a certain level of attraction is possibly a prerequisite for substitution, since synonymous expressions may need to share at least some functional features for replacement to take place. De Smet et al. (2018: 204, 217) explain the process of attraction by means of analogical change, the process whereby one form becomes more similar to another which it already resembled, usually from a formal perspective (Trask 2007: 15–16). However, analogy is here understood in an extended sense to include also expressions which are functionally rather than formally or structurally similar (see also Nuyts & Byloo 2015: 36). De Smet et al. (2018) thus argue that expressions exhibiting functional similarity parallel each other’s behavior through an interchange of characteristics. This theory of semantic parallelism is also evinced in the fact that various synonymous lexical expressions can occur in the same metaphorical mappings and thus come to share more senses over time (e.g. Stefanowitsch 2008; Turkkila 2014). For instance, as argued by Stefanowitsch (2008: 96–99), the nouns joy and happiness, which share 3 Synonymy 87 the literal sense ‘the emotion or state of being highly pleased or delighted’, have developed similar metaphorical mappings in which the emotion is described as a liquid (e.g. river of joy/happiness) or as a source of heat (e.g. sparks of joy/happiness and burn with joy/happiness). Another example concerns the near-synonymous adjectives from the domain of TASTE, for example, delicious,delectable,luscious, and tasty, which besides referring to ‘foods of pleasant flavor’, have all been metaphorically extended to ‘very attractive’ when used for people (e.g. OED, s.v. tasty adj. 1a and 1b). These cases suggest that with time near-synonyms can come to develop the same figurative senses, thus becoming more similar in meaning. Additionally, De Smet et al. (2018: 205) claim that the processes of differentiation and attraction are mutually exclusive. While this seems to be the case when pairs of synonyms are considered, the argument does not necessarily hold if larger groups are taken into account. This is so since different members of a synonym set may hold diverse and sometimes opposite relations to one another. Consider a hypothetical set of near-synonyms with three members, A, B, and C. If we focus only on the diachronic development of Aand B, over time they can become either more similar (i.e. attraction) or more dissimilar (i.e. differentiation), but not both processes at the same time.41 However, if we take all the members of the set into account, A and Bmay become attracted, while simultaneously Band Cmay become differentiated. Therefore, by adopting a broader perspective and including larger groups of near-synonyms in the analysis, both types of changes can be found to be at work in one and the same synonym set. This is an issue that has, to my knowledge, not yet been explored, and that can be of crucial importance for lexical synonymy, where, as discussed in Section 3.2.1, near-synonyms tend to come in larger groups rather than in pairs. 3.4 SUMMARY The present chapter has traced the semantic phenomenon of synonymy, dating back to the structuralist school, which proposes various classifications of this semantic relation. Although different types of synonymy were postulated, only the two-fold distinction between absolute and non-absolute or near-synonymy is relevant for our purposes in this dissertation. Section 3.2 presented an overview of the existing synchronic research on the internal semantic structure of pairs and sets of near-synonyms that has been conducted from a distributional corpus-based 41 A third possibility is that of stability, that is, when no changes in the functional profiles of the near-synonyms occur. DANIELA PETTERSSON-TRABA 88 perspective. Two waves of research were distinguished, which point to a constant evolution of the field, both concerning the number of near-synonyms and the factors included in the analyses, but also regarding other methodological aspects, such as the particular techniques used to measure semantic (dis)similarity. Then, an in-depth review of studies dealing with adjectival near-synonyms was provided, focusing on the types of factors that are typically considered and the most important findings of such studies. In the majority of cases, semantic factors emerged as the most significant ones to explain the choice between synonymous adjectives, outperforming morphosyntactic variables. Lastly, Section 3.3 was devoted to the few diachronic studies on lexical synonymy with a corpus-based distributional orientation, which demonstrate that the relation between particular near-synonyms is often unstable over time. Additionally, typical semantic changes of a semasiological nature such as specialization, generalization, amelioration, and pejoration were shown to be equally relevant in onomasiology, as these processes are also often resorted to in order to explain changes in the internal structure of synonym sets. It was seen that the prototypical view in the literature is that synonyms become differentiated with time if one is not displaced by another that takes over its functions. Nevertheless, recent research has shown that another possible outcome is for synonyms to progressively become more similar instead of more dissimilar, thus undergoing attraction. The shortage of diachronic research on lexical synonymy shows that there is a clear need for more analyses of this nature. In this context, the present dissertation aims at partially filling this existing gap by analyzing in the recent history of AmE a set of adjectival near-synonyms from the olfactory domain, namely fragrant,perfumed,scented,sweet-scented,and sweetsmelling. Chapters 5–7 present the results of corpus-based analyses of this synonym set. First, however, Chapter 4 introduces the synonyms object of study, as well as the reasons for the selection of this particular near-synonym set. Moreover, the data retrieval and annotation processes of the databases employed in the subsequent corpus-based analyses in Chapters 5–7 are also explained in detail. 4 THE CONCEPT PLEASANT SMELLING: DATA SELECTION AND ANNOTATION As became evident in the previous chapter, although absolute synonyms are virtually nonexistent, languages abound in near-synonyms, that is, words with similar, though not identical meanings. This is particularly true for English, which, due to its long history of borrowing from other languages, such as French and Latin in ME and eModE, displays an exceptionally large number of roughly synonymous expressions (Kay & Allan 2015: 12, 31, 88). Consequently, the task of selecting one specific synonym for the analysis is not an easy one. As seen in Sections 3.3 and 3.4, where previous studies of individual pairs and sets of adjectives were reviewed, certain tendencies are observed, with some semantic domains and even specific synonyms receiving particular attention. Thus, for instance, several analyses have focused on -ic and -ical adjectives (i.e. Gries 2001; 2003; Kaunisto 2001) and others have paid attention to basic descriptive adjectives from the domains of SIZE and AMOUNT, namely Biber, Conrad & Reppen (1998: Section 2.6) on big,large, and great, Taylor (2003) on high and tall, and Gries & Otani (2010) on the following two sets: big,large, and great, on the one hand, and little,small, and tiny, on the other. This dissertation, however, sets out to examine a set of adjectival near-synonyms from the relatively understudied semantic domain of SMELL: fragrant,perfumed,scented,sweet-scented, and sweet-smelling, which designate the concept PLEASANT SMELLING. Section 4.1 is devoted to a description of this synonym set, drawing on data from different dictionaries and thesauri, and to an explanation of the motivations for selecting this particular group of adjectives. Different analyses conducted on this synonym set are presented in subsequent chapters (i.e. Chapters 5–7) on the basis of corpus data. Corpora are the mainstream data sources for research in basically all fields of contemporary linguistics, semantics and historical linguistics constituting no exceptions. For the purposes of the present dissertation and due to the relatively low frequency of some of the items in the synonym set object of study (cf. Section 5.1), a rather DANIELA PETTERSSON-TRABA 90 large corpus is desirable in order to achieve a representative sample. The corpus used here is COHA (Davies 2010–), which contains more than 400 million words. A description of this corpus and its suitability for the analyses in this dissertation is provided in Section 4.2, together with the data retrieval process of the occurrences of the near-synonyms employed in the case studies in Chapters 5, 6, and 7. The resulting databases were annotated for a series of languageexternal and language-internal variables pertaining to the contextual features of each attestation of the five near-synonyms. This annotation process is described in depth in Section 4.3. 4.1 INTRODUCING THE SYNONYM SET Five adjectival near-synonyms from the olfactory domain are the object of study of the analyses conducted in the present dissertation, namely fragrant,perfumed,scented,sweet-scented, and sweet-smelling. These five items are exemplified in (11)–(15): (11) Her feet, broad and solid from years of walking, easily passed over the tricky terrain of low shrubs, dead leaves, fallen trees, and trailing vines. It had rained a little last night, and the moist earth was fragrant. (COHA, 2009, FIC, WifeGodsNovel) (12) Jake touched warm, breathing woman, inhaled her freshly bathed scent and found her primal essence beneath the perfumed soap and body lotion. (COHA, 2007, FIC, WolfTalesIII) (13) Smoked Pork Chops with Apple-Red Chile Compote # You can think of this smoky entre from W. Park Kerr’s “Burning Desires: Salsa, Smoke and Sizzle From Down by the Rio Grande,” as pork chops with red-hot scented applesauce all grown. (COHA, 2007, NEWS, Denver) (14) As his eyes adjusted ever so slowly to the gloom, he saw looming shadows, blurred shapes like enormous trees that stirred not at all in the gentle sweet-scented wind. (COHA, 2005, FIC, FantasySciFi) (15) Tapping one slim cigarette out of the pack, she brought the sweet-smelling tobacco to her lips and searched blindly for her lighter, then leaned back in her chaise lounge on the patio of the villa she’d rented for the summer on the island of Saint Martin. (COHA, 2003, FIC, RoomService) 4. The concept PLEASANT SMELLING: Data selection and annotation 91 The main motivations for selecting these particular lexical items are explained in Section 4.1.1. Section 4.1.2, in turn, offers a review of the existing descriptions of these adjectives as provided in dictionaries and thesauri, concentrating specifically on their etymology, different senses, and examples of use. 4.1.1 Main motivations for the selection of the concept PLEASANT SMELLING As mentioned in the introduction to this chapter, the semantic domain of SMELL has hitherto been relatively understudied in semantic research (but cf. Ibarretxe-Antuñano 1999; Digonnet 2018 for metaphor analyses on words referring to the sense of smell). A common claim in the literature is that humans, despite being able to identify and discriminate a wide range of smells, experience difficulties when having to name them (cf. Ibarretxe-Antuñano 1999: 36; Lorig 1999: 392;; Yeshurun & Sobel 2010: 216, among others). This can be one of the reasons why the odor vocabulary has not been as widely researched as other domains such as those of COLOR (e.g. Berlin & Kay 1969; Biggam 2010; 2012; Anderson & Bramwell 2014) and COOKERY (e.g. Lehrer 1969; 1974), or the lexicon used to describe other senses (i.e. hearing, vision, taste, and touch), which have been argued to be much more extensive and precise in English (e.g. Sperber 1975: 115–116; Digonnet 2018: 178–179). However, the HTOED, which provides a semantic classification of all the senses of words, lists a whole 904 word senses under the semantic category SMELL AND ODOUR,42 which have been and/or are still used in English, a fact that appears to contradict previous claims about the extensiveness and precision of this semantic domain in English. Of these words, a great amount are adjectives, some of which are similar to those selected in the present study with some minor variations in definition or use, making this a particular interesting semantic field for the study of near-synonymy. Examples of these semantically related adjectives include aromatic (OED, s.v. aromatic adj. 1), balmy (OED, s.v. balmy adj. 3), odoriferous (OED, s.v. odoriferous adj. 1), odorous (OED, s.v. odorous adj.), perfumy (OED, s.v. perfumy adj), redolent (OED, s.v. redolent adj. 1 and 2), and sweet (OED, s.v. sweet adj. 2), among others, many of which are loanwords from French and/or Latin that entered English between the 15th and the 17th centuries (cf. also Durkin 2014: 413–414). 42 Although AmE orthographic conventions are used in the present dissertation, the British spelling of the labels of semantic categories in the HTOED is retained in order to maintain the original nomenclature. DANIELA PETTERSSON-TRABA 92 Second, as demonstrated in previous pilot studies on this synonym set (Pettersson-Traba 2015; 2018; 2019; 2020a; 2020b; forthcoming), there seem to be some intriguing diachronic developments regarding the relationship between the selected adjectives in the latter part of Late Modern and Present-Day AmE. By analyzing the noun collocates of the adjectives in attributive function, Pettersson-Traba discovers in these works a tendency for the adjectives to be increasingly used to modify artificial, as opposed to natural, smells. Additionally, fragrant and perfumed, which were initially the most frequent adjectives, are gradually replaced by scented, thus reflecting a change, which still seems to be ongoing, in the relation between the near-synonyms over time. The concept PLEASANT SMELLING therefore appears to be a particularly interesting and relevant object of study for the purposes of the present dissertation, which aims at uncovering the semantic processes that specific lexical synonyms undergo with time, as well as the underlying motivations for such processes. Finally, another crucial reason for selecting this synonym set is its rather low degree of polysemy, which is advantageous for the innovative approach adopted here. Given that the main focus is on the semantic relation of synonymy, examining a set of lexical items that, besides being synonymous, are also highly polysemous would result in an extremely complex and almost unfeasible study. For example, analyzing a highly polysemous synonym set would require a great amount of manual pruning in order to ensure that the final datset contained exclusively those instances which are truly synonymous and therefore interchangeable. This would entail a great deal of “donkey work” as, to my knowledge, no semantically tagged diachronic corpora exist to date.43 Nevertheless, the semantic structure of the selected five nearsynonymous adjectives is not devoid of complexity, with at least three senses being shared by all of them, including a figurative one. This issue is further discussed in the next section. 4.1.2 Revising reference works The information about the near-synonyms reviewed in this section is based on different English dictionaries and thesauri, both historical and present-day. The list of present-day dictionaries 43 It is even difficult to find semantically tagged synchronic corpora, although a few do exist. Examples are the English SemCor Corpus created by the WordNet project research team (Landes, Leacock & Fellbaum 1998), which includes semantic annotations for sense, and the BBN Pronoun Coreference and Entity Type Corpus (Weischedel & Brunstein 2005), which provides information about the co-indexation of pronouns and their antecedents, as well as a semantic categorization into entity types (e.g. PERSON,PRODUCT,PLANT,TIME,QUANTITY, and EVENT). 4. The concept PLEASANT SMELLING: Data selection and annotation 93 includes the following: American Heritage Dictionary of the English Language (AHDOE), Cambridge Dictionary (CD), Collins Dictionary (CoD), Lexico,Longman Dictionary of Contemporary English (LDOCE), MacMillan Dictionary (MD), Merriam-Webster (MW), and Newbury House Dictionary of American English (NHDAE). All of them offer information about the current senses of the selected adjectives and examples of use. Moreover, several of the dictionaries contain a thesaurus section, thus providing synonyms and antonyms of words. CoD, which uses data from the COBUILD Corpus to exemplify the different meanings of words, also gives information about their usage frequency, not only in PDE but also in earlier periods of the language. The reference work therefore serves to shed light on the different senses and uses of the near-synonyms at issue in contemporary English. In addition, historical information has been drawn from the OED, which is considered to be the largest and most comprehensive dictionary of English (e.g. Hoffmann 2004: 18). First published in 1884 and now on its third edition, available online, the OED constitutes an unparalleled database when it comes to tracking the history of English words and their meanings from the time when they entered the language until the present-day, thus serving as an essential guide for research on historical English. As such, it includes earlier meanings and etymology alongside current meanings. Consequently, the OED offers more illustrations of word usage than most other dictionaries, as it contains quotations from different periods of English. Each entry in the OED provides information about the pronunciation, spelling, frequency band in PDE, etymology, POS, and senses of a word. By making use of the examples of usage one can ascertain the first known attestation of the different meanings of a word and, in many cases, it is also possible to establish the relation between senses. The OED makes evident that fragrant,perfumed, and scented are loanwords that were borrowed from French in eModE, specifically during the first half of the 16th century (cf. examples (16)–(18) below), as is the case of many words in this particular semantic domain (Durkin 2014: 414; cf. Section 4.1.1 above). However, despite this common origin, the way they entered the English language differs. Fragrant, which comes from the French adjective fragrant, ultimately dates back to Latin IUƗJUDQW-em, the present participle of IUƗJUƗUH‘to smell sweetly.’ However, although it derives from a participial form, when fragrant became an English word, its participial nature was no longer recognizable. Both perfumed and scented, in turn, are derivatives formed by the suffix -ed (OED, s.v. -ed suffix1) attached to the verbal base DANIELA PETTERSSON-TRABA 94 forms perfume and scent, which were also borrowed indirectly from Latin via French.44 Contrary to fragrant,perfumed and scented were already formed in English by derivation, thus retaining their participial shape, which was used to form the past participles of weak verbs. It is, therefore, probably correct to assume that perfumed and scented still retain part of their verbal semantic functions in PDE, namely that of ‘denoting a process or action’ (Biber et al. 1999: 63). This may also be the case of the adjective sweet-scented, first attested at the end of the 16th century (cf. example (19)), given that it is a compound formed by the loanword scented and the OE adjective sweet. On the other hand, sweet-smelling is a composite form first attested in the early 15th century that combines the adjective sweet and the gerund form of smell,also of OE origin (cf. example (20)).Therefore, as on many other occasions in the history of English, near-synonymy arises in this specific case as a result of borrowing from Latin and French in the eModE period, with loanwords —fragrant,perfumed,scented, and, to a certain extent, sweet-scented— coming to be established in the language alongside an already existing native expression with the same meaning, that is, sweet-smelling. (16) The fragraunt odour, and oyntment of swete flour (c.1530. OED, s.v. fragrant adj.) (17) Suffitus, perfumed. (1538. OED, s.v. perfumed adj. 1) (18) Many here smell strong but none so ranke as he A stronger sented knaue [knave] then he was cannot bee. (?c.1562. OED, s.v. scented adj. 1) (19) Sweet sented Roe (1591. OED, s.v. sweet-scented adj. a) (20) A place..Y-set aboute with floures so swete smellyng. (c.1400. OED, s.v. sweetsmelling adj. 1) Although not all the dictionaries consulted distinguish the same senses and subsenses for the five near-synonyms, they all provide the same basic meaning ‘having a sweet pleasant smell’ for the five adjectives. This general sense is in fact the only definition provided for sweetscented and sweet-smelling in the OED,Lexico, and CoD, which include an independent entry 44 Another possibility, according to the OED, is that perfumed and scented derive from the nominal forms perfume and scent plus the suffix -ed (OED, s.v. -ed suffix2). 4. The concept PLEASANT SMELLING: Data selection and annotation 95 for these two compound words.45 Similarly, fragrant occurs only with this sense in all eight PDE dictionaries, while the OED points out that this adjective can also be used figuratively to mean simply ‘sweet or pleasant’, as in example (21), where fragrant modifies the noun memory, which refers to an abstract entity and can therefore not emit an odor: (21) Their fragrant mem’ry will out last their tomb (1782. OED, s.v. fragrant adj.) Whereas a few dictionaries also define perfumed and scented exclusively in the general and figurative senses just mentioned, others make a distinction between two senses, which will here be labeled ‘natural’ and ‘artificial’.46 This is the case of, for instance, MD, which lists the following two definitions for perfumed and thus serves to illustrate these two senses: (i) ‘[P]leasant to smell because of natural qualities’ (MD, s.v. perfumed adj.) (ii) ‘[P]leasant to smell because perfume has been added or used’ (MD, s.v. perfumed adj.) Therefore, according to the dictionaries that draw this distinction, there is a difference in nuance of meaning depending on whether the source of the smell is natural, in which case modified nouns designate entities which can release a smell on their own, such as wallflowers in (22), or artificial, where the adjectives collocate with nouns referring to entities which can acquire a 45 MW does not provide an entry for sweet-smelling, only for sweet-scented.MW and the OED also offer an additional subsense for sweet-scented, namely ‘in names of species or varieties of plants having sweet-smelling flowers, leaves, etc.’ (OED, s.v. sweet-scented adj. b), as in sweet-scented pea or sweet-scented geranium. Examples exhibiting this sense are excluded from the present study as the adjective in this case forms part of the vernacular name of the plants. None of the other four adjectives seem to be used in this sense according to the reference material consulted in this dissertation. 46 Moreover, two additional senses are provided for scented in the OED: (i) ‘Of tea, tobacco, etc.: flavoured with an aromatic ingredient. Also: having a fragrant taste as if flavoured by an aromatic ingredient. […]’ (OED,s.v. scented adj. 2b) (ii) ‘With modifying word. Chiefly of an animal: having a sense of smell of the specified kind. […]’ (OED,s.v. scented adj. 3) The former meaning is classified in the OED as a subsense of the artificial sense and, since the other adjectives have also been found to collocate with teaand tobacco-related nouns, for instance sweet-smelling tobacco in example (15) above, examples featuring this sense are included in the subsequent corpus analyses. Meaning (ii), on the contrary, is exclusive to scented and is therefore excluded here. DANIELA PETTERSSON-TRABA 102 from 0 to 0.4 instances per million words. Interestingly, most changes seem to have taken place from approximately the 19th century onwards, precisely the time span covered in the present dissertation. Despite this valuable frequency information for PDE and earlier periods of English, neither the OED nor CoD shed any light on the distribution of the adjectives in different senses and more specific contexts of use. The dictionaries and thesauri reviewed in the present section offer some important insights into the relation between the five near-synonymous adjectives fragrant,perfumed,scented, sweet-scented, and sweet-smelling regarding both their similarities and possible differences in terms of frequency and semantic characteristics. However, the information provided by these reference works is still insufficient in order to acquire a full understanding of the structure of the concept PLEASANT SMELLING and of the degree to which the five adjectival near-synonyms are interchangeable. As such, some questions are yet left unanswered: (i) Are perfumed and scented the only adjectives of the set that are used in the artificial sense, i.e. ‘pleasant to smell because perfume has been added or used’, or do fragrant, sweet-scented, and sweet-smelling also appear in such semantic contexts? (ii) If fragrant,sweet-scented, and sweet-smelling are also used in the artificial sense, is there any difference in their distribution, both in this and in the other senses (i.e. natural and figurative) of the adjectives? (iii) Does the distribution of the five adjectives across senses remain stable over time or does it fluctuate? In other words, are the apparent changes in frequency of the adjectives revealed in CoD specific to one or several contexts of use or do they, on the contrary, occur equally across all contexts? (iv) The reference works point to some semantic (dis)similarities among the adjectives, but do differences concerning, for instance, connotation, style, and/or morphosyntactic features also exist? The analyses in the present dissertation are designed in order to try to answer these research questions. However, before moving on to the results of these analyses, which are provided in Chapters 5–7, Sections 4.2 and 4.3 offer a description of the corpus used, as well as the data retrieval and annotation processes. 4. The concept PLEASANT SMELLING: Data selection and annotation 103 4.2 CORPUS DESCRIPTION AND DATA RETRIEVAL Given the relatively low frequency of the five near-synonymous adjectives object of study, a large corpus is necessary to carry out the analyses in Chapters 5–7. For this reason, the data used was drawn from COHA. Released in 2010, COHA is the largest historical corpus of English, containing about 400 million words of running text from more than 100,000 individual texts, which are divided into four different genres or text-types: fiction, popular magazines, newspapers, and non-fiction.52 COHA covers the period 1810–2009 in AmE, and this time span of 200 years is split into 20 decades. Table 3 displays the number of words across the different decades and text-types. As shown in the table, COHA is not a balanced corpus in terms of number of words, neither across decades nor across genres, that is, it contains considerably more data for the later decades than for the earlier ones, and fiction accounts for a much higher percentage of the data than the other three genres. For instance, until the 1850s there are less than 17 million words per decade, whereas the last two decades, 1990s and 2000s, include nearly double the amount of running text, almost 28 and 30 million words, respectively. Despite the unequal distribution of the data, an effort was made by the corpus compilers to keep the distribution of the different genres even through the 20 decades. For instance, fiction totals between 48–55% of the data in each decade (see the last column in Table 3). Similarly, roughly the same percentage of popular magazines, non-fiction, and newspapers is maintained over time, with the exception of newspapers in the first six decades. This distribution ensures that possible changes are not just a byproduct of fluctuations in the number of words per genre over time, but that they reflect actual changes. 52 Within the four genres there are further subclassifications. For more details about the main sources used for each of the four text-types, see the information available at http://www.helsinki.fi/varieng/CoRD/corpora/COHA/basic.html. DANIELA PETTERSSON-TRABA 104 Table 3. Number of words in COHA across decade and genre DECADE FICTION POPULAR MAGAZINES NEWSPAPERS NONFICTION BOOKS TOTAL % FICTION 1810s 641,164 88,316 – 451,542 1,181,022 54% 1820s 3,751,204 1,714,789 – 1,461,012 6,927,005 54% 1830s 7,590,350 3,145,575 – 3,038,062 13,773,987 55% 1840s 8,850,886 3,554,534 – 3,641,434 16,046,854 55% 1850s 9,094,346 4,220,558 – 3,178,922 16,493,826 55% 1860s 9,450,562 4,437,941 262,198 2,974,401 17,125,102 55% 1870s 10,291,968 4,452,192 1,030,560 2,835,440 18,610,160 55% 1880s 11,215,065 4,481,568 1,355,456 3,820,766 20,872,855 54% 1890s 11,212,219 4,679,486 1,383,948 3,907,730 21,183,383 53% 1900s 12,029,439 5,062,650 1,433,576 4,015,567 22,541,232 53% 1910s 11,935,701 5,694,710 1,489,942 3,534,899 22,655,252 53% 1920s 12,539,681 5,841,678 3,552,699 3,698,353 25,632,411 49% 1930s 11,876,996 5,910,095 3,545,527 3,080,629 24,413,247 49% 1940s 11,946,743 5,644,216 3,497,509 3,056,010 24,144,478 49% 1950s 11,986,437 5,796,823 3,522,545 3,092,375 24,398,180 49% 1960s 11,578,880 5,803,276 3,404,244 3,141,582 23,927,982 48% 1970s 11,626,911 5,755,537 3,383,924 3,002,933 23,769,305 49% 1980s 12,152,603 5,804,320 4,113,254 3,108,775 25,178,952 48% 1990s 13,272,162 7,440,305 4,060,570 3,104,303 27,877,340 48% 2000s 14,590,078 7,678,830 4,088,704 3,121,839 29,479,451 49% TOTAL 207,633,395 97,207,399 40,124,656 61,266,574 406,232,024 51% The size and scope of COHA make it the perfect resource for the purposes of the present dissertation, as it allows us to examine low frequency items such as the near-synonyms under study with greater confidence than smaller corpora. Additionally, COHA is freely available and easily accessible via its inbuilt online interface, which provides many convenient user-friendly search tools. In particular, given that COHA is POS-tagged, it is possible to search only for the lexical items of interest here, and thus exclude the verbal uses of perfumed and scented. The interface also makes it possible to retrieve the collocates of a given word. By means of this option, one can select a search window of up to nine words to the left and nine words to the right of the node word (i.e. L9–R9). Moreover, the search for collocates can be restricted to specific POS and, for instance, retrieve only noun collocates. The interface also permits the researcher to establish a frequency and/or PMI threshold for the collocates, thereby limiting the search to include only those which occur relatively frequently with a node. Two separate databases were created from the data in COHA, given that the aims of the case studies in Chapters 5 and 6, on the one hand, and Chapter 7, on the other, differ and thus impose distinct requirements on the data. For the first database the following queries were made by using COHA’s online interface: ‘fragrant_j*’, ‘perfumed_j*’, ‘scented_j*’, ‘sweet- 4. The concept PLEASANT SMELLING: Data selection and annotation 105 scented_j*’, ‘sweetscented_j*’, ‘sweet scented’, ‘sweet-smelling_j*’, ‘sweetsmelling_j*’, and ‘sweet smelling’, where the POS-tag ‘_j*’ indicates that only adjectives were searched for. The strings ‘sweet scented’ and ‘sweet smelling’, written as two separate words without hyphenation, were not specified for POS because it is not possible to add one POS-tag to two separate words, as the tag would only affect one of them. In addition, the instances of perfumed and scented tagged as verbs in COHA were also retrieved, that is, ‘perfumed_v*’ and ‘scented_v*’, where ‘_v*’ stands for verb. This step was taken in order to identify possible adjectival uses of these lexical items which had been erroneously tagged as verbs. The instances of the adjectives sweet-scented and sweet-smelling were retrieved by means of three distinct search strings given that, although the two compounds are mostly spelt with a hyphen, alternative spellings exist, as in many other cases of compounds (Huddleston & Pullum et al. 2002: 451).53 It is a well-known fact that it is often very difficult to disambiguate between compound words and syntactic constructions formed by two independent lexical items (e.g. a blackbird vs. a black bird), given that there is not a clear-cut division between the two, but rather a cline (e.g. Biber et al. 1999: 589–590; Huddleston & Pullum et al. 2002: 1644). This problem is further aggravated when texts from different historical periods are considered, as in the present dissertation, because as pointed out by Quirk et al. (1985: 1569), it is usually the case that compounds first occur as individual words and only once they become more established and lexicalized are they written as hyphenated or even as single words. Although alternative interpretations may exist for such ambiguous cases, the decision was taken to include in the analysis the examples corresponding to the search strings ‘sweet scented’ and ‘sweet smelling’ because they seem to behave exactly as those instances in which the adjectives are written with a hyphen or as one word. Consider in this connection examples (30)–(32): (30) Before it lay a neat smooth little court, surrounded by a close hedge, of a sweet scented red and white flower, resembling the honeysuckle in shape. (COHA, 1828, MAG, NorthAmRev) 53 For the sake of convenience, throughout the present dissertation, these two compounds are always referred to by means of the hyphenated alternative, that is, sweet-scented and sweet-smelling, as the vast majority of examples retrieved from COHA correspond to these spellings (cf. Table 4 below). DANIELA PETTERSSON-TRABA 106 (31) It has white, sweetscented flowers, arranged in the same manner as the last; stem without spots. Leaves ovate lanceolate, quite smooth. (COHA, 1851, NF, FlowerGardenBrecks) (32) Bladder Senna. Colutea, an ancient name of a bush with sweet-scented flowers. The genus includes a number of species of shrubs, with yellow or orange, pea-shaped flowers, which are succeeded by seed-vessels like bladders. (COHA, 1851, NF, FlowerGardenBrecks) In all three examples, sweet(-)scented modifies the noun flower(s) without any discernible difference in meaning. It is worth mentioning here that a series of criteria, both syntactic and non-syntactic, have been used to distinguish between a true compound and a mere combination of two independent words (Biber et al. 1999: 589–590; Huddleston & Pullum et al. 2002: 448– 451). Among syntactic criteria, combinations of two words can be individually coordinated with and modified by a different word, whereas this is not possible in the case of compounds. As to non-syntactic criteria, compounds are typically stressed on the first element, are spelt as one word (either without or with a hyphen), and have a non-transparent meaning. Combinations of two independent words, in turn, are typically stressed on the second element, are written as two separate words, and have a transparent meaning. These syntactic and non-syntactic criteria, however, were not easily applicable to the current dataset for various reasons. First, given that COHA contains only written language, it is not possible to know the stress patterns of the relevant examples. Second, orthography is not a reliable clue, as many compounds have alternative spellings and, as mentioned above, compounds are typically first written as separate words and later on as hyphenated or single words as they become more accepted and established. Moreover, hyphenation is claimed to be less common in AmE (Quirk et al. 1985: 1569). Finally, meaning is also problematic in this case, given that the semantics of the compounds sweet-smelling and sweet-scented are highly transparent in the sense that they are the sum of their components, namely ‘that smells sweet’ and ‘that has a sweet scent’, respectively. Therefore, their meaning is the same if they are spelt as two separate units. Considering the difficulty of applying these criteria to the data from COHA, the decision to include in the analyses all instances corresponding to the search queries ‘sweet scented’ and ‘sweet smelling’ was further reinforced. 4. The concept PLEASANT SMELLING: Data selection and annotation 107 Table 4 summarizes the data retrieval process of the first database, relevant to the analyses in Chapters 5 and 6, and the initial number of examples extracted by means of each query. Table 4. Number of instances retrieved per query QUERY NUMBER OF INSTANCES RETRIEVED fra g rant _j *3,374 perfumed _j *792 scented _j *792 sweet-scented 173 sweetscented _j *5 sweet scented 37 sweet-smellin g_j *207 sweetsmellin g_j *4 sweet smellin g 37 perfumed _ v* 340 scented _ v* 577 Total 6,338 The second database, relevant for the discussion in Chapter 7, does not include the instances of the adjectives, but their noun collocates. Therefore, the collocates option in COHA was used to retrieve the lemmas of the collocates of fragrant,perfumed,scented,sweet-scented, and sweet-smelling by means of the POS-tag _nn* in an L5–R5 context window. An L5–R5 context window was selected in this case since it has been shown that tighter windows, such as one or two words, often lead to data sparseness, especially if low frequency items are considered (Sahlgren 2006). Moreover, such tight windows are often more appropriate to retrieve semantically (dis)similar terms, such as synonyms or antonyms of the target word (Peirsman, Heylen & Geeraerts 2008: 40). On the contrary, if one is interested in typical collocates of the target, it is convenient to loosen the context window somewhat, for instance, to L5–R5. Another option would be to consider only collocates which are syntactically connected to the target, in the present case either syntactically —or semantically— modified nouns of the adjectives (e.g. the fragrant flower;the flower is fragrant). However, this would also lead to a lower number of retrieved collocates, and thus again to data sparseness, as in the case of tighter context windows. In addition, only the noun collocates of the adjectives were considered because, as argued by, for example, Geeraerts (1986), Justeson & Katz (1995), and Gries (2001; 2003), nouns are more informative than other word types when it comes to the semantics of adjectives (cf. Section 3.2.2). As mentioned earlier (cf. Section 2.3.2), Geeraerts’ (1986) study on the DANIELA PETTERSSON-TRABA 108 Dutch adjective vers ‘fresh’ demonstrates that the fine-grained aspects of meaning of polysemous adjectives can be discovered by examining the nouns they modify. The second database did not require further annotation due to the fact that it is used only in the analyses of individual collocates in Chapter 7. On the contrary, the instances retrieved for the first database were manually pruned and subsequently annotated for a number of variables, both language-external and language-internal ones. The pruning and annotation processes of the first database are the focus of the next section. 4.3 DATA ANNOTATION Prior to the annotation process, the instances in the first database were revised one by one to exclude false positives, as well as other problematic examples. First, in the case of duplicated occurrences, that is, when the same instance appeared more than once in the corpus, only one of them was maintained in the database. Second, three cases of perfumed from the same source (examples (33) and (34) below) were used in a metalinguistic context and were therefore excluded, as they are not actual instances of the adjective: (33) A new alliterative bombardment is inaugurated by “poop” with its doubled “p’s.” closely followed by the doubled “p’s” in “purple,” and then hammered home by another “p” in “perfumed.” (COHA, 1909, MAG, Harpers) (34) Note also the magical use of the “f’s” and “v’s” in “perfumed,” “love,” and “silver,” and of the “m’s” in “perfumed,” “them,” “made,” and “amorous.” (COHA, 1909, MAG, Harpers) Third, two instances of perfumed and two of fragrant were also removed given that they were part of the name of a brand (35) and of the title of a book (36): (35) She burned your translation of The Perfumed Garden, claiming you would not have wanted to publish it unless you needed the money for it, and you didn’t need it, of course, because you were now dead. “Burton was speechless for one of the few times in his life. |p109 Frigate looked out of the corner of his eyes at Burton and grinned. 4. The concept PLEASANT SMELLING: Data selection and annotation 109 He seemed to be enjoying Burton’s distress.” Burning The Perfumed Garden wasn’t so bad, though bad enough. (COHA, 1971, FIC, YourScatteredBodies) (36) It wasn’t Sabine who had mixed up some poisonous sludge as a joke to mock at him and his obsessions. Why hadn’t he believed her? She had given him a bottle of Arkady Ferson’s Fragrant Goddess, prepared by him and described in his 1969 treatise, the veracity of which Jeremy had always rejected out of hand. (COHA, 2007, FIC, FantasySciFi) Fourth, during the course of the manual revision, some false positives of ‘perfumed_j*’and ‘scented_j*’ were identified, that is, examples which were erroneously tagged as adjectives, but in fact corresponded to verbal uses. This is the case of, for instance, examples (37) and (38): (37) He was determined to have a luxurious bath, to be shaved and perfumed, to leave behind him the very dust of his past life. (COHA, 1893, FIC, SingerFromSea) (38) The cops had heard that one before, and scented trouble. One of them began to advance. (COHA, 1970, MAG, Time) In (37), perfumed is a past participle in a verb phrase in the passive voice, with the meaning ‘[t]o impregnate with a (usually pleasant) odour; to impart a (sweet) smell to; to apply perfume to’ (OED, s.v. perfume v. 2). In (38), scented is the past participle or past simple of the verb scent, used figuratively to mean ‘[t]o perceive or discover as if by smell; to detect or discern instinctively or from subtle indications; to intuit’ (OED, s.v. scent v. I. 2b). Among the instances of perfumed and scented tagged as verbs, that is, those retrieved by means of the queries ‘perfumed_v*’ and ‘scented_v*’, some were in fact adjectival rather than verbal uses of the items in question. As is well-known, the distinction between past participles and participial adjectives is not always straightforward, and several such ambiguous instances were identified in the database. Therefore, a decision had to be made regarding which of these cases counted as adjectival uses and which, on the contrary, counted as verbal ones. A number of criteria have been proposed in the literature (e.g. Granger 1983; Andersson 1985; Quirk et al. 1985: Section 7.16; Biber et al. 1999: 505–506; Huddleston & Pullum et al. 2002: 78–79) to classify these ambiguous examples as being located closer to the adjectival end or closer to the DANIELA PETTERSSON-TRABA 110 verbal end of the continuum. Even though participial adjectives are not considered prototypical instances of the adjectival class, a participle is regarded more adjective-like than verb-like if:54 (i) the focus is more on the state resulting from an action rather than on the action itself, as in example (39), where the emphasis is on the look of the man, not on the process whereby he acquired his barbered and perfumed appearance; (ii) it can occur with lexical copular verbs such as appear,feel,orseem (cf. example (40)); (iii) it can be modified by adjectival intensifiers such as very (cf. example (41)); (iv) it can be inflected for degree, either comparative or superlative (cf. examples (39) and (42)); and (v) it can be coordinated with central or true adjectives (cf. examples (40) and (43)). (39) And, oddly enough, this look of premature senility was not masculine but feminine. Though no more barbered and perfumed than the next Italian man, he evoked the black mass of the dressing-table and the hand-mirror (COHA, 1950, FIC, CastColdEye) (40) Our voices seemed low and scented with fragrant incense, hidden behind the staticcracked passion of the Baptists, the Pentecostal brimstone and fury, the preacher’s calling from the pulpit, the congregation’s response from the pews. (COHA, 1998, FIC, MassachRev) (41) They did, however, keep up a suspicious intimacy with a brilliantly lighted, though not very fragrantly scented, saloon on the left. (COHA, 1858, FIC, LifeAdventures) (42) The breezes that bore them were more exquisitely scented than those which had soothed my doubts, bad stimulated my imagination on her feast-days. (COHA, 1922, MAG, Atlantic) 54 It is worth noting here that these criteria are postulated to disambiguate between past participles and participial adjectives in PDE. Therefore, it is not entirely clear whether they can also be safely applied to data from earlier periods in the history of English. However, these criteria are often the only way of reaching a relatively objective decision regarding such ambiguous cases. 4. The concept PLEASANT SMELLING: Data selection and annotation 111 (43) The rest of him, which was Spanish, was indolent and arrogant and perfumed. (COHA, 1933, FIC, Harpers) Examples of perfumed and scented which satisfied any of these five criteria were consequently kept in the database and subsequently included in the analyses. All other occurrences of these two items tagged as verbs were excluded from the count. Finally, some occurrences of the five near-synonyms were counted more than once. This is the case of examples in which the adjectives modify two or more coordinated nouns, as in (44): (44) For years Italy led in perfuming; it supplied the rest of Europe with sweet bags, perfume cakes for throwing on fires, fragrant candles and cosmetics, scented gloves and pomanders (COHA, 1922, MAG, Mentor) Four entries in the database were derived from example (44): two instances of fragrant and two of scented. This was so because both adjectives here modify more than one noun, which appear in coordination, to wit, candles and cosmetics, in the case of fragrant, and gloves and pomanders, in the case of scented. The reasons for taking this decision were two. First, in many such cases, the coordinated nouns belong to different semantic domains and even instantiate different senses of the adjectives, which poses problems for the annotation of the variables Sense (cf. Section 4.3.1.1 below) and Semantic category (cf. Section 4.3.1.2 below). In (44), for instance, the nouns modified by fragrant, namely, candles and cosmetics, were classified as belonging to the semantic categories OBJECT and COSMETICS, respectively. Similarly, the nouns modified by scented,gloves and pomanders, were categorized as TEXTILE AND CLOTHING and COSMETICS, respectively. Second, given that one of the variables for which the data was annotated included the specific nouns modified by the adjectives (cf. Section 4.3.2 below), coordinated nouns had to be kept separate. As mentioned in Section 4.2, the data retrieval process resulted in a total of 6,338 instances (cf. Table 4 above). Following the manual pruning discussed in the preceding paragraphs, 5,764 occurrences of the five adjectives remained in the first database. These examples were subsequently annotated for eleven variables of different types, including language-internal DANIELA PETTERSSON-TRABA 118 Both USAS and the HTOED provide detailed hierarchies of semantic classes and domains, which in fact overlap to a large degree. For instance, both sources distinguish categories for body parts, people, plants, time, and weather. However, some important differences between them are worth highlighting, especially concerning their respective advantages and disadvantages. On the one hand, whereas USAS automatically tags words according to their semantic category —albeit with a certain degree of manual editing—, the HTOED is less automatic, as the researcher has to select manually the correct sense for the lexical items of interest in the OED before being able to retrieve its semantic classification.59 On the other hand, while the HTOED is extremely sense specific, that is, the same word typically occurs in many different semantic classes, this is often not the case of USAS, which commonly provides only one or two tags for each word. In fact, many examples of metonymies and metaphors are not captured by USAS, while they do appear in the HTOED.60 To illustrate this point, consider example (50): (50) Down in the vale where cowslips are growing, Where violets breathe thro’ sweet scented lips, Where brook o’er the bright pebbly bottom is flowing, And bee of the nectar of columbine sips. A monarch it stands of regnative power, In a graceful symmetrical pose; (COHA, 1895, FIC, OurProfessionOther) In (50), the noun lips is a typical example of metaphorical sense extension, in which it is used to refer to a part of a flower, rather than to a body part, thus being used with the meaning ‘[s]omething resembling the lips of the mouth’, specifically ‘[o]ne of the two divisions of a bilabiate corolla or calyx’ (OED, s.v. lip n. II. 5c). The HTOED provides an entry for this sense of the noun, whereas USAS classifies this instance of lip as a body part (i.e. B1: ANATOMY AND 59 One can also use the HTOED directly by searching for all the items included in a given semantic class (for such use of the HTOED, cf., for instance, Allan 2015 on EDUCATION). However, in order to do so, it is first necessary to know which specific category or categories to examine, which is not the case in the present dissertation, whose goal is to discover to which category the modified nouns belong in a particular context of use. 60 The definitions of metaphor and metonymy followed in the present dissertation are those of conceptual approaches within cognitive semantics (cf., for instance, Lakoff 1987; 1993; Lakoff & Johnson 2003 for metaphor and Panther & Radden 1999; Kövecses 2010 for metonymy; for more information, see Section 2.4). 4. The concept PLEASANT SMELLING: Data selection and annotation 119 PHYSIOLOGY), even when access is given to the whole context in (50) and even though this is a fairly well-established metaphor.61 The classification of nouns modified by fragrant,perfumed,scented,sweet-scented, and sweet-smelling resulted in a total of twelve semantic categories: (i) ABSTRACT; (ii) BODY AND PEOPLE;(iii)CLEANING;(iv)COSMETICS;(v)EARTH,ATMOSPHERE,AND WEATHER;(vi)FOOD AND DRINK;(vii)OBJECT; (viii) PLANTS AND FLOWERS;(ix)SENSATION;(x)SPACE; (xi) SUBSTANCE AND MATERIAL; and (xii) TEXTILE AND CLOTHING. The category ABSTRACT groups together under one single heading all abstract nouns, such as those denoting actions, processes, beliefs, states, and temporal terms. These are all entities which cannot emit an odor and thus the near-synonymous adjectives, when modifying such nouns, must be interpreted in the figurative sense ‘sweet or agreeable’. An example is provided in (51): (51) When she slips away in the dusk to-night I shall put a period to my thought of Maria de Guadalupe Rosalia Merced Castello. I want to keep this fragrant memory of her. (COHA, 1922, FIC, JaneJourneysOn) The category BODY AND PEOPLE contains nouns designating human beings and human body parts, together with proper nouns and pronouns that refer to people. Both USAS and the HTOED propose separate classes for body (i.e. B1: ANATOMY AND PHYSIOLOGY and THE BODY, respectively), on the one hand, and for people (i.e. S2: PEOPLE in USAS and PEOPLE in the HTOED), on the other. However, in the present study the decision was made to keep the two together due to the similarities of the examples, which all refer to the pleasant smell people exhibit, either coming from a specific body part, as in (52), or in general, as in (53). (52) The Spencer woman was there with us -- before us -- all around us. “I am Armand Dalberg’s wife” was pounding in my brain. Then I felt a soft little hand slip into mine; a perfumed hair tress touched my cheek; and the sweetest voice, to me, on earth whispered in my ear. (COHA, 1906, FIC, ColonelRedHuzzars) 61 In fact, body parts constitute a very common source of conventional metaphors, as they represent familiar concepts, e.g. head/foot of a mountain or mouth of a cave (Kay & Allan 2015: 159–160; cf. also Kövecses 2010: 18). DANIELA PETTERSSON-TRABA 120 (53) Tea and toast unobserved before them, music drifting unheard about them, furred and fragrant women coming and going; all this was but the vague setting for their own thrilling drama of love and confidence. (COHA, 1921, FIC, BelovedWoman) A few instances of occasional meanings of clothing nouns which present extensions of meaning via the metonymical pattern PIECE OF CLOTHING FOR PERSON (e.g. Waag 1908; Nyrop 1913; Paul 1920; Esnault 1925; cf. also Geeraerts 2010: 33) or OBJECT USED FOR USER (Lakoff & Johnson 2003: 38; Kövecses 2010: 172), previously mentioned in Section 2.1, were also categorized as BODY AND PEOPLE. This is the case of (54), where bodice does not refer to the piece of clothing, but to the person wearing it: (54) One member of that body was not present. Well-nigh all Williamsburgh knew by now that Mr. Marmaduke Haward lay at Marot’s ordinary, ill of a raging fever. Hooped petticoat and fragrant bodice found reason for whispering to laced coat and periwig; significant glances traveled from every quarter of the building toward the tall pew where, collected but somewhat palely smiling, sat Mistress Evelyn Byrd beside her father. (COHA, 1902, FIC, Audrey) Nouns referring to cleaning and personal care products were classified into two separate categories, following the organization in the HTOED, but contrary to that in USAS, which keeps the two together under the label CLEANING AND PERSONAL CARE (i.e. class B4). These two categories are (i) CLEANING, which encompasses those products whose essential purpose is that of cleansing, either the body or areas and objects which people inhabit or use on a daily basis, and (ii) COSMETICS, which includes those products used mainly to enhance or improve physical appearance, that is, for beautification purposes. Examples of perfumed and sweet-smelling modifying a CLEANING and a COSMETICS noun are provided in (55) and (56), respectively: (55) He watched her empty the leather pouch, daintily clean it out, fill the washbowl again and sink into it to her upper arms. The soap was perfumed. She inhaled deeply.” Luxury is everywhere.” (COHA, 1973, FIC, EagleEye) 4. The concept PLEASANT SMELLING: Data selection and annotation 121 (56) “This smoke is getting thick. Some of this toilet water might help if I sprinkled it about.” One whiff of the sweet-smelling cologne was enough for Bragdon and he bolted up the companionway, leaving the stateroom door wide open and the prisoner free to go where he pleased. (COHA, 1902, FIC, BrewstersMillions) Apart from instances like (56), CLEANING also comprises cases such as (57), where tub is an example of an extension of meaning through the conceptual metonymy CONTAINER FOR CONTENTS (e.g. Kay & Allan 2015: 165). Here what is fragrant is not the tub itself, but the water it contains, which has been impregnated with a cleansing agent of some type. (57) Sherry, always an early riser, had already breakfasted in her chamber and was, Adam knew, even now sunk deep in a hot, fragrant tub. Washing herself clean of him, arming herself for another day of battle. (COHA, 2000, FIC, ComeNearMe) Nouns referring to natural geographical terms, expressions related to the weather, and atmospheric conditions were coded as EARTH,ATMOSPHERE,AND WEATHER. A great majority of the nouns belonging to this category are grouped into the rather general class THE EARTH in the HTOED, particularly within the subclasses LAND,WATER, and WEATHER AND THE ATMOSPHERE. In USAS nouns of these types are included in the category W:THE WORLD AND OUR ENVIRONMENT, particularly in W3: GEOGRAPHICAL TERMS and W4: WEATHER. Thus, the decision of maintaining such nouns together in one individual category is well justified, as they in fact belong to the same general semantic domain in both the HTOED and USAS. An example of this category is given in (58): (58) Turning from the mountain scenes we have described, let us back once more to Constantinople, and direct our footsteps up the fragrant valley where the Barbyses threads its meandering course. (COHA, 1851, FIC, CircassianSlave) In this category we also find a wide range of metonymical or metaphorical extensions of meaning in which temporal expressions and states are used to refer to the surrounding environment and atmospheric conditions of an outdoor location, and sometimes even to the DANIELA PETTERSSON-TRABA 122 characteristic actions and processes of the natural entities present in those locations. Consider (59) and (60): (59) And how we used to drive the sleepy old cows through the wood and meadow land, wading knee deep in the clover that seemed to be holding up its red lips to be kissed, while the wild flowers like gossamer chalices filled with dew, made the rosy morning fragrant.(COHA, 1873, FIC, WhiteSlave) (60) Accordingly, she had attired herself in a becoming negligee, and had spent the fore part of the night somewhat restlessly, occasionally emerging on the veranda and gazing down into the perfumed gloom of the garden.(COHA, 1892, FIC, GoldenFleeceRomance) In (59), the noun morning does not denote ‘[…] the early part of the day, esp. from sunrise until noon or lunchtime’ (OED, s.v. morning n. A 1a), but instead refers to ‘[…] the early part of a day as characterized by the particular weather, conditions, sentiments, etc., prevailing or experienced during that time’ (OED, s.v. morning n. A 2). Therefore, morning is in this instance not just an abstract temporal expression, but it also refers to the concrete surroundings and atmospheric conditions of the wood and meadow land through which the narrator and other individuals drive the cows. In fact, it is evident that the smell stems from the wild flowers mentioned in the example. Consequently, instances of the type in (59) would correspond to the metonymical pattern TIME PERIOD FOR A CHARACTERISTIC ACTIVITY IN THAT PERIOD (Kövecses 2010: 258), in which the characteristic activity could be that of the flowers blooming again in the morning after closing at night. Note that TIME is a particularly complex concept. In fact, it is one of the most common target domains in metaphor and metonymy and thus appears in a wide range of different mappings (Kövecses 2010: 26), which makes it particularly difficult to pinpoint the exact pattern of such examples. Similarly, in (60), gloom does not just refer to the darkness of the garden, but also to the garden itself, which is most probably filled by flowers and plants. The conceptualization of the state as a space is reinforced by the preposition into (into the perfumed gloom of the garden), which expresses motion or direction. As in (59), difficulties arise when trying to characterize the mapping, given that SPACE is also a common target domain. However, one possibility could be that of the metaphorical pattern STATES ARE 4. The concept PLEASANT SMELLING: Data selection and annotation 123 LOCATIONS, where a more abstract aspect is understood in terms of physical location (Kövecses 2010: 163). Be that as it may, the importance here lies not in identifying the specific figurative pattern at stake, which is beyond the purposes of the present dissertation, but in the fact that when the adjectives modify such nouns, these nouns are often not used literally but figuratively. Within the category FOOD AND DRINK we find nouns corresponding to the semantic class with the same label in the HTOED and to the F1: FOOD and F2: DRINKS USAS categories, which include edible and drinkable products, as well as nouns referring to different meals of the day (e.g. breakfast,lunch,anddinner). Example (61) was, therefore, grouped into this category: (61) Every morning believers came to offer Han-shan a bowl of fragrant soup that steamed his face as if to make him sweat. (COHA, 1970, NF, BuddhistLeader) Moreover, the category FOOD AND DRINK also comprises several metonymic examples corresponding to the conceptual metonymy CONTAINER FOR CONTENTS (e.g. Kay & Allan 2015: 165), already discussed in relation to CLEANING,in which the near-synonymous adjectives occur with nouns denoting containers (e.g. bowl,cup,andglass) but modify the contents within them, as in (62): (62) They enjoy afternoon tea precisely as other women enjoy it, and over a fragrant cup they relax, talk and laugh and tell stories and have what they need more than anything else in their lives, a really good time. (COHA, 1905, NF, RadiantMotherhood) A relatively general category is that of OBJECT, as it contains a mixture of concrete nouns referring to material things which can be seen and touched, including furniture (e.g. bed,couch, and drawer), paper documents (e.g. letter,envelope, and book), and other objects (e.g. box, candle,andlamp). Many of these nouns belong to the semantic classes MATTER in the HTOED and O2: OBJECTS GENERALLY,Q1.2: PAPER DOCUMENTS AND WRITING, and Q4: THE MEDIA in USAS, among others. The decision was made to group these more specific classes into a more general one due to data sparseness, as retaining the more specific categories would yield very low figures, which would make it difficult to obtain significant results. An example of OBJECT is shown in (63): DANIELA PETTERSSON-TRABA 124 (63) We who produce the world where Heirston finds John Robshaw pillows along with Calypso’s own Pom Pom cotton throws and scented candles ($38 each) in eleven fragrances. (COHA, 2009, MAG, COCA) Categories L3: PLANTS in USAS and PLANTS in the HTOED were here classified under the label PLANTS AND FLOWERS. This category is illustrated in (64): (64) Then the angel went down to the earth, and he came to a beautiful rose-bush upon which bloomed a rose lovelier and more fragrant than any of her kind. (COHA, 1896, FIC, SecondBookTales) Again, we find several metonymical instances in this semantic category. In this case, the metonymical extension follows the pattern CHARACTERISTIC FOR CHARACTERIZED ENTITY (Waag 1908; Nyrop 1913; Paul 1920; and Esnault 1925; cf. also Geeraerts 2010: 32), in which the near-synonymous adjectives modify a noun that denotes an attribute of a plant or a flower. Such an example is illustrated in (65): (65) The woman sprang back from the flowers as though a poisonous serpent, hidden in their fragrant beauty, had struck her. (COHA, 1912, FIC, TheirYesterdays) In this example, even though fragrant modifies the noun beauty, the pleasant and sweet smell applies to the flowers, since beauty here is clearly a characteristic of the flowers. The ninth category, SENSATION, includes nouns referring to any of the five physical senses, that is, smell (e.g. aroma,odor,andscent), taste (e.g. flavor and taste), hearing (e.g. music, sound,andvoice), sight (e.g. look), and touch (e.g. kiss). Evidently, the great majority of the nouns in my data correspond to the sense of smell, since the adjectives under analysis belong to this particular semantic domain, as in example (66), where fragrant collocates with smell: (66) He spread the faded army blanket, which was old but clean, on the straw of the second loft. The smell was fragrant and sweet. (COHA, 1979, FIC, SilverGhost) 4. The concept PLEASANT SMELLING: Data selection and annotation 125 Nevertheless, several instances of nouns belonging to one of the other four senses were also found in the data and were therefore included in this category, which corresponds to the broad semantic classes PHYSICAL SENSATION in the HTOED and X3: SENSORY in USAS. In general, when the adjectives at issue here modify any of the senses apart from smell, they are used in the figurative sense ‘pleasant or agreeable’, as in example (67), where perfumed modifies the noun murmur, which refers to ‘[a] word or sentence spoken softly or indistinctly; faint or barely audible speech […]’ (OED, s.v. murmur n. 4): (67) He picked up one of the small books. “This is The Book of Everything.” There was a rustle among the girls and a perfumed murmur of “Everything?” “Everything.” he said firmly. (COHA, 1958, FIC, Maggie-Now) Examples such as (67) are instances of synesthetic metaphors, which describe one sensory domain on the basis of another one (e.g. Geeraerts 2010: 35). Here, an adjective from the olfactory domain is used to describe a noun from the auditory domain, and thus perfumed is used figuratively to mean simply ‘sweet or agreeable’. While the category EARTH,ATMOSPHERE,AND WEATHER included, as seen above, natural geographical terms, among other types of nouns, the category SPACE contains nouns referring to enclosed or indoor locations, such as room,boudoir, and kitchen, as well as those denoting human —as opposed to natural— geographical terms, for instance, avenue,road, and city. The principal reason for maintaining these two categories separate is that whereas the category EARTH,ATMOSPHERE,AND WEATHER contains instances of the adjectives solely in the natural sense, SPACE comprises also occurrences of the adjectives in the artificial sense, as well as indeterminate cases. Additionally, while the great majority of nouns in the former group are classified as THE EARTH in the HTOED, especially within the subclasses LAND,WATER, and WEATHER AND THE ATMOSPHERE, and in W3: GEOGRAPHICAL TERMS and W4: WEATHER in USAS, this is not the case of nouns in the category SPACE. Instead, these nouns appear mainly in the general category SOCIETY, especially within the more specific classes INHABITING AND DWELLING and TRAVEL, in the HTOED, and in H:ARCHITECTURE,BUILDINGS,HOUSES AND THE HOME and in M3: MOVEMENT/TRANSPORTATION:LAND and M7: PLACES in USAS. An example of the category SPACE is given in (68), where scented modifies the noun room: DANIELA PETTERSSON-TRABA 126 (68) As he and Jock left the warm, scented room behind them, and faced the white, still cold of an apparently dead St. Ange, the boy turned a drawn face upon Jock, and cried tremblingly, “Say, you”. (COHA, 1911, FIC, JoyceNorthWoods) As in the case of EARTH,ATMOSPHERE,AND WEATHER, the category SPACE also contains instances of metaphors following the pattern STATES ARE LOCATIONS, but in which the space is an indoor location, as in example (69): (69) Look, there are orchids and camellias and oleanders.’ He was silent so long that she thought he hadn’t heard her. She looked up inquiringly and saw that he was staring not into the perfumed dampness of the hall but at her. (COHA, 1944, FIC, Dragonwyck) Similarly to example (60) above, dampness does not just refer here to the condition or state of the hall, but also to the hall itself, that is, an indoor space, where there are several flowers (i.e. orchids, camellias, and oleanders). The category SUBSTANCE AND MATERIAL includes nouns referring to liquid, solid, or gaseous substances (e.g. ash,ether,liquid,andsteam), and to materials from which objects are made (e.g. amber and wood). Moreover, this category also contains nouns referring to drugs, especially tobacco (e.g. cigar,cigarette,andtobacco). These drug-related nouns, very few in number, were included here as they are in fact substances of a kind. In USAS most of the nouns belonging to SUBSTANCE AND MATERIAL are grouped into the semantic class O1: SUBSTANCES AND MATERIALS GENERALLY, whereas they mainly correspond with the general category MATTER in the HTOED. Example (70) illustrates this semantic category: (70) Most wood gives off a pleasant aroma when it’s cut, but wood dust is more than just fragrant - it’s also hazardous to your health. (COHA, 2003, MAG, MotherEarth) Finally, the last category, TEXTILE AND CLOTHING, comprises nouns referring either to textile material (e.g. linen,silk,andwool) or to different types of clothing (e.g. clothes,dress, 4. The concept PLEASANT SMELLING: Data selection and annotation 127 garment,andglove). These nouns all belong to the categories TEXTILES AND CLOTHING in the HTOED and B5: CLOTHES AND PERSONAL BELONGINGS, in the case of clothing, and O1.1: SUBSTANCES AND MATERIALS GENERALLY:SOLID, in the case of textile material, in USAS. An example is provided in (71), where perfumed collocates with the noun gloves: (71) […] and lay piled in heaps beneath the Sand Hill fort-many youthful gallants from Spain and Italy among them, noble volunteers recognized by their perfumed gloves and golden chains. (COHA, 1868, NF, HistoryUnitedNetherlands) The twelve semantic categories distinguished and illustrative examples of nouns belonging to each of them are provided in Table 7. Table 7. Classification of nouns into semantic categories SEMANTIC CATEGORY EXAMPLES OF NOUNS ABSTRACT (ABS)charm, darkness, fear, knowledge, memory, personalit y BODY AND PEOPLE (B&P)arm, bodice, cheek, g irl, hair, lock, woman, wrist CLEANING (CL)bath, deodorant, disinfectant, shampoo, soap, tub, water COSMETICS (COS)colo g ne, cosmetics, cream, g loss, oil EARTH,ATMOSPHERE,AND WEATHER (EAW)air, breeze, darkness, gloom, hill, mist, morning, rain, summer, valle y , water FOOD AND DRINK (F&D)apple, beverage, bowl, chicken, coffee, cup, glass, rice OBJECT (OBJ)book, box, candle, couch, lamp, letter, notepaper, stationar y PLANTS AND FLOWERS (P&F)beauty, bloom, bud, flower, leaf, loveliness, pine, rose, shrub, tree SENSATION (SEN)aroma, flavor, kiss, look, music, odor, scent, smell, taste, tone SPACE (SPA)avenue, bakery, boudoir, chamber, city, dampness, house, room SUBSTANCE AND MATERIAL (S&M)amber, dust, fume, liquid, oil, smoke, steam, vapor, water, wood TEXTILE AND CLOTHING (T&C)cambric, cloth, dress, g love, linen, pillow Since the classification was made on the basis of the referents of the nouns, which may vary from instance to instance, some of the nouns, especially those which are highly polysemic, such as breath,darkness,lip,oil,andwater, were classified into more than one semantic category (cf. Table 7). First, the adjectives at issue here can function as modifiers of several senses of DANIELA PETTERSSON-TRABA 230 In fact, it could even be argued that the combination scented candle is in the early stages of a lexicalization process, understood as the change whereby in certain linguistic contexts, speakers use a syntactic construction or word formation as a new contentful form with formal and semantic properties that are not completely derivable or predictable from the constituents of the construction or word formation pattern (Brinton & Traugott 2005: 96). This is so because, while a candle is employed primarily as a source of artificial light (OED, s.v. candle n. 1a), the main use of a scented candle nowadays is that of refreshing or perfuming the air of a place, much like air fresheners. This can be deduced from examples such as (117), where scented candle is referred to as a type of air freshener delivery method: (117) Many other air freshener delivery methods have become popular since, including scented candles, reed diffusers, potpourri, and heat release products. (CD, s.v. candle n. collocations) The remaining two prominent collocates of scented are gale and wind, two related words, as the former is a hyponym of the latter, which belong to the semantic category EARTH, ATMOSPHERE,AND WEATHER. 6.3 DISCUSSION The analyses conducted in this chapter proved the adequacy of the set of variables selected for examination since they all emerged as significant predictors of the choice between the nearsynonymous adjectives fragrant,perfumed, and scented. This indicates that variables of a varied nature should be considered in order to reach a better understanding of the variation between lexical near-synonyms. However, not all predictors play an equally important role according to the statistical methods employed here. In Section 6.2.1, both the multinomial regression models and the random forest analysis pointed to the conclusion that the intralinguistic semantic variables were stronger determinants for the particular synonym set under analysis, especially the two related predictors Semantic category and Sense, but also Countability and Concreteness. Similarly, the mixed-effects analyses computed in Section 6.2.2 indicated that the individual noun collocates considerably improved the performance of the 6. In-depth onomasiological analysis of the synonym set: A multivariate approach 231 regression models, as idiosyncratic collocational patterns explained a great amount of the variation existing between the near-synonyms. This is in line with previous research (cf. Section 3.2.2), which demonstrates that zooming in on the nouns that adjectives modify is a particularly effective approach to uncover their internal semantic structure (Geeraerts 1986: 282–284). On the other hand, the morphosyntactic variables Syntactic function and Degree, although significant, played a relatively minor role, particularly the latter (cf. Figure 16). This finding highlights the importance of considering the POS of the near-synonyms object of study when selecting the variables to be included in the analysis: while morphosyntactic variables are indeed crucial for near-synonymous verbs (e.g. Divjak 2010: 183–193), their effect size is rather negligible in the case of adjectives, at least in the set examined here. This reminds us of Hank’s (1996: 92; 96) and Liu’s (2010: 61) arguments that, despite the similar goals of usage-based corpus research on near-synonymy, the micro-procedures that need to be employed in each study vary, since the ways in and the dimensions on which near-synonyms differ also vary from one set to another. Additionally, one of the two extralinguistic variables, to wit, Period, also exerted a great influence on the distribution of the near-synonyms at stake, being ranked third in order of importance according to the random forest analysis. The fact that Period is a more powerful predictor than most of the other variables, both semantic and morphosyntactic ones, suggests that the diachronic dimension of lexical variation cannot be disregarded. As shown in Section 3.3, this issue has not received the attention it clearly deserves, as only a handful of diachronic distributional corpus-based studies on lexical near-synonyms exist to date, especially if compared to the relatively large amount of historical research on other linguistic phenomena, including polysemy and constructional synonymy. Concerning the specific effects of the predictors, most of the diachronic onomasiological tendencies of Sense and Semantic category uncovered in Section 5.4 are confirmed by the regression analyses, as many of these trends reach statistical significance. Once again, fragrant emerges as the default choice in the natural sense and all five prototypically ‘natural’ semantic categories (i.e. EARTH,ATMOSPHERE,AND WEATHER,FOOD AND DRINK,PLANTS AND FLOWERS, SENSATION, and SPACE). We witness, however, a slight decrease in the probability of this adjective in most of these contexts, mainly in favor of scented, which increases significantly over time, although it does not threaten the prominence of fragrant with these types of nouns. On the other hand, some substantial changes do occur in the artificial sense and with most of the related semantic categories, to wit, CLEANING,COSMETICS,OBJECT, and SUBSTANCE AND DANIELA PETTERSSON-TRABA 232 MATERIAL. In fact, we observe a reorganization of the near-synonyms in many of these contexts of use. (i) The probability of fragrant decreases somewhat in the artificial sense, and that of perfumed does so more drastically. Scented, in turn, gains ground at the expense of the other two adjectives, particularly from P2 to P3, and becomes the default choice in this sense in P4. (ii) With respect to specific semantic categories, the probability of scented increases particularly in CLEANING,COSMETICS, and OBJECT, but also somewhat in SUBSTANCE AND MATERIAL. In fact, in CLEANING and in OBJECT scented becomes the default choice, ousting perfumed in the former and both fragrant and perfumed in the latter. In COSMETICS and SUBSTANCE AND MATERIAL,scented converges in probabilities with perfumed and fragrant, respectively, although it does not surpass them. Similarly, the internal structure of the set is also reorganized in the figurative sense and in the category ABSTRACT, where fragrant loses its dominance over time; by P4, perfumed and fragrant are almost equally likely. Finally, in the value ‘indeterminate’, fragrant also loses ground in favor of scented, while perfumed remains stable. This change seems to be restricted to BODY AND PEOPLE, where scented also displays a substantial upward tendency. As in Chapter 5, the results of this chapter suggest the existence of a probably still ongoing process of substitution within the synonym set object of study, whereby scented gains ground at the expense of fragrant and perfumed. Figure 27 depicts this process in a more visual manner, where both the overall frequencies of each adjective, as well as those within the artificial, indeterminate, and natural senses, are displayed.105 Here, the vertical axis represents the percentages of use of the adjectives in each sense, whereas the horizontal axis shows the four historical periods. Within the plot, a different pattern is used to signal the frequencies of each adjective: squares are used to indicate those of fragrant and black diamonds those of scented, while the absence of a pattern in the middle area of the graph corresponds to perfumed. Additionally, a different color signals the frequencies of each sense: purple stands for the artificial sense, white for indeterminate uses, and green for the natural sense. 105 The figurative sense is not included in this graph as the frequencies of the three adjectives in this sense are almost negligible. 6. In-depth onomasiological analysis of the synonym set: A multivariate approach 233 Figure 27. The use of fragrant, perfumed, and scented in the artificial, indeterminate, and natural senses over time Figure 27 combines the patterns and colors to portray the percentages of use of the three adjectives in each sense over time. By focusing exclusively on the patterns, the ongoing process of replacement can clearly be observed. Overall, fragrant —located at the bottom of the figure— decreases considerably in frequency throughout the period 1810–2009, while the opposite tendency is true for scented —located at the top of the figure. However, while fragrant becomes less and less commonly selected over time in favor of scented, it still remains the dominant choice overall in P4. Perfumed, in turn, does neither increase nor decrease substantially. In P4, the situation seems to be one of a greater degree of competition between the near-synonyms, with all three adjectives being more equally distributed than in P1. Nevertheless, this greater degree of competition between the variants does not appear to result in an increasingly higher degree of differentiation between the synonyms as time passes. In fact, the contrary tendency seems to hold in this particular case, namely attraction. This can be observed by focusing on the colors in Figure 27, that is, by zooming in on the frequencies of each adjective in the different senses. By P4, all three adjectives are much more frequently used in the artificial sense 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% 100% P1 P2 P3 P4 Fragrant Artificial Fragrant Indeterminate Fragrant Natural Perfumed Artificial Perfumed Indeterminate Perfumed Natural Scented Artificial Scented Indeterminate Scented Natural DANIELA PETTERSSON-TRABA 234 —the purple areas— than in earlier periods. On the contrary, none of the adjectives exhibit a substantial upward tendency in the natural sense —the green areas—: fragrant and perfumed both clearly decrease in this sense over time, and scented shows only a minor increase. Similarly to the development of the adjectives in the artificial sense, all of them also increase in the value ‘indeterminate’ —the white areas— from P1 to P4. Interestingly, the diachronic development of the adjectives in so-called indeterminate contexts is very similar to their evolution in the artificial sense. This could be taken as an indication that a relatively large amount of such ambiguous cases is, in fact, ‘artificial’, although this cannot be discerned in the actual instances in COHA, not even if their extended contexts are analyzed.Moreover, more than half of the examples (56.96%) which are classified as indeterminate belong to the semantic category BODY AND PEOPLE. The reader may recall that a considerable increase over time of this particular semantic category emerged from the analysis (cf. Section 5.2.2), and that this upward tendency may have been the result of the growing availability and use of cosmetics and other personal care products applied to the body (cf. Section 5.2.3). Therefore, it may well be the case that the general rise of the indeterminate sense is, at least in part, due to the increase of the category BODY AND PEOPLE and that of potential artificial uses of the adjectives which had to be categorized as ‘indeterminate’ due to a lack of sufficient evidence. The end result of the frequency fluctuations described in the previous paragraph is a more even distribution of the adjectives in P4, both generally and across the different senses. In general terms, the overall frequencies of the adjectives are more similar in P4, although fragrant remains the dominant variant of the set. Additionally, their individual distribution across senses is also more similar at the end of the time span analyzed, with all three adjectives being more commonly used than before in both the artificial sense and in indeterminate cases, and less so in the natural sense, except in the case of scented, which does not recede from natural contexts of use. Therefore, the findings point to the conclusion reached in Chapter 5 that both processes of substitution and attraction are at play in the history of this synonym set. De Smet et al. (2018: 217) postulate that attraction results from analogical change, whereby functionally similar words or structures parallel each other’s behavior through an interchange of characteristics, an explanation that would account, for instance, for near-synonymous expressions developing the same metaphorical mappings. However, a more likely explanation in this particular case, as argued in Section 5.2.3, is that the tendency of all three adjectives to denote more and more artificial aromas derives from extralinguistic forces. In particular, the underlying motivations 6. In-depth onomasiological analysis of the synonym set: A multivariate approach 235 for change may be related here to the socio-cultural and technological advances of American society over the time span 1810–2009, whereby these specific types of aromas have become increasingly prominent in the day-to-day life of the American population. Turning now to the significant predictors which do not interact with Period, namely the semantic factors Concreteness rating and Countability and the morphosyntatic variables Degree and Syntactic function, fragrant is the default choice in all levels of the four predictors throughout the entire time span analyzed. This is not surprising, given the much higher overall frequency of fragrant in the corpus as compared to the other two adjectives. However, there are some contexts of use in which both perfumed and scented are more likely than expected, whereas the reverse is often true for fragrant: (i) Perfumed is more likely with highly abstract nouns, i.e. those with a low concreteness rating, than with more concrete nouns, while the opposite tendency holds for scented. (ii) As concerns Countability, fragrant is dominant in all three levels (count, non-count, and other), but less so with non-count nouns, where scented is more probable than expected, and with pronouns and proper nouns (‘other’), where perfumed is marginally more likely than expected. (iii) In the case of the variable Degree, fragrant is again salient in all levels, but much more so in the comparative and superlative degrees, because perfumed and scented are hardly ever used in either comparative or superlative form in COHA. (iv) Finally, as regards Syntactic function, fragrant outperforms the other two adjectives in all functions, especially as a predicative complement. However, in the postpositive use and in other minor functions its probability is not as pronounced: in the former, scented exhibits a higher probability than expected, whereas the same is true for perfumed when serving as a predicative adjunct, fused-modifier head, or predeterminer. These results entail that the overall prominence of fragrant is not threatened in any of these morphosyntactic or semantic contexts, and that this situation is maintained diachronically. However, some interesting differences are identified between perfumed and scented, especially DANIELA PETTERSSON-TRABA 236 as regards Concreteness rating, Countability, and Syntactic function, which suggests that a division of labor may exist between these two adjectives in some of these contexts. Some stylistic differences between the three near-synonyms also emerged from the regression analyses. Crucially, some diachronic changes can be said to exist as regards the probability of the adjectives in some text-types in COHA. An overall increase of scented is attested in all three text-types, but it is in non-fictional texts where this upward tendency is most pronounced. In fact, scented, the least probable adjective in non-fiction in P1, becomes the most likely choice in P4. In fiction and periodicals scented also increases and outperforms perfumed in P3 and P4, but it does not reach the probability of fragrant in any of these two text-types. Therefore, the overall increase of scented in frequency uncovered in Section 5.1 seems to occur across all text-types, though the increase is sharper in non-fictional texts. Thus, as in previous studies on the distribution of near-synonyms, stylistic differences also play a role on the choice between the three adjectives examined here (e.g. Biber, Conrad & Reppen 1998: Section 2.6; Kjellmer 2003; Liu & Espino 2012). However, it must be noted here that the categorization of text-types used in the present dissertation (fiction, non-fiction, and periodicals) is relatively coarse-grained and could be further refined in future research to pinpoint exactly in which particular subgenres each adjective is preferred (e.g. poetry, novels, movie and play scripts, editorials, advertisements, academic writing from different sciences). Lastly, the goodness-of-fit indexes of the multinomial regression models provided in Section 6.2.1 signal that there is still room for improvement. As previously explained, the accuracy and the R2values of both Models A and B indicate that the set of predictors included in the analyses and their interactions are insufficient to explain all the variation in the data. Three possible reasons why these models are not optimal are postulated here. First, although a rather large number of predictors of different types has been examined, other variables may also play a role in the choice between the near-synonyms. Two factors which have been shown to influence linguistic variation in different dimensions of language are the opposing forces of priming (Gries 2005; Szmrecsanyi 2005; 2006) and the so-called horror aequi principle (e.g. Rohdenburg 2003; Vosberg 2003). The former consists in the (explicit) repetition or reiteration of linguistic elements precisely due to the fact that they have been employed in a recently produced utterance. The latter, in turn, implies the avoidance of the repeated use of “formally (near-) identical and (near-) adjacent […]” linguistic elements (Rohdenburg 2003: 236). For instance, in example (118) scented, as opposed to fragrant or perfumed, may have been chosen 6. In-depth onomasiological analysis of the synonym set: A multivariate approach 237 as a modifier of the noun fruit precisely because it has already been used earlier in the same sentence (i.e. scented bloom), thus potentially constituting a case of priming. On the contrary, in (119) fragrant may have been selected as a modifier of the COSMETICS noun oils because scented, which is a more likely choice with this particular semantic category, has already been employed in the previous noun phrase (i.e. scented soap), hence a possible instance of horror aequi. Similarly, the avoidance of perfumed occurring in the neighboring contexts of the noun perfume or of the verb perfume may illustrate cases of the horror aequi principle.106 (118) About me were miles on miles of apple orchards, and in our own I worked and played from the time of pink and scented bloom to that of scented and yellow fruit. (COHA, 1930, MAG, Atlantic) (119) “Everybody is waiting for it to happen down here, hoping everything will jell,” said Sarah Laight, an ex-pat Brit who sells scented soaps, fragrant oils and summer-oflove dresses in her Life boutique. (COHA, 1994, NEWS, SanFran) Moreover, the non-optimal values of the multinomial models may result from interactions between variables which were not tested for in the present analyses. The reader may recall that only the interactions between the extralinguistic variable Period and the other predictors were included here, as the main aim of this dissertation is to examine the diachronic development of the near-synonymous adjectives. However, interesting findings may be revealed if a larger number of predictors, for instance, Text-type and Sense or Semantic category, are cross-cut so as to see whether functional differences exist between the near-synonyms in different genres. Nevertheless, the inclusion of additional variables and variable interactions would require a larger dataset than the one used in the present dissertation. The second reason why the models with only fixed effects are not optimal could be that some of the variation in the data may be relatively random. After all, we are here dealing with the phenomenon of lexical synonymy and semantically related terms need to exhibit a sufficiently high degree of interchangeability without a change in meaning in order to be considered near-synonyms. If this was not the case, language users would not be able to resort to near-synonyms for the purposes of avoidance of repetition and coherence in text production, among others (Murphy 2013: 289–290). 106 The same holds for the combinations scented scent or fragrant fragrance. DANIELA PETTERSSON-TRABA 238 Finally, and more importantly, the mixed-effects regression analyses presented in Section 6.2.2 highlighted the great importance of taking into consideration the idiosyncratic behavior of the near-synonyms. By including the specific noun collocates of the adjectives in the model, a substantial improvement in model fit was achieved. The resulting models with both fixed and random effects (cf. Section 6.2.2) explained almost double the amount of variation than a model with only fixed effects (cf. Section 6.2.1). This means that, in this particular case, what really distinguishes between the adjectives are the specific nouns that they modify. Thus, Firth’s (1957a: 11) famous quote “you shall know a word by the company it keeps” (cf. Section 2.3.3), by means of which he emphasizes the importance of collocational patterns in lexical semantics, seems to be an adequate description of the synonym set at issue. However, as shown by the models with fixed effects only presented in Section 6.2.1, collocations do not tell the whole story, since the fixed effects are also significant determinants of the variation between fragrant, perfumed, and scented. 6.4 SUMMARY This chapter has focused on the onomasiological structure of the synonym set under analysis, thus emphasizing the competition over time between the three adjectives fragrant,perfumed, and scented in different contexts of use. A wider range of variables than in Chapter 5 was examined by means of more powerful statistical tools, which enabled us to confirm several of the findings obtained previously and also to reach additional conclusions. In particular, it was shown that variables of a semantic nature, specifically those pertaining to the noun collocates of the adjectives, had a much stronger explanatory power than morphosyntactic and stylistic ones, although the latter two variable types also contributed significantly to the models. Moreover, the effect of the extralinguistic variable Period emerged as highly significant in this particular synonym set, thus highlighting the importance of adding a diachronic dimension to the study of near-synonyms. However, models containing only fixed effects proved to be insufficient to explain the variation in the data and had to be complemented by random effects capable of capturing the idiosyncratic collocational behavior of the lexical items object of study. Given the great importance of the individual noun collocates of the near-synonymous adjectives, the next chapter zooms in on their idiosyncratic collocational behavior throughout the period 1810–2009 by using different methodological approaches that are specifically geared towards analyses of this kind. 7 IDIOSYNCRATIC COLLOCATIONAL PREFERENCES OF THE NEAR-SYNONYMS As seen in Section 2.3.2, The notion of collocation constitutes one of the central tenets of the distributional approach from its earliest stages, being emphasized already in several of Firth’s most important works (e.g. 1935; 1957a). With the advent of corpora, the notion became even more essential in lexico-semantic analysis and, therefore, research on collocational behavior flourished in this field of study. Similarly, advancements in the methods and statistical techniques used both to measure the degree of association between nodes and their collocates and to represent such associations led to an increasing interest in this phenomenon. In particular, the notion of collocation has figured prominently in research on near-synonymy as one of the types of contextual patterns used to quantify the degree of (dis)similarity between potentially near-synonymous expressions. In fact, many scholars argue that the more contextual ground two lexemes share, including specific collocates, the more similar they are (e.g. Rubenstein & Goodenough 1965; Miller & Charles 1991; Gries 2001; 2003; Gries & Otani 2010; Liu 2010). For instance, Gries (2001; 2003) examines the degree of collocate overlap between pairs of near-synonymous -ic/-ical adjectives by accounting for the amount of significant collocates they share, as well as the amount of significant collocates one near-synonym exhibits, but not the other. Due to the broad application of the notion of collocation and the boost it received from the distributional corpus-based approach, it is not surprising that this term has been interpreted in various ways. For instance, Gries (2013: 138–139), in his overview of collocational research over the last 50 years, identifies a series of parameters on which the notion has varied. Among these, he includes the nature of the linguistic items examined (e.g. POS), the frequency threshold for a co-occurring item to be considered a collocate, the distance between the items making up the collocation (i.e. context window), and the relation between the node and the collocate, that is, whether mere adjacency between the items in question suffices or whether the DANIELA PETTERSSON-TRABA 246 Table 28. Distance matrix for the data Fragrant Perfumed Scented P1 P2 P3 P4 P1 P2 P3 P4 P1 P2 P3 P4 Fragrant P1 0 0.24 0.41 0.47 0.17 0.53 0.66 0.63 0.29 0.42 0.47 0.53 P2 0.24 0 0.29 0.58 0.57 0.43 0.66 0.64 0.52 0.52 0.61 0.64 P3 0.41 0.29 0 0.21 0.68 0.55 0.61 0.60 0.57 0.58 0.30 0.51 P4 0.47 0.58 0.21 0 0.75 0.65 0.57 0.44 0.80 0.60 0.58 0.55 Perfumed P1 0.17 0.57 0.68 0.75 0 0.13 0.50 0.49 0.05 0.36 0.51 0.77 P2 0.53 0.43 0.55 0.65 0.13 0 0.24 0.40 0.17 0.16 0.34 0.59 P3 0.66 0.66 0.61 0.57 0.50 0.24 0 0.15 0.60 0.70 0.17 0.67 P4 0.63 0.64 0.60 0.44 0.49 0.40 0.15 0 0.65 0.46 0.05 0.27 Scented P1 0.29 0.52 0.57 0.80 0.05 0.17 0.60 0.65 0 0.22 0.39 0.78 P2 0.42 0.52 0.58 0.60 0.35 0.16 0.70 0.46 0.22 0 0.41 0.58 P3 0.47 0.61 0.30 0.58 0.51 0.34 0.17 0.05 0.39 0.41 0 0.34 P4 0.53 0.64 0.51 0.55 0.77 0.59 0.67 0.27 0.78 0.58 0.34 0 Cluster analysis is a group of methods that helps to identify groups or clusters of objects on the basis of their (dis)similarities (Levshina 2015: 301). The objective of such methods is therefore to create groups comprising objects that are maximally similar to one another and maximally dissimilar to those objects in other groups, all this on the basis of the data at hand, in this case the distance matrix displayed in Table 28. Here, Hierarchical Agglomerative Clustering (HAC) analysis is used, which organizes the data into clusters and subclusters in a tree-like structure by first considering each individual node (here, ‘fragrant P1’, ‘fragrant P2’, etc.) as one separate cluster, then continually merging those two which are most similar until there is only one macro cluster containing all nodes (Steinbach, Tan & Kumar 2005: Chapter 8). MDS, in turn, is an approach that depicts differences between two or more objects as distances in a twoor three-dimensional plot (Levshina 2015: 336–337). The input for this test is again the distance matrix shown in Table 28 above and the output consists in a spatial representation of the nodes so that the more similar two nodes are, the closer they are located to each other in this spatial configuration. The goal of MDS is to shed light on the structure of the data in an easily understandable and visual manner (Kruskal & Wish 1978: 7; Wickelmaier 2003: 4). The two methods, HAC and MDS, are used here since they complement each other: HAC is useful to identify coherent and delimited groups in the data, while MDS can be said to be semantically more realistic in the sense that it does not imply a categorical either/or split 7. Idiosyncratic collocational preferences of the near-synonyms 247 between individual clusters, but instead displays a more precise and continuous cline of semantic (dis)similarity (Jansegers & Gries 2017: 14). The second type of analysis to which the data was submitted was collocational network analysis, which supplements the findings of SVS modeling. This is so because, as argued by Heylen et al. (2015: 154–155), “the technique [SVS] is too much of a black box to be suitable for in-depth lexical analysis […]”. In other words, SVS results in a series of distance values, but it averages over a great amount of individual collocates so that it is not always easy to determine exactly where the (dis)similarities in the data reside. The idea of visualizing the associations between nodes and their collocates by means of networks was first proposed by Philips (1985; 1989) and has since then been put into practice and further refined by numerous scholars (e.g. Baker 2005; 2017: 95–101; McEnery 2006a; 2006b; Brezina, McEnery & Wattam 2015). Put simply, collocational networks are built by automatically selecting the collocates of the node words, usually the most prominent ones according to some statistical measure of collocational strength (e.g. PMI, ǻ3, or log-likelihood), and then visualizing them in the form of a network. In this network, collocates are connected to their nodes by means of arrows, with the length of the arrow indicating the strength of the association between them: the shorter the arrow, the stronger the association. Collocational networks have been used for various purposes in the specialized literature, ranging from discourse analysis (e.g. McEnery 2006a; 2006b; Brezina, McEnery & Wattam 2015) to historical linguistics (Baker 2017: 95–101). In particular, Baker’s study on the competing prepositional forms on/upon and round/around described in Section 3.3 is of special relevance for the purposes of the present dissertation, as it constitutes the first application of this method to analyze the development of the collocational behavior of near-synonyms over time. Baker demonstrates that, through collocational networks, one can observe increases or decreases in the frequency of near-synonyms, as well as the competition between them: if one near-synonym draws away collocates from another, this most probably means that it is taking over some of its semantic space. Here, four separate collocational networks were built, one per period, hence enabling us to determine which collocates the near-synonymous adjectives under analysis share at different points in time, but also to identify potential variations in the internal semantic structure of the synonym set. The association measure employed was again PMI and, following the methodology in Baker (2017: 98–100) for low-frequency words, two thresholds were established for a collocate to be considered significant and thus be included in the DANIELA PETTERSSON-TRABA 248 networks: a minimum raw frequency of 5 with the node words and a PMI value of 3 or higher in each of the four periods.111 These thresholds drastically decreased the number of collocates of the near-synonyms in the dataset if compared to the SVS analysis, thus allowing us to conduct a more qualitative and careful examination of the collocational profiles of the adjectives at issue. At this point, it became evident that sweet-scented and sweet-smelling could not be included in the collocational network analysis, as almost none of their collocates emerged as prominent after applying these two thresholds: only the noun flower complied with these criteria as a collocate of sweet-scented in P1 and sweet-smelling in P2. For this reason, the two adjectives were excluded from the collocational networks and, for the sake of comparability, also from the SVS analysis.112 Establishing lower thresholds, that is a raw frequency lower than 5 and a PMI lower than 3, would not be a particularly effective solution, as such networks would include collocates which may not be especially prominent. Additionally, as mentioned earlier (cf. Section 2.3.2), a PMI of 3 or higher is often considered to indicate that two items co-occur significantly more often than expected by chance (e.g. Church & Hanks 1990; Church et al. 1991; Church et al. 1994; Liu 2010), so that lowering this threshold would result in nonsignificant collocates entering the analyses. On the contrary, higher thresholds would lead to even greater data sparseness, particularly in the case of perfumed and scented, which are, as seen in Section 5.1, not as frequent in COHA as fragrant (cf. Table 9). Therefore, the thresholds proposed by Baker (2017: 96) for words of a higher frequency, namely a minimum raw frequency of 10 and a PMI of 6 were here discarded. Table 29 shows the settings used to identify collocates, namely the PMI and frequency cut-off values, the context window, and some additional filters. The additional filter column makes it explicit that only noun collocates were considered, thus discarding other content (e.g. adjectives and verbs) and function words (e.g. prepositions and pronouns). 111 Brezina, McEnery & Wattam (2015: 159) argue that it is useful to apply a frequency threshold when selecting collocates on the basis of PMI scores, given that these scores have “a propensity to highlight unusual combinations […] that occur only once or twice in the corpus” and which are thus not particularly informative. 112 The most frequent noun collocates of sweet-scented and sweet-smelling in each period are provided in Tables C.2 and C.3 in Appendix C. 7. Idiosyncratic collocational preferences of the near-synonyms 249 Table 29. Settings for the construction of collocational networks PMI threshold Frequency threshold Context window Additional filters 35L5–R5 Function words and content words other than nouns removed Nowadays different tools are available in order to build collocational networks, including CONE (Gullick et al. 2010) and GraphColl (Brezina, McEnery & Wattam 2015). However, in this dissertation, the networks were created in Rby means of the igraph package (Csardi & Nepusz 2006). The reasons for this decision were twofold: (i) loading COHA into a program like GraphColl required a lot of computation time due to its large size and (ii) Rallows greater flexibility in the customization of the visual parameters of collocational networks. 7.2 RESULTS As mentioned in the introductory section to this chapter, the results of the two methods employed to examine the idiosyncratic collocational behavior of the near-synonyms fragrant, perfumed, and scented, to wit, SVS modeling and collocational network analysis, are presented separately in what follows. First, in Section 7.2.1 the findings of SVS are provided through the two different visualizations techniques selected here, i.e. HAC and MDS. Then, Section 7.2.2 presents the most prominent collocates of each adjective by means of collocational networks. 7.2.1 Semantic Vector Space modeling Figure 28 displays the distance between the near-synonymous adjectives across periods according to HAC in the form of a dendrogram: the higher two nodes are merged, the more dissimilar they are. The grey rectangles indicate the optimal number of clusters on the basis of the average silhouette width of this solution, an index which serves as an indicator of the internal coherence of the clusters. This index ranges from 0 to 1, with values higher than 0.2 signaling that clusters are well-formed (Levshina 2015: 311). DANIELA PETTERSSON-TRABA 250 Figure 28. Dendrogram of the collocational preferences of fragrant, perfumed, and scented throughout the period 1810–2009 A HAC solution with five clusters obtains the highest average silhouette width, namely 0.49. The five clusters are the following (from left to right in Figure 28): (i) Cluster 1: perfumed and scented in P1 and P2. (ii) Cluster 2: scented in P4. (iii) Cluster 3: perfumed in P3 and P4 and scented in P3. (iv) Cluster 4: fragrant in P3 and P4. (v) Cluster 5: fragrant in P1 and P2. The stability of the clusters obtained can be estimated by computing their Approximately Unbiased (AU) p-values, which determine how well supported they are by the data and thus also the chances of achieving the same clusters if a new data sample was to be analyzed. AU pvalues range from 0 to 1 and the closer they are to 1, the more stable the clusters (Levshina 2015: 315–317). The AU p-values for the five clusters in Figure 28 are shown in Table 30: 7. Idiosyncratic collocational preferences of the near-synonyms 251 Table 30. AU p-values for the five-cluster solution in Figure 28 Cluster AU p-value 1 (perfumed and scented in P1 and P2) 0.89 2(scented in P4) 0.90 3(perfumed in P3 and P4 and scented in P3) 0.93 4(fragrant in P3 and P4) 0.88 5(fragrant in P1 and P2) 0.78 Of the five clusters obtained, none reaches the traditional 0.95 significance threshold, but most of them are close to this value. The least stable cluster is that containing fragrant in P1 and P2 (i.e. cluster 5), which implies that if a new data sample was to be examined, these two nodes would possibly not be part of the same cluster. On the other hand, the cluster containing perfumed in P3 and P4 and that of scented in P3 (i.e. cluster 3) is almost statistically significant and therefore represents the most stable cluster. The remaining three clusters exhibit values that range from those obtained for cluster 5 to those for cluster 3, but they are closer to the latter. Figure 28 above shows that the collocational profiles of the three adjectives seem to pattern according to period and near-synonym. First, the collocational profiles of the near-synonyms in P1 and P2 are grouped together (i.e. clusters 1 and 5), as well as those in P3 and P4 (i.e. clusters 3 and 4). Second, the collocational profiles of fragrant (i.e. clusters 4 and 5) differ from those of perfumed and scented, which are part of the same clusters (i.e. clusters 1 and 3). The only exception to this pattern is the collocational behavior of scented in P4, which forms its own cluster (i.e. cluster 2) and is thus kept apart from the nodes ‘perfumed P3’, ‘perfumed P4’, and ‘scented P3’ (i.e. cluster 3). A semantically more realistic picture with no categorical either/or splits is provided by the MDS analysis of the data. Figure 29 plots a two-dimensional MDS solution of the collocational preferences of fragrant,perfumed, and scented across periods. This figure should be interpreted in the following way: each separate node in the graph represents the collocational behavior of one of the near-synonyms in a specific period, and the closer two nodes are located on the MDS map, the more similar their collocational preferences. The quality of the MDS solution can be examined by computing the so-called stress value (Levshina 2015: 341): the smaller the stress, the better the performance of the model. As a rule of thumb, values higher than 0.2 indicate that the model can be further improved. The stress of the MDS solution visualized in Figure 29 is 0.21, indicating an almost acceptable fit. A slightly better fit is obtained by a three-dimensional DANIELA PETTERSSON-TRABA 252 MDS solution (i.e. stress = 0.16). However, only a two-dimensional solution is discussed here, since it allows for a much clearer and easily interpretable visualization of (dis)similarities in the collocational behavior of the three adjectives (cf. Szmrecsanyi & Kortmann 2009: 1650; Szmrecsanyi 2017: 351 for a similar approach).113 Figure 29. Two-dimensional MDS map of the collocational preferences of fragrant, perfumed, and scented throughout the period 1810–2009 The vertical axis (Dimension 2) in Figure 29 organizes the nodes according to nearsynonym, with perfumed being located at the top of the figure and fragrant at the bottom. Scented, in turn, occupies the middle ground but, overall, it is much closer to perfumed than to fragrant, with the exception of P4. The horizontal axis (Dimension 1), on the contrary, organizes the nodes on a temporal continuum from left to right, with P1 and P2 exhibiting negative values and P3 and P4 exhibiting positive ones. Independent-samples t-tests on the MDS coordinates, that is, the values of the nodes on Dimensions 1 and 2, confirm the visual interpretation of Figure 29.114 This is so because fragrant differs significantly from both perfumed and scented with respect to their positions on the vertical axis (t= -6.8, df = 4.49, p < 0.01 and t= -4.98, df = 5.12, p< 0.01, respectively), but the difference between perfumed and scented is only marginally significant (t= 2.27, df = 5.72, p= 0.06). However, the 113 Although only a slightly better fit is achieved with a three-dimensional MDS, more variation is of course captured by this solution. For the three-dimensional MDS map, cf. Figure C.1 in Appendix C. 114 This test is used to compare the means of two groups of datapoints, here the mean MDS coordinates of the near-synonyms across periods (e.g. fragrant vs. perfumed), on the one hand, and the mean MDS coordinates of the periods across near-synonym (e.g. P1 vs. P2), on the other (Levshina 2015: 87). 7. Idiosyncratic collocational preferences of the near-synonyms 253 differences between the near-synonyms along the horizontal axis are not significant. In fact, on this axis, where the nodes are organized according to period, P1 diverges significantly from P3 and P4 (t= 4.27, df = 2.78, p< 0.05 and t= 5.81, df = 3.17, p< 0.01, respectively), P2 diverges significantly from P3 and P4 (t= 4.28, df = 3.89, p< 0.05 and t= 6.4, df = 3.98, p< 0.01, respectively), but the differences between P1 and P2, on the one hand, and between P3 and P4, on the other, are not significant (t= 1.50, df = 3.06, p= 0.23 and t= 2.74, df = 3.81, p= 0.055, respectively). In turn, the differences among the periods on the vertical axis do not reach statistical significance. This translates into the following: concerning their collocational profiles, whereas the near-synonyms seem to be relatively similar in P1 and P2, on the one hand, and in P3 and P4, on the other, there is a substantial change from P1/P2 to P3/P4, that is, at the turn of the 19th to the 20th century. This finding is in line with those obtained in the previous two chapters, where important changes were also discovered at this particular point in time. The two-dimensional MDS solution of the SVS analysis displayed in Figure 29 allows us to identify differences in collocational behavior between, on the one hand, fragrant and the other two adjectives and, on the other, all three near-synonyms in the 19th century versus the 20th and 21st centuries. However, as mentioned above, Heylen et al. (2015: 154–155) argue that SVS is “too much of a black box” to clarify the nature of the (dis)similarities uncovered. Therefore, in order to shed some light on such (dis)similarities, further analyses are required. Before turning to the more qualitative analysis carried out in Section 7.2.2 by means of collocational networks, in the remainder of the present section, quantitative techniques are applied to the results of the SVS to identify what the Dimensions (i.e. axes) in the MDS plot in Figure 29 stand for. To this end, the 2,682 types of noun collocates identified in Section 7.1 were classified into semantic categories by resorting to USAS (cf. Section 4.3).115 Following USAS semantic classification for each near-synonym in each period (e.g. ‘fragrant P1’, ‘fragrant P2’), a PMI score signaling their collocational strength with each semantic category and subcategory in USAS was calculated. This was done with the aim of determining whether correlations exist between the MDS coordinates of the nodes and their PMI scores with the 115 Given the data-driven nature of SVS analysis, the semantic classification of the noun collocates in this chapter was carried out automatically using USAS, instead of the HTOED. As mentioned in Section 4.3, despite being less precise, USAS allows researchers to automatically analyze strings of words according to their semantics and thus check the domains to which particular words belong DANIELA PETTERSSON-TRABA 254 semantic categories and sub-categories. The significance of correlations was assessed by means of Pearson’s product-moment correlation tests (Levshina 2015: 116–130). Besides a significance value, this test provides a correlation coefficient, represented by the letter r, which ranges from 0 to 1 and indicates the strength of the correlation: the higher the value, the stronger the correlation. Positive coefficients signal a positive correlation, which means that the values of the tested variables, here the MDS coordinates of the nodes and their PMI scores with the USAS semantic categories, jointly decrease and increase. In other words, if the values of one variable increase, so do the values of the other variable, and vice versa. On the contrary, negative coefficients signal a negative or inverse correlation, implying that the values go in opposite directions: if the values of one variable increase, the values of the other decrease, and vice versa. The eight correlations shown in Table 31 emerged as significant: Table 31. Significant correlations between the dimensions in Figure 29 and semantic (sub-)categories in USAS MDS dimension USAS semantic (sub-)category Direction of correlation Correlation coefficient (r) p-value Dimension 1 F4: FARMING AND HORTICULTURE negative -0.63 < 0.05 Dimension 1 W4: WEATHER negative -0.58 < 0.05 Dimension 2 B:THE BODY AND THE INDIVIDUAL positive 0.88 < 0.001 Dimension 2 L3: PLANTS negative -0.54 0.06 Dimension 2 M3: MOVEMENT/TRANS PORTATION (LAND) negative -0.53 0.07 Dimension 2 Q1: COMMUNICATION positive 0.77 < 0.01 Dimension 2 Q4: THE MEDIA positive 0.69 < 0.05 Dimension 2 S2: PEOPLE positive 0.80 < 0.01 Given that some of the semantic categories in Table 31 differ from those used in previous chapters, examples of nouns belonging to these categories that exhibit a significant correlation with any of the MDS dimensions are shown in Table 32: 7. Idiosyncratic collocational preferences of the near-synonyms 255 Table 32. Examples of types of noun collocates in the semantic (sub-)categories in USAS Semantic (sub-)category Examples of types of noun collocates B:THE BODY AND THE INDIVIDUAL B1: hand, hair, neck, shoulder, wrist B4: detergent, perfume, lotion, shampoo, soap B5: clothes, handkerchief, garment, glove, lace F4: FARMING AND HORTICULTURE crop, farm, farmhouse, field, mow,pasture, plantation, vineyard L3: PLANTS cactus, garden, magnolia, pollen, pumpkin, root, vine M3: MOVEMENT/TRANSPORTATION (LAND)avenue, path, pathway, road, wayside Q1: COMMUNICATION draft, leaflet, letter, message, note, notepaper, stationery Q4: THE MEDIA article, book, edition, newspaper, paper S2: PEOPLE boy, child, gentleman, madam, maiden, man, people, person, woman W4: WEATHER breeze, flood, fog, haze, mist, rain, snowfall, storm, wind First, regarding the behavior of the three near-synonymous adjectives across periods, that is, Dimension 1 (horizontal axis in Figure 29), two semantic categories exhibit significant negative correlations with the MDS coordinates, namely F4: FARMING AND HORTICULTURE and W4: WEATHER. This means that the higher the value of a node on Dimension 1, the lower its PMI score with these categories. Therefore, in P3 and P4 the adjectives occur with F4 and W4 nouns significantly less frequently than in P1 and P2. This change could again perhaps constitute a reflection of the transformation of American society, evolving from a preindustrial and mainly rural society in the 19th century to the world’s major industrial and commercial power in the 20th and 21st centuries (cf. Section 5.2.3). As a result of this process of modernization, the importance of farming and horticultural activities in the daily life of most American citizens has surely drastically declined over time, particularly after the First and Second Industrial Revolutions of the 19th century, which resulted in mass migration from the countryside to the cities. In a more indirect way, the importance of weather phenomena may also have decreased with time, as farming and horticultural activities rely greatly on the climate. Second, the MDS coordinates on Dimension 2 (vertical axis in Figure 29) are negatively correlated with the PMI scores of the near-synonyms and the semantic categories L3: PLANTS and M3: MOVEMENT/TRANSPORTATION (LAND). This implies that the adjectives collocate with nouns belonging to these two categories more often when their values on MDS Dimension 2 DANIELA PETTERSSON-TRABA 262 Figure 32. Collocational network of fragrant, perfumed, and scented in P2 (1860–1909) Both perfumed and scented increase in number of collocates with respect to P1, with 16 and 10 collocates, respectively, but they are still much less productive than fragrant.As in P1,fragrant again exhibits a clear natural orientation, with 54 ‘natural’ noun collocates (e.g. clover,foliage, herb,leaf, and vine). As in the previous period, there is a strong correspondence between the noun collocates of this adjective and the (sub-)categories L3 (23 types), F(8 types), and W4(4 types). In fact, we witness an increase in nouns belonging to L3 in this period, with collocates such as pine,lily, and clover emerging also as prominent collocates of fragrant. Additionally, its collocates now seem to be somewhat more varied, including also nouns belonging to other (sub-)categories, such as X3.5: SENSORY (SMELL) (e.g. scent and smell) and O1: SUBSTANCES AND MATERIALS GENERALLY (e.g. steam and smoke). Nouns belonging to the artificial sense, for instance, B4: CLEANING AND PERSONAL CARE and B5: CLOTHES AND PERSONAL BELONGINGS, on the contrary, do not often co-occur with fragrant. Similarly, scented seems to display a 7. Idiosyncratic collocational preferences of the near-synonyms 263 preference for entities denoting a natural smell, collocating with nouns such as bloom,blossom, breeze,andgrass, which mainly belong to (sub-)categories L3 and W4. Contrariwise, perfumed seems to be the most common adjective to modify artificial smells, as suggested by collocates such as hair,handkerchief,lace,pocket, and silk, that is, nouns in (sub-)categories B1: ANATOMY AND PHYSIOLOGY (1 collocate) and B5 (4 collocates), but also in O1: SUBSTANCES AND MATERIALS GENERALLY (3 collocates, namely oil,water, and smoke). The five most prominent collocates of each adjective in this period —in decreasing order according to their PMI scores (cf. Tables C.4, C.5, and C.6 in Appendix C)— are the following: (i) fragrant:odor,balsam,honeysuckle,geranium,andpetal. (ii) perfumed:handkerchief,lace,blossom,oil, and silk. (iii) scented:soap,grass,blossom,bloom, and breeze. This list confirms the general tendencies described above, with fragrant and scented occurring mostly with ‘natural’ collocates, and perfumed with ‘artificial’ ones. Nevertheless, in the case of scented, the most prominent collocate is soap, which belongs to the category B4: CLEANING AND PERSONAL CARE, and is thus a noun that occurs with the adjectives at issue when they are used in the artificial sense. Figure 33 displays the collocational network for P3. DANIELA PETTERSSON-TRABA 264 Figure 33. Collocational network of fragrant, perfumed, and scented in P3 (1910–1959) With 39 collocates, fragrant is still the most productive adjective of the synonym set in this period, but it now seems to start losing some ground if compared with previous stages. Similarly, perfumed decreases somewhat in comparison to P2, with 11 nouns, i.e. five less than in P2. Scented, in turn, appears with one more collocate (11 in all) than in the preceding period. As regards the semantics of the collocates of the near-synonyms, the results point to a certain degree of specialization of the adjectives: fragrant still dominates in the natural sense, with 25 collocates clearly depicting natural smells (e.g. blossom,forest,lily,pine, and rose), while perfumed favors the artificial sense (e.g. garment,handkerchief,powder, and soap). As becomes evident from the collocational network in Figure 33, the noun collocates of fragrant and perfumed mainly correspond with the same semantic categories as those discussed for P2, namely L3, F,W4, X3.5for fragrant, and O1 and B1, B5, and O1for perfumed. On the other hand, scented does not seem to exhibit a clear preference for any of the two senses, occurring both in 7. Idiosyncratic collocational preferences of the near-synonyms 265 the natural sense, especially with nouns in category L3 (e.g. flower,garden, and tree), and in the artificial sense, with nouns from different categories, especially O1 (e.g. powder and oil), but also B4 (e.g. soap and handkerchief). Again, the five most prominent collocates of each adjective reveal that fragrant is mostly used to denote natural smells and perfumed artificial ones: (i) fragrant:aroma,incense,blossom,perfume,and clover. (ii) perfumed:soap,handkerchief,garment,bath,andpowder. (iii) scented:soap,powder,envelope,handkerchief, and cake. In the case of scented, all top five nouns clearly belong to the artificial domain.117 Interestingly, in the case of fragrant,L3 nouns seem to become less dominant, with only 2 items (blossom and clover) in its top five, as opposed to 4 in previous periods. Also worthy of notice is the fact that perfumed and scented share several of their most prominent collocates in this period, namely soap,handkerchief, and powder, which may indicate that their prototypical uses are more similar than to those of fragrant. Figure 34 shows the collocational network for P4. 117 The noun cake can be used either in the semantic category F:FOOD AND FARMING, as in example (i), thus referring to a natural aroma, or in B4: CLEANING AND PERSONAL CARE, as in (ii), thus referring to an artificial one. (i) Half the bread compartment was filled with dainty sandwiches of bread and butter sprinkled with the yolk of egg and the remainder with three large slices of the most fragrant spice cake imaginable. (COHA, 1909, FIC, GirlLimberlost) (ii) “I put a cake of scented soap among your handkerchiefs,” she said, rather breathlessly. (COHA, 1922, FIC, BreakingPoint) In the data, scented always collocates with cake in the latter category (B4: CLEANING AND PERSONAL CARE), as opposed to fragrant, which, in all cases but one, co-occurs with cake in the former category (F:FOOD AND FARMING). DANIELA PETTERSSON-TRABA 266 Figure 34. Collocational network of fragrant, perfumed, and scented in P4 (1960–2009) As shown here, fragrant continues losing ground, occurring now with only 33 collocates. The number of collocates of perfumed also decreases slightly in contrast with P3, from 11 to 9. Contrariwise, the number of collocates of scented continues rising (13 collocates). As in the previous periods, many of the collocates of scented are shared with one or both of the other adjectives, but many of them are now more strongly associated withscented. In order words, the PMI scores of these nouns are now higher with scented than with fragrant or perfumed. This is the case of hair,perfume, and smoke for fragrant, and of bath,oil, and water for perfumed (cf. Tables C.4–C.6 in Appendix C). This finding seems to be in line with those of previous chapters: scented comes to occupy some of the semantic space previously belonging to fragrant and/or perfumed (cf. Sections 5.4 and 6.2.1). As concerns the semantics of the nouns, those related to natural smells still collocate more commonly with fragrant (21 natural collocates, e.g. bloom,flesh,jasmine, and shrub), while perfumed remains as the preferred 7. Idiosyncratic collocational preferences of the near-synonyms 267 choice with artificial smells (e.g. bath,handkerchief,oil, and soap). Again, the semantic categories of nouns with which the adjectives collocate are basically the same: mainly categories F,L3, W4, X3.5, and O1 in the case of fragrant, and B1, B4, B5, and O1 in the case of perfumed. As in the previous period, scented seems to occupy an intermediate position in this respect, being common with both ‘natural’ (mainly category L3; e.g. geranium and flower) and ‘artificial’ collocates (mainly categories B1, B4, B5, and O1; e.g. hair,perfume,soap,sheet, and oil). The top five noun collocates of the three adjectives in P4 are the following: (i) fragrant:aroma,jasmine,shrub,goddess,andblossom. (ii) perfumed:soap,handkerchief,bath,smell,andoil. (iii) scented:geranium,soap,candle,perfume, and bath. The five most prominent collocates of scented are, as in P3, mostly ‘artificial’, with the exception of geranium, which as we saw in Section 6.2.2, occurs with scented only in one specific text in COHA. All top noun collocates of perfumed except smell are also clearly ‘artificial’, while the opposite is true of those of fragrant. Interestingly, two of the artificial noun collocates of scented are, once more, shared with perfumed, namely soap and bath, again indicating a higher degree of similarity between these two adjectives. A different but complementary perspective of the collocational preferences of the adjectives can be achieved by classifying their most prominent collocates visualized in the collocational networks in Figures 31–34 into the types proposed by McEnery & Baker (2017: 25–30), to wit, consistent, initiating, terminating, and transient. A summary of the absolute number and percentages of the different types of collocates of the three near-synonyms is provided in Table 33. The detailed classification of noun collocates into types is shown in Tables 34–36. DANIELA PETTERSSON-TRABA 268 Table 33. Number and percentage of consistent, initiating, terminating, and transient collocates of fragrant, perfumed, and scented Consistent Initiating Terminating Transient TOTAL Fragrant N27 12 22 35 96 %28.13 12.51 22.91 36.46 100 Perfumed N6441125 %24.00 16.00 16.00 44.00 100 Scented N3911326 %11.54 34.62 3.85 50.00 100 Here, the focus is mainly on consistent, initiating, and terminating noun collocates, particularly on the last two types, as they are the most informative ones regarding possible diachronic fluctuations in collocational preferences. Transient collocates are disregarded, as they are basically indicative of transient or punctual changes (McEnery & Baker 2017: 28), while the focus here is on more stable diachronic tendencies or the lack of such trends. As displayed in Table 33, fragrant and perfumed seem to be the most stable adjectives, with a much higher amount of consistent collocates, 28.13% and 24%, respectively. Scented, in turn, exhibits a lower rate of such collocates, namely 11.54%. If we consider the share of initiating and terminating collocates of the adjectives, an interesting picture emerges. Fragrant displays the largest amount of terminating collocates (22.91%) and the lowest amount of initiating ones (12.51%), which means that it loses more collocates over time than it gains. This finding replicates those of Section 5.1 (cf. Table 9 and Figure 3), where an overall decrease in frequency of this adjective was identified during the time span analyzed. Interestingly, some patterns can be deduced from the semantics of the noun collocates that fragrant gains and loses over time (cf. Table 34).118 118 In Tables 34–36, the following abbreviations are used to refer to consistent, initiating, terminating, and transient collocates, respectively: cons., init., ter., and tran. 7. Idiosyncratic collocational preferences of the near-synonyms 269 Table 34. Consistent, initiating, terminating, and transient noun collocates of fragrant COLLOCATE P1 P2 P3 P4 TYPE COLLOCATE P1 P2 P3 P4 TYPE air + + + + Cons. hay + + + Cons. apple +Init. heat +Init. aroma ++Init. herb + + + Cons. atmosphere + + Tran. hillside +Tran. balsam +Tran.honeysuckle ++ Ter. basket +Tran.incense + + + Cons. berry +Tran.jasmine +Init. beverage +Tran.June +Tran. bloom + + + + Cons. leaf ++++Cons. blossom + + + + Cons. lily + + Tran. blush +Ter.liquid +Tran. bough + + Tran. load +Tran. bouquet + + Tran. magnolia +Tran. bower ++ Ter. meadow ++ Ter. bowl +Tran.melody +Tran. bread +Init memory +Tran. breath + + + Cons. minute +Init. breeze + + + + Cons. odor ++++Cons. bud ++ Ter. oil + + Tran. bunch +Tran.orchard +Tran. cedar +Tran.perfume ++++Cons. chicken +Init. petal ++ Ter. cigar + + + Cons. pile +Tran. cigarette +Tran.pine + + + Cons. cloud + + Tran. plant +Ter. clover + + Tran. porch +Tran. cluster ++ Ter. rose ++++Cons. coffee + + + Cons. scent + + + Cons. cup + + + + Cons. shade ++ Ter. curl +Ter.shrub + + + Cons. dew + + + Cons. sight +Ter. dusk +Tran.smell + + + Cons. fern +Tran.smoke + + + Cons. fir +Tran.spice ++ Ter. flesh +Init. spring +Tran. DANIELA PETTERSSON-TRABA 270 Table 34. Continued COLLOCATE P1 P2 P3 P4 TYPE COLLOCATE P1 P2 P3 P4 TYPE flower ++++Cons. steam + + + Cons. foliage ++ Ter. summer + + + Cons. forest ++Init. sunshine +Tran. fruit ++ Ter. tea + + + Cons. gale +Ter.tobacco + + + Cons. garden ++++Cons. tree + + Tran. garland +Ter. verdure +Tran. garlic +Init. vine ++ Ter. geranium +Tran.weed +Tran. goddess +Init. wind +Ter. grass ++++Cons. wine +Ter. grove ++ Ter. wood ++ Ter. gum +Tran.wreath ++ Ter. hair ++Init. Terminating collocates correspond mostly with nouns belonging to the semantic categories L3: PLANTS (e.g. bower,foliage,honeysuckle,petal, and plant) and W4: WEATHER (e.g. gale, meadow,andwind), whereas initiating ones are concentrated mainly in categories F:FOOD AND FARMING (e.g. apple,bread,chicken,andgarlic) and B1: ANATOMY AND PHYSIOLOGY (e.g. flesh [of body] and hair). As seen in section 5.3.2, fragrant increased in P4 with nouns belonging to the semantic category FOOD AND DRINK;the findings obtained here point to a similar conclusion. However, it is worth mentioning that fragrant also gains some new L3 collocates (e.g. forest and jasmine) and loses some Fcollocates (e.g. fruit,spice,andwine), which suggests that its prototypical uses remain largely the same over time. This can also be observed by zooming in on the consistent collocates of this adjective, most of which are L3 nouns, especially basic level terms such as bloom,blossom,flower,garden,grass, and leaf and W4 nouns, such as breeze, dew, and summer. Many of the consistent collocates of fragrant correspond with those that exhibited a high degree of association with this adjective in the mixed-effects regression analysis carried out in Section 6.2.2. These are the three tobacco-related nouns cigar,smoke, and tobacco. In turn, the noun weed, which emerged as a prominent collocate of fragrant in the mixed-effects regression analyses, is here categorized as a transient collocate, occurring only in P2. This shows the importance of conducting a fine-grained qualitative analysis of the individual collocates in different periods, since some collocates are more informative than 7. Idiosyncratic collocational preferences of the near-synonyms 271 others, occurring continually over time. Finally, two other consistent collocates of fragrant are also in line with the results of the mixed-effects regression model. These are the two nouns odor and perfume, which belong to the sensory domain. While both lemmas appear as prominent collocates in all four periods, odor displays a downward tendency with fragrant in P4: out of its 6 occurrences in this period, 5 are dated before the 1980s, with only one example of this noun co-occurring with fragrant in the last 30 years in COHA. The decrease of odor as a collocate of fragrant may be related to the fact that this noun has undergone a process of pejoration in recent years, as explained in Section 3.3, being nowadays used mainly with a negative connotation to depict ‘unpleasant smells’, much as many of its related forms (e.g. odorous; cf. Section 4.1.2). On the contrary, fragrant and the other adjectives analyzed in the present dissertation are explicitly used to denote pleasant smells, and no indication has been found in the corpus of their use to designate disagreeable ones. Therefore, it is not at all surprising that odor, which progressively becomes more strongly associated with negative smells, almost vanishes in the last decades in COHA as a collocate of fragrant. Additionally, in the examination of the last decade in COCA (2010s), odor appears only once in the surrounding context of fragrant (i.e. L5–35 window), a fact which reinforces this hypothesis. As regards perfumed, its relative stability in terms of its overall frequency (cf. Section 5.1) seems to be reflected also in its proportion of initiating and of terminating noun collocates. As can be seen in Table 33 above, these two types of collocates are evenly distributed, namely 16%. This indicates that, over time, perfumed has gained as many new prominent collocates as it has lost. As concerns the semantics of these collocates (cf. Table 35), 3 out of 4 initiating ones correspond to uses in the artificial sense and in indeterminate uses, namely bath and soap, i.e. category B4: CLEANING AND PERSONAL CARE, and body. i.e. category B1: ANATOMY AND PHYSIOLOGY. On the contrary, two of its terminating collocates belong to L3: PLANTS (flower and tree) and another one is the sensation noun flavor, which has already been commented on at several points in this dissertation, as the use of perfumed with this noun is restricted to just a few particular texts in COHA in P1 (cf. Sections 5.3.2 and 6.2.2). [Document text truncated for crawler view.]