Co-varying verbs and adjectives of it-extraposed constructions with to-infinitive clauses in academic discourse: a quantitative corpus-driven study
Abstract
7
Full text
Co-varying verbs and adjectives of it-extraposed constructions with to-infinitive clauses in academic discourse: aquantitative corpus-driven study Jarosław Wiliński (Siedlce) ABSTRACT This paper employs the background assumptions of usage-based Construction Grammar (Goldberg 1995, 2006, 2013), Frame Semantics (Fillmore 1982), and aquantitative corpus-driven method for investigating the reciprocal interaction between lexical items occurring in two different slots of agrammatical construction. The method, referred to as co-varying collexeme analysis (Stefanowitsch and Gries 2005; Stefanowitsch 2013; Hilpert 2014), is applied to the determination of strongly attracted and repelled pairs of adjectives and verbs occurring in the extraposition construction with to-infinitive clauses in American English. Using the data extracted from the academic sub-corpus of COCA, the author seeks to indicate that some pairs of adjectives and verbs co-occur significantly more frequently than expected in the it is ADJ to V-construction. Furthermore, the results of the analysis of the co-variation of collexemes in two different slots of the same construction seem to suggest that such strong correlations between these slots can be determined by frame-semantic knowledge and/or discourse-functional properties of the construction under study. KEYWORDS Construction Grammar, Frame Semantics, extraposition, co-varying collexeme analysis, COCA DOI https://doi.org/10.14712/18059635.2019.1.1 1. INTRODUCTION Extraposition in English has received much attention over the last three decades (Quirk et al. 1985; Seppänen, Engström and Seppänen 1990; Seppänen 1999; Kaltenböck 2000, 2003, 2004; Kaatari 2010). Some research studies have compared extraposed constructions with clefts (Pérez-Guerra 1998; Calude 2008) as well as extraposition to right dislocation (McCawley 1988; Collins 1994), while others have focused on discourse functions of different extraposed structures (Mair 1990; Herriman 2000a,b; Hoey 2000; Hewings and Hewings 2002). Some researchers have also explored the occurrence of epistemic, deontic, or evaluative adjectives in avariety of extraposed constructions (Biber et al. 1999; Van Linden 2012), their clausal complementation (Mindt 2011), and various valency properties (Herbst et al. 2004). Thus far, however, scant attention has been paid to the quantitative evaluation of adjectives and verbs in the extraposed construction complemented by to-clauses, the statistical validation of their occurrence in academic discourse, or the empirical confirmation of previous assumptions and speculations about their use. Hilpert’s (2014)
8 LINGUISTICA PRAGENSIA 1/2019 case study of different pairs of adjectives and verbs occurring preferentially in the it is ADJ to V-construction and Wiliński’s (2017) quantitative investigation of adjectives complemented by to-clauses in academic discourse are notable exceptions. Using the data retrieved from the BNC corpus, Hilpert established that there are certain pairs of verbs and adjectives closely related semantically that display astrong preference for the construction in question. On the basis of the data extracted from the academic sub-corpus of COCA, Wiliński in turn found that some adjectives are more strongly attracted to this construction than others, and that the occurrence of certain adjectives in this construction is more significant than their use in different types of extraposed constructions. Given that Hilpert’s study solely concerned the use of the it is ADJ to V-construction in British English and Wiliński’s study was not specifically designed to capture interdependencies between adjectives and verbs, there is still aneed for the quantitative determination of strongly attracted and repelled combinations of verbs and adjectives in the it is ADJ to V-construction in American English and for the qualitative analysis of their usage in this kind of extraposition, in view of the widespread occurrence of the construction in question in academic discourse. Thus, using data extracted from the academic section of the Corpus of Contemporary American English, the author seeks to determine pairs of verbs and adjectives strongly and loosely associated with the pattern under investigation. The remainder of this paper is structured in the following fashion. Section 2 explains both the theory fundamental to the semantic explanation of pairs of verbs and adjectives occurring in the it is ADJ to V-construction, and the methodology underlying the quantitative analysis of these combinations. Section 3 describes the corpus, the data, and the tools. Section 4 outlines the statistical procedure employed in this study. Section 5 defines the construction under scrutiny and discusses its function and usage. Section 6 combines the findings of the quantitative analysis with asemantic description of adjectives and verbs and elucidates the contribution of various semantic frames to the constructional meaning. Section 7 assesses the findings and formulates some proposals for future research. 2. THEORY AND METHOD This study rests on the theoretical foundation provided by the usage-based Construction Grammar (Goldberg 1995, 2006, 2013) and Frame Semantics (Fillmore 1982; Fillmore and Atkins 1992; Fillmore and Baker 2010). Aconstructional approach to grammar assumes that there is no strict division between grammar and lexicon. Grammar consists of constructions or symbolic units, pairings of aform and ameaning/function, i.e. conventionalized associations of aphonological structure and asemantic/ conceptual structure. Crucially for the current study, the notion of construction is not restricted to morphemes or words, but encompasses more complex and schematic constructions such as partially filled structures, lexically unspecified patterns, argument structure constructions, extraposed constructions, idioms, etc. This theory places special emphasis on actual frequencies of usage or occurrence and hence it is
JAROSłAW WILIńSKI 9 explicitly usage-based (see Bybee 2010, 2013) in the sense that exposure to, or use of, constructions is deemed to influence the linguistic system of speakers and hearers, while sufficient frequency is anecessary condition for the entrenchment (Langacker 1987) and the achievement of construction status of alinguistic expression. Frame Semantics is atheory of linguistic meaning formulated by the American linguist Charles Fillmore that aims to explain the meaning of lexical units and grammatical constructions in terms of frames or prototypical scenes relating to structured— but to some extent individually and culturally varying— background knowledge (Fontenelle 2003). For example, the word sell cannot be understood without access to all the essential knowledge associated with the commercial transaction frame (cf. Petruck 1996), i.e. the situation of commercial transfer involving, among other things, aseller, abuyer, goods, money, and the relations between them. Thus, words and constructions activate, or evoke, frames of encyclopedic knowledge providing the background and motivation for their meanings. Frame Semantics has been put into practice in the Berkeley FrameNet project (Fillmore, Johnson and Petruck 2003; Fillmore and Baker 2010). The primary aim of FrameNet is to systematically describe syntactic and semantic valency patterns of lexical units based on extensive corpus annotation. In this respect, the syntactic properties of lexical units in corpora are systematically aligned with the semantic frames evoked by the words. The description of each frame in the FrameNet database includes the following components: the name of the frame, adefinition of the situation the frame is supposed to represent, the set of core and non-core frame elements (semantic roles) associated with the frame, and the corresponding word senses (lexical units) that activate the frame. Practically all semantic frames and their modified definitions discussed in this study are taken from the Berkeley FrameNet project: importance, difficulty, mental stimulation, statement, becoming aware, remembering information, grasp, cogitation, awareness, coming to believe, separating, mental property attribution, expectation, fairness evaluation, usefulness evaluation, correctness evaluation, frequency evaluation, likelihood, and prediction. The remaining semantic frames along with their descriptions, implemented in the description of semantic properties of verbs and adjectives, are created by the author himself: in other words, risk evaluation and realism evaluation. The methodology of quantitative corpus linguistics is applied in this study. The method called co-varying collexeme analysis (Stefanowitsch and Gries 2005; Stefanowitsch 2013; Hilpert 2014) is aimed at determining combinations of elements that occur more often than would be expected by chance in the it is ADJ to V-construction, i.e. identifying pairs of adjectives and verbs that are significantly attracted to, or biased towards, the investigated pattern in the academic section of COCA through astatistical evaluation of the observed frequencies of the lexemes in question in relation to the overall frequency of the construction in the corpus. The output is aranking list of the so-called co-varying collexemes, i.e. of those pairs of lexemes that exhibit astronger preference for the investigated construction than others. Although this technique is quantitative, the results of this analysis are evaluated qualitatively and subjectively. For example, the meanings of the lexemes that are strongly associated
10 LINGUISTICA PRAGENSIA 1/2019 with the construction can be interpreted with respect to the semantic frames to which they are relativised. 3. CORPUS, DATA, AND TOOLS The data were gathered from the downloadable version of the academic part of the Corpus of Contemporary American English (COCA), i.e. the full-text data corpus purchased from Mark Davies. This academic part contains approximately 81 million words coming from nearly 100 different peer-reviewed journals. These encompass awide range of academic disciplines: e.g. acertain percentage from philosophy, psychology, religion, world history, education, and technology. The observed frequencies were retrieved from the corpus by means of MonoConc Pro, aconcordance program. This tool was used to search through the corpus for all the occurrences of adjectives and verbs in the construction under study. Each concordance line was manually inspected to determine the frequencies of all pairs of adjectives and verbs co-occurring in the relevant pattern. Then, all these frequencies required for the computation of the mutual association between combinations of elements in the it is ADJ to V-construction were entered in a2-by-2 table and submitted to the Fisher exact test. The p-value provided by this test was used to gauge the strength of association, i.e. the degree of attraction to or repulsion from the it is ADJ to V-construction: the smaller the p-value, the higher the probability that the observed distribution is not due to chance and the higher the strength of the association between two slots of the same construction (cf. Schmid and Küchenhoff 2013). This calculation of statistical significance was performed by means of an online Fisher’s exact test calculator for two-by-two contingency tables. The rest of the values and expected frequencies were calculated by means of Microsoft Excel spreadsheets. The resulting frequency lists then provided the input to the co-varying collexeme analysis. Generally, p-values are so low that their significance lies only in the number of decimal places. These values are precisely expressed in numbers of the type “1.31E–10” (see for example the result provided for the combination difficult to ascertain in Table 2 below), which reads “1.31 times 10 to the power of minus 10”, i.e. 0.000000000131. To simplify things, alog transformation of these p-values is frequently given in some studies (e.g. Hilpert 2014; Perek 2014) which apply Coll.analysis 3, an R script written by Stefan Gries. This script uses alog transformation to the p-values yielded by the Fisher exact test, and turns the sign into aplus if the association is one of attraction (i.e. the actual frequency of occurrence of averb and an adjective exceeds the expected frequency) and into aminus in the case of repulsion (i.e. the actual frequency of the combination is lower than the expected frequency). This provides amore readable value than p-values, expressed in powers of ten (cf. Hilpert 2014: 402; Perek 2014: 69). The values of the strength of association between adjectives and verbs larger than 1.301 mean that particular combinations are significantly attracted to the construction, whereas the values lower than –1.301 mean that combinations are significantly repelled by the construction (cf. Hilpert 2014: 402).
JAROSłAW WILIńSKI 11 It is worth noting that the use of the p-value as asignificance measure has come under heavy criticism from Schmid and Küchenhoff in their last publication (2013: 539). This criticism centres on the issue of whether or not the Fisher exact p-value incorporates an effect size. Gries (2015: 520) argues that although “p-values are not effect sizes, p-values by their very nature reflect acombination of different things including the size of the sample(s), the variability of the sample(s), and the effect size.” The rationale for the use of the Fisher exact as asignificance test is that, in comparison to other statistical tests, this measure can be used to assess the interaction among variables when data is very unevenly distributed and/or infrequent (cf. e.g. Stefanowitsch and Gries 2003: 9; Gries and Stefanowitsch 2004a: 101; see also Gries 2012 and Gries 2015: 508 for further arguments). 4. STATISTICAL PROCEDURE The procedure followed in this study consisted of four stages. This procedure can be illustrated with the aid of the adjective important in the adjective slot along with the verb note in the verb slot of the it is ADJ to V-construction. Important in adjective slot of the it is ADJ to V-construction All other adjectives in adjective slot of the it is ADJ to V-construction Total Note in verb slot of the it is ADJ to V-construction a: Frequency of adjective (important) and verb (note) in the it is ADJ to V-construction 318 (97.91) b: Frequency of all other adjectives and verb (note) in the it is ADJ to V-construction 168 x: Total frequency of verb (note) in the it is ADJ to V-construction 486 Other verbs in verb slot of the it is ADJ to V-construction c: Frequency of adjective (important) and other verbs in the it is ADJ to V-construction 1628 d: Frequency of all other adjectives and other verbs in the it is ADJ to V-construction 7545 y: Total frequency of all other verbs in the it is ADJ to V-construction 9173 Total e: Total frequency of adjective (important) in the it is ADJ to V-construction 1946 f: Total frequency of all other adjectives the it is ADJ to V-construction 7713 z: Total frequency of the it is ADJ to V-construction 9659 Table 1.Co-occurrence table for aco-varying-collexeme analysis At the initial stage of this procedure, the observed frequencies were calculated on the basis of the data extracted from the corpus. First, all occurrences of the construction under study were identified from the corpus: 9659. Second, the frequency of the adjective important in the adjective slot was determined: 1946. Third, the frequency of the verb note in the verb slot was calculated: 486. Finally, the frequency of the adjec-
12 LINGUISTICA PRAGENSIA 1/2019 tive important and the verb note appearing together was counted: 318. These four values were derived from the corpus directly, while the remaining ones resulted from subtraction in Table 1.For example, in order to calculate the frequency of all other adjectives and the verb note in the it is ADJ to V-construction, the frequency of the adjective important and the verb note in the same construction (318) was subtracted from the total frequency of the verb note in the it is ADJ to V-construction (486), giving the result (168). Table 1 above displays the actual frequencies necessary to carry out aco-varying collexeme analysis of the adjective important and the verb note in the construction under scrutiny (for expository purposes, it also gives the expected frequencies for the adjective important and the verb note in parentheses). At the second stage, these observed values were used to calculate the expected frequency of the adjective (important) and the verb (note) in the investigated construction. This calculation was performed in Microsoft Excel in the following fashion. For this combination of elements, its column total was multiplied by its row total, and the result was divided by the overall table total. For example, for the top cell containing the figure 318, its column total (1946) was multiplied by the row total (486), giving the figure (945756). Then this figure was divided by the table total (9659), yielding the result (97.91). If the observed frequency of the adjective (important) and the verb (note) together in the construction is significantly higher or lower than expected, the relation between this pair of lexemes is one of attraction or repulsion respectively (the adjective important and the verb note are then assumed to be significantly attracted or repelled collexemes of this construction). At the third stage, the degree of attraction, or the association strength, between the adjective (important) and the verb note was estimated by means of the Fisher exact test. To this end, the following four frequencies were employed: the frequency of the adjective (important) and the verb (note) in the it is ADJ to V-construction, the frequency of all other adjectives and the verb (note) in this construction, the frequency of the adjective (important) and other verbs in the pattern in question, and the frequency of all other adjectives and other verbs in the investigated construction. The p-value resulting from the computation of the Fisher exact test for this combination is exceptionally small: 5.21E–111. This means that the adjective (important) and the verb note share strong mutual attraction in the investigated construction, but this can only be determined by comparing the observed frequencies of the adjective (important) and the verb note with the expected ones. As this comparison indicates, the adjective (important) occurs more frequently than expected with the verb note in the construction. In other words, important and note are highly significant, very strongly attracted collexemes of the construction. This procedure was employed for all pairs of adjectives and verbs in the investigated construction. At the next stage, the results were arranged, first, according to the direction of association (attracted or repelled), and second, according to their association strength. Finally, the data were interpreted qualitatively and subjectively. More specifically, the results of the quantitative analysis were integrated with asemantic description of the most strongly attracted pairs of adjectives and verbs, and the contribution of frame-semantic knowledge to the meaning of the pattern under study was explained.
JAROSłAW WILIńSKI 13 5. IT-EXTRAPOSITION WITH TO-INFINITIVE CLAUSES It-extraposition with to-infinitive clauses refers to asyntactic process by which ato-infinitive clause is shifted (extraposed) from its initial position (i.e. its subject position) to the end of asentence. This usually involves the use of the dummy pronoun it as asubject. Classic examples of this type of extraposition are given in (1), with the non-extraposed counterparts being provided in (2): (1) it-extraposition with to-infinitive clauses a. It is impossible to buy aflat here b. It is difficult to find agood wife c. It is important to be able to speak English (2) non-extraposition a. To buy aflat here is impossible b. To find agood wife is difficult c. To be able to speak English is important Given that it-extraposition is similar to non-extraposition in respect to its structure and logico-semantic properties, many researchers from the formal school of generative linguistics (e.g. Rosenbaum 1967; Huddleston 1971; Emonds 1972) treated the sentences in (1) as syntactic derivations or transformations of the sentences in (2). Recently, however, abody of empirical evidence (e.g. Francis 1993; Biber et al. 1999; Kaltenböck 2000) obtained from naturally occurring data in corpora has suggested that examples of non-extraposition are extremely rare in corpora. For example, Kaltenböck’s research (2000: 158) revealed that instances of it-extraposition considerably outnumber those of its non-extraposed counterpart with aratio of 1:7.8 in the British section of the International Corpus of English. Biber et al. (1999: 676, 724), in turn, noticed that occurrences of that-clauses and to-clauses in non-extraposition are extremely infrequent in spoken language, and that it-extraposition is much more preferred. In addition, as noted by Quirk et al. (1985: 964–965), some extraposed examples do not allow for reversion to non-extraposed constructions, in either writing or speech. Hence, it is debatable whether sentences such as those in (1) are indeed transformations of the sentences in (2), and it seems to be more acceptable to treat It is impossible to buy aflat here as aconstruction (apairing of form and meaning/function) in its own right, and to examine it accordingly, rather than consider it as aversion of something that is used extremely infrequently. In this study, therefore, examples such as the ones in (1) are assumed to be atype of the English it-extraposition construction, apartially lexically-filled pattern consisting of three fixed lexical items (it is […] to […]) and two flexible slots that can be filled by adjectives and verbs. This pattern can be represented structurally and schematically as [it is ADJ to-infinitive clause], where adummy subject it is followed by the third person singular form of the verb be, apredicative adjective, and ato-infinitive clause. The use of this construction can be exemplified by the following sentences retrieved from the corpus:
14 LINGUISTICA PRAGENSIA 1/2019 (3) It is important to note that this assessment is criterion based and leveled by grade (4) It is hard to imagine that nearly 300,000 men died or were wounded here almost acentury ago (5) It is reasonable to expect asignificant reacceleration of inflation in the near future Regarding discourse-functional properties of extraposed constructions with to-infinitive clauses, much research (Huddleston 1984; Collins 1994; Gómez-González 1997; Herriman 2000a; Hoey 2000; Hewings and Hewings 2002; Rowley-Jolivet and Carter-Thomas 2005) has shown that it-extraposition in examples such as those in (3), (4) and (5) can serve two crucial functions. First, it is commonly used in both speech and writing to avoid long and heavy subject clauses because they sound awkward, thereby placing them at the end of the sentence, in accordance with the principles of end-weight and end-focus. This function allows speakers to convey new pieces of information in away that is easier to process (cf. Huddleston 1984: 453; Quirk et al. 1985: 863; Erdmann 1990: 137–8; Collins 1994: 15–16). Aslightly different view is expressed by Mair (1990: 39), Miller (2001), and Kaltenböck (2005), who found that extraposed clauses convey not only new but also given information. Second, it-extraposition allows aspeaker/writer to express subjective opinions about some state-of-affairs by presenting them as if they were generally accepted views rather than his/her personal judgement, hence introducing evaluative comments at the beginning of asentence (cf. Herriman 2000b: 211; Gómez-González 2001: 272; Rowley-Jolivet and Carter-Thomas 2005: 51; Kaltenböck 2005: 137). Although discourse-functional and structural properties of various types of extraposed constructions have received systematic treatment in the literature, the role of the adjectives and verbs in extraposed patterns with to-infinitive clauses has largely been ignored and neglected. Hence, research into interdependences between adjectives and verbs in this kind of it-extraposition deserves more attention. The ratio nale for undertaking such an empirical study is that the meanings of adjectives and verbs enormously influence the constructional meaning. For example, the combination of elements (important to note, hard to imagine, reasonable to expect) in (3), (4), and (5) contribute substantially to the understanding of the illustrative sentences by assigning different meanings to the constructions under scrutiny. In these cases, adjectives denoting importance, difficulty, or aspecific mental property co-occur with verbs that can be used to introduce astatement, to denote awareness or knowledge about afact, or to express the belief that some phenomenon will take place in the future. Thus, the quantitative investigation of such pairs and their semantic description with respect to the semantic frames they activate may enable us to find subtle distributional differences in their use and understand their role in the investigated pattern, as well as to broaden our knowledge and understanding of the meaning and function of the construction. Given the semantic and discourse-functional properties of the examples mentioned above and the results of the study conducted by Hilpert (2014), it is possible to predict roughly what adjectives and verbs are likely to occur in both slots of this construction in academic discourse. The adjectival slot should prefer adjectives expressing the speaker’s or writer’s evaluative judgement, whereas the verbal slot should
JAROSłAW WILIńSKI 15 prefer verbs denoting cognitive processes, introducing astatement, and/or conveying new facts and states of affairs. These predictions will be tested below. It is important, however, to note that even the detailed description of the construction’s semantics does not allow us to predict whether adjectival and verbal slots in this pattern are related semantically and in what way. It follows from the principle of semantic compatibility (Stefanowitsch and Gries 2005: 11) that co-occurrences of adjectives and verbs are expected to be semantically coherent, but it does not specify what kind of semantic coherence we could expect for the construction under investigation. In this study, it is assumed that the semantic coherence between two different slots of this construction can be determined by frame-semantic knowledge, i.e. arelationship between semantic frames evoked by adjectives and verbs. 6. FINDINGS AND DISCUSSION The concordancer extracted 9659 occurrences of the it is ADJ to V-construction containing 3956 different combinations of adjectives and verbs, out of which 2713 occurred only once in the investigated construction. Because of the limitation of space, however, this section will only interpret the findings for the 30 most strongly attracted and repelled co-varying collexemes of this pattern. Table 2 provides the results of aco-varying collexeme analysis (PFisher exact) for the 30 most strongly attracted combinations of adjectives and verbs. It also displays the observed and the expected frequencies for each pair of lexical items occurring in two different slots of the investigated construction. The figures (a, e, x, z) were derived from the corpus directly, while the remaining figures (b, c, d, f, y) result from addition and subtraction. The results support the prediction that the semantic coherence between two different slots of the construction under study is based on arelationship between two frames, i.e. on frame-semantic knowledge. Furthermore, the specific suggestions concerning the meaning of this construction are also confirmed. For this construction, we find that combinations of lexemes evoking the importance frame and difficulty frame constitute the bulk of the most strongly associated pairs of co-varying collexemes in the ranking list. The former frame is evoked by combinations such as important to note, important to recognize, important to remember, important to understand, important to keep in mind, important to acknowledge, and important to consider in ranks 1, 9, 10, 15, 19, 20 and 23. The p-values resulting from the calculation of the Fisher exact test for these collexeme pairs are exceptionally small: 5.21E–111, 2.11E–36, 2.43E–34, 8.04E–16, 4.77E–13, 7.16E–13, 7.16E–13, respectively. Acomparison of the observed and the expected frequencies of these pairs of lexical items occurring in two different slots indicates that these pairs occur more frequently than expected in the construction. In other words, they are highly significant and very strongly attracted to each other in this construction. Note also that important to note is the most strongly associated co-varying-collexeme pair in the construction, since its p-value is exceptionally small (5.21E–111) and the observed frequency is much higher than the expected one. This combination is
22 LINGUISTICA PRAGENSIA 1/2019 It is useful [to recall that, at its start, Poland’s powerful Solidarity movement lacked clear and cohesive leadership] ACTION, some action or desirable state of affairs (in this case, the recollection of facts) is considered by aspeaker as useful for abenefiting party. True to say, ranked number 21, can be described with reference to the background knowledge associated with the correctness evaluation frame and the statement frame. For example, the speaker of the sentence It is true [to say, however, that even his earliest horses and riders had an unsettled or unsatisfactory partnership] INFORMATION introduces the statement by means of the verb say and judges apiece of information to be correct or true. The pair rare to find in rank 25 is aconcrete instance representing arelationship between the frequency evaluation frame and the becoming aware frame. The first frame refers to an action, event or salient entity evaluated by acognizer as being frequent or rare, while the second one concerns acognizer becoming aware of some phenomenon, an entity, or asituation in the world. Both frames are evoked by the combination of rare and find in the sentence It is rare [to find astein entirely made of ivory] EVENT. Finally, the bottom of the ranking list contains impossible to predict, apair of lexemes invoking the likelihood frame and the prediction frame. The first frame is concerned with the likelihood of ahypothetical event (the state of affairs or occurrence) being evaluated by ajudge. The second one, in turn, has something to do with an event or state that is predicted by aspeaker to occur or hold true at afuture time. The co-occurrence of impossible and predict in the sentence It is impossible [to predict which students will be future bullies, victims, or bystanders] HYPOTHETICAL EVENT is evidence of astrong correlation between these two frames, i.e. amutual connection affecting the semantic coherence between the two slots of the construction under investigation. At the last stage of the interpretation, it is also worth pointing out pairs of adjectives and verbs that are not significantly attracted to the construction in academic discourse: that is, co-varying collexemes that occur less frequently than expected in the investigated pattern. The results of acollexeme analysis for the 30 most strongly repelled pairs of the it is ADJ to V-construction are shown in Table 3. The top of the ranking list in this table is dominated by pairs of adjectives and verbs such as difficult to be, important to say, difficult to consider, important to imagine, possible to note, necessary to note, possible to be, important to be, important to find, difficult to have that are not strongly attracted lexemes, since their p-values resulting from the calculation of the Fisher exact test are very high: 1, in all of these cases. In addition, acomparison of the observed and the expected frequencies for each of this pair shows us that these collexemes occur less frequently than expected in this construction and hence they are loosely associated with the pattern under scrutiny in academic discourse. Acursory look at Table 3 already reveals that adjectives such as difficult, possible, important, impossible, hard, and reasonable demonstrate aloose association with the verb be, that the adjectives important, necessary and possible occur less frequently than expected with the verb say, that the adjectives possible, necessary and easy are loosely associated with the verb note, and that difficult and possible co-occur extremely rarely with the verb remember in the investigated pattern. In addition, some of these adjectives have aweak correlation with verbs such as have, do, go, and take, since these verbs occur extremely infrequently in the construction in academic discourse. Apos-
JAROSłAW WILIńSKI 23 sible explanation for their loose association in the investigated pattern may lie in the function and usage of the it is ADJ to V-construction in academic discourse. The results of the analysis for the 30 most strongly attracted combinations have revealed that the verbal slot of this construction exhibits astrong preference for verbs conveying new facts and information, i.e. verbs introducing astatement, denoting awareness and expectation, and evoking semantic frames such as grasp, remembering inform ation, coming to believe, or becoming aware. These verbs, in turn, have astronger tendency to occur with particular types of adjectives than with others. For example, the verb say co-occurs more frequently with the adjectives fair, safe and, true than with important, necessary and possible, the verb note tends to collocate more often with important and interesting than with possible, necessary and easy, and the verb remember prefers the adjective important to difficult and possible. The interdependence between these adjectives and verbs is strongly determined by specific semantic frames, the construction’s function, the speaker’s or writer’s communicative intention, and the context in which such constructions are used. 7. CONCLUDING REMARKS In conclusion, the findings of this investigation have indicated that the semantic coherence between the most strongly associated co-varying-collexeme pairs of the it is ADJ to V-construction is based on frame-semantic knowledge, i.e. arelationship between semantic frames evoked by adjectives and verbs co-occurring in the construction under study in aspecific situational, discourse, and conceptual-cognitive context. The co-varying collexeme analysis has revealed not only the high degree of semantic coherence that exists between different adjectival and verbal slots of the pattern in question, but also systematic relationships between semantic frames that determine this semantic coherence. These relationships are clearly not the exception, but the rule for this construction. It has been found, for example, that the interaction between the adjective important and the verbs note, recognize, remember, understand, keep in mind, acknowledge, and consider in academic discourse is determined by areciprocal relationship between the importance frame and several other frames activated by these verbs, e.g. statement (note, acknowledge), becoming aware (recognize), remembering information (remember, keep in mind), grasp (understand), and cogitation (consider). The interdependence between the pairs hard to imagine, hard to believe, difficult to imagine, and difficult to know, in turn, is based on aclose relationship between the difficulty frame and the awareness frame. The results also confirm previous predictions about types of adjectives and verbs preferred by both slots of this construction in academic discourse. The adjectival slot seems to show amarked preference for adjectives expressing the speaker’s or writer’s evaluative judgement, whereas the verbal slot prefers verbs denoting cognitive processes, introducing astatement, and/or conveying new facts and states of affairs. For example, adjectives denoting importance (e.g. important), difficulty (hard), or aspecific mental property (e.g. reasonable) co-occur with verbs that can be used to
24 LINGUISTICA PRAGENSIA 1/2019 introduce astatement (important to note), to denote awareness of or knowledge about afact (hard to imagine), or to express the belief that some phenomenon will take place in the future (reasonable to expect). Alogical explanation as to why such combinations are preferred by writers and speakers may lie in the nature and specificity of academic discourse. In this kind of register, researchers aim to present, interpret and comment on the findings of their studies. To this end, they seek to convey new factual information about the current state of their research by expressing their evaluative opinions on the importance of their results, the difficulties encountered in the process of their interpretation, the correctness of their predictions, the meaning of an idea, the occurrence of aphenomenon, the likelihood of ahypothetical event, etc. All these findings support the specific suggestions concerning the semantic and discourse-functional properties of the it is ADJ to V-construction. The illustrative examples, discussed in section 6, show two main and partially related functions of this type of extraposition in academic discourse. First, aspeaker or awriter attempts to express his/her evaluative opinion in an indirect way by introducing the evaluative comments in the form of the dummy it, the verb form is, and adjectives such as important, difficult, hard, reasonable, fair, or true at the beginning of asentence, as in It is important [to note that data suggest mixed results for the success of anti-bullying programs] UNDERTAKING. Second, aspeaker or awriter aims to introduce acompletely new idea into the discourse, anew topic that is linked to the previous context or has no direct link with the preceding context. This new idea is introduced at the end of asentence, as in It is interesting [to note that they correspond to different stabilizing control laws] STIMULUS. These findings are in agreement with earlier studies into the discourse function of extraposition (e.g. Collins 1994; Herriman 2000a; Hoey 2000; Hewings and Hewings 2002; Kaltenböck 2005; Rowley-Jolivet and Carter-Thomas 2005), and with the results of Hilpert’s (2014) co-varying-collexeme analysis of adjectives and verbs in the it is ADJ to V-construction. Using the data extracted from the BNC corpus, Hilpert found, for example, that adjectives denoting ease and difficulty (difficult, easy, hard) co-occur with verbs pertaining to cognitive processes (see, imagine, believe), while the adjective important co-occurs with verbs introducing astatement (note, remember). He also states that combinations such as it is interesting to note or important to remember “do not carry focal information in themselves and are usually less prominently stressed than the material that follows”, thereby “setting the stage for anew piece of information in discourse” (Hilpert 2014: 402). Hilpert’s analysis, however, was restricted to the indication of the 20 most strongly attracted combinations, as its primary aim was to demonstrate the application of the quantitative method for asemantic analysis of the it is ADJ to V-construction. Surprisingly, apart from pairs such as reasonable to suppose, important to realise, important to stress, interesting to compare, and good to be, the ranking list of the most strongly associated co-varying-collexeme pairs in the current study contains the same fifteen combinations interpreted by Hilpert (cf. 2014: 402) as the most significant pairs of the investigated construction. These combinations, however, hold various positions in both lists. In Hilpert’s (see 2014: 402) ranking the top nine positions are occupied by interesting to note, fair to say, important to remember, true to say, reasonable to assume,
JAROSłAW WILIńSKI 25 hard to believe, hard to imagine, important to note, and unrealistic to expect, while in the present study the nine most strongly attracted pairs are important to note, interesting to note, hard to imagine, fair to say, reasonable to assume, reasonable to expect, safe to say, unrealistic to expect, and important to recognize. This suggests, for example, that pairs reflecting the relationship between the importance frame and the statement frame, between the mental stimulation frame and the statement frame, between the difficulty frame and the awareness frame, between the fairness evaluation frame and the statement frame, and between the mental property attribution frame and the statement frame are the five most strongly attracted pairs of this construction in the academic section of COCA, while the combinations instantiating the relationship between the mental stimulation frame and the statement frame, between the fairness evaluation frame and the statement frame, between the importance frame and the remembering information frame, between the correctness evaluation frame and the statement frame, and between the mental property attribution frame and the statement frame co-occur more frequently with this pattern in the corpus of general British English. In other words, out of the five relationships between frames listed above, three occur in both studies at the top of the ranking list. Five notable exceptions, listed by Hilpert but not included in the ranking list of the current research, are reasonable to suppose, important to realise, important to stress, interesting to compare, and good to be. Reasonable to suppose (ranked number 10 in Hilpert’s list) and important to stress (ranked 13 in Hilpert’s table) are also among the most attracted pairs of this construction in this study but occupy lower positions: reasonable to suppose, with 7 occurrences, is in rank 46, while important to stress, with 15occurrences, is in rank 33.The remaining three combinations occur very rarely in the academic register, thus being among the least strongly associated pairs of the construction in academic discourse: important to realise (1 occurrence, in rank 2503), interesting to compare (4 occurrences, in rank 531), and good to be (2 occurrences, in rank 928). Apossible explanation for their loose association in the construction under scrutiny may lie in the influence of academic discourse on the preferred combinations of semantic frames. For example, this kind of register allows speakers or writers to present their evaluative opinions about the importance of the current state of their studies by introducing the adjective importance, evoking the importance frame, and verbs activating several other frames, e.g. statement (note, acknowledge), becoming aware (recognize), remembering information (remember, keep in mind), grasp (understand), and cogitation (consider), rather than by introducing the adjective important and the verb realise, activating the importance frame and the coming to believe frame, or the combination important to be, reflecting the relationship between the importance frame and the existence frame. The interdependence between the adjective important and these verbs is strongly determined by the speaker’s or writer’s communicative intention and the academic context in which such combinations are used. The co-varying collexeme analysis applied in this study has proved to be an effective technique for the determination of the most strongly associated co-varying-collexeme pairs of the it is ADJ to V-construction, and hence may be employed for the identification of the most significant pairs of lexemes co-occurring in other types of it-ex-
26 LINGUISTICA PRAGENSIA 1/2019 traposed constructions. Future research, for example, might focus on determining interdependencies between adverbs and adjectives found in two different slots of itextra posed constructions complemented by to-infinitive clauses or that-clauses. Such aquantitative analysis could reveal those combinations that occur more often than would be expected by chance, considering the respective frequencies of their participating elements. This in turn may be accompanied and supported by an analysis of semantic frames associated with these participating elements. Given that the current research was confined to the academic register, it would also be interesting to explore the distribution of adjectives and verbs in the investigated pattern across different types of both written and spoken registers, in view of the possible existence of slight variations in their occurrence. Future research, therefore, may determine the most strongly attracted co-varying collexeme pairs of the construction in other sections of COCA. REFERENCES Biber, D., S.Johansson, G.Leech, S.Conrad and E.Finegan (1999) Longman Grammar of Spoken and Written English. Harlow: Pearson Education. Bybee, J.J. (2010) Language, Usage, and Cognition. Cambridge: Cambridge University Press. Bybee, J.J. (2013) Usage-based theory and exemplar representations of constructions. In: Hoffmann, T. and G.Trousdale (eds) Handbook of Construction Grammar, 49–92. Oxford: Oxford University Press. Calude, A.S. (2008) Clefting and extraposition in English. ICAME Journal 32, 7–34. Collins, P. (1994) Extraposition in English. Functions of Language 1, 7–24. Emonds, J. (1972) Areformulation of certain syntactic transformations. In: Peters, S. (ed.) Goals of Linguistic Theory, 21–62. Englewood Cliffs, NJ: Prentice-Hall. Erdmann, P. (1990) Discourse and Grammar: Focusing and Defocusing in English. Tübingen: Niemeyer. Fillmore, Ch. J. (1982) Frame semantics. In: The Linguistic Society of Korea (eds) Linguistics in the Morning Calm, 111–137. Seoul: Hanshin Publishing Company. Fillmore, Ch. J. and B.T. Atkins (1992) Toward aframe-based lexicon: The semantics of RISK and its neighbors. In: Lehrer, A. and E.F.Kittay (eds) Frames, Fields and Contrasts, 75–102. Hillsdale, New Jersey: Lawrence Erlbaum Assoc. Fillmore, Ch. J. and C.Baker (2010) Aframes approach to semantic analysis. In: Heine, B. and H.Narrog (eds) The Oxford Handbook of Linguistic Analysis, 313–340. Oxford: Oxford University Press. Fillmore, C.J., C.R.Johnson and M.R.L.Petruck (2003) Background to FrameNet. International Journal of Lexicography 16(3), 235–250. Fontenelle, T. (ed.) (2003) Special issue on FrameNet and frame semantics. International Journal of Lexicography 16(3), 231–385. Francis, G. (1993) Acorpus-driven approach to grammar: principles, methods and examples. In: Baker, M., G.Francis and E.TogininiBognelli (eds) Text and Technology: In Honour of John Sinclair, 137–156. Amsterdam: John Benjamins. Goldberg, A. (1995) Constructions: AConstruction Grammar Approach to Argument Structure. Chicago: Chicago University Press. Goldberg, A. (2006) Constructions at Work. The Nature of Generalization in Language. Oxford: Oxford University Press. Goldberg, A. (2013) Constructionist approaches to language. In: Hoffmann, T. and G.Trousdale (eds) Handbook of Construction Grammar, 15–48. Oxford: Oxford University Press. Gómez-González, M.A. (2001) The Theme-Topic Interface. Evidence from English. Amsterdam and Philadelphia: John Benjamins.
JAROSłAW WILIńSKI 27 Gómez-González, M.A. (1997) On subject itextrapositions: Evidence from present-day English. Revista Alicante de Estudios Ingleses 10, 95–107. Gries, S.Th. and A.Stefanowitsch (2004a) Extending collostructional analysis: Acorpus-based perspective on alternations. International Journal of Corpus Linguistics 9 (1), 97–129. Gries, S.Th. (2012) Frequencies, probabilities, association measures in usage-/exemplarbased linguistics: Some necessary clarifications. Studies in Language 36, 477–510. Gries, S.Th. (2015) More (old and new) misunderstandings of collostructional analysis: on Schmid and Küchenhoff (2013). Cognitive Linguistics 26 (3), 505–536. Herbst, T., D.Heath., I.F.Roe and D.Götz (2004) AValency Dictionary of English: ACorpus-Based Analysis of the Complementation Patterns of English Verbs, Nouns and Adjectives. Berlin: Mouton de Gruyter. Herriman, J. (2000a) Extraposition in English: AStudy of the Interaction between the Matrix Predicate and the Type of Extraposed Clause. English Studies 6, 582–599. Herriman, J. (2000b) The Functions of Extraposition in English Texts. Functions of Language 7(2), 203–230. Hewings, M. and A.Hewings (2002) It is interesting to note that …: Acomparative study of anticipatory ‘it’ in student and published writing. English for Specific Purposes 21, 367–383. Hilpert, M. (2014) Collostructional analysis: Measuring associations between constructions and lexical elements. In: Glynn, D. and J.Robinson (eds) Corpus Methods for Semantics: Quantitative Studies in Polysemy and Synonymy, 7–38. Amsterdam & Philadelphia: John Benjamins Publishing Company. Hoey, M. (2000) Persuasive rhetoric in linguistics: Astylistic study of some features of the language of Noam Chomsky. In: Hunston, S. and G.Thompson (eds) Evaluation in Text. Authorial Stance and the Construction of Discourse, 28–37. Oxford: Oxford University Press. Huddleston, R. (1971) The Sentence in Written English. Cambridge: Cambridge University Press. Huddleston, R. (1984) Introduction to the Grammar of English. Cambridge: Cambridge University Press. Kaatari, H. (2010) Complementation of Adjectives: ACorpus-Based Study of Adjectival Complementation by thatand to-Clauses. Unpublished MA thesis. Uppsala University, Department of English. Kaltenböck, G. (2000) It-extraposition and nonextraposition in English discourse. In: Mair, C. and M.Hundt (eds) Corpus linguistics and linguistic theory, 157–175. Amsterdam and Atlanta: Rodopi. Kaltenböck, G. (2003) On the syntactic and semantic status of anticipatory it. English Language and Linguistics 7(2), 235–255. Kaltenböck, G. (2004) It-Extraposition and Nonextraposition in English: AStudy of Syntax in Spoken and Written Texts. Wien: Braumüller. Kaltenböck, G. (2005) It-extraposition in English: Afunctional view. International Journal of Corpus Linguistics 10 (2), 119–159. Langacker, R. (1987) Foundations of Cognitive Grammar. Theoretical Prerequisites. Volume I.Stanford, CA: Stanford University Press. Mair, C. (1990) Infinitival Complement Clauses in English. Cambridge: Cambridge University Press. McCawley, J.D. (1988) The Syntactic Phenomena of English. Volume II. Chicago: The University of Chicago Press. Miller, P.H. (2001) Discourse constraints on (non)extraposition from subject in English. Linguistics 39 (4), 683–701. Mindt, I. (2011) Adjective Complementation: An Empirical Analysis of Adjectives Followed by thatClauses. Philadelphia: John Benjamins. Pérez-Guerra, J. (1998) Integrating rightdislocated constituents: Astudy on cleaving and extraposition in the recent history of the English language. Folia Linguistica Historica XIX, 7–25. Perek, F. (2014) Rethinking constructional polysemy: The case of the English conative construction. In: Glynn, D. and J.Robinson
28 LINGUISTICA PRAGENSIA 1/2019 (eds) Corpus Methods for Semantics: Quantitative Studies in Polysemy and Synonymy, 61–86. Amsterdam & Philadelphia: John Benjamins Publishing Company. Petruck, M.R.L. (1996) Frame semantics. In: Verschueren, J., J.-O. Östman, J.Blommaert and C.Bulcaen (eds) Handbook of Pragmatics, 1–13. Philadelphia: John Benjamins. Quirk, R., S.Greenbaum, G.Leech and J.Svartvik (1985) AComprehensive Grammar of the English Language. New York and London: Longman. Rosenbaum, P.S. (1967) The Grammar of English Predicate Complement Constructions. Cambridge, MA: M.I.T. Press. Rowley-Jolivet, E. and S.Carter-Thomas (2005) Genre awareness and rhetorical appropriacy: Manipulation of information structure by NS and NNS scientists in the international conference setting. English for Specific Purposes 24, 41–64. Schmid, H.-J. and H.Küchenhoff (2013) Collostructional analysis and other ways of measuring lexicogrammatical attraction: Theoretical premises, practical problems and cognitive underpinnings. Cognitive Linguistics 24(3), 531–577. Seppänen, A., C.G.Engström and R.Seppänen (1990) On the so-called anticipatory It. Zeitschrift für Phonetik, Sprachwissenschaft und Kommunikationsforschung 43, 748–776. Seppänen, A. (1999) Extraposition in English revisited. Neuphilologische Mitteilungen 100: 51–66. Stefanowitsch, A. and S.Th. Gries (2003) Collostructions: Investigating the interaction between words and constructions. International Journal of Corpus Linguistics 8, 209–243. Stefanowitsch, A. and S.Th. Gries (2005) Covarying collexemes. Corpus Linguistics and Linguistic Theory 1 (1), 1–43. Stefanowitsch, A. (2013) Collostructional analysis. In: Hoffmann, T. and G.Trousdale (eds) Handbook of Construction Grammar, 290–306. Oxford University Press. Van linden, A. (2012) Modal Adjectives: English Deontic and Evaluative Constructions in Synchrony and Diachrony. Berlin and Boston: Walter de Gruyter. Wiliński, J. (2017) Normal and anomalous occurrences of adjectives in extraposed constructions with to-infinitive clauses: Aquantitative corpus-based study. In: Wiliński, J. and J.Stolarek (eds) Norm and Anomaly in Language, Literature and Culture, 89–106. Frankfurt am Main: Peter Lang. SOURCES AND TOOLS The Corpus of Contemporary American English (COCA). The full-text data (1990–2012). Available from https://www.corpusdata.org/ purchase.asp The FrameNet project. Available from https:// framenet.icsi.berkeley.edu/fndrupal/ MonoConc Pro (MP 2.2). Available from http:// www.athel.com/mono.html Fisher’s Exact Test. Available from http://www. langsrud.com/fisher.htm Jarosław Wiliński Siedlce University of Natural Sciences and Humanities Wydział Humanistyczny Uniwersytetu Przyrodniczo-Humanistycznego w Siedlcach ul. Żytnia 39, 08-110 Siedlce ORCID ID: 0000-0002-3136-6529 e-mail: [email protected]