scieee AI-readable full text Open interactive document viewer

Lexical Enrichment and Structural Complexity in the Acquisition of English as a First Language: Ditransitive, Causative and Relative Constructions

Proroković, Jakov

Abstract

This doctoral dissertation investigates various developmental parameters in the context of first language acquisition through a comparative triconstructional analysis (ditransitive, causative, and relative constructions) in children whose first language is English. The research is based on corpus analysis, during which the mentioned constructions were extracted and analysed from the English part of the CHILDES corpora. For a comprehensive interpretation and understanding of first language acquisition, intra-linguistic evidence (i.e., language phenomena related to syntax and lexicon) is observed and compared across three age-defined groups (0-3, 4-6, and 18+). Given the specificities of the three observed constructions, the analyses are tailored to their characteristics (such as frequency, complexity, and variability of certain lexemes and structural patterns appearing in these constructions). Simultaneously, they are conducted to allow not only inter-age group comparisons but also inter-constructional comparisons (such as the degree of similarity in the representation and frequency rankings of certain lexical units, as well as measures like structural and lexical complexity, including the mean length of utterance, lexical density, and type-token ratios). The main contribution of this research is the methodologically unique way of monitoring the parallel development of children's speech at the lexical, morphological, and syntactic levels, which analytically and interpretatively unifies the mentioned constructions through the lens of cognitive theories of language acquisition that are primarily based on the ‘usage-based’ model. Among other things, the research results show what can be described as a linguistic version of Pareto's principle, where all three age groups in the production of the observed constructions largely rely on a few lexical units in certain construction slots, and where the majority of their linguistic expression when it comes to said constructions is occupied by a relatively small number of structural patterns. Significant lexical and structural similarities among the same constructions were confirmed at several levels across the observed age groups, along with a more “conservative” language use and a more pronounced concentration of certain items and syntactic patterns with the declining of age. Furthermore, the ‘item-based’ claims about “simpler” constructions being more lexically restricted production-wise than more “complex” constructions was not confirmed; rather, an examination of certain lexical items used in specific construction slots showed that item-based tendencies are present in all constructions across all three age groups. In conclusion, the data on frequency and lexical and syntactic complexity obtained through comparative analysis of child and adult speech contribute to existing body of research on the nature of language acquisition and development.

Full text

SVEUČILIŠTE U ZADRU POSLIJEDIPLOMSKI SVEUČILIŠNI STUDIJ HUMANISTIČKE ZNANOSTI Jakov Proroković LEXICAL ENRICHMENT AND STRUCTURAL COMPLEXITY IN THE ACQUISITION OF ENGLISH AS A FIRST LANGUAGE: DITRANSITIVE, CAUSATIVE AND RELATIVE CONSTRUCTIONS Doktorski rad Zadar, 2024. SVEUČILIŠTE U ZADRU POSLIJEDIPLOMSKI SVEUČILIŠNI STUDIJ HUMANISTIČKE ZNANOSTI Jakov Proroković LEXICAL ENRICHMENT AND STRUCTURAL COMPLEXITY IN THE ACQUISITION OF ENGLISH AS A FIRST LANGUAGE: DITRANSITIVE, CAUSATIVE AND RELATIVE CONSTRUCTIONS Doktorski rad Mentor izv. prof. dr. sc. Marco Angster Zadar, 2024. SVEUČILIŠTE U ZADRU TEMELJNA DOKUMENTACIJSKA KARTICA I. Autor i studij Ime i prezime: Jakov Proroković Naziv studijskog programa: Poslijediplomski sveučilišni studij Humanističke znanosti Mentor/Mentorica: izv. prof. dr. sc. Marco Angster Datum obrane: 11. listopada 2024. Znanstveno područje i polje u kojem je postignut doktorat znanosti: humanističke znanosti, filologija II. Doktorski rad Naslov: Lexical Enrichment and Structural Complexity in the Acquisition of English as a First Language: Ditransitive, Causative and Relative Constructions UDK oznaka: 811.111'232 Broj stranica: 265 Broj slika/tablica: 21/47 Broj bilježaka: 106 Broj korištenih bibliografskih jedinica i izvora: 396 Broj priloga: 2 Jezik rada: Engleski III. Stručna povjerenstva Stručno povjerenstvo za ocjenu doktorskog rada: 1. doc. dr. sc. Frane Malenica, predsjednik/predsjednica 2. prof. dr. sc. Marijan Palmović, član/ica 3. izv. prof. dr. sc. Mojca Kompara Lukančič, član/ica Stručno povjerenstvo za obranu doktorskog rada: 1. doc. dr. sc. Frane Malenica, predsjednik/predsjednica 2. prof. dr. sc. Marijan Palmović, član/članica 3. izv. prof. dr. sc. Mojca Kompara Lukančič, član/članica UNIVERSITY OF ZADAR BASIC DOCUMENTATION CARD I. Author and study Name and surname: Jakov Proroković Name of the study programme: Postgraduate doctoral study programme in Humanities Mentor: Associate Professor Marco Angster, PhD Date of the defence: 11 October 2024 Scientific area and field in which the PhD is obtained: Humanities, Philology II. Doctoral dissertation Title: Lexical Enrichment and Structural Complexity in the Acquisition of English as a First Language: Ditransitive, Causative and Relative Constructions UDC mark: 811.111'232 Number of pages: 265 Number of figures/tables: 21/47 Number of notes: 106 Number of used bibliographic units and sources: 396 Number of appendices: 2 Language of the doctoral dissertation: English III. Expert committees Expert committee for the evaluation of the doctoral dissertation: 1. Assistant Professor Frane Malenica, PhD, chair 2. Full Professor Marijan Palmović, PhD, member 3. Associate Professor Mojca Kompara Lukančič, PhD, member Expert committee for the defence of the doctoral dissertation: 1. Assistant Professor Frane Malenica, PhD, chair 2. Full Professor Marijan Palmović, PhD, member 3. Associate Professor Mojca Kompara Lukančič, PhD, member Izjava o akademskoj čestitosti Ja, Jakov Proroković, ovime izjavljujem da je moj doktorski rad pod naslovom Lexical Enrichment and Structural Complexity in the Acquisition of English as a First Language: Ditransitive, Causative and Relative Constructions rezultat mojega vlastitog rada, da se temelji na mojim istraživanjima te da se oslanja na izvore i radove navedene u bilješkama i popisu literature. Ni jedan dio mojega rada nije napisan na nedopušten način, odnosno nije prepisan iz necitiranih radova i ne krši bilo čija autorska prava. Izjavljujem da ni jedan dio ovoga rada nije iskorišten u kojem drugom radu pri bilo kojoj drugoj visokoškolskoj, znanstvenoj, obrazovnoj ili inoj ustanovi. Sadržaj mojega rada u potpunosti odgovara sadržaju obranjenoga i nakon obrane uređenoga rada. Zadar, 18. listopada 2024. In Appreciation I would like to express my heartfelt gratitude to my family—my wife, my parents, and my friends—who stood by me throughout the journey of creating this dissertation. Their patience and understanding during the more challenging periods of this work meant the world to me. I am also deeply thankful to my Department, which has been a steadfast source of support since I embarked on this campaign, to my former professors who inspired me long before I considered pursuing a PhD, and, finally, to my mentor, whose expertise and guidance were invaluable throughout this process. To my baby boy! Project Acknowledgment This work has been realised under the institutional project Multilingualism: between theory and empirical knowledge (ViTE – Višejezičnost: između teorije i empirije) (IP-01-2023-14), fully financed by the University of Zadar. Table of contents 1. INTRODUCTION: CORPUS AND COGNITIVE LINGUISTICS................................... 1 1.1. General introduction .................................................................................................... 1 1.2. Cognitive Linguistics ................................................................................................. 10 1.2.1. The framework of Cognitive Linguistics ......................................................... 10 1.2.2. Construction Grammar ..................................................................................... 13 1.2.3. Usage-based linguistics .................................................................................... 15 1.3. Corpus linguistics ...................................................................................................... 18 1.3.1. Corpus-based and corpus-driven methods ....................................................... 18 1.3.2. Corpus linguistics and Cognitive Linguistics ................................................... 21 1.3.3. Extracting the data – corpus as a method ......................................................... 22 1.4. Three constructions – the increase in complexity ..................................................... 24 2. RESEARCH PURVIEW AND METHODOLOGY ......................................................... 31 2.1. Exploratory scope of the research ............................................................................. 31 2.1.1. Structural similarity and complexity ................................................................ 31 2.1.2. Animacy ........................................................................................................... 32 2.1.3. Lexical richness and similarity ......................................................................... 33 2.2. Research questions and hypotheses ........................................................................... 42 2.2.1. Research questions ........................................................................................... 42 2.2.2. Hypotheses ....................................................................................................... 44 2.3. Method ....................................................................................................................... 46 2.3.1. Approaching the corpus ................................................................................... 46 2.3.2. Sample .............................................................................................................. 47 3. DITRANSITIVE CONSTRUCTION: CORPUS-BASED COMPARISON OF CHILD AND ADULT SPEECH ........................................................................................................... 50 3.1. Operationalization and previous research .................................................................. 50 3.1.1. Defining ditransitives ....................................................................................... 50 3.1.2. Ditransitive verbs vs. ditransitive constructions .............................................. 55 3.1.3. Previous research .............................................................................................. 59 3.2. Ditransitive constructions – results............................................................................ 65 3.2.1. Extraction method and the examples of ditransitive constructions .................. 65 3.2.2. Structural complexity and similarity ................................................................ 68 3.2.3. Lexical arrangement and overlap ..................................................................... 74 3.2.4. Animacy ........................................................................................................... 84 3.2.5. Lexical richness ................................................................................................ 87 3.3. In sum ........................................................................................................................ 93 4. PERIPHRASTIC CAUSATIVE CONSTRUCTION: CORPUS-BASED COMPARISON OF CHILD AND ADULT SPEECH ....................................................................................... 96 4.1. Operationalization and previous research .................................................................. 96 4.1.1. Defining causation ............................................................................................ 96 4.1.2. Analytic/Periphrastic Causatives .................................................................... 100 4.1.3. Previous Research .......................................................................................... 109 4.1.3.1. Causative alternation and periphrastic causatives ...................................... 109 4.1.3.2. The acquisition of periphrastic causatives ................................................. 111 4.2. Periphrastic causative constructions – results ......................................................... 117 4.2.1. Extraction method and the examples of periphrastic causatives .................... 117 4.2.2. Structural complexity and similarity .............................................................. 121 4.2.3. Lexical arrangement and overlap ................................................................... 129 4.2.4. Animacy ......................................................................................................... 136 4.2.5. Lexical richness .............................................................................................. 140 4.3. In sum ...................................................................................................................... 143 5. RELATIVE CONSTRUCTION: CORPUS-BASED COMPARISON OF CHILD AND ADULT SPEECH .................................................................................................................. 146 5.1. Operationalization and previous research ................................................................ 146 5.1.1. Defining relative clauses and sentences ......................................................... 146 5.1.2. Defining the relative construction .................................................................. 154 5.1.3. Previous research ............................................................................................ 158 5.2. Relative constructions – results ............................................................................... 166 5.2.1. Extraction method and the examples of relative constructions ...................... 166 5.2.2. Structural complexity and similarity .............................................................. 169 5.2.3. Lexical arrangement and overlap ................................................................... 181 5.2.4. Animacy ......................................................................................................... 187 5.2.5. Lexical richness .............................................................................................. 190 5.3. In sum ...................................................................................................................... 194 6. CONVERGENCE OF DATA AND FINAL DISCUSSION.......................................... 197 6.1. Inter-constructional data .......................................................................................... 197 6.2. Final discussion ....................................................................................................... 214 7. CONCLUSION ............................................................................................................... 221 8. BIBLIOGRAPHY ........................................................................................................... 223 8.1. References: .............................................................................................................. 223 8.2. CHILDES corpora included in the research: ........................................................... 246 9. ABSTRACT .................................................................................................................... 251 9.1. Abstract (in English) ................................................................................................ 251 9.2. Abstract (in Croatian) .............................................................................................. 253 10. APPENDIX ................................................................................................................. 255 10.1. Queries used to target constructions in the corpus ............................................. 255 10.2. Additional data on constructions ........................................................................ 262 List of tables Table 1. Measures of lexical richness considered in this study ............................................................. 40 Table 2. Examples of (candidate) ditransitive constructions found in the CHILDES corpora ............. 66 Table 3. Distribution of different ditransitive patterns across age groups (relative frequencies normalised per million tokens and rounded to the closest unit) ............................................................................... 69 Table 4. Examples of the produced ditransitive constructions in the CHILDES corpora (patterns from Table 3 indicated in square brackets) .................................................................................................... 69 Table 5. Spearman’s rank correlation of the observed ditransitive patterns across age groups ............ 70 Table 6. Distribution of different ditransitive patterns across age groups (relative frequencies normalised per million tokens; some patterns are merged together)........................................................................ 72 Table 7. Distribution of ditransitive verbs across age groups ............................................................... 75 Table 8. Spearman’s rank correlation of the most frequent ditransitive verbs across age groups......... 79 Table 9. Spearman’s rank correlation of the most frequent ditransitive verbs across age groups (without the verbs ‘tell’ and ‘show’) ................................................................................................................... 80 Table 10. Most frequent first (indirect) objects in give-NP-NP ............................................................ 81 Table 11. Most frequent second (direct) objects in give-NP-NP .......................................................... 81 Table 12. First (indirect) object animacy .............................................................................................. 86 Table 13. General overview of the sentences containing give-NP-NP across age groups .................... 88 Table 14. Lexical richness for the ditransitive give-NP-NP .................................................................. 91 Table 15. Examples of (candidate) periphrastic causative constructions found in the CHILDES corpora ............................................................................................................................................................. 119 Table 16. Distribution of causative constructions across age groups (relative frequencies normalised per million tokens and rounded to the closest unit) ................................................................................... 123 Table 17. Examples of the produced periphrastic causative constructions in the CHILDES corpora (patterns from Table 16 indicated in square brackets) ........................................................................ 123 Table 18. Spearman’s rank correlation of the observed causative patterns across age groups ........... 124 Table 19. Observed frequencies of particular verbs in the pattern V NP to V (excluding get and help) ............................................................................................................................................................. 126 Table 20. Distribution of causative constructions across age groups (relative frequencies normalised per million tokens; some patterns are merged together) ............................................................................ 128 Table 21. Main verbs in the periphrastic causatives ............................................................................ 129 5 which constructions best accommodate certain verbs: (1) there are specific semantic constraints that apply to certain English constructions and verbs (see Pinker, 1989), which children learn at some point, although it is not completely clear how; (2) the more frequently children hear a verb used in a particular construction, the less likely it will be for them to use it in a novel construction (Clark, 1987; Bates & MacWhinney, 1989 ; Braine & Brooks, 1995 and others) and (3) if children hear a verb used in a construction that serves the same purpose as a possible generalisation, they may recognize the non-canonical nature of that particular generalisation. 8 This thesis does not necessarily put the focus entirely on verbs, but rather on the syntactic patterns within constructions along with distributional data pertaining to the lexical items in those construction, whether nouns or verbs. Nevertheless, similar logic can be applied to both nouns and constructions, that is, it is possible to rephrase the Tomasello's observation to encapsulate not only verbs, but nouns as well, if not even constructions in their entirety (as they represent and behave like items in their own respect). Aside from verbs, there may be constraints that children learn concerning the use of nouns and patterns, the frequency of these may facilitate the acquisitional process, and they may avoid using certain items and patterns which would otherwise appear sound to them – but due to lack thereof in the input – resort to those more frequent, even when they do not necessarily represent the “easier” route. In this study, most of the data is interpreted through the lens of frequency and inter-group correspondence, 9 but this does not mean that the interpretation of the results should solely be restricted to the most frequent items in the input and their relevance; that is, it takes into account structures and items that are overrepresented, as well as those that are underrepresented or marked. 10 A quote by Ibbotson and Tomasello (2009, p. 80) wonderfully outlines the way in 8 In some sense, the last claim implies that input trumps generalisation; e.g. a child might expect a construction as He danced his friend but after hearing He made his friend dance, infer that the verb ‘dance’ does not occur in transitive constructions – particularly in view of the fact that the adult makes an effort to circumvent the “irregular” use of the verb by using the more marked (periphrastic causative) alternative (see Tomasello, 2011). 9 The data obtained in this research can be observed through several different lenses: intra-constructional, interconstructional, intra-age group and inter-age group observations. Inter-age group comparions are no doubt at the focus of this research, but they often cannot be interpreted in isolation from other construction and age-specific observations. In this sense, intra-constructional remarks delve into lexical and structural properties of the constructions observed primarily in relation to patterns occuring in them and the corresponding frequency trends, whereas inter-constructional observations concern differences across constructions, and not patterns occuring within them (see Footnote 1). Intra-age group data is a term used to capture all construction-specific results but within one age group, whereas inter-age group comparison takes all of the aforementioned into account to gain a better insight into developmental differences between the observed age groups in the language they produced. 10 In linguistics, markedness as a concept has been referred to different phenomena in linguistics, but recently it tends to be used as the designation of cases which go against our expectations of what they should look like. A phenomenon is more marked than another if our general expectations regarding its possibility of occurring are 6 which the language acquisition process should be approached through the lens of frequency, and serves as an instruction for interpreting the results of this study: “And so our proposal is actually that frequency and maximal marking play very different roles in acquisition. Children will acquire first the instantiation of the construction they hear most frequently. But then they will bring together different instantiations of the construction on the basis of similar function, such that a prototype is formed including all of the marking options. And so we might actually propose two different routes children might use to get to their own prototype representation of a linguistic construction (as well as various possible mixtures of these strategies).” Some approaches to language acquisition focus on the end result, while others emphasize the process of how children acquire language. Linguists tend to focus on the end result, such as the grammatical patterns, while psycholinguists are more interested in the ways that speakers access the pertinent structures and produce the utterance in a stepwise procedure by which the utterance is produced (Clark, 2009, pp. 10-12). In other words, psycholinguistic approaches to language acquisition tend to focus on the internal, cognitive processes behind using particular constructions such as the relative construction. These aspects of language acquisition remain beyond the scope of this study, although they are sometimes referred to as possible interpretation tools for the data retrieved. Instead, the focus of this study is on the actual language use in the context of language development and acquisition. Given that the emphasis is on the three constructions in terms of how they are realized, the approach undertaken in the thesis no doubt belongs more to the linguistic domain, though it is difficult to isolate the issues completely; the questions addressed often (at least to a degree) concern the pragmatics of language use, the need to use the observed patterns, the process of accessing and forming the constructions in question. Nevertheless, the linguistic approach is evident in the fact that the final product is in the focus of the study, i.e. the representation of the observed constructions in everyday speech with respect to their internal syntactic and semantic features. This is also the point where the thesis steps into an explorative domain where several different frameworks converge into one; that is, it mostly remains under the scope of Cognitive Linguistics, making lesser, and thus closely correlates with frequency and distribution, but also conceptual and structural complexity (Gaeta, 2017). 7 use of the ideas such as the ‘usage-based’ and Construction Grammar (CxG) 11 approaches when it comes to interpreting language patterns and language production, but it also introduces corpus-based methodology as a tool to observe actual language use, and makes claim about the value of such findings. In solely observing the language produced, the corpus-based approach emphasizes the linguistic aspect of acquisition, while taking into consideration the social setting in which the language is produced, as well as the cognitive foundations that the observed age groups build upon (cf. Núñez, 2007). The comparison of early and adult language allows for conjectures about the processing of external linguistic input and input reproduction. Aa a method, a corpus-based research entails the collection and analysis of large amounts of naturally occurring language data, or 'corpora', in order to study patterns of language use and meaning. One of the main advantages of corpus data is that it reflects how language is actually used in real-world contexts, rather than being based on artificially constructed examples or experimental tasks (see Section 1.3 for more details). On the other hand, the limitations are obvious in the sense that the researcher always has to bear in mind the difference between language competence and language production. 12 Still, the strongest argument in favour of corpus data is that it allows researchers to study a large and diverse sample of language, which is potentially useful for studying language development. Children's language development can be difficult to study experimentally, as it involves a wide range of factors and is influenced by a child's individual experiences and environment, whereas corpus data provides a rich source of naturally occurring language data that can be analysed to investigate how children's language develops over time, and how it is influenced by different factors. In the context of this study, the research places considerable emphasis on the comparative data, while emphasizing (dis)similarity and correlation solely in linguistic terms as children move towards adult-like utterances. Overall, the study also argues for the position that such approaches to actual 11 Note that the term Construction Grammar tends to orthographically vary across different papers and studies, and though it tends to be written with initial capital letters, it can also sometimes take the uncapitalized form – especially when the author wants to emphasize the fact that it represents a family of approaches (sometimes dubbed “flavours”) which share some core underlying principles, rather than a unified monolithic theoretical strain (cf. Goldberg & Suttle, 2010). It has also been referred to as a framework or a field in its own respect (for a recent example, see Ungerer & Hartmann, 2023). 12 It needs to be noted that the claims on the usefulness of actual language use data, and the ability to generalise on it (no matter how restricted) is not necessarily a radical proposition. For example, Stefanowitsch (2020, p. 4) has argued that if we believe linguistic competence to be at least broadly reflected in linguistic performance, then it should be possible to model linguistic knowledge based on observations of language use. In this sense, corpus linguistics is in the same situation as any other empirical science with respect to deducing underlying principles from specific manifestations influenced by other factors. 8 language data can serve as a valuable tool for investigating language development because they provide a comprehensive and realistic view of language use, but also allow researchers to study a large and diverse sample of language data. This study is based on the family of corpora called CHILDES (Child Language Data Exchange System), that is, on those studies which were conducted among children whose first language is English (see the full list of CHILDES corpora sources used under Section 8.2). CHILDES is currently defined as an organized database that contains written records of spoken language (see Section 2.3.2 for more) which were mostly contributed by researchers who collected the data for their own research purposes (cf. Corrigan, 2012; MacWhinney, 2015). Nonetheless, the type of research such as the one employed for the purposes of this thesis is not a complete novelty, with several studies having already utilized CHILDES corpora as a secondhand source of data to compare parental input and child language (see Liu et al., 2008; Goodman et al., 2008; Li & Fang, 2011; Shin, 2022). For example, a study by Li and Fang (2011), who utilized the Manchester corpus data (part of CHILDES), which consists of transcripts of English-speaking children aged between 1 year and 8 months to 2 years and 25 days, aimed to explore the relationship between the frequency of words in the language children hear from their mothers (input) and the words they produce themselves (output). Apart from having noted that children use more words with concrete and easily visualized meanings, the researchers also observed that the mere quantity of words spoken to children by their mothers was not significantly correlated with the number of words acquired by the children, which according to them, indicated that the comprehensibility of the input is crucial children’s language acquisition; i.e. it is essential to make sure they understand what is being said to them, rather than just having them mimic sounds or words. A considerable overlap in caregiver input and child production when it comes to construction types expressing transitive events was observed by Shin (2022), while Goodman et al. (2008) found higher parental frequency to be associated with earlier acquisition within specific lexical categories, but also noted that frequency interacts with category, modality and developmental stage in its impact of the acquisition of vocabulary. Liu, Zhao and Li (2008) conducted corpus research that examined lexical composition patterns within the vocabularies of children and their caregivers in three different language groups (English, Mandarin, and Cantonese) in CHILDES family of corpora on age groups spanning from 13 to 60 months. They highlighted that language-specific differences in the linguistic input significantly influenced children's language output. This influence was reflected in the varying 9 percentages of nouns, verbs, and adjectives present in a child’s productive lexicon at different developmental stages, indicating also that, as children continued to grow and develop, their lexical composition patterns gradually became more similar to those of their parents. Again, any evidence based on such corpora needs to be interpreted with care since it represents data collected from spoken records. There may be a considerable disparity between what speakers can produce, and what they do produce, given their language abilities. Nevertheless, the primary advantage of CHILDES is twofold: it contains transcriptions of natural conversations between children and parents, and the congregated corpora of numerous child language studies is compelling in its size. The English corpora alone comprise more than 22 million word tokens, which enables reliable generalisations to be made based on the data. The sheer size of such corpora, and the improbability of collecting reproducible and retestable findings from the accidental data, allow us to make reasonable inferences on linguistic competence and not just performance. All in all, the primary aim of the study is to investigate the lexical and structural similarities among different age groups in terms of their preferred lexical and syntactic choices within observed constructions. This primarily includes determining the extent of similarity in verb and noun frequency across these constructions, examining whether there is a discernible difference in similarity between the 0-3 age group and the 4-6 age group, and specifically assessing if the 0-3 age group exhibits greater similarity to child-directed speech than the 4-6 age group. The thesis also intends to identify any differences in terms of animacy concerning key lexical items (nouns) used within the constructions. Measures of lexical richness and complexity are also included for the purposes of reflecting differences in language use across age groups. Finally, the data should also shed light on the variation in the degree of item-based character across said constructions. The first chapter of the thesis serves as an introduction to the general framework of the research, providing key points about the theoretical background exploited in the study design and data interpretation (see incoming subsections of Chapter 1). Chapter 2 contains several important subsections that delve into the exploratory scope of the research, outlining the types of analyses to be conducted on the corpus data, the research questions and hypothesis, corpus description and the method by which the constructions were recovered from the transcriptions. The research results on the three constructions observed – ditransitives, periphrastic causatives, and relatives – are outlined in chapters 3, 4, and 5, respectively. Each of them deals with the 10 constructions separately, but in a rather systematic fashion; each construction-specific chapter can be further subdivided into sections that operationalize the construction observed, recapitulate the previous research on the topic, and lay out the results of the corpus analysis. For each of the constructions, there is an additional section which summarizes the main findings of the examination and reflects on the stated research questions and hypotheses. Chapter 6 moves away from intra-constructional data towards inter-constructional data and discusses the implications of these findings on the stated research questions and the overall understanding of this heterogenous and comprehensive dataset. Finally, it presents study limitations that are to be considered by the reader of this thesis, after which the conclusion (Chapter 7) once again responds to the stated research objectives. 1.2. Cognitive Linguistics 1.2.1. The framework of Cognitive Linguistics Cognitive Linguistics is a field that emerged in the 1970s as response to the premises and inferences of both Generative Grammar and Montague Grammar (Evans, 2012, p. 130). The framework of Cognitive Linguistics started to form and evolve around the names such as George Lakoff, Charles Filmore, Leonard Talmy, Ronald Langacker and others (cf. Lakoff, 1982; Lakoff & Thompson, 1975; Filmore, 1975; Talmy, 1978; Langacker, 1987; etc.). Hence, Cognitive Linguistics emerged as not as a “single, closely articulated, theory”, but rather as a “broad theoretical and methodological enterprise” (Evans, 2012, p. 130). Similarly to Evans, Geeraerts (2006, p. 2) also calls it a “flexible framework rather than a single theory of language”, further stating that “it constitutes a cluster of many partially overlapping approaches rather than a single well-defined theory that identifies in an all-or-none fashion whether something belongs to Cognitive Linguistics or not.” Perhaps his most apt description of Cognitive Linguistics is that of a “theoretical conglomerate”, consisting of theories which diverge from traditional approaches primarily in their treatment of the interrelationship between grammar and meaning. Although this multiplicity and intermixture of various emerging approaches at the time might be confused with inconsistency, this is where, in fact, lies the strength of this theory. Despite the fact that a wide range of claims has been made within this framework, they are typically oriented towards different subsections of the language phenomenon, and they mostly complement one another. This thesis exploits the accounts that 11 position themselves well on the theoretical and methodological line between cognitive and corpus linguistics. Sometimes, terminological confusions might arise in addressing of the terms cognitive and Cognitive Linguistics, and Geeraerts (2006) provide us with a clear distinction between the capitalized Cognitive Linguistics and the uncapitalized cognitive linguistics. According to him, uncapitalized cognitive linguistics marks any approach in linguistics which observes language as a cognitive phenomenon. In that sense, generative grammar also belongs to this broader category of cognitive linguistics because it “attributes a mental status to the language” as well, regardless of the fact that it views language as an autonomous system, described through formalized syntactic structures and rules. Cognitive Linguistics is taxonomically on the same level as generative grammar and represents only one approach within cognitive linguistics. Nevertheless, given that the term is now more frequently associated with the narrower definition of Cognitive Linguistics, and in the rest of the thesis, the references to the term ‘cognitive’ will be the ones relating to this ‘theoretical conglomerate’ that represents a move away from the previous approaches to grammar. 13 Considering the multiplicity of ideas that exist with the cognitive framework, there is a need to establish what exactly these ideas have in common. According to Evans (2012), we need to start with two commitments: the Cognitive Commitment and the Generalisation Commitment. The former represents the idea that cognitive linguistics deals with descriptions and models of language by taking into account other cognitive and brain sciences. The latter is a commitment of linguists to describe the knowledge as a by-product or a result of general cognitive abilities. In other words, in order to earn the label ‘cognitive’, the theory must be in sync with what is known about the way in which the brain functions, and keep track of the new discoveries made by the fields of psychology, neuroscience and even studies on artificial intelligence. As already stated, another important common trait of approaches within Cognitive Linguistics is the way in which meaning is approached. Language is primarily about meaning 13 Here, “previous” primarily refers to the generative-transformational approaches to grammar. Although there are shared points of agreement between certain theoretical strains within the Cognitive framework (such as Construction Grammar) and generative-transformational approaches to grammar (both recognize the language as part of the cognitive system), the most robust points of departure are that the advocates of the latter view formal structures as independent from semantic and pragmatic aspects of discourse, and that they argue for a factual necessity of hard-wired knowledge of otherwise (according to them) unlearnable grammar (see Goldberg, 2006, pp. 4-5). 12 and the information conveyed by a particular language expression is at the core of linguistic analysis, Geeraerts (2006) outlines the treatment of meaning in Cognitive Linguistics through four key assertions: (1) linguistic meaning is perspectival, (2) linguistic meaning is dynamic and flexible, (3) linguistic meaning is encyclopaedic and non-autonomous, and (4) linguistic meaning is based on usage and experience. That meaning is perspectival is evident in the fact that it both reflects and shapes the outside world. The way in which we express something is constrained by the conditions in which we find ourselves at the moment of speaking, and the interpretation of that expression must take context into account. That linguistic meaning is dynamic and flexible is evident in the way meaning tends to change and shape depending on the social or temporal context. The argument that it is encyclopaedic and non-autonomous is inspired by the idea that it cannot be isolated from the world in which we live, or from our cognitive capacities for that matter. It is encyclopaedic in that it depends on other knowledge that we have acquired thus far in the world. Finally, the idea that linguistic meaning is based on usage and experience is rooted in the way Cognitive Linguistics perceives the emergence of meaning – meaning of a particular expression or a linguistic item is attainable only through recurring experience of how it is used, in which context, for what purposes etc. In the initial stages of these new theoretical assumptions, the term frequently encountered was that of Cognitive Grammar (CG), which somewhat constitutes a bedrock of Cognitive Linguistics. Nevertheless, given its uniformity and specificity, CG separates itself as only one model of grammar within Cognitive Linguistics (there is also Construction Grammar and Word Grammar; see Geeraerts & Cuyckens, 2007); in other words, CG represents a particular linguistic theory within Cognitive Linguistics, that can be regarded as radical due to the claims such as that of grammar being completely symbolic (i.e. filled with meaning) or that fundamental grammatical notions such as ‘verb’, ‘noun’, or even ‘subject’, “have unified conceptual characterizations” (Langacker, 2010). Although the emphasis is often on the term ‘cognitive’, notion which has become entrenched in linguistics as an epitome for the claims of language being an essential manifestation of cognition, CG designates a theory that bridges the gap between cognition and interaction. Langacker states that the relationship between cognition and language is evident in the fact that it emerges from more general phenomena, such as perception, attention and categorization, and consequently cannot be autonomous: 13 “The concern with cognition—not shared by all strands of functionalism—is fully compatible with the latter’s emphasis on social interaction. It is only through interaction in a sociocultural context that language and cognition are able to develop. By the same token, an account of linguistic interaction cannot ignore the assessment by each interlocutor of the knowledge, intentions, and mental state of the other, as well as their apprehension of the overall discourse context. In its basic principles, CG (despite its name) strikes what is arguably a proper balance between cognitive and interactive factors.” (2010, p. 89) Cognitive Grammar is thus only a small part of this theoretical conglomerate; Geeraerts (2006) states that there are, in fact, 12 major parts that create the basis for Cognitive Linguistics. Apart from Cognitive Grammar, there is also Construction Grammar, usage-based linguistics, schematic network, prototype theory, conceptual metaphor, grammatical construal, radial network, image schema, metonymy, mental spaces and frame semantics. In the context of this thesis, the part related to theoretical background, operationalization and interpretation, mostly draws from studies in Construction Grammar and usage-based linguistics. Let us then briefly define what these two terms relate to and why they are relevant for this research. 1.2.2. Construction Grammar Much like Cognitive Grammar, the term Construction Grammar (CxG) is generally used as an umbrella term for a number of theoretical strains that share some key underlying assumptions pertaining to the lexico-syntactic continuum. One of the most prominent is the Goldbergian one, later known as Cognitive Construction Grammar (CCxG), which only represents one ‘type’ (see Boas, 2013, 2021). Other prominent accounts include Unification Construction Grammar (UCxG), Cognitive Grammar (CG) and Radical Construction Grammar (RCxG) (for an overview, see Goldberg, 2006, pp. 213-214). As can be assumed from the name, the priority is given to the term ‘construction’, or its more elaborate counterpart ‘grammatical construction’. Hence, the first radical idea of Construction Grammar, that distinguishes it from its competing theoretical frames, is a move in the perception of grammar, and its interrelationship with semantics. First, the construction as a term that represents a particular syntactic pattern, started to be treated as a primary unit of a language. The crucial notion here 14 is that of a ‘unit’, because it immediately hints towards the existence of something that is inherent solely to the ‘unit’ in question, i.e. there is something that can be ascribed to the construction as a whole. If, in fact, we start to treat a construction as an entity itself, we must immediately recognize that there are characteristics that cannot be deduced from its components, and thus the contribution of the construction to meaning (or something else) can be detected on a greater level than on that of its constituents. In other words, this is where Construction Grammar differs from generative so-called componential models, where words are typically interpreted as associations between meaning, syntactic category and form (Croft, 2007). In the early efforts to capture the distinctions between the emerging Construction Grammar and its rival theories, Charles Filmore (1988) laid out several arguments which he found in common with other syntacticians working within CxG: (1) Construction Grammar does not have transformations like transformational grammars, i.e. relationships that used to be captured by derivations of sentences and their structure are now captured by grammar as a whole; (2) Construction Grammar differs from simple-phrase structure grammar in that its structural categories may store complex bundles of information, and perhaps most importantly, it allows for a particular linguistic expression to be able to simultaneously instantiate more than one grammatical construction at the same level; (3) Construction Grammar differs from the generativist tradition in that it attempts to simultaneously describe syntactic patterns and the semantic and pragmatic purposes to which these patterns are dedicated. Probably, the most important distinguishing feature of Construction Grammar is its treatment of syntax and lexicon. Similarly to what has already been stated about the interrelationship between grammar and meaning within the framework of Cognitive Linguistics, lexicon is not treated separately from syntax. While in the generative tradition, lexicon is viewed as a component in itself, in the Construction Grammar it is viewed as a part of the continuum which encompasses both lexicon and syntax. The inevitable conclusion is then that the only difference between lexicon and constructions is in degree (Croft, 2007, p. 470-471). The difference between words and constructions is in syntactic complexity; constructions may be made up of words or phrases, while words are syntactically simple. According to Croft (ibid.), if we take an example of morphologically complex words, the interpretation of Construction Grammar 21 investigate (dis)similarities between adult and child language use, whilst laying groundwork for future investigations into the phenomenon of language acquisition. 1.3.2. Corpus linguistics and Cognitive Linguistics Given that this thesis mainly relies on the framework of Cognitive Linguistics, it appears useful to explain how this framework fits alongside corpus linguistics. First, as already noted, this study methodologically relies on corpus linguistics, i.e. although the term has been used to represent more than just mere methodology, this research builds upon theories surrounding first language acquisition and treats corpus approaches primarily as a means to an end. Although nowadays, the fact that corpora play a significant part in the research of linguists working within the cognitive framework might seem ordinary, this was not necessarily always the case. Cognitive Linguistics started to emerge around 1980s, and although it differed significantly in terms of its theoretical foundations and postulates from the traditional generative accounts, it still continued with the similar methodological restraints towards the linguistic data that is to be included in the research (Stefanowitsch 2011). It is often said that Cognitive Linguistics elaborates on the linguistic categories before engaging the data, that is, it defines categories beforehand, while corpus-driven linguistics induces categories from the data itself (Teubert, 2008). At first glance, these two seem to be the opposites that might not unify together well, but it is precisely the combination of the two that gives credit to both. The absence of data is complemented by a comprehensive theoretical overview in one, while the absence of theory is complemented by the abundance of data in the other. In her explanation of why these two work well together, Gilquin (2010, p. 15) claims: “It is true that corpus linguistics extracts categories from the data and cognitive linguistics posits them beforehand, but this is no different from saying that corpus linguistics starts with the data, while cognitive linguistics starts with the theory. What is crucial is that both paradigms recognise the existence of categories. More fundamentally, they both see category membership as a matter of degree rather than a yes-or-no question. In cognitive linguistics, this principle is expressed by means of the concept of prototypicality, originally a concept from psychology, according to which categories are organised around a maximally representative example (the prototype). Depending on the similarity they exhibit with the prototype, the other members occupy 22 a more or less central position within the category. As a consequence, categories have no clear boundaries and overlap with one another. It is the same fuzziness that corpus linguistics has revealed.” The idea of unifying cognitive framework and corpus studies has probably been best represented by the ‘usage-based’ models in modern linguistics (see Barlow & Kemmer, 2000; Mukherjee, 2004; Tummers et al., 2005; Gries & Stefanowitsch, 2006; Heylen et al., 2008 etc.). In this context, the term ‘usage-based’ seems to be the common denominator between the two theoretical and methodological frameworks. In fact, in their definition of Cognitive Linguistics some argue that Cognitive Linguistics is usage-based by default; for instance, Glynn (2010, p. 89) states that: “Cognitive Linguistics is, by definition, a usage-based approach to language. Its model of language places usage at the very foundations of linguistic structure with a linguistic sign, the form-meaning pair, argued to become entrenched through repeated successful use. It is this entrenchment that renders symbolic gestures linguistic rather than merely incidental and represents the key to structure in language. Patterns of language usage across many individuals can be argued to be indices of shared entrenchment. When large numbers of language users possess the same or similar entrenchment, we can talk about grammar, that is, linguistic structure.” Corpora thus serves to attest the use of language in various settings, contexts and manners, relying on the shared properties between language representation and language manifestation. The usage-based model in many ways represents a bridge between the tools and methodology of corpus linguistics and the theoretical account of Cognitive Linguistics. 1.3.3. Extracting the data – corpus as a method Considering that the three constructions targeted by this research are structurally and lexically different, different methods were used for their extraction from the corpora. The extraction of ditransitives, periphrastic causatives and relative constructions, was conducted performing queries which exploited either the exact word anchors or POS anchors. Given the flexibility and openness of ditransitive construction to accept new verbs (Goldberg, 1989, 1992), the construction was extracted solely through the use of its syntactic pattern. For periphrastic causatives and relative constructions, the extraction relied on both specific lexical 23 items and part of speech category. The latter approach is less inductively prolific (especially when it comes to the periphrastic causative), since it does not reveal new lexemes as heads of these constructions. Nonetheless, the mixture of the syntactic pattern along with the particular lexemes in the extraction process, although not ideal, helps to reduce the redundant patterns that make the filtering operation considerably more difficult and time consuming. This leads us back to the differences between the use of corpus-driven and corpus-based research, and detecting where this thesis positions itself on this continuum. In terms of data extraction from the corpora, Biber (2010, p. 162) states that “corpus-driven analysis assumes only the existence of word forms”. In other words, such research would rely on targeting particular lexical bundles or word sequences (for instance, one may try to discover the frequency of the sequence never have I ever in a particular corpus). Corpus-driven approaches are thus primarily defined by their focus on particular word sequences, that are rarely structurally complete or idiomatic, and on the analysis that is based solely on the rates of particular occurrences and their distribution across different texts. In terms of just data extraction, it would appear that this research does not start out as a corpus-driven research, but rather belongs to a more hybrid type of form that is also referred to as ‘pattern grammar studies’: “The pattern grammar studies might actually be considered hybrids, combining corpusbased and corpus-driven methodologies. They are corpus-based in that they assume the existence (and definition) of basic part-of-speech categories and some syntactic constructions, but they are corpus-driven in that they focus primarily on the construct of the grammatical pattern…” (Biber, 2010, p. 175) Basically, these studies are corpus-driven because they focus on the linguistic units that emerge from corpus analysis, but start from the position of corpus-based studies as they take into consideration the pre-defined constructions that lead the extraction process. This hybrid approach to corpus is also that one that bridges the gap between corpus linguistics and constructional approach in the cognitive framework. The research must, at times, employ particular lexical items because relying solely on syntax would be impractical to the point of inadequacy, but there also needs to be room for syntax and its contribution to the meaning of the construction, particularly from the constructional perspective. This is perhaps best illustrated by the quote of Hunston and Francis (2000, p. 3): “Patterns and lexis are mutually dependent, in that each pattern occurs with a restricted set of lexical items, and each lexical item occurs with a restricted set of patterns. In 24 addition, patterns are closely associated with meaning, firstly because in many cases different senses of words are distinguished by their typical occurrence in different patterns; and secondly because words which share a given pattern tend also to share an aspect of meaning.” This methodological position is inspired by the idea that particular grammatical patterns are tied to particular meanings. For example, the ditransitive pattern is that of a subject – verb – object – object, and by targeting the word classes that occupy such positions (for instance, PRO – V – NP – NP) we ought to retrieve the instance with particular verbs that do not usually carry ditransitive meaning. The role of such research becomes partly to confirm that syntax can dictate the meaning, and not necessarily just the lexical items. While the concerns around syntax-semantics interface are not the primary focus of this investigation, it remains to be noted that such approach to corpora is at the methodological core of this thesis. 1.4. Three constructions – the increase in complexity As already stated, in order to tackle the acquisition of the ditransitive, causative and relative patterns, the thesis approaches the matter by using evidence from corpora whilst interpreting the extracted data with the help of cognitive theoretical framework, meaning that the aforementioned patterns will thus be referred to as constructions (cf. Section 1.2.2). As already discussed in the introduction, the move away from the traditional perception of meaning being restricted to lexicon started to change with the first observations that syntax can influence meaning. Although the idea was popularized by Goldberg (1995), there were many before her arguing for the interrelatedness between syntax and meaning (Givon, 1985; Langacker, 1985; Clark, 1987; MacWhinney, 1989 etc.). Even back in 1968, Bolinger observed that any change in the syntactic form must imply a change in the meaning (p. 127), which ultimately drove Goldberg to reflect on their findings and postulate a “Principle of no synonymy” (1995, p. 67): “If two constructions are syntactically distinct, they must be semantically or pragmatically distinct”. Because of this, the constructions targeted by this thesis will be analysed primarily based on their different syntactic patterns. In a certain way, this principle leads us to assume that even minute differences in syntax can in fact influence the interpretation of these constructions, i.e. different patterns with slightly altered syntax that appear within ditransitive, causative or relative constructions can be regarded as mini- 25 constructions in themselves. Although the starting point of this work, these assumptions are not the object of this study in themselves, nor are the taxonomic differences between the patterns that we find among the family of constructions targeted in this research. Given that this thesis does not deal directly with detailed theoretical and syntactic accounts of a singular construction, or solely with its interpretation, it needs to be noted that the emphasis will be on the extracted data and the possible conclusions that might be drawn from it, and not on the presuppositions of any theoretical account about a particular construction and what its variants might entail. Nevertheless, the fact that certain syntactic patterns are more frequent than others within certain age groups reveals something about the processing difficulty of the same. In the research, the term ‘construction’ is used for all three grammatical patterns observed in this study. While the term is more frequently used in the Goldbergian sense with ditransitives and periphrastic causatives, especially when it comes to research in Construction Grammar, relatives tend to be referred to as clauses (relative clauses, however, only constitute a part of the relative construction). Before proceeding, it is necessary to define what exactly constitutes a construction in this particular context. As already stated before, even words are conceived as constructions from the point of view of Construction Grammar (see Croft, 2007, p. 471). Any syntactic pattern which has adopted a “conventional function in a language” can be regarded as a grammatical construction, i.e. the fact that its structure is associated with a particular meaning and that this association has become linguistically conventionalized makes it a construction (Filmore 1988). This ‘construction’ can be described through terms such as external and internal syntax. The former signifies speakers’ knowledge about the properties of a particular construction, and how they might be accommodated by a wider syntactic contexts. The latter term (internal syntax) signifies the construction’s make-up, that is, the description of the construction in terms of its grammatical pattern (ibid.). Considering that constructions are viewed primarily as “form-meaning correspondences” (Goldberg, 1995), the label construction is more often reserved for the syntactic patterns which obviously contribute to the meaning of a particular sentence, that is, even when the participants of the particular construction are replaced by non existing words (or even verbs for that matter), there is still a meaning that can be expected from that particular construction. In other words, the non-compositionality principle suggests that the construction can be regarded as a construction if its form or meaning cannot be predicted solely by the component parts which constitute it (Goldberg, 1995, p. 13). This phenomenon has particularly been investigated and corroborated on the ditransitive 26 construction (Goldberg, 1989, 1992), while the periphrastic causative construction has also been similarly approached (Gilquin, 2006, 2017). One of the reasons why these are often referred to as constructions is partly because one can imagine that a learner might store these as units, considering their degree of prototypicality and structural coherence, whereas relative constructions are on a different syntactic level from ditransitives and are often referred to as complex sentences (see Diessel, 2004). On the other hand, if we take a sentence such as The man gave the boy a ball which was black, it is hard to determine which construction is superior to another in terms of syntactic complexity. If anything, it appears that it is the ditransitive construction which accommodates the relative clause, and in the same sense accommodates a relative construction (a ball which was black). This happens because relative clauses act similarly to modifiers in a sentence, allowing themselves to be attached to any noun phrase. In terms of Cognitive Linguistics, relative clauses exhibit high schematicity in that they can be accommodated by various schemas, and yet their presence transforms the entire construction into a relative one. Similarly, if we look at the sentence The man forced the boy to throw the ball which was black, the question that arises is whether we should interpret this sentence primarily through the lens of a periphrastic causative or of relative construction. This however, does not represent a problem for Construction Grammar, which, in allowing polysemy within a family of constructions, must also allow for constructional polirepresentations at the same time. In this sense, it is also important to distinguish relative clauses from relative constructions. While the relative clause is an integral part of the relative construction, it is only in its entirety that we can argue for an advancement in syntactic complexity when compared to ditransitive or periphrastic causative constructions. Although this thesis addresses the suppositions concerning gradual complexification of the children’s grammar, the benefit of the study will be the additional evidence pertaining to the differences in complexity between the observed constructions. The term ‘complexity’ itself has proven problematic to clearly define, with various interpretations and uses by scholars. For example, ‘linguistic complexity’ has been used in different contexts for different purposes, some of which saw the construct as being intricately linked with sentence processing (i.e. the degree of computational resources a structure requires) (see Gibson, 1998), while others used it as a “cover term” for a variety of factors that include syntactic, thematic and semantic complexity (see Friederici et al., 2006). More straightforward operationalisations of the term moved away from associations with difficulty, and instead saw 27 the amount of information needed to describe it as an objective property of the structure in question (Dahl, 2004, p. 2). Building upon that, Odlin (2012) used a similar approach by referring to it as “descriptive complexity”, and later concluded that “linguistic complexity” is best defined as the level of descriptive detail required in accounting for all phenomena that may present a challenge to learning. On the other hand, Trudgill (2017) used the term in the context of sociolinguistic typology where he defined it as a multifaceted concept, arguing that complexification is best defined as an opposite process of simplification, i.e. opposite of regularization, reduction of redundancy, and an increase in transparency. Factors contributing to linguistic complexity thus include greater irregularity, heightened syntagmatic redundancy such as repetition of information, increased use of morphological categories, elevated levels of allomorphy, and a higher degree of fusion. When it comes to the analysis of intra-constructional differences and the specific syntactic patterns within each of them, the use of the term in this thesis is closest to the one presented by Pallotti (2015), who defined it purely in structural terms, where complexity stems directly from the quantity of linguistic elements and their interconnections (see also Arnold et al., 2000, where the term has been operationalized based on the length of constituents). Such operationalization deliberately excludes considerations of cognitive cost (difficulty) and developmental dynamics (acquisition), and instead takes into account the length of phrases, the number of phrases within a clause, the quantity of clauses per unit, the number of word-order patterns, and the mean length of clauses. 14 For instance, the use of the term ‘complexity’ in ditransitive constructions is mainly linked to the nature and length of the construction components, such as nominal elements, 15 and it is referred to as ‘structural complexity’. In this context, the intricacy of individual elements is closely connected to the overall complexity of the construction, and the same logic is applied for other constructions. The reference to the complexity of the entire construction builds upon the complexity of its nominal elements, and it is not to be conflated with processing difficulty. For example, in relatives, the studies of relative constructions tend to concentrate on that particular aspect (see Section 5.1.3), i.e. 14 It is essential to acknowledge, however, that even seemingly simple approaches like word count-based definitions can pose challenges, as exemplified in cases where ambiguity arises regarding whether a word should be counted as one or two, such as in the instances of "sports car" or "sportscar" (Pallotti, 2015, p. 123). In English, the somewhat common discrepancy between phonological and orthographic complexity/length also needs to be recognized as a factor. 15 Note that research has demonstrated how the length of arguments influences the selection of dative realization (see Bresnan et al., 2007). 28 employing a pronoun in a construction instead of a modified noun does not necessarily make it less complex in terms of processing, as seen in the contrast between the cognitive load required to process the relativized noun the car which was driven and the relativized pronoun it which was driven. On the other hand, the thesis also makes use of term ‘lexical complexity’ as described by Lu (2012), which is further elaborated in Section 2.1.3. Though processing difficulty remains beyond the scope of this study, and though the term is primarily used to denote an increase in linguistic material when it comes to specific intra-constructional comparisons of pattern frequencies, the results provide a solid foundation to further one’s inspection into the reasons for certain trends in production of specific patterns and the related implications on developmental dynamics and cognitive cost in general. As far as inter-constructional differences and the use of the term complexity are concerned, the fact that there may be more to it than simply an increase in (or complexification of) linguistic material calls for consideration. For instance, the periphrastic causative construction is assumed to be more complex than the ditransitive construction because the predicate is more complex. On the level of meaning, the periphrastic causative also implies a greater number of agents in the event; given that the action is governed by two entities in the process, we may even assume the entailment of two events on the metaphoric level of the construction. Whether with a full infinitive or a bare one, the periphrastic causative is a biverbal construction and it entails two events – one which precedes and causes the other to start. The syntactic complexification is followed by the complexification in meaning, and the comparison might remind of Diessel’s claims about the complexification of the nonfinite clauses in early children’s speech. Diessel (2004) distinguishes between children’s early nonfinite complement clauses based on the roles of the subject and the direct object in these constructions. He concludes that the infinitival extension of the nonfinite clause can be interpreted as a complexification of the initial expression, and that in the process of acquisition, children tend to learn this incrementally. “Moreover, while children’s early nonfinite complement clauses are always controlled by the matrix clause subject (e.g. I wanna sing), later nonfinite complements are often controlled by the direct object, which occurs between the finite verb and the infinitive or participle (e.g. He told him to leave). Based on these findings, I argue that the development of nonfinite complement constructions can be seen as a process of clause expansion. Starting from structures that denote a single situation, children gradually 29 learn the use of complex sentences in which a nonfinite complement clause and a matrix clause express a specific relationship between two states of affairs.” (Diessel, 2004, p. 49) Diessel’s inference about the increase in complexity is applied to constructions where one clearly behaves as an extension to the other. This is evident in their interpretation, the increase in syntactic complexity in terms of argument and predicate structure and the fact that one can be a direct extension of the other (He told him to leave instead of just saying He told him). Although there is no direct implication of the extension, the argument of the increase in complexity is similar to Diessel’s. This is not to say that there can be no overlaps between the ditransitive construction and the periphrastic causative. If we take Diessel’s example He told him to leave, we can argue that the construction shares features with both ditransitives and periphrastic causatives. Similarly to ditransitives, there is something (a message) that is being transferred between the agent and the recipient. Similarly to the periphrastic causative, there is an implication (although causation is not guaranteed) that the occurrence of one event leads to the occurrence of another (‘his suggestion leads to a departure’). Clearly, there are points of overlap where it becomes difficult to assess the difference in complexity, but it is still justifiable to argue that the aforementioned ditransitive construction (if we grant it such a label) is more complex than the prototypical construction with give (as in He gave him the ball). The point is that these discrepancies can be captured by the corpus-based approach, and that they may reveal interesting new data about language acquisition and processing. Figure 1 illustrates the operationalization of the increase in complexity in the three constructions targeted by this research. Figure 1. Graphical representation of the increase in complexity across the three constructions Legend: NP – noun phrase, V – verb, C – clause, REL – relativizer 30 Thus far, the research in language acquisition has primarily been focused on the differences in the acquisition difficulty between similar patterns within a family of constructions; for instance, with different types of ditransitives (see Gropen et al., 1989; Campbell & Tomasello, 2001; Snyder & Stromswold, 1997¸ Fischer, 1972 etc.) or relatives (see Diessel & Tomasello, 2005; Diessel, 2009a). Such studies tend to observe differences in the speed of mapping, the relevance of the input, the acquisitional time frames for these patterns etc. Although this thesis mainly investigates the range of frequencies between different syntactic patterns within each of the construction families observed, the chapter following the intra-constructional (i.e. construction-specific) analysis evaluates the discrepancies between all three observed constructions (for inter-constructional data, see Chapter 6). 37 process, as opposed to those words which occur less frequently. Although this modification may provide more fine-grained results, the problem may arise with ‘drawing the line’ between highly frequent and lowly frequent lexical items. Following Engber (1995) and Lu (2012), the LD measure in this research is calculated as the ratio between content open-class words and the total number of words in a sample text. Lexical sophistication is a measure that ties closely to lexical density because, much like density, it relies on the ratios between specific content words and, depending on different variations, the overall numbers of content words, the number of word types etc. In the analysis, in the thesis I will use the index of lexical sophistication as conceptualized by Linnarud (1986) or Hyltenstam (1988) that divides the number of sophisticated lexical words by the total number of lexical words. The index directly concerns the frequency of one’s usage of lexical items that are not part of their narrowest and commonest language environment, but without delving into diversity. Although it may initially seem so, the basic measure itself does not reveal much about the range of lexical items used by the speaker. Theoretically speaking, it is possible for a person to have a poorly diverse vocabulary despite employing lexical items found beyond the most common ones, i.e. one could theoretically have good knowledge of the sophisticated lexicon without necessarily knowing the elementary terms. In this context, sophistication is to be regarded in theory as a separate concept from diversity, despite often corresponding to diversity in practice (for example, see Durrant & Brenchley, 2019). In this research, following Lu (2012), sophisticated lexical words were considered to be those not found amongst the most frequent 2000 words derived from the British National Corpus (Leech et al., 2014). The measures of sophistication have since been extended to specific parts of speech and have centred around verbs in particular. For example, one such measure has been suggested by Wolfe-Quintero et al. (1998), who proposed the measure of sophisticated verbs (CVS1) in the text to be altered along the lines of corrected type–token ratio (Carroll, 1964) so that the sample size effect is reduced. As opposed to just looking at the number of sophisticated verbs in a text, it takes into account word types; the index is calculated by dividing the number of sophisticated verb types by the square root of the doubled number of verbs in the sample text. Like with lexical sophistication, the goal is to assess how far the speaker goes beyond the most frequent words in the vocabulary, but with the focus on verbs, albeit with the introduction of diversity aspect. As opposed to lexical density which concerns variation only in relation to parts of speech, lexical diversity provides an estimate of the range of lexical items used in the evaluated 38 text. High levels of lexical diversity would indicate more developed vocabulary levels, or at least a demonstration of the same. In order to achieve relatively high scores of lexical diversity, the evaluated corpora would have to represent speech or writing that employs different lexical items and does not repeat them that often in the text once they have been used. It is no surprise that, as lexical density, lexical diversity also tends to be of higher values in written texts as opposed to spoken ones. In the context of this work, the term is used as a cover term for the straightforward measures of range such as the number of different words in the text, as well as type-token ratios which look into variation with respect to size of the sample. As Duran et al. (2004) note, the number of different words (NDW) is the most straightforward way to tackle lexical diversity in one's speech, but in order for it to work, the researcher has to standardize the size of the samples (for example, balance them either by the number of utterances or the number of tokens). For obvious reasons, the conclusions deduced from samples of different sizes are inherently wrong and defeat the purpose of the measure; if the size of one sample is bigger than the other, chances are that it will contain more different words. Both methods of sample standardization (either by the number of tokens or utterances) have their own advantages depending on which developmental aspects are emphasized, but the goal should be for researchers to agree on a single method so that cross-research comparisons could be carried out. Other methods that help with reducing sample size issues look at random subsets selected from the sample and take the mean number of types of words found across those subsets. One such measure used in this thesis (adopted from Lu, 2012) calculates the mean number of word types of 10-random 50 word samples (NDW-ER50). In both cases of NDW measures, the important thing to note is that they are to be understood as measures of ‘range’ (Malvern et al., 2004). Type-token ratio (TTR) is in some sense a natural extension of NDW. The basic value is calculated by dividing the number of different words in a text by the number of tokens in a text. The term ‘type’ refers to unique words in a text, whereas ‘tokens’ represent each individual occurrence with a particular corpus position (i.e. each string of characters divided by space). In a text that contains three words, out of which two are repeated, the number of types will be two and the number of tokens three. Consequently, the measure is always expressed through values between 0 and 1, where the higher values represent greater diversity (Malvern et al., 2004). The obvious difference between TTR and NDW is that the former takes the size of the language sample into account, but the size can also become an issue. The problem is that the increase in 39 the number of tokens after a certain point becomes counterproductive for TTR values because maintaining lexical diversity in spite of the output volume produced would require extremely rich (or unlimited) vocabulary that is paired with little or no repetition. The shorter samples are more likely to (unjustly) benefit in comparison to longer ones. This is especially evident in the context of language acquisition and development, where it has already been stated and observed that linguistically more advanced children produce longer utterances (both verbally and morphologically richer) and they also produce them more frequently in given time frames (see Richards, 1987). The implication of the varying lexis being held constant is that TTR will negatively correlate with the number of tokens in the compared samples, and any increase in size will necessarily lead to a reduction in the value of TTR, only seemingly exemplifying a decrease in lexical diversity (ibid., p. 206). One of the solutions for TTR being largely affected by the sample size is to calculate the average TTR based on the subset of samples derived from the observed sample. This socalled Mean segmental TTR method was proposed by Johnson (1944), who suggested that each of the samples gets divided into segments of the designated length and TTR calculated for each, with the final index being the result of average subsample TTRs. As Malvern et al. (2004) note, the measure is an equivalent of TTR for the expected number of words in the sample. Other transformations which aimed to achieve a constant value of TTR for samples regardless of their size include Root TTR (RTTR) by Guiraud (1960) and Corrected TTR (CTTR) by Carroll (1964). While the former divides the number of types by the square root of the number of tokens (RTTR), the latter does the same but by double the number of tokens (CTTR). The latest and perhaps the most precise solution has been the so-called Moving-average TTR (MATTR) by Covington & McFall (2010), also based on the mean TTRs, but with different method of capturing the subset, i.e. first a designated window length of words is chosen and then TTRs are computed for each of the ranges between the first and every subsequent word and the last (e.g. if the window length is 500, than MATTR is calculated by computing TTRs for 1-500, 2500, 3-500 until the end), and the final value amounts to the mean TTR of the calculated subset TTRs. Since the compared samples in the research are already balanced for the random number of utterances (albeit not for the number of tokens), elemental values of type-token ratios such as MATTR and RTTR will suffice for the evaluation. 22 The exploration of this corpus data in terms of diversity is further supplemented by the transformed measure of verb variety SVV1 22 For additional values on lexical richness, see the Appendix 10.2. 40 (as suggested by Wolfe-Quintero et al. (1998) following the original measure by Harley and King (1989)), which reduces sample-size effect. The idea behind the indices of verb-variation is the same as with lexical diversity, only instead of looking at the ratios between word types and their overall number, the indices take into account solely verb types and the total number of verbs in the sample. Table 1. Measures of lexical richness considered in this study 23 Index Formula Lexical Density (LD) Nlex/N Lexical Sophistication (LS1) Nslex/Nlex Corrected Verb Sophistication (CVS1) Tsverb /√2Nverb Number of Different Words (NDW) T NDW (expected random 50) (NDW-ER50) Mean T of 10 random 50-word samples Mean Segmental TTR (MSTTR-50) Mean TTR of all 50-word segments Root TTR (RTTR) T/ √𝑁 Corrected TTR (CTTR) T/√2N Squared Verb Variation (SVV1) Tverb2/Nverb Legend: the number of tokens of words (N), lexical words (Nlex), sophisticated lexical words (Nslex), and verbs (Nverb), the number of types of words (T), sophisticated verbs (Tsverb), adjectives (Tadj) In order to analyse lexical richness of the samples across age groups, in the thesis I cross-examined the average scores calculated on sentences containing the prototypical construction patterns explored in this research. The average scores for the listed measures were calculated on random samples of 200 utterances (exceeding 2000 tokens on average) for each observed age group. 24 The analysis of lexical richness will be provided for utterances that centre around the observed constructions, and which consequently often extend beyond them. One reason for observing entire utterances and not isolated constructions (although there will be references to the same lexical indices pertaining solely to constructions as if they were isolated from the rest 23 For additional explanation of indices, consult Lu (2012). 24 The scores were rechecked with different database samples and they showed relative consistency in numbers. The choice of that sample size (random samples of 200 utterances) is additionally supported by the results of Zenker and Kyle (2021), who found many lexical diversity indices to be stable even on samples with less than 200 tokens, except for some measures of type-token ratios. 41 of the sentence) is because a more fine-grained description will already be provided in terms of head lexical elements and syntactic patterns that characterize these constructions. More importantly, it gives an information about the context in which the constructions are usually employed. A great part of analysis in language development and acquisition is sometimes too concentrated on the patterns as if they were isolated from the rest of the sentence just because there are no inherent constraints on their immediate linguistic environment, and yet many stable phrases (if not entire utterances) often go beyond the basic constructions. Furthermore, it may be interesting to observe how do entire utterances that contain these constructions differ amongst themselves, what role do larger language units such as sentences bear on the children’s processing of the embedded constructions, and whether corpus-derived data would reveal something about the interplay of these structures in relation to age-differences. Such approach to lexical richness in general is not something usually applied in language research; it is not common to observe measures of lexical richness for one specific construction (or the utterance that contains that specific construction, as is the case in this analysis) because the usual aim is to capture differences in view of density and diversity across samples that represent linguistic skills/competence in its entirety (whether written or spoken). Nevertheless, the lexical richness-based comparison of utterances that centres around one particular construction across age groups should additionally shed a light on some more specific construction-related developmental aspects. For example, the difference in lexical density demonstrated within a particular construction across age groups tells us whether the information uttered by these age groups in that particular construction is more concentrated. The increase in lexical diversity within a particular construction suggests that a particular age group is more lexically versatile in producing that construction. Moreover, given that three constructions are observed in this research, we might speculate as to why there might be a drop (or an increase) in the measure of lexical richness in one construction, as opposed to others. Finally, the measures can serve to confirm that the speculated differences in complexity between the observed constructions indeed exist when it comes to real language use, or even that a particular construction is more challenging to master from the perspective of language acquisition. The data may reveal that constructions themselves bear striking resemblances across age groups, but the question would remain whether the same goes for the rest of the language, even the segment which is closely related to the observed phenomena. 42 2.2. Research questions and hypotheses So far, the introduction has outlined some of the key theoretical considerations for a corpusbased study that aims to tackle the issues surround language acquisition and development. In usage-based linguistics, the degree to which the input plays a role in language acquisition has been discussed and explored in a number of studies (see Section 1.2.3). The productivity of specific language patterns that uniquely intertwine with specific semantic features has been in the focus of Construction Grammar (see Section 1.2.2), while Cognitive Linguistics served to provide a more comprehensive theory of mind that allows for assertions about linguistic structures emerging from use (see Section 1.2.1). The research questions and hypotheses in the following subsections tie together the aforementioned theoretical and practical research in the hope of providing additional data on some issues relatively underinvestigated in this manner thus far. 2.2.1. Research questions As stated in the introduction, the main goal of the thesis is to explore the structural and lexical similarities across the observed age groups (0-3, 4-6 and adults) using corpus-based analysis to target three specific constructions. Aside from examining inter-age and intraconstructional data, the study also addresses inter-constructional differences in the data retrieved, as well as the degree of ‘item-based’ character in each. The utilization of CHILDES in this manner is not entirely novel, with various studies, such as those by Liu et al. (2008), Goodman et al. (2008), Li & Fang (2011), and Shin (2022), having already employed CHILDES corpora to investigate parental input and child language, and having found significant relations between parental input and language produced by children in several respects (see Section 1.1). These studies predominantly concentrated on lexical aspects and similarities, whereas the contribution of this thesis is that it scrutinizes correspondence between specific construction slots and aspects associated to both lexical and structural complexity of the said constructions. The research question pertaining to lexical and structural similarity builds upon usagebased accounts of language acquisition (see Section 1.2.3). Though some lexical and structural similarity is obviously expected between adult input and child output in any language 43 acquisition theory, the degree to which it occurs with specific sub-patterns of the constructions observed and the corresponding construction slots is a valuable contribution to the discussion on input-output correspondence. Furthermore, the research question pertaining to animacy builds on the idea that the imitative tendencies of early child language should affect the lexical character of items occupying the corresponding construction positions (cf. Buckle et al., 2017), but with a more dominant representation of specific animacy values which would otherwise be considered prototypical in the observed construction slot. The idea that children consider the semantics of items heard most frequently in the observed construction slots and consequently overproduce words which fit the dominant semantic frame is another extension of item-based suppositions about early language development that transposes to a semantic aspect such as animacy. Furthermore, the analysis of lexical richness and complexity has been used in various fields of linguistic research, with language acquisition being one of them (see Blake et al., 1993; Malvern et al., 2004; Yoder and Stone, 2006; Yoder, 2006; Li & Fang, 2011 etc.). However, the measures are usually employed on developmental data obtained longitudinally, whereas the inter-group comparison characteristic of this research has been underrepresented so far. Finally, the unique addition of this thesis concerns the inter-constructional comparison in terms of itembased tendencies, exploiting the idea that simpler constructions (see Section 1.4) should resemble the input more than others (cf. Tomasello & Brooks, 1999). The research questions and hypotheses that exemplify the aforementioned topics are the following: 1. Is there any lexical and structural similarity between age groups in terms of preferred lexical and syntactic choices in the observed constructions? 1a. To what extent are the constructions similar with respect to verb frequency? 1b. To what extent are the constructions similar with respect to noun frequency? 1c. Is there a difference between the observed age groups of children in terms of the degree to which they are similar? Does the 0-3 age group exhibit greater similarity to child-directed speech than the 4-6 age group? 2. Are there any differences in terms of animacy with respect to key lexical items (nouns) in the constructions? 3. Will the measures of lexical richness and complexity reflect the differences in language use across age groups with respect to the observed constructions? 4. Is there a difference across constructions in the degree of their item-based character? 44 2.2.2. Hypotheses 1. There will be a considerable lexical similarity in constructions across age groups in the observed constructions. This assumption is anchored in the fact that children of the earliest age ought to reproduce the most frequent patterns as they occur in their immediate linguistic environment. Having in mind the item-based tendencies in early language, the following hypotheses are then derived from this: 1.1. The selection of certain verbs in constructions will reveal a considerable similarity across age groups. 1.2. The selection of certain nouns in constructions will reveal a considerable similarity across age groups (albeit smaller than when it comes to verbs). 1.3. In comparison to other two, the youngest age group’s production of constructions will be characterized by a greater disproportion between the most frequent items/patterns and the rest (most frequent ones will be more pronounced than in other age groups). 1.4. The language of 0-3 age groups will be more similar to child-directed adult speech than the language of age groups 4-6 when it comes to the ratio of the use of certain structures and lexemes. As children grow up, their language should be more similar to the adult one in terms of how competent and productive they are, but not in terms of reiterating lexical choices that appear in specific construction slots or even the degree to which they reproduce the same patterns in the input. If this assumption turns out to be correct, in a way it agrees with the following: child-directed speech does help at the earliest ages, but its reach is limited in the sense that children learn language imitatively up to a certain age, but as their cognitive systems develop, they begin to rely on other mechanisms such as distributional/statistical learning and ‘pattern finding’ skills. At that moment, the speech aimed at children becomes less relevant, and other spontaneous speech takes precedence, from which the slightly older children begin to derive their own conclusions about what language use allows and disallows in terms of context, grammaticality, logic etc. This fact may additionally be supported by more equivalent ratios in children and 45 adults regarding certain lexemes and structures, as well as a simultaneous explosion in terms of the frequency of constructions after reaching the age of four. 2. The language of 0-3 age groups should be least flexible when it comes to the animacy of key lexical nouns in the constructions; meaning that the category of animacy considered prototypical for particular lexemes in the observed constructions will also be the most saturated in the youngest age groups when compared to others. This hypothesis reflects the idea that the children’s learning of constructions is not just imitative in character, but also distributional; that is, by taking into account the linguistic context and remembering various instances in which a particular item (or a phrase) can be used, children should be able to infer on some lexical properties of these items, such as animacy. In turn, their early suppositions concerning the use of these items (along with imitative tendencies) should reflect on the character of these items; they will favour the prototypical (most frequent) uses found in the input. In other words, there is a reason to assume that children's representation of specific language patterns closely ties to the semantic character of the same construction which they encounter most often in the input (for experimental data, see Buckle et al., 2017). 3. The measures of lexical richness and complexity will differ across age groups despite the potential lexical similarity. Despite the item-based character of language development, the measures of lexical richness and similarity should still reflect differences in the syntactic and lexical elaborateness of older age groups in comparison to younger ones. 4. Simpler constructions will be more item-based in character than more complex constructions; that is, ditransitive constructions should be more item-based than causative constructions, and causative constructions should be more item-based than relative constructions. This prediction builds on the perspective explained in Tomasello & Brooks (1999), which assumes that the learning of the earliest constructions does not take place in the context of abstract syntactic patterns that children have access to from an early age, but rests on the emulative learning of constructions around its central words. In this sense, 46 it could be expected that the learning of patterns that justify the label of ‘construction’ to a greater degree (because they have a more stable syntax of an almost inherent meaning; e.g. ditransitives compared to relatives) will be anchored to their central elements to a higher degree as well. On the other hand, constructions with greater ‘combinatorial’ potential in children should be less similar to the input itself. 2.3. Method 2.3.1. Approaching the corpus In this construction-aimed corpus-based research, a key aspect of the methodology involves operationalizing the constructions we are examining. In this sense, each section that looks into the data on particular constructions also includes an introduction covering the theoretical background of the phenomenon that is going to be observed, the overview of previous research, and the operationalization of the observed phenomena. Essentially, the undertaken research process can be outlined in the following way: the so called “scouting” phase which included characterizing the phenomenon and reviewing the relevant literature, followed by the observation of the linguistic phenomenon in the primary sources, and finally the gathering of new information and applying deductive reasoning (as described by Gries, 2013, pp. 8-9). In other words, the more general theoretical overview provided in Chapter 1 is complemented with construction-specific research overview that also serves as a background for the proper methodology (the extraction of the constructions from the corpus) – the sections operationalizing individual linguistic phenomena observed serve as introductions to chapters analysing them (see Section 3.1, Section 4.1, and Section 5.1). The most time-consuming process in this research was the data curation process, i.e. extraction of data from the corpora and the removal of “false positives” from the final dataset. In order to retrieve the target constructions from the corpus, it was first necessary to locate them in the transcriptions (enabled by the Sketch Engine-integrated part-of-speech automated tracking tools) and to filter out those instances that actually represent the constructions targeted by the research from those that, for one reason or another, manifest the same part-of-speech pattern but do not qualify for said constructions in terms of meaning or even structure (for example, due to repetitions, interjections, redundancies or non-registered sentence breaks, an isolated syntactic layout of the speech segment may accidentally correspond to target 53 Analogous semantic alternations have been observed on ditransitive construction denoting less overt ‘transfers’, such as with cognitive transfer verbs: (3) (a) John taught the students French (b) John taught French to the students Example (a) entails that the students have mastered French, while in the example (3b), it is only obvious that the students were indeed educated, but it is not conclusive that his process was successful in the end (Green 1974). Similarly, ditransitive alternations with the verb throw also vary in interpretation: (4) (a) Dick threw John a ball (b) Dick threw a ball to John In sentence (4a), it is implied that John has caught the ball he was thrown. However, in the second example, John is perceived more as a target than a recipient; i.e. in the prepositional dative alternation, John could have easily been dead or asleep while being thrown a ball, whereas this is not conceivable in the first sentence (Pinker, 2013, pp. 97-98). Furthermore, a distinction between caused motion and caused possession was explored on the verbs send and throw by Rappaport Hovav and Levin (2008), who argued that both may take either of the aforementioned meanings (either the recipient has come into the possession of the theme or they have not). Such distinction is important, because it not only allows us to distinguish between the two types of construction, but it also make it easier to define and operationalize a ‘prototypical’ ditransitive, which we can then define as only the one that takes an implied ‘possessor’ as its recipient. In many cases, the meaning of a direct object is closely related to that of a recipient. This was noted by Larson (1988) on the prepositional dative alternation but the same logic applies to double-object constructions. Consider the following: (5) (a) Beethoven gave the world the Fifth Symphony, (b) Beethoven gave his patron the Fifth Symphony In (5a) we assume that the “transfer of possession is metaphorical”, and the Fifth Symphony refers to a composition in general. In sentence (5b), the implication is that Beethoven gave a physical copy where the composition is written to his patron. In other words, the nature of the 54 recipient governs the exact semantic role that can be assigned to the direct object in a ditransitive construction. Goldberg (2002) has argued that that the ditransitive construction needs to be distinguished from other prepositional alternations that share similar semantics, i.e. she considered it wrong to view ditransitives as derivations from its “corresponding paraphrases”, and instead suggested that we treat these as constructions in and for themselves (ibid., p. 336). In other words, she calls the ‘double-object construction’ a ditransitive construction, while the dative construction with the preposition to is to be called a ‘caused-motion construction’ (e.g. Jack gave a gift to John), and the prepositional ‘for alternation’ should be named something along the lines of ‘transitive construction + benefactive adjunct construction’ (e.g. Jack bought a gift for John). When it comes to the complexity of arguments, it has been noted that the canonical double object ditransitive favours ‘simpler’ recipients (or generally indirect objects) when compared to its dative alternation. For example, speaker will tend to choose a construction such as He gave his son a couple of CDs more frequently than He gave a couple of CDs to his son, whereas when it comes to more complex recipients, the preferred construction is the one with the prepositional complement – e.g. He gave the spare copy to one of his colleagues is likely to be preferred over He gave one of his colleagues the spare copy. Correspondingly, the construction with a pronoun as a direct object (e.g. I gave Kim it) seems awkward or even grammatically ‘inadmissible’ for English language speakers (examples from Huddleston & Pullum, 2002, pp. 309-310). In terms of semantics, some scholars have argued that the ditransitive construction shares some features with the causative one. For instance, the verb give in a ditransitive construction can supposedly be decomposed into ‘cause-to-have’ predicate (see Richards, 2001; Harley, 2002; Jung & Miyagawa, 2004). In other words, many of the double-object patterns in English language can be decomposed into a causative predicate. Another restriction on the double-object ditransitive construction has been observed in terms of argument animacy (see Gropen et al., 1989; Harley, 2002), which suggests that the referent in the first object position (indirect object in a double-object construction) must be an animate one. The following examples serve to illustrate this claim (taken from Harley, 2002, p. 35): 55 (6) (a) The editor sent the article to Sue. (b) The editor sent the article to Philadelphia. (c) The editor sent Sue the article. (d) The editor sent Philadelphia the article. While it is harder to observe how examples (6a) and (6c) diverge in interpretation, the nature of the noun Philadelphia calls for a different reading between a double-object construction and the prepositional dative variant. The prepositional dative allows for an interpretation where Philadelphia is not necessarily an animate being, but more of a location, whereas this type of interpretative flexibility is not allowed in the double-object construction. In a double-object construction, Philadelphia stands for some sort of organization (a group of people) and as such, represents an animate recipient. This rule derives from a semantic criterion stating that the referent in the first object position in a double-object ditransitive must be a “prospective possessor of the referent of the second object” (Gropen et al., 1989, p. 207), and considering that a prospective possessor is perceived as an animate being, the conclusion is that only animate recipients occur in double-object ditransitive constructions. On the other hand, the prepositional dative allows for the indirect object to be interpreted as a location and not necessarily a possessor or a recipient. Indeed, there is a lot of variability depending on the choice of verbs in a given construction. For instance, it has been observed that with verbs such as show, give and tell, the theme tends to be an inanimate object, while the recipient is consistently animate. On the other hand, verbs such as bring, send and throw allow for the recipient to be an inanimate thing, which is in line with the interpretation that these verbs also more readily accept more ‘goal-like’, as opposed to always requiring ‘possessor-like’ recipients (Haspelmath, 2015, p. 37). 3.1.2. Ditransitive verbs vs. ditransitive constructions The constructional approaches to ditransitive constructions treat this syntactic pattern as a symbolic unit, meaning that there is a deeper semantic characterization almost inherent to this particular argument structure construction. This is an interpretation that moves away from lexically-rooted perspectives, where the verb in a construction dictates the meaning of it, i.e. every variation in the meaning between these sequences is explained primarily via verb senses. 56 This dispute relates to one’s outlook on a range of meanings that can be attributed to a particular construction; the question is whether ditransitive constructions cover a wide range of similar but nevertheless distinct meanings, or whether these meanings are enough alike to be treated as one. In other words, are these semantic alterations explicated by verb senses, or do we attribute the construction one underlying meaning that encompasses all semantic realizations for each of the instantiations. For instance, the following examples illustrate different semantic entailments realized in the same construction (examples from Gisborne & Trousdale, 2008): (7) (a) Jane gave Peter a cake. (b) Jane sent Peter a cake. (c) Jane faxed Peter a letter. The first sentence (7a) entails that the Peter has received the cake, while this is not the case in second sentence (7b), where it is only understood that the cake was indeed sent, but that it did not necessarily reach its destination. In the third sentence (7c), Peter also has not received the physical copy of the letter, but was instead sent only the contents of it. In some cases, the differences in interpretation seem to exist quite clearly and it is necessary to recognize them; however, aside from recognizing these different semantic entailments, we may also recognize the commonalities. Instead of arguing that the verb ‘send’ has multiple word senses that depend on the syntactic frame which it occupies, it is possible to argue that there is an underlying meaning inherent to the construction, in which we can fit a more general “underspecified sense” of the verb send (Gisborne & Trousdale, 2008, pp. 2-3). Therefore, in the constructional approach, there is a meaning that can be associated with the construction itself to the point where even if you fit nonce words in it, there is a high probability that the speaker could infer the semantic constraints of the syntactic pattern in question. In order to distinguish between ditransitive verbs and ditransitive constructions (at least in terms of how they are conceived in Construction Grammar), we have to recognize the fact that some transitive verbs and intransitive verbs can appear in a ditransitive pattern, as well as the fact that ditransitive verbs can occur in other types of syntactic environment. For instance, a ditransitive verb such as send can take on the intransitive form with a prepositional complement (e.g. He sent for him) or in a transitive prepositional pattern (e.g. He sent him into the comma). This is not exclusive to the verb 'send' but it also happens with prototypical ditransitives such as give (consider the intransitive prepositional construction such as Mary 57 gave freely to the poor). More importantly, verbs which are not usually perceived as ditransitive verbs occur in ditransitive constructions in the context of 'transfer' (e.g. intransitive blow in Mary blew John a kiss or a transitive throw in Chris threw Pat the ball) (examples from Stefanowitsch & Gries, 2003, p. 227). Examples such as these may be treated as evidence for the existence of inherent meaning to a ditransitive syntax. In the example He gave it all he could, there is no real recipient, at least not one that we could locate in a physical word, i.e. the ‘recipient’ in this case is typically something abstract, in a form of a goal that the agent has to attain. One could argue the construction to be idiomatic, but the idea of ‘transfer’ still exist to a certain degree – in order to attain a goal, a person will sacrifice their time and will, and thus in the process, the attainment of a particular goal becomes the recipient of a person's effort. Clearly, there are examples in English where a ditransitive construction does not convey the usual thematic roles of agent-recipient-theme, but from a comprehensive viewpoint, the interpretation of the said example falls under the scope of the semantic extension allowed by the ditransitive construction itself. Furthermore, in the aforementioned example, the theme is represented by a clause, but it may be replaced by a noun phrase at any time. In fact, the construction is rather productive; apart from the clausal complement, we can use an adjectival one (e.g. He gave it his best) or a noun phrase instead (e.g. He gave it a shot). Trousdale (2008) discusses composite predicates that resemble ditransitive constructions in terms of lexicalization and grammaticalization, i.e. according to him, the construction as a whole can also be subject to lexicalization and grammaticalization processes. This interpretation is supported by the construction-based account and it can be observed on the examples of constructions resembling the canonical ditransitive construction. Most give composite predicates share the same semantic and syntactic properties; a ‘give-gerund’ composite predicate (e.g. I gave him a kicking) resembles the ‘give ditransitive construction’ in many aspects, but the partial idiomaticity of the expression can sometimes dissociate it from the canonical ditransitive. For instance, I’ll give her a seeing to is even more idiomatic, in that it typically means I’ll have sex with her or I’ll physically harm her, and only the second meaning is evident from its prepositional variant (I’ll see to her) (examples taken from Trousdale 2008; 35). The point is that not all ‘give composite predicates’ entail some type of transfer regardless of their shared syntactic characteristics, and part of the reasons for this are the processes of grammaticalization and lexicalization, and this does not necessarily conflict with the 58 construction-based account and the assumed shared semantic alikeness of particular grammatical patterns (the verb give developing a telic aspect in the ‘give-gerund’ composite predicate is interpreted here as an example of grammaticalization, while its idiomatic meaning is viewed as an example of a lexicalization process) (ibid., pp. 35-36) . Likewise, some reframing effects of the verb give, and consequently some ditransitive forms, have been observed in English. The verb give in a ditransitive construction is said to influence the semantics of the expression it paraphrases. For instance, the expressions I kissed Gwen and I gave Gwen a kiss differ in that the verb give in the ditransitive construction turned an atelic event into a telic one; the first expression does not imply any temporal boundaries, while the second has the connotation of the event taking place instantaneously (Trousdale, 2008, p. 40). Although this research adopts the Construction Grammar’s perspective on constructions, this does not imply a rigorous ‘structure-above all’ approach to the studied constructions, or the children’s language for that matter. As Goldberg stated herself, despite the fact that we observe shared properties of certain linguistic pattern, we still have to acknowledge the existence of same patterns with independent meanings, which means that we cannot go completely beyond the verbs as such in our interpretation of these structures: “That is, there do exist instances of constructional ambiguity: a single surface form having unrelated meanings. It must be emphasized that it is not being claimed that meaning is simply read off surface form. What is being suggested here is simply that by putting aside rough paraphrases and considering all instances with a formal and semantic similarity, broader generalizations can be attained. In order to identify which argument structure construction is involved in cases of constructional ambiguity, attention must be paid to individual verb classes. In fact, in order to arrive at a full interpretation of any clause, the meaning of the main verb and the individual arguments must be taken into account.” (Goldberg, 2002, p. 335) When it comes to verbs used in ditransitive construction, the collostructional analysis 27 done on the International Corpus of English (ICE-GB) revealed that the verb give is by far the most 27 Collostructional analysis is a set of quantitative corpus-linguistic methods enabling researchers to quantify the strength of the relationship between word constructions and the grammatical structures they appear in. The primary goal is to identify the lexical items that commonly collocate with a specific grammatical construction, which in 59 frequent one (Stefanowitsch & Gries, 2003, p. 229). Other salient collocates were also tell, send, offer, show, cost, teach, followed by award, allow, lend, deny, owe, promise, earn, grant, allocate, wish, accord, pay, hand, guarantee, buy, assign, charge, cause, ask, afford, cook, spare, drop. In terms of Construction Grammar, such analysis might suggest a connection between the inherent meaning of construction and the usually assumed meaning of the collocate verb. In other words, the semantic properties of the most significant collexems tend to correlate with the perception of what constitutes a ditransitive verb, i.e. the more salient the verb used in the ditransitive construction, the more likely it is for speakers to conceive it as a ditransitive verb (ibid., p. 229). If anything, lexical analysis of frequent collocates allows us to claim that a ditransitive construction with the verb give can be regarded as a prototypical one, or, from the perspective of Construction Grammar, that the speakers most frequently associate it with the meaning and syntax of the ditransitive construction. The dominance of the double-object ditransitive construction over the prepositional dative alternation was observed by a number of studies and authors (see Biber et al., 1999; Campbell & Tomasello, 2001; Stefanowitsch & Gries, 2003; Siewierska & Hollmann, 2007 etc.). A corpus-based research on the Lancashire dialect texts (that constitute a part of the British National Corpus) has revealed the vast majority of ditransitive clauses to be double object ones (a total of 83%) and featuring the following verbs: blow, bring, buy, cook, draw, get, fetch, give, hire, make, offer, owe, pay, post, read, save, send, sell, show, take, teach, tell and write (Siewierska & Hollmann, 2007). 3.1.3. Previous research When it comes to language acquisition research, ditransitive constructions represent an interesting object of study due to a number of reasons, some relating to the complexification of simpler structures (such as transitives) and their developmental trajectory in child language, and the other relating to divergent interpretations and operationalizations of this syntactic pattern (e.g. the ‘lexical rules’ account vs. the constructivist approaches). Different suppositions regarding the interplay between syntax and semantics, how they are decoded and explicated, turn facilitates semantic analysis of the grammatical constructions observed (see Stefanowitsch, 2013 and Hilpert, 2014). 60 present important concerns to language acquisition researchers for the same reason language processing research relates to the field of language acquisition (see Gropen et al., 1989; Snyder & Stromswold, 1997; Campbell & Tomasello, 2001; Goldberg, 2006). The research in first language acquisition has found ditransitive constructions interesting for a number of reasons, the most important ones being the fact that they are acquired and demonstrated relatively early in children’s speech, that they share relatively stable semantic properties (some sort of ‘transfer’ being the recurrent theme) and that they are cognitively complex (the action they denote involves three participants) (Dixon, 1991). Many studies that have dealt with the acquisition of the ditransitive construction have looked into the order in which datives are acquired. The results of these studies have consistently shown that the acquisition of (ditransitive) double-object construction arises earlier in children’s speech than the prepositional ‘to/for’ dative alternations (see Bowerman, 1990; Pinker, 1984; Snyder and Stromswold, 1997; Campbell and Tomasello, 2001). The argument concerning the order of acquisition, which has also been the subject of research in second language acquisition, has been used to support the claims of what constitutes a derived construction and which formal pattern can be regarded as the ‘source’ (see Sánchez-Calderón & Fernández-Fuertes, 2018); for instance, if double-object constructions are performed earlier than ‘to/for’ datives, one could use this finding as a part of the argument for the claim that ‘to/for’ datives are in fact the ones that are derived (and not the other way around). The research on the acquisition of ditransitive construction has revealed that children master such patterns at a relatively early age – the age range at which the children master the double-object construction tends to vary between 1;8 (years; months) to 2;11, whereas when it comes to the prepositional alternation or the ‘caused-motion construction’, children tend to master it sometime between 2;0 and 3;4 years of age. In some cases, they master these constructions simultaneously, but the tendency seems to be that the prepositional pattern is acquired later, and the temporal gap between the production of these constructions in early language may be as long as 12 months (Snyder & Stromswold, 1997). The hypothesis that children produce double-object constructions before the prepositional to-dative was corroborated by Campbell and Tomasello (2001), although the overall age at which the children acquire datives was later than posited before; for example, one child first produced a dative construction with 2:9 years of age. Their conclusion with regards to the age of acquisition was 61 that the strong correlation between the points at which children start to produce these two constructions suggests a grammatical relatedness between the two. One of the first major papers on the acquisition of datives was supplied by Gropen et al. in 1989, where they argued against the claims that input plays a major role in children’s acquisition of dative constructions. Similarly to this one, an important finding that raised a lot of questions regarding the acquisition of ditransitive construction was the one by Snyder and Stromswold (1997), whose comparison of input and output led them to infer that the frequency of certain ditransitive patterns in the input children receive did not play a role in terms of their output. Although their focus was on the age of acquisition and whether the increased frequency of ditransitive constructions with the verb give might facilitate the process of acquisition, the absence of correlation between the recurring patterns in input and the age of acquisition was indicative of the input’s overall role in the early language development. Such findings were challenged by Campbell and Tomasello (2001, p. 257), who also expressed some methodological concerns regarding the previous works, arguing that it would be far more beneficial to approach the matter by numerically separating instantiations according to the verbs used and applying a cross-examination of dative-construction frequencies between child and adult language. Their results pointed in a different direction to that of previous research, i.e. the general similarity between the frequency of use when it comes to specific verbs revealed a strong correlation between the adult input and children’s output. The data was interpreted within the framework of usage-based models, which predict exactly that the verbs heard most often by the children would reflect on their language use, and in turn produce similar tendencies in terms of frequency of use– children did, indeed, produce most frequently those verbs which they heard most frequently. Moreover, this lead them to believe that children indeed learn double-object dative first precisely because it was the construction they tended to hear most often in their immediate linguistic environment. As stated, Gropen et al. (1989) argued that the frequency of a particular construction in the input (the language which they hear from adults in their immediate environment) does not affect children’s acquisition of the double-object dative. Their analysis was also conducted on the data from CHILDES, or more specifically, on the data from Brown’s and MacWhninney’s corpora. Their results suggested that neither of ditransitive alternations (double-object and the prepositional dative) consistently emerged first in the acquisition, despite suppositions that one precedes the other in the acquisition order. In the end, they reported that “about 95% of the 62 children's double-object sentences (tokens) and about 86% of the verbs (types) they use in double-object sentences could have been based on argument structures acquired conservatively from adult speech” (1989, p. 220). However, they emphasized precisely the smaller portion of the double-object uses which did not correspond to the input, thereby rejecting “strict” conservativism in favour of the “weak” version; although not frequently, children did demonstrate to use verbs in the double-object construction that they have not encountered previously in the same construction (in the adult input). When Campbell and Tomasello (2001, p. 259) made an objection to their research, stating that a more beneficial way to explore the role of the input would be to look at lexically specific relationships (rather than combine all verbs), they opted for a more sensitive analysis, observing each verb used in double-object datives and prepositional datives, comparing child and adult frequencies for each of the verbs separately. Their analysis revealed that children and adults had the same preference for 21 verbs (out of the total 26 they observed) when it comes to the choice of alternation in which they tend to use it (whether it is double-object dative or the prepositional dative). Their preferred choices in ditransitive alternations differed with 5 verbs, 3 of which the child preferred to use in the prepositional dative, whereas the parent did not (for the other 2, it was the other way around). Finally, their conclusion was that children indeed tend to learn the double-object dative first because this was what they heard in the input most often (also, the input that they received from their parents seemed to favour the double-object construction in the ration of 2 to 1). When it comes to previous research on verbs that appear in ditransitive constructions, such as give, bring, and show, Bowerman (1990) looked into longitudinal data on the spontaneous speech produced by her two English speaking daughters in order to see whether there are any innate linking rules that might facilitate the acquisitional process. Her results supported the claim that linking rules were learned rather than inherited innately; when it comes to mapping thematic roles onto syntactic functions, the children had difficulties with verbs that one would presume were easier to link than some others. Instead, she proposed that children clearly learn linking rules and that their early production of errors ought to be interpreted as a side-effect of “overregularizations of a statistically predominant linking pattern to which they had become sensitive through linguistic experience” (ibid.). 28 Overall, Bowerman's diary-based 28 The proposition that linking rules might be innate was famously suggested by Pinker (1984; later revised in 1996), who hypothesized that children somehow seem to have a pre-set understanding of lexical and syntactic categories such as N, V, NP and VP), whereby they are able to recognize the major thematic roles in the sentence. According to this account of grammar acquisition, the child’s first challenge in the learning process is to locate 69 Table 3. Distribution of different ditransitive patterns across age groups (relative frequencies normalised per million tokens and rounded to the closest unit) 29 0-3 4-6 18+ V-PRO-MOD-N 333 627 1068 V-PRO-N 90 197 135 V-N-MOD-N 28 43 133 V-PRO-PRO 62 90 48 V-MOD-N-MOD-N 8 30 38 V-N-N 22 18 21 V-MOD-N-N 2 5 7 V-N-PRO 4 5 5 V-MOD-N-PRO 0 2 0 TOTAL 549 1017 1455 Table 4. Examples of the produced ditransitive constructions in the CHILDES corpora (patterns from Table 3 indicated in square brackets) Construction pattern Example Corpus Age V-PRO-MOD-N My mommy [gave me the candles from Charlotte] cause those candles were going down all was going to melt. MACWHINNEY 3;3;15 V-PRO-N Okay then you'll have to [earn me money]. KUCZAJ 4;9;12 V-N-MOD-N I got ta [show Gil some of my pictures]. BROWN 4;2;17 V-PRO-PRO Will you [read me it]? MACWHINNEY 3;8;3 V-MOD-N-MOD-N And a and a little woman [giving a little child a mouse]. MACBATES 5;0 V-N-N Hulk's girl friend we [call Hulk Woman]. MACWHINNEY 3;10 V-MOD-N-N I [give big tiger a bracelet]. MACWHINNEY 2;6;17 V-N-PRO You [give mummy one]? BELFAST 18+ V-MOD-N-PRO Daddy [give the ball me]. MANCHESTER 2;1;7 Furthermore, as expected given the initial outline of the specific pattern frequencies, the Spearman's rank correlation reveals very high values for inter-group similarities (N=9). The value is slightly higher for the similarity between 4-6 age group and adults, though the margin 29 The data on frequencies concerning the distribution of ditransitive construction patterns across age groups has already been published in Proroković and Malenica (2023). Also, note that the zeros in the table do not necessarily represent absolute zeros (i.e. a complete absence of the pattern), but they do indicate that the pattern has been observed less than 0.5 times across million tokens. The same is true for other tables concerning relative frequencies of different construction patterns in the thesis. 70 is small (see Table 5). The highest similarity in terms of rank ordering is between 0-3 and 4-6 age groups, but given the very high values for all, the overall take away is that all three age groups exhibit more or less the same hierarchy of use in terms of specific ditransitive patterns with respect to object elaborateness. Table 5. Spearman’s rank correlation of the observed ditransitive patterns across age groups 0-3 4-6 18+ 0-3 1 0.983** 0.950** 4-6 1 0.967** 18+ 1 *ρ=0.683 p=0.05 **ρ=0.833 p=0.01 The analysis of ditransitive constructions across age groups reveals similar tendencies in terms of their argument structure and complexity. To begin with, the most frequent ditransitive pattern is the one with the pronoun as its first object (indirect one) and the complex nominal as its second (direct) object. Considering that the recipient in the ditransitive construction tends to imply an animate being, it is not surprising that the most frequent selection for the argument takes on the form of a pronoun. Even in the non-prototypical cases of ditransitives with proverbial character, it is not rare to find the neuter, third-person singular pronoun it (e.g. (…) so I feed it this giant piece of meat; HALL, 4;6), and sometimes this may be the case due to inverted word order (e.g. Miriam gave it me; SUPPES, 2;0;24). When it comes to the theme, the preferred choice in the most frequent pattern is the complex nominal. 30 There is no doubt that even the youngest age groups have mastered the use of all double-object ditransitive subtypes, at least constituency-wise, and that the trends are remarkably similar between the age groups. However, there are just enough differences that would account for the expectations of children’s language being structurally simpler, while at the same time allowing us to maintain the position that children’s language relies on the input to a degree where it becomes almost imitative in character. The increase in the overall complexity (see Section 1.4 for a construction-specific use of the term) is observable as the age progresses. The increase can be observed in the most frequent ditransitive pattern V-PRO-MOD-N (e.g. (…) he's the one who buyed me cotton candy; 30 Note that a complex nominal does not necessarily denote a huge degree of complexity; in fact, the most frequent variant is a single noun simply preceded by an article (here designated as MOD, which includes both determiners and adjectives in this case). 71 MACWHINNEY, 4;4;1), but also in other variations—when the constructions containing ‘complex nominals’ are grouped together and compared the spike in the complexity becomes even more evident. In the 0-3 age group, the percentage of constructions which incorporated ‘complex nominals’ in some of the slots equals to 67.46% of the overall uses, in the 4-6 age group it amounts to 69.49%, and in the adult corpora it takes up 85.7%. As can be expected, children prefer to use shorter and simpler constructions, at least in terms of constituents that make up their phrases. In many cases, the variation may be explained simply by the fact that children are more likely to drop determiners before nouns, as well as the fact that adults use more clauses with verbs such as tell and show. Furthermore, one needs to be careful when ascribing the syntactic complexity to the construction as a whole. In other words, it is possible to suggest that the measurement of complexity applies solely to the nominal elements included in the observed pattern, and not to the construction itself. Children may very well acquire the construction and only later progress in terms of more complex nominal elements (arguments) that they use as a part of the ditransitive construction. And yet, given that the ditransitive construction tends to be affected by the length of the argument, or at least the choice to realize either the double-object or the prepositional dative construction (see Bresnan et al., 2007), we may safely assume that the complexity of arguments associates closely with the structural complexity of the overall ditransitive construction. Naturally, the data from Table 3 indicates that the complexity of the nominal elements in the ditransitive construction increases with the progression of age. For example, the results indicate that the position of the second (direct) object in the ditransitive constructions tends to be occupied most by a complex nominal in the adult speech, regardless of the fact that it is child directed. The first look at the data indicates a rise in the total use of ditransitive constructions. Clearly, the amount of overall use of ditransitive constructions incrementally rises with age – first, it almost doubles, and between the 4-6 age group and the adult group it increases by approximately 30%. Ditransitive constructions are evidently mastered by children younger than 4, but the overall use is still significantly smaller than the one in older age groups. What does this tell us about the use of ditransitives and the children’s earliest utterances? Given that we are looking at relative frequencies (normalised per 1 million tokens), the data illustrates the fact that the overall language use by either age groups is occupied with a particular portion of produced ditransitive instances. In other words, the more ditransitives there are, the less of other 72 types of utterances are left in the overall language uttered by the speaker. This in turn means that this type of data reveals also something about what is not included in this analysis, and yet was produced by the speakers. It could be posited that such discordance in the data can be explained by the fact that earliest age groups prefer (or require) to use other types of constructions in their everyday communication. A more accurate assumption would probably be that the overall speech of children recorded in various CHILDES subcorpora contains a lot of incoherent children’s talk, such as single, double, or triple-word utterances with little or no meaning. The distinction between the last two assumptions is particularly important because it leads to two different conclusions. For instance, the first assumption (that children may simply be using more of other types of constructions which then occupy a great part of their overall language production) does not really help in terms of inferring about the overall language complexity of early age groups; this is a part of the reason why this thesis observes the use of other constructions as well. The second assumption (that there is a lot of incoherent children’s talk occupying the rest of the produced language), which is probably more accurate, helps put the data into perspective and shows the degree of overall linguistic prowess that we ascribe to particular age groups. Clearly, there is nothing groundbreaking about finding that the language of children younger than 4 might be simpler than that of adults, or that of 4-6 age group. However, it does reveal to what degree this might be the case, and when it comes to ditransitive constructions, it appears that the use of them increases dramatically and progressively between the studied age groups. Table 6. Distribution of different ditransitive patterns across age groups (relative frequencies normalised per million tokens; some patterns are merged together) Normalised frequencies per million tokens 0-3 4-6 18+ VPRON/MOD-N 423 824 1203 V-N/MOD-N-N/MOD-N 60 96 199 V-PRO-PRO 62 90 48 V-N/MOD-N-PRO 4 7 5 The data from children’s speech agrees with arguments about pronominality 31 as being one of the factors for predicting the realization of the ditransitive (see Collins, 1995, Bresnan & 31 The term is used to dissociate phrases headed by pronouns of any kind from those headed by nouns. 73 Nikitina, 2003, Bresnan et al., 2007) that is, the double-object ditransitive is inclined to incorporate the pronoun in the position of the first (indirect) object, but prefers to avoid the pronouns in the final position. When the production rates of ditransitive construction patterns are merged together to better reflect the disparity in the use of (modified) nominals and pronouns with respect to the object positions they occupy (see Table 6), the Chi-square test shows that the difference in ditransitive production rates is significant, χ2 (6, N = 4) = 62.9231, p < 0.05. When the patterns are merged into two groups depending on the part-ofspeech of the second object (V-NP-N/MOD-N and V-NP-PRO), the Chi-square test again shows that the difference in ditransitive production rates is significant (χ2=55.32, df=2, p<0.01), with the ratios between the two construction patterns increasing significantly towards the older age groups (this is especially visible in the adult group, where the pattern which takes a noun as its second object is represented to the greatest extent). This is why it is particularly interesting to observe the difference between the age groups in the pattern V-PRO-PRO, which is also suggested to be a non-canonical ditransitive that sounds awkward (cf. Haspelmath, 2015, p. 10); in relative terms, the overall proportion of said pattern is highest in the 0-3 age group. When it comes to 0-3 age group, it takes up around 11% of their overall ditransitive use, in 4-6 it takes up around 8%, while adults only use it in 3% of the cases. One obvious possibility might be related with the lexical richness in the early age groups, where it may have happened that the child could not retrieve the adequate lexeme needed in the situation at hand, and opted for the pronoun instead; instead of saying give me the teddy, the child may have said give me it/him (e.g. I want you to give me it; BELFAST, 2;6;16). However, this explanation is not necessarily sufficient considering the number of arguments even the youngest age groups have been observed to use in the ditransitive construction (see Proroković & Malenica, 2023). An important point regarding the corrective function of the frequency itself, suggested within the usage-based tradition (Lieven, 2010), might be further challenged here, i.e. if we imagine the construction to be a non-canonical one which needs grammatical rectification, 32 we have to think about the reasons which moderate this drop in the overall portion of such construction 32 The fact that something is rather unorthodox in grammatical terms is not integral for the claim about the corrective purpose of the input, and in fact, it even reduces the strength of the argument to a degree. In other words, if the pattern is a non-canonical one grammatically speaking (e.g. there may be more “elegant” syntactic alternatives for the same expression), it would require an even greater reinforcement in the input in order to be produced by children. The situation in this case is the opposite one – the supposedly non-canonical pattern which is seldom found in the input is produced more in children’s speech in relative terms. This can also indicate that the pattern is not so syntactically unorthodox, or that the noun/pronoun distinction is not that relevant in the early developmental context. 74 use. The reason may very well be that the sole frequency in the input bears the corrective function; low frequency in the input may account for the progressive drop in the use of such ditransitive pattern – the children are able to observe that such use of the ditransitive construction is not common and infer that the pattern V-PRO-PRO is generally best avoided. However, the results still remain counterintuitive, given the greater relative production of the pattern in the younger age groups. It is important to note that a moderate anti-nominal inclination has also been observed in some periphrastic causative construction patterns (see Section 4.2.2), though never to the degree where the use of pronouns overshadows the use of nominals in the younger age groups – especially if such relation is not attested in the input. 3.2.3. Lexical arrangement and overlap The following part of the analysis explored the most frequent verb choices within the ditransitive construction (see Table 7). As heads of phrases, verbs have been traditionally interpreted as focal points in the acquisition of language and constructions (see Goldberg, 2006). Both ideas of the so-called ‘pathbreaking verbs’ and ‘verb-islands’ recognise the importance of prototypical items frequently found in certain constructions, which are acquired first (and if not first, they show a tendency to appear early on) and which tend to have a stimulating role in terms of subsequent language development. The ‘pathbreaking verbs’ hypothesis (see Ninio 1999, 2003) is, in some respect, similar to Tomasello’s early ‘verb-island’ hypothesis (1992), later rephrased and altered into ‘item-based’ hypothesis. The two hypotheses differ with regards to whether the relationship between different verbs is fluid within and across constructions; in some sense, the verb island hypothesis views the syntax of a particular construction as initially learned for each verb separately, meaning that the language of children consists of item-specific patterns initially treated separately throughout grammatical development, and only later generalised and understood in terms of constructions. While the ‘verb island hypothesis’ agrees with rote learning principles, the ‘pathbreaking verb’ hypothesis resonates more with what is known as the ‘Subset Principle’ (see MacWhinney, 2005, p. 59) where the grammatical forms are envisioned as parts of the slightly more complex grammars, and where the child generalises the use of a particular verb across competing constructions (e.g. across alternating dative constructions) if the same has been attested in the input (see Fodor & Crain, 1987). However, the two theories agree on the fact that the learning of a particular 75 construction begins with prototypical verbs heard frequently in the input, which later guide and facilitate the acquisition process of other lexical variations within and across constructions. These results do not directly test the ‘pathbreaking verbs’ hypothesis, but they certainly show an incredible disparity in the use of certain verbs in ditransitive constructions. The data further corroborates the claim of the verb give being the most representative example in the ditransitive construction, and probably the ‘pathbreaking’ one in most of the cases. More importantly, the data also shows an important role of the following 4 most frequent verbs in the production of ditransitive construction. This type of comparative analysis provides valuable data relating to the item-based language learning hypothesis (Tomasello, 2000b, 2001c; MacWhinney, 2014), and contributes additional evidence on the relationship between children’s produced utterances (output) and the language they heard in their environment (input). 33 In terms of synchronic comparative data, the claim transposes to expectation that certain verbs will cover most of the ditransitive uses, as well as that overall use of ditransitive constructions might be headed by fewer number of verbs in early speech. Table 7. Distribution of ditransitive verbs across age groups 34 Verbs/ Age 0-3 4-6 18+ give 1073 42.65% 798 41.85% 11162 37.25% tell 233 9.26% 367 19.24% 6295 21.01% get 374 14.86% 207 10.85% 2654 8.86% show 157 6.24% 124 6.50% 1885 6.29% make 186 7.39% 71 3.72% 1405 4.69% buy 91 3.62% 92 4.82% 950 3.17% bring 75 2.98% 40 2.10% 867 2.89% read 68 2.70% 37 1.94% 600 2.00% find 22 0.87% 5 0.26% 407 1.36% ask 5 0.20% 20 1.05% 385 1.28% 33 Tomasello (2001c) himself stated that these claims are best challenged by negative evidence, that is, by observing which items do not appear in certain patterns, despite their otherwise common presence in the input. This type of analysis unfortunately falls beyond the scope of this research as well. However, while it is valid to maintain the usefulness of such data, the method would not be without clear limitations and drawbacks. In the context of CHILDES, considering that we do not possess data on the entire linguistic input that the child has received throughout their lifetime, we cannot derive conclusions, not with absolute certainty at least, about their atypical overgeneralisations or supposedly creative language uses in terms of deviations from what they have heard. It would surely be possible to make a leap of faith based on some justified assumptions, but there is no doubt that it presents methodological issues from a rigid perspective on what qualifies for data comprehensive enough for such testing. 34 The table on frequencies concerning the distribution of ditransitive verbs across age groups has already been published in Proroković and Malenica (2023). 76 Again, the overall resemblance in the use of particular verbs in ditransitive constructions was evident. By far, give is found to be the most frequently used verb in double-object ditransitives, accounting for more than a third of ditransitive constructions in each of the studied age groups. Frequency-wise, the verb is followed by tell, get, show, make and other, with relatively small discrepancies in terms of their frequency-based hierarchy of usage. The frequency order by which these verbs are used in ditransitive constructions is remarkably similar between age groups. What is more, on a broader scale, the results reflect the exact usagebased expectations of how output would resemble the input: (1) verbs which are most frequent in the adult language should also be the ones most frequent in children’s language and (2) despite overall similarity in distribution, children’s early language should still focus around fewer verbs than adult language. For instance, if the verb give is the most present one in the input, the same should be reflected in the children’s output when it comes to ditransitive constructions; furthermore, if we assume that the verbs most frequent in the input tend to influence and guide the acquisitional process of the constructions in which they appear, we should also expect its representation to be slightly more frequent in the early language. This is exactly what the data shows when it comes to the verb give. The situation is the same when it comes to verbs get and make, but slightly different when it comes to tell and show. This is to be expected, partly because they are not the most frequent ones and the aforementioned reasoning tends to weaken the further you go down the input frequency lane, but the major reasons might be the fact that tell and show do not constitute a canonical ditransitive construction. In other words, constructions containing these verbs have been described as ditransitives implying a ‘mental’, rather than physical transfer (see Malchukov et al., 2010). Furthermore, if usage-based assumptions were incorrect, one could expect at least some lexical differences across age groups within constructions. For example, why would children so frequently choose the verb give over the verb get in ditransitive constructions. These two verbs share many characteristics in terms of meaning and are frequently used in the same sense. What is more, give tends to have a very restricted meaning, whereas get appears more flexible meaning-wise, implying a wider range of situations aside giving, including the grammaticalized ones (such as the periphrastic causative one; see Section 4.1.2). If one were to completely ignore the effects of the input, one should expect more competition between the two. Indeed, the getgive ratio is larger (the difference in frequency is smaller despite seemingly corresponding 77 disparity in overall proportion between the two) in the 0-3 age group, but it does not seem nearly enough for us to infer any sort of competition between the two. If anything, it seems to be explained by the fact that adults are simply more lexically diverse in their speech. If the five most frequent verbs (see Figure 2) in ditransitive constructions are added together, they constitute approximately 80% of the total use across all three age groups; more precisely, the verbs give, tell, get, show and make, take up 80.40% in the 0-3 age group, 82.16% in the 4-6 age group and 78.10% in the adult group. In some sense, the distribution of verbs in the produced utterances reflects the Pareto principle, although with slightly different rations; for instance, in 0-3 age group, the first 5 verbs take up approximately 10% of different verbs found in ditransitive construction, but in terms of overall use, they take up approximately 80%. The disparity is even more pronounced when it comes to the ‘vital few’ and their productive potential, as they take up approximately 5% of ditransitive uses but also constitute close to 80% of the overall use of ditransitives. Figure 2. Most common verbs used in ditransitive constructions Perhaps the disparity in the usage of tell and show seems to be one of the most interesting elements of the data, for they are not more salient in the youngest category as one might expect (if one supposes that the most frequent verbs would also be most salient in the youngest category because the older ones tend to increase the range of verbs used in ditransitive 0.00% 5.00% 10.00% 15.00% 20.00% 25.00% 30.00% 35.00% 40.00% 45.00% give tell get show make Verbs used in ditransitive constructions 0-3 4-6 18+ 78 construction). Tomasello (2000a) has hypothesised that language tends to be imitatively learned first, which may entail simultaneous or subsequent efforts to comprehend the utterance. 35 Here, Tomasello carefully elaborates the concept of ‘imitative learning’ and adds the element of comprehension, claiming that ‘imitative learning’ is unjustifiably stigmatised in language acquisition research, and that it needs not be interpreted in the context of rigid repetition of verbatim, but rather elaborated with the role of ‘understanding’ and the ability to execute ‘functionally based distributional analysis’. In other words, children need to reproduce the surface linguistic form so that it reflects its conventional communicative intent, thus exhibiting the fact that they understand its underlying function (see Tomasello 1998, 2009). Seemingly, these results counter the idea that certain linguistic patterns are first learned and understood later, given that the frequency of the two verbs in the 0-3 age group clearly does not follow the same frequency-based order of usage. If such hypothesis were true, we might expect that the order will be the same, regardless of the ‘understanding’ element. But given that understating can play an important role simultaneously to imitation, the results actually confirm some of Tomasello’s conjectures regarding early verb usage and language character. The evidence clearly shows that the youngest age group has clearly produced/mastered ditransitive uses of the two verbs, but also that their use is rather restricted when compared to other verbs in the same constructions and the distributions in the other two age groups. The important conclusions that we might draw from this data are the following: (1) The level of abstraction and cognitive capacity surely plays a role in acquisition—in the usage-based canon there traditionally exist an insistence on the mechanisms of reading cues in environment and there is no doubt that more abstract verbs rely on cues less than others, meaning that children should find it harder to acquire such verbs; (2) The two verbs (tell and show) often imply more complex syntax of the direct object (such as clauses), which has partly skewed the results in favour of adults and (3) Children’s language also reflects the variety of needs and situations in which they find 35 Quote from Tomasello (2000, p. 73): “In either case, the main point is that young children begin by imitatively learning specific pieces of language in order to express their communicative intentions, for example, in holophrases and other fixed expressions. As they attempt to comprehend and reproduce the utterances produced by mature speakers – along with the internal constituents of those utterances – they come to discern certain patterns of language use (including patterns of token and type frequency), and these patterns lead them to construct a number of different kinds of (at first very local) linguistic categories and schemas. As with all kinds of categories and schemas in cognitive development, the conceptual “glue” that holds them together is function…” 85 Figure 4. First (indirect) object animacy A more detailed look into first object animacy tells us that the recipients are most commonly represented by names (ANp) in 0-3 age group and adults, whereas the most frequent choice in 4-6 age group are general animates (ANc; see Table 12). Animate toys (ANt) occupy the third position in 0-3 age group, while in 4-6 age group and adult group, they share the third spot (in terms of frequency) with inanimates. Body parts (ANb) are almost non-existent in the first object position. Again, 0-3 age group resembles the adult one more in terms of their most frequent choices, as predicted by the first hypothesis, while at the same time they remain reluctant to employ inanimates (INAN) to a greater degree. The results are in line with the usage-based framework, whereby children remain reluctant to produce patterns seldom encountered in their input – or at least – they tend to produce it even more rarely than these would otherwise be encountered in the input. Early language is primarily guided by the most frequent prototypical instances, and as the children grow older, the range of patterns to which their language use opens up expands significantly. 0-3 4-6 18+ Inherent Animates 79% 83% 70% Contextual Animates 16% 9% 15% Inanimates 5% 9% 14% 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% Proportion of use 1st object animacy (type-based) 86 Table 12. First (indirect) object animacy 0-3 4-6 18+ ANc 39% 54% 24% ANp 40% 28% 47% ANt 16% 9% 14% ANb 0% 0% 1% INAN 5% 9% 14% As far as animacy of the second object is concerned in the ditransitive give-NP-NP, the cross-age group comparison surprisingly indicates a more constrained use of object types with respect to their animacy. While the recipients were mostly represented by inherent animates, the themes of ditransitive constructions were overwhelmingly represented by inanimates. In the 0-3 age group, they appear to take 89% of overall use, whereas in the 4-6 and adult age groups, they surpass the 90% margin. In contrast to other findings so far, the 0-3 age group does not show the greatest saturation in the most frequent category, but instead demonstrates greater dispersion (albeit to a small degree) between the two remaining categories – inherent and contextual animates. Figure 5. Second (indirect) object animacy 0-3 4-6 18+ Inherent Animates 6% 3% 3% Contextual Animates 4% 4% 3% Inanimates 89% 94% 94% 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% 100% Proportion of use 2nd object animacy (type-based) 87 One possibility is that, prior to 4 years of age, children remain more flexible with second object realizations for some reason; what can be transferred becomes more constrained with the progression of age (which seems opposite to the data on first object animacy). The less conservative use of second object animacy suggests that children possibly go beyond the input and that their perception of what qualifies for a theme in the ditransitive give-NP-NP is possibly determined by other factors. To be fair, the resemblance between the two groups is striking, and there is no particular reason to suspect considerable differences in the treatment of items between the three age groups. Nevertheless, the small differences could be triggered by some linguistic, or even extralinguistic factors. For example, it is possible that children early on believe that anything can be given, regardless of whether the object of transfer agrees with the transfer itself; that is, the case where animate themes are employed are relatively rare in English language, and when they are employed – they require some compliance by the person who is being given, or a greater effort (some sort of consideration) at the hand of the person who does the giving, and potentially a greater willingness to receive what is being given at the hand of the recipient (e.g. I gave my husband a son/a dog). This usually assumes either shared knowledge between the speakers (cf. Kecskes & Zhang, 2009) or children's ability to read other people's intentions (cf. Tomasello, 2001b), and children’s egocentrism and lack of common ground comprehension may result in the effect observed (no matter how small). 3.2.5. Lexical richness Mean length of utterance 38 and lexical density 39 were calculated in order to provide more information on the structural and lexical complexity of the sentences containing ditransitive constructions. Although the measures are directly related with the ditransitive constructions, they capture the sentences in their entirety (as they are marked in the CHILDES corpora) and many of them contain other types of constructions given that they may consist of more than one clause, or have ditransitive construction embedded in more complex syntactic 38 Since the intent was to capture all of the sentences used in the ditransitive pattern (as is the case with subsequent analysis of other constructions studied), MLU was calculated with the help of the CLAN program that is not restricted to syllables or characters as most automated text analysis tools available online, but is able to morphologically parse the text once the MOR grammar is integrated and calculate the required measures. 39 Lexical density is calculated as the ratio between content or lexical words (nouns, verbs, adjectives, adverbs) and the total number of words. 88 patterns. Additional constraints on the results are imposed due to the fact that the corpora represent transcripts of spoken language, and the choices where to end a particular sentence (to put a full stop) in the process of transcription may not have always been consistent or carefully considered. Nevertheless, this type of information ought to provide context in terms of ditransitive constructions’ immediate syntactic and lexical environment. Table 13. General overview of the sentences containing give-NP-NP across age groups 40 0-3 4-6 18+ Token count 6975 6960 99654 Word type count 682 799 3283 Character count (without spaces) 25215 24960 365272 Utterance count 1014 798 11205 Open class words 3008 2963 43926 Closed class words 3860 3868 53936 MLU (morpheme-based) 7,16 9,21 9,30 The results indicate a rise when it comes to the MLUs of the sentences containing ditransitive constructions with the progression of age. When it comes to the spoken language of children aged from 2;6 to 4;0, the previous data shows an increase in average MLU (calculated in words per sentences) going from 2.9 up to 4.1, and when it comes to children aged 4;0 to 6;11, it progressively increases from 4.1 to 4.7 (see Rice et al., 2010). Naturally, the average MLU in our data is considerably higher (morpheme-based MLU is 7.16 for 0-3 age group, 9.21 for 4-6 age group and 9.3 for the adult group), partly because the limitation on the type of sentences included (those containing ditransitive construction) and partly because of the aforementioned reasons (the occasional problematic identification of actual sentence-endings). Despite the amount of nonce and unintelligible words, which were recorded in CHILDES and undoubtedly more frequent in younger age groups (although CLAN eliminates most of them in the analysis), and which prolong the sentences at least in terms of their length, the results still indicate an increase with the progression of age. In other words, children produce shorter sentences than adults when it comes to those utterances containing ditransitive constructions, and those older than 4 produce longer sentences than the younger ones. The increase is much 40 Note that the token and word type counts do not include repetitions or revisions (for this table, and the subsequent tables for the corresponding constructions), as set by default in the CLAN program and the applied 'eval' function (for detailed explanation, see MacWhinney, 2023a, pp.132-133) Also, the function excludes gesture-only utterances from the total utterance count, while the morpheme-based MLU evaluation excludes morphemes in utterances representing unintelligible words transcribed as 'xxx', 'yyy' or 'www' codes. 89 smaller between the spoken language in the 4-6 age group and that of adults, but it is apparent between the two younger age groups observed; when it comes to word-based MLU, the difference almost amounts up to entire two words per sentence. This is a significant increase in the context of the data provided by Rice et al. (2010) which revealed how long it takes for children to increase their sentences by approximately two words, and how this progression somewhat slows down after 4 years of age. The results of this research are consistent with their findings, but in light of additional evidence from the adult group, they also suggest a potential approach to MLU’s growth limit; that is, there are three possible explanations: (1) 4-6 age group comes very close to the adult language in terms of how elaborate their spoken utterances tend to be, (2) the increase might be bigger if the adult speech was not mostly child-directed in character, (3) MLU is not a particularly reliable measure of language development. The last argument appears least credible considering the bulk of research that suggests otherwise (see Rondal et al., 1987; Rice et al., 2006; Wieczorek, 2010; Pavelko & Owens, 2019 etc.). So the question is, where exactly do these sentences differ across age groups? Do children produce shorter sentences because they find it harder to retrieve crucial content words? Is it because their language is poorer in adjectives or adverbs, or is it the function words that they tend to omit? The data on lexical density provides certain clues on where exactly their spoken language appears to differ. The similarity is not to be disregarded considering that lexical density can oscillate substantially, albeit more observably in terms of spoken-written language (Ure, 1971; Halliday, 1985). When it comes to lexical density on its own, some research has shown that it does not tend to be sensitive to age as expected, or that is beneficial to observe it on its own when examining cross-age group differences (Johansson, 2008). However, if this data is considered in the light of MLU, it can indicate where the differences between the studied age groups tend to occur. In other words, what can we infer from the fact that the utterances between age groups differ in length but do not differ in lexical density? First, it means that the expansion of sentences in the 4-6 age group and adults does not necessarily contribute to the number of content words – given that the ratio stays similar (see Table 14), we could conclude that function and lexical words participate equally in the sentence-expansion process. Moreover, the lack of difference between the observed age groups might suggest that children tend to drop some of the function words from their speech (which would agree with one earlier assumption regarding article dropping); for example, if we imagine the adult language to be saturated more with 90 modifiers such as adjectives and adverbs, we should also expect a higher level of lexical density in those age groups. Considering that this does not happen, the data might suggest that the reason for this might come at the expense of functional words such as articles, which might be dropped from children speech, and which could in turn inflate the final numbers on lexical density. In fact, this is not surprising at all – despite their significantly poorer vocabulary, children tend to orient themselves on content words in their speech (especially early on), elaborating them with gestures to convey the final message (e.g. daddy look, doggy). If anything, when “gibberish” is removed from early language, content words would be expected to dominate (especially in the youngest age groups). On the one hand, we have a rather poor lexicon in the youngest age groups which negatively affects lexical density as they find it hard at times to retrieve the content words they require, while on the other hand, we have a tendency of early language to orient itself on the heads of phrases that seem to carry most of the meaning. In the adult language, the situation is somewhat opposite, where the lexical prowess is maximised along with the syntactical one, resulting in sentences which are both well-formed and well-informed. Data on both MLU and lexical density seem to agree well with such view on the opposing processes which participate in determining lexical and structural complexities of adult and early speech. When it comes to the basic measure of lexical density, there is little or no variation between all age groups in utterances that contain (or centre around) the double-object give ditransitive constructions (Table 14). The values seem to be somewhat surprising, given that an isolated ditransitive construction should be more favourable to content words in the LD ratio. For example, a sentence such as give the man a job would give an LD of 0.6 (three content words and two grammatical words). In fact, the LD values for isolated give ditransitive constructions in 0-3 age group was approximately 0.66. This means that ditransitive give constructions tend to take up only a segment of the utterances produced, and that the rest of the utterance compensates for lack of function words. What is more, we could expect lexical densities in this case to be even higher in the younger age groups because they should produce shorter sentences without going beyond the parameters of the construction at hand. In that sense, a greater number of such short but concentrated utterances would produce higher values of LD. 91 Table 14. Lexical richness for the ditransitive give-NP-NP Lexical density and sophistication Lexical diversity: range and variation LD LS1 CVS1 NDW NDW-ER50 MSTTR-50 RTTR SVV1 0-3 0.44 0.25 0.04 257 29.70 0.56 7.02 2.12 4-6 0.45 0.23 0.24 360 31.10 0.61 8.77 6.62 18+ 0.45 0.26 0.38 379 33.80 0.64 8.89 8.24 When it comes to lexical sophistication (LS1), the numbers are again similar across age groups. Approximately, there seems to be one relatively rare word (lexical word that does not belong in the top 2000 most frequent words according to the BNC) per 4 lexical words in total across all age groups, that is, one sophisticated lexical word and three common lexical words on average in utterances containing ditransitive constructions. Although it may be redundant to scan for the amount of sophisticated verbs because the utterances centre around a constant ditransitive give, the analysis still reveals a spike in sophisticated words with the progression of age (CVS1). Such results imply that with the progression of age, other elements (including sophisticated verbs) are added to ditransitive constructions in the produced utterances (e.g. give the man a ball or else I’ll force you). Although a measure of diversity rather than sophistication, the SVV1 index concerning verb variation also shows an increase with the progression of age. 41 Not only do 4-6 age groups and adults use less common verbs more often, but they also see to use a wider range of the verbs overall even in sentences revolving around ditransitive constructions. The question is what role does the complexification of a larger language unit such as a sentence bear on the children’s processing of the embedded construction, or in this case, what does utterance complexification in the input bears on the children’s processing of the embedded ditransitive? When compared to lexical density, lexical diversity seems to have a somewhat steeper curve as far as utterances containing ditransitive constructions are concerned (see Figure 6). This is not necessarily surprising given that the two indices do not always demonstrate stable correlations amongst themselves and that they measure very different concepts. However, when interpreted together, they tell us where exactly the changes occurred in the samples. Considering that the density remained consistent across age groups and diversity grew, we can 41 Once again, note that the effects are not brought about by the disproportion in the number of utterances produced by age groups because the values in question were calculated on a random sample of utterances that was equalized across age groups. 92 assume that diversity was not elevated due to prevalence of either lexical or grammatical words. Instead, the utterances in the speech of children aged 4-6, and then in adult speech, have either grown proportionately in the number of lexical and grammatical words or they have stayed the same in terms of sheer volume, but with an increase in the number of distinct word forms in both of those cases. When we add the calculation of MLUs to this assessment, which tells us that the average length of the utterance did increase with the progression of age, we begin to understand that both length and diversity increased, but the change in lexical and grammatical words-ratio was absent in the case of utterances with embedded ditransitive constructions. An increase similar to the one expressed by MSTTR-50 can also be observed with RTTR. Figure 6. Density and diversity in the ditransitive give-NP-NP One of the most reliable indices pertaining to the range of lexical items employed by the speaker has been the number of different words in the sample, especially in the context of child language development (cf. Klee, 1992; Miller, 1991; Moyle et al., 2007; Hadley et al., 2018 etc.). When it comes to utterances centred around ditransitive constructions, both the total number of different words (NDW) and the averaged NDW of 10 random subsamples also show an increase in the variety of lexical items employed with the progression of age. The size of the increase is not particularly large, but along with the measure of sophistication, it indicates where 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0-3 4-6 18+ density (LD) diversity (MSTTR-50) 93 the increase in variation occurs. Given that the indices of general lexical sophistication were relatively stable (even lexical sophistication which takes into account the types; see the Appendix 10.2), whereas verbal sophistication slightly increased, the data suggests that the ageincreased variety (regardless of its size) in utterances containing ditransitives was anchored in some common general lexical items, as well as some less frequent verbs (at least in terms of the entire sentence). In turn, the nature and volume of differences in view of lexical richness suggest that adults (at least as far as the language surrounding ditransitive constructions is concerned) moderate their language use around children. In order to be understood better, the gravity of increase between 0-3 and 4-6 age groups has to be additionally assessed in relation to the other constructions observed in this research. 3.3. In sum • There is a significant rise in the use of ditransitive constructions with the progression of age. The amount of overall use of ditransitives almost doubles between the 0-3 and the 4-6 age groups, and the increase continues with difference of approximately 30% between the 4-6 age group and adults. • When it comes to ditransitive constructions, the data shows a significant degree of similarity across age groups in the distribution of both syntactic patterns and lexical items. Although there is a considerable overlap in the proportion of particular ditransitives with respect to their argument structure and complexity, it can also be noted that the overall complexity increases with the progression of age. Children use shorter and simpler constructions, at least in terms of constituents that make up the constructions they produce. • The overall similarity in the use of ditransitive verbs was noticeable; the verb give was found to be the most commonly used, accounting for over a third of all two-object constructions in each age group studied. The verbs tell, get, show, make, and others follow give in terms of frequency, with only minor differences in their order of usage. The results align with the expected usage patterns based on input and output: (1) the verbs most frequently used by adults are also the most frequently used by children, and (2) although the overall distribution is similar, children's early production of ditransitives tends to concentrate on a smaller set of verbs compared to adult language 94 (however, it cannot be said that the 0-3 age group concentrates on fewer items in their usage than the 4-6 age group). • The frequency order of these verbs in ditransitive constructions showed a remarkable degree of similarity across different age groups. The results of the Spearman's rank correlation analysis reveal a strong correlation between the 0-3 age group and the other two age groups, indicating a high and statistically significant similarity in the frequencybased ranking of the most frequent 22 ditransitive verbs used. Interestingly, there appears to be a stronger correlation between the 4-6 age group and the adult group compared to the correlation between the 0-3 age group and the adult group, which contradicts the initial hypothesis. • When considering the entire range of nouns, the analysis indicates a higher level of similarity between the 0-3 and adult group than between 4-6 and adult group. This disparity is particularly noticeable in the second object position, which is surprising given that the wider range of possible items can occupy the second object slot when compared to the first object. • The first hypothesis is confirmed across several levels; there is a considerable similarity across age groups with respect to specific nouns, verbs, and syntactic patterns in terms of frequency rankings. However, the hypothesis predicting a greater similarity between the 0-3 age group and adults in comparison to 4-6 and adults is not consistently proven across both verbs and nouns, rendering the data inconclusive in ditransitive constructions. • The analysis of the data regarding the animacy of the first object in the give-NP-NP ditransitive indicates that all age groups predominantly use inherent animates as recipients in this position. The second and third positions, however, show some differences between the age groups in terms of the presence of contextual animates and inanimates. In particular, inanimates are less commonly found in the 0-3 age group, while they are most frequently used in the adult group, with their usage rate doubling as the age groups progress. These proportions suggest that adults are more flexible in their use of inanimates in the first object position compared to children, especially those in the youngest age group. The expectation that the youngest age group would use animate items considered prototypical in designated slots frequently is confirmed as well, but not necessairly to the point where this trend is more pronounced than in 4-6 age group. 101 fact, mentioned as a separate type of causative: “Syntactic causatives are distinguished by many authors from constructions which refer to causative situations but do not represent cohesive units, thus being biclausal sentences” (2001, p. 886). Nevertheless, as they state later on, the distinction between the syntactic causatives and biclausal causatives is not “clear-cut” and it depends on the degree of fusion (for instance, the ‘same predicate’ causative make someone go constitutes a syntactic causative, while the force someone to go represents a biclausal causative). 43 Similar typology can be found in Dixon (2000), as he distinguishes between morphological causatives, causatives making up a single predicate, syntactic causatives (periphrastic or analytic causatives), lexical pairs in causative relation and ambitransitive verbs that can be regarded as causatives, and finally the realization of the causative effect through the exchange of auxiliaries which has been observed in some languages spoken by indigenous Australians (Mangarayi and Ngan'gityemerri). In (English) grammar, the term periphrastic is used for expressions employing two or more words to denote what can usually be replaced by sole inflection (Collins English Dictionary, 2021). In English, the periphrastic causative construction can also be described as consisting of a noun phrase followed by a causative verb and another noun phrase followed by an infinitive, whereby the first noun phrase is the subject of the main causative verb and the second noun phrase the subject of the infinitive (Hollmann, 2003). Causative verbs such as have and make can take a bare infinitive (have/make someone do something), while the verbs get, cause, force and others take a full infinitive form (get/cause/force someone to do something). Naturally, there are other distinctions to consider within the periphrastic causative, but these are at the lower semantic level. For instance, when it comes to analytic/periphrastic causatives, Stefanowitsch (2001) distinguishes between different types of causal links encoded (i.e. thematic roles of the participants). For instance, make-causative can encode both agent-like and stimulus-like participants when it comes to the subject of the main verb (the causer) (ibid., 2001, p. 3): “Agent: Several hours later she saw police arrive and make villagers dig the graves. 43 This research is not interested in the strict separation between the two, and takes both into account when it comes to their acquisition and processing. This distinction between biclausal and monoclausal syntactic causatives is not without its problems, as it seems to rest on the fact that one verb takes a bare infinitive, while the other takes a full infinitive form (for example, consider the help-construction which can take both). For a detailed discussion on differences between monoclausal and biclausal structures, and the ways in which these can be tested, see Batinić Angster (in press). 102 Stimulus: The semi-nude picture of O.J. and Nicole Simpson on the cover of People magazine made me shudder.” Comrie (1976) argues that the causative sentence is characterized by an underlying structure that consist of a matrix and an embedded sentence and that these are fused together in some languages to form a derived structure, where there is no obvious separation of the clauses. The structure is derived from a matrix sentence that expresses the causer of the action in the position of the subject and the embedded sentence with the subject that takes on the role of the one who carries out the action. Nevertheless, Comrie later on argues that English hardly exhibits structures, where there is a “compression of an underlying complex structure with embedding into a derived-structure simplex sentence” (1976, p. 303). However, biverbal causative constructions in English surely exhibit a considerable degree of fusion as well. For instance, in the derived construction John made Mary give the book to John, we have two underlying predicates which can be paraphrased as: John forcing Mary into something and Mary doing that particular something. As noted by Kulikov et al. (2001), patterns such as make NP give can be regarded in many ways as ‘same predicate’ causative construction. There might be some hesitation from Comrie to use English as an example, partly because sentences such as I opened the door qualify for a causative sentence just as I made John go (Shibatani, 1976). To a certain extent, one could even argue that the monoverbal I opened the door might contain a more complex underlying structure, at least on the conceptual level (thus merging the sentences such as I did something and the door opened). A subset of verbs which exhibit this kind of causative sense in the transitive clausal pattern are called ‘labile verbs’, with approximately 800 of them existing in English language (see McMillion, 2006). They can occur in both transitive and intransitive patterns (I opened the door vs. The door opened), In the transitive clausal pattern, these verbs express causative events where the subject initiates a change in the direct object, acting as an external agent, whereas in intransitive patterns, the subjects, typically inanimate entities, are either interpreted as causing the action themselves or are affected without any external cause being relevant (ibid.). Naturally, it is hard to justify this reasoning in terms of language processing, and this is just a part of the reason why such constructions are excluded from the research. In order to distinguish between the lexical and periphrastic causatives (for which he uses the label ‘analytical’, Lemmens (1998, pp. 21-28) uses the example of the verb kill which had been previously referred to as a lexical causative. The lexical causatives and the periphrastic 103 causatives are equated on the basis of semantics – for instance, the expression cause to die and kill were assumed as semantically equal. The fact that ‘kill’ derives from the ‘cause to die’ stems from the generative perspective on semantics – through lexicalization, the verb is thus a derived form of the aforementioned underlying structure. Lemmens claims otherwise even in terms of semantics and claims that lexical and periphrastic causatives “encode different conceptualizations of a given situation” (1998, p. 21). The general definition of the causatives encapsulates verbs that refer to a situation where there is a causal relation between two events, whereby one of the events is believed by a speaker to be caused by another, i.e. causatives refer to verbal constructions that mean cause to V or make V where the V stands for the embedded verb (Nedjalkov & Sil’nickij, 1973; Kastovsky, 1973; Kulikov, 2001). This general definition encompasses sentences such as John opened the door, Peter made John go and the morphological causative suffix in Turkish: Ali Hasan-ı ol-dur-du (Ali:NOM Hasan-ACC dieCAUS-PAST) meaning Ali killed Hasan (Kulikov, 2001; Comrie, 1976). One of the reasons as to why so much focus as been put on defining causative constructions is in many ways connected to cognition and the way linguists assume how we process language, i.e. the fact that we distinguish between the synthetic causatives and periphrastic constructions is in relation to the history of why a verb kill may be, according to some (Lakoff, 1966), treated the same as ‘cause to die’. In order to illustrate why the synthetic causative differs from the periphrastic one, Fodor (1970) presents the case of John caused Mary to die, analysing the interplay between syntax and semantics via the ‘do so’ transformation of it. In the transformation examples from Fodor (1970, p. 431), the contention is that the example (8d) is not possible, and yet it should be for the verb ‘kill’ to be equated with ‘cause to die’ (8) (a) John caused Mary to die and it surprised me that he did so. (b) John caused Mary to die and it surprised me that she did so. (c) John killed Mary and it surprised me that he did so. (d) *John killed Mary and it surprised me that she did so. The argument about the lack of equivalence between the two can be observed in the comparison between (8b) and (8d). Sentence (8b) is possible in English language, and there is no confusion that the speaker will be aware what the ‘do so’ construction is referring to. However, this is not the case in sentence (8d), i.e. it is not clear that ‘do so’ refers to Mary’s dying. Another observation by Fodor when it comes to the difference between the monoclausal and the 104 biclausal causative has to do with the temporal aspect of it. The examples from Fodor (1970, pp. 432-433) illustrate this point: (9) (a) Floyd caused the glass to melt on Sunday by heating it on Saturday. (b) *Floyd melted the glass on Sunday by heating it on Saturday. The point that he makes is that the implication of the periphrastic causative is that one can cause a certain event, or bring about a certain effect, by doing something at a time which is distinct from the particular time of the caused event. If we look at the monoclausal synthetic causative Floyd melted the glass, it is not possible to assume this temporal gap between the cause and effect (i.e. if you do melt something, the action of melting happens simultaneously with the melting of the object in question). As Lemmens (1998, p. 25) puts it: “At the risk of arguing the obvious, I maintain that lexical causatives such as kill almost invariably code events in which the agent's deed and the patient's affectedness fail to coincide referentially. However, unlike analytical causatives (like the cause toconstruction), lexical causatives neutralize this difference and code the components as concurrent and integrated.” Many problems that arise from the definitions of constructions relates to the distinction between the so-called deep and surface structure. For instance, this is at the root of the monoclausal vs periphrastic causative discussion. 44 In terms of deep structure, it is safe to assume that the periphrastic causative is a biclausal causative construction, i.e. there is a subject for each of the events entailed by the construction. However, at the surface level, the biclausal construction can be interpreted as a monoclausal one (especially with causative constructions where the causee takes an intransitive verb) containing the complement phrase that may be paraphrased as a clause. For instance, in the sentence Her friend made Mary cry, the action chain slightly changes as there is no final patient we usually expect to encounter in the periphrastic causative; Mary is thus the patient and the final recipient of the energy initially brought about by the causer. The fact that this is a monoclausal structure is evident in the fact that Mary is the object of the (arguably) phrasal verb make cry. These concerns about the structure of the periphrastic causative may be resolved by the passivization test (Gilquin, 2010, p. 65), which shows that some verbs such as have allow for different interpretations of arguments in terms of the subjectobject relations. Mary was made to cry by her friend differs from the passivized form of the 44 While the discussion on definitions is mentioned here, in the context of this research, it is more important to recognize the difference between the surface forms, partly due to the focus on language acquisition, but mostly because of the corpus methodology and the necessity for consistency and clarity in the analysis. 105 sentence Her friend had Mary kill Jack, where even in the passivized Her friend had Jack killed by Mary, Mary remains the subject of the verb kill. According to Gilquin (2010, p. 65), this can be seen as the proof of the biclausal nature of the causative construction with have. By now, it is important to once again highlight that the thesis focuses exclusively on periphrastic/analytic causatives that take either full infinitives (e.g. get him to do) or bare ones (e.g. have him do). Another biverbal causative construction is the so-called passivized syntactic causative (whereby two words make up a single predicate with the auxiliaries have and get; e.g. get this done), but this is beyond the scope of this research. 45 Syntactically speaking, the ordering of their constituents corresponds to the form NPX→ VCAUSE → NPY → VPEFFECT, where the NPX is the subject of the main verb (i.e. VCAUSE). In the periphrastic variant, NPY is the object of the main verb (i.e. VCAUSE) and the subject of the subordinate verb (VPEFFECT), while in the single predicate variant (passivized causative), NPY is generally the object of a biverbal predicate. In the periphrastic variant NPY thematically corresponds to the agent of the embedded verb, while in the passivized syntactic causatives it corresponds to a patient. In the periphrastic variant NPY is the causee, while in the passivized causative construction, the causee is external. In the passivized construction, the causee is removed from the sentence (it can be introduced as the optional by-complement; e.g. John had him killed by professional assassins) despite the fact that the causer remains the subject of the sentence. In other words, John may be the causer in the sentence John had him killed but not necessarily the agent of the embedded verb (causee) (as in: the killing was done by professional assassins) and in the sentence John got him killed, again, the source of the killing might be unintentionally related to John (as in: John did something that unintentionally led to the killing of one man). Some linguists (see Dik, 1997, p. 9) have argued that the periphrastic causative construction can be explained in terms of valency extension, meaning that the number of arguments in the source construction can be increased with this additional causativisation; e.g. the transposition of the sentence The dog ate the food to The owner made the dog eat its food. Gilquin (2010) suggests that within the framework of Cognitive Linguistics, the periphrastic causative construction may be best explained via the addition of the link to the so-called action 45 The syntax of the passivized causative significantly differs from the periphrastic causative. It exhibits a single predicate as opposed to a complex one in the periphrastic causative construction, with the verbs get or have taking on the role of the auxiliary verb. Moreover, while there is a distinction in using have and get as the auxiliary in the passivized causative, the role of the causative verb in a periphrastic causative is pivotal to the meaning of the construction (consider the degree of semantic difference between force him to do and cause him to do in comparison to get this done and have this done). 106 chain. Gilquin (2010) borrows the notion of the action chain from the theory of Cognitive Linguistics, or more precisely, she adapts Langacker’s diagram (Figure 7; 1990, p. 221) for the causative construction. The diagram of the periphrastic causative construction graphically illustrates the addition of a link to the beginning of the action chain, which in turn leads to the inclusion of the action source in the predicate of the construction. In such operationalization of the causative construction, the causer becomes the initiator of the action chain, the causee is the first recipient of this energy, and the final recipient is called the patient. In between the causer, cause and the patient, there are two subevents. First is the so-called causing event, formed by the causer and the causative verb. After the energy has been transferred to the causee, another event is initiated – this event is called a caused event in which the causer/action chain initiator needs not be directly involved. Figure 7. Schematic action chain of a periphrastic causative construction (adapted from Langacker, 1990, p. 221) Not all verbs with causative implications are necessarily treated as periphrastic causatives, or as verbs constituting causative constructions in general terms. Often, the disagreement about what constitutes the causative stems from the semantics of the verb, and consequently the construction. For instance, Dixon (2000, pp. 22-23) claims that the narrow interpretation of the prototypical causative construction would include a sentence such as Mary made John go, but it would not include the sentence Mary ordered John to go. The wider interpretation, which is adopted in this research, permits verbs such as order, persuade, convince and others to be treated as non-prototypical causatives. However, despite the fact that 107 some constructions doubtlessly convey a causative relation, the line has to be drawn somewhere. For instance, Song (1996, p. 36) regards the sentence such as I speak and child eats as a causative construction, but although we can semantically map out a causative relation similar to the one of ordering someone to do something, the matter is approached from the perspective of language processing and acquisition which warrants a coherent and precise definition of the studied phenomena, The research for this thesis is thus restricted to a selection of verbs, partly due to methodological concerns and the difficulties related to the extraction of the targeted constructions from the corpus. Rather than assuming a monoverbal, biverbal or even a biclausal description of the periphrastic causative construction, from the general perspective of our research, it may be best to define it as a “two-part configuration” where a non-finite complement clause is controlled by the causative verb, where the relationship between the two implies a causal relation that brings about a certain effect in the end (Gilquin, 2010, p. 1). Wolff & Song (2003) argue that beside the straightforward implication of the verbs that are consistently treated as causatives (e.g. cause, make, force) precisely because they encode causation in its purest form, the periphrastic causatives encompass a wider category of semantic restrictions. Here, the authors add the concepts of ‘enable’ and ‘prevent’; for instance, the verbs such as enable, let and allow would also constitute a periphrastic causative that resonates with the semantics of ‘enable’, while the verbs prevent, block and restrain belong to the semantic category of ‘prevent’. This definition of the periphrastic causative is the one where the “embedded clause encodes a particular result”, i.e. there is a strong implication that the encoded result will occur, or a result in a sense of intervention to prevent a particular change that would otherwise occur had there not been an intervention (as in the case of ‘prevent’) (Wolff & Song, 2003, pp. 285-286). Some authors studied the periphrastic causative construction, limiting themselves only to verbs such as make, have, get, cause and let (Ammon, 1980; Baron, 1977; Shibatani, 1976). The usual verbs that are mentioned in the periphrastic causative construction have been expanded by Wolff & Song (2003), who came up with 49 periphrastic causative verbs in their analysis of causation based on both syntactic and semantic criteria. They divide the verbs into two groups depending on whether they can be directed towards sentient and nonsentient causees (1), or exclusively sentient causees (2): (1) allow, block, cause, enable, force, get, help, hinder, hold, impede, keep, leave, let, make, permit, prevent, protect, restrain, save, set, start, stimulate, and stop 108 (2) aid, bar, bribe, compel, constrain, convince, deter, discourage, dissuade, drive, have, hamper, impel, incite, induce, influence, inspire, lead, move, persuade, prompt, push, restrict, rouse, send, and spur When it comes to the causative to-infinitive verbs in English, Sugawara (2016) has observed their differences and subdivided them according to semantics. Depending on the semantic perspectives of implicativity and affectedness of objects, this lexically-oriented division yields four types of the causative to-infinitive pattern. The constructions with implicative verbs imply that the action denoted by the infinitival complement has actually been carried out, whereas the affectedness of the object reveals whether the influenced participant undergoes a psychological change. If the verb is an implicative one, it means that the action has indeed been carried out. Finally, the distinction is between the implicative verbs which induce a psychological change (1), the implicative verbs which do not necessarily induce a psychological change (2), the nonimplicative verbs which induce a psychological change (3) and the non-implicative verbs which do not necessarily induce a psychological change (4): (1) induce-type: induce, influence, etc. (2) force-type: force, compel, oblige, drive, cause, etc. (3) persuade-type: persuade, convince, decide, determine, etc. (4) order-type: order, command, instruct, direct, etc. (Sugawara, 2016, p. 15) To sum up, the present study delves into periphrastic causative constructions. These constructions are specifically defined as those where causation is expressed through complex predicates that embed either full infinitive or bare infinitive forms of the caused event verb. To elaborate further, syntactically, these constructions follow a specific order of constituents (NPX→ VCAUSE → NPY → VPEFFECT): the subject of the main verb (referred to as NPX) followed by the causative verb (VCAUSE), then the subject of the embedded verb (NPY), and finally the embedded verb itself (VPEFFECT). It is worth noting that this research adopts a broader definition of periphrastic causatives, which not only includes the prototypical cases but also encompasses borderline instances where the causer simply facilitates the action of the embedded verb rather than directly causing it in a traditional sense (such as help and let-causative constructions among others). 109 4.1.3. Previous Research 4.1.3.1. Causative alternation and periphrastic causatives Much of the research in the area of causative verbs has been focused on the interpretation of causative alternation (Levin & Rappaport Hovav, 1994; Pye & Levy, 1994; Marcotte, 2005; Schäfer, 2009; among others). The phenomenon has also been under the scrutiny of those working in the field of early language acquisition and psycholinguistics (Pinker 1989). The alternation can be illustrated by the transitive-intransitive sentence pair: Pat opened the door / The door opened. Some find the alteration interesting in terms of crosslinguistic existence of the phenomenon and its implications. Other are more focused on the theoretical issues around discerning which verbs are inherently dyadic and which are monadic. More precisely, the question is often “whether all intransitive verbs have transitive counterparts with the paraphrase appropriate to the causative alternation” (Levin & Rappaport Hovav, 1994, p. 42), and the results seem to suggest that the intransitivity of a verb is not necessarily sufficient to ensure the verb’s participation in the alternation. Some suggest that agentivity is one of the crucial factors in the causative alternation and that the agentive rules might not license the alternation despite the existence of a number of exceptions to the rule (ibid., p. 39). Instead of agentivity, the related notions of ‘internally’ and ‘externally caused’ eventualities may help with the understanding of the causative alternation mechanism, and according to Levin & Rappaport Hovav (1994, p. 49), the concept of internal cause subsumes agency. As they further explain: “For agentive verbs such as play, speak, or work, the inherent property responsible for the eventuality is the will or volition of the agent who performs the activity. However, an internally caused eventuality need not be agentive. For example, the verbs blush and tremble are not agentive, but they, nevertheless, can be considered to denote internally caused eventualities, because these eventualities arise from internal properties of the arguments, typically an emotional reaction.” (1994, p. 49). Moreover, when it comes to the lack of the causative alternation for an internally caused verb, the answer resides in the properties of the argument structure, meaning that – in order to introduce an external cause – one has to “express the causative use of internally caused verbs periphrastically” (instead of saying The clown laughed me, the expression will be The clown made me laugh; more examples can be found in Levin & Rappaport Hovav, 1994, p. 55). 110 Particularly detailed work on periphrastic causatives came into prominence in the last two decades with researchers such as Stefanowitsch (2001), Hollmann (2003, 2006), Gilquin (2004, 2010) and Lauer (2010). Gilquin’s study (2004) shares a part of the methodological framework with this thesis. In her study of the main English causative verbs, Gilquin (2004) also dealt with the periphrastic causative constructions with the verbs get, have, make and cause and based her study on the corpus data (more precisely, a subset of the British National Corpus), thus orienting herself towards a usage-based cognitive approach to the interpretation of the structures and choices of particular causative verbs. Gilquin (2006) also argued that periphrastic causative constructions accommodate fewer number of secondary verbs than typically assumed, that is, verbs heading the periphrastic causative constructions are apparently restricted in number, but those occurring alongside as collexemes are not as obviously confined. The data she extracted from the British National Corpus revealed several things: the periphrastic causative constructions with make as its head verb and bare infinitive form for its secondary verb (“X make Y Vinf”) were clearly the most frequent ones when compared to alternatives with get, have and cause as their head verbs. On the other hand, when it comes to coreferential causative constructions, 46 get was the most frequent causative verb heading the periphrastic causative (see Gilquin, 2007). Furthermore, Gilquin (2006) found that, for each of the constructions with the aforementioned head verbs – there are different collexemes that tend to be the preferred choice by the speakers; for example, in “X make Y Vinf”, the most distinctive verb-collexemes were feel, laugh, look, think, wonder, appear, seem, want, sound, jump etc. Gilquin further concluded that the variation in preferences towards certain verbs in the secondary non-finite position suggest that causative constructions should be treated as separate entities, rather than “(near) synonyms” (ibid., p. 40), with respect to their syntax and causative verbs heading the construction. Similarly to Gilquin, Hollmann (2006) also used The British National Corpus in order to obtain the results for the active and passive periphrastic variations of the causative make (also dubbed the most general causative in terms of semantics). Being interested more in the processes behind passivization, Hollmann’s aim was to gather the quantitative evidence demonstrating correlations between the properties of transitivity and passivisability. Due to his background in the typological work, he was interested in the correlations that can be stated as 46 Coreferential causative constructions are those where the causer and the causee constitute a single entity, or two parts of a single entity (e.g. He got himself to start training again). 117 has been provided by Ambridge and Lieven (2015). As they succinctly put it, children make errors in language acquisition because of the competition between different forms of language, from single words to abstract constructions; in the same way, the mechanism of competition that is responsible for errors is also responsible for why children stop making errors (ibid., pp. 494-495). Early on, children may use the wrong form because it is more frequent or because they have not yet learned an appropriate alternative. However, as they learn new schemas and constructions that better fit the items in their intended message, they retreat from error. This process is influenced by two factors: statistics and semantics. Children learn the probabilistic links between items and constructions, causing certain constructions to be activated more frequently. Additionally, they learn the fine-grained semantic properties of different construction slots, making certain verbs more appropriate for certain constructions over others. As a result, the likelihood of children accepting and comprehending verbs in non-attested constructions decreases as verb frequency increases and semantic compatibility with the target construction improves. 4.2. Periphrastic causative constructions – results 4.2.1. Extraction method and the examples of periphrastic causatives Given that I am working on the already existent corpora, the methodology of the extraction process relies on the POS (Part-Of-Speech) tagging provided by the platform Sketch Engine. Essentially, the methodological process comes down to clearly defining the targeted construction and employing the syntax that will locate the constructions of interest in the corpora. Drawing from the definitions of the periphrastic causative construction stated above, I looked at the standard periphrastic causatives with the verbs get, make and have, but I have also expanded the selection of these verbs with the ones that take full infinitives, and carry a compelling causative meaning such as cause, force, compel, coerce, lead, oblige, drag, drive, press, pressure, blackmail, strong-arm. The category is further expanded with the verbs order, command, appoint, direct, require, request, demand, dictate, ordain, convince, motivate, urge, ask, beg, bid, tell, instruct, authorize, empower, entitle, endorse, sanction, support, insist, talk; these are verbs which do not necessarily imply the realization of the outcome (the final outcome of the caused event), but there is a strong implication of the realization (whereby the causee 118 does what the causer led them to do). Further considered are the verbs in the category of allowing someone to do something – such as allow, permit, enable, grant, license, approve. Furthermore, the productive let and help were also considered separately due to their productivity and frequency. They do not include the prototypical semantic information conveyed by the causative, but they do share a considerable portion of syntax and semantics with causative verbs. Examples of negative causation were also considered, but ultimately were not included in the extraction process and these examples might prove revealing in some further research. These constructions can be described as ‘Vx someone/thing from V-ing something’ (where the Vx can be replaced by verbs such as forbid, prevent, stop, ban, block, deny, cancel, censor, disallow, halt, hinder, impede, inhibit, oppose, restrain, restrict, avert, forestall, prohibit, thwart, limit, suspend, disrupt, stall, obstruct). All in all, the analysis aimed to cover the majority of periphrastic causative constructions, even those which are semantically not clear-cut in terms of causation, i.e. they all imply it to a higher or lesser degree, or at least imply that the action of the subject (causer) is somehow related to the realization of the secondary event/action (for example, help and let constructions). These constructions were included primarily because they share the syntax of the rest of canonical periphrastic causatives (just as benefactives were included in Section 3, given that they share a non-prepositional pattern with canonical ditransitives). Some of the periphrastic causative constructions extracted after filtering out the corpora can be found in Table 15. The examples include a variety of realizations, differing in terms of causative verbs, animacy of the causee, infinitival alternations (to vs. bare infinitive) etc. The simplest constructions are restricted solely to the causative verb, the adjacent causee and the effect expressed by the secondary verb (such as Let me kick). Others are more complex and, aside from the aforementioned, include the causer and the patient as well (e.g. Somebody's helping him hold it because the mine's real deep). As already explained, various verbs have been included as potential verbs expressing causative relationship, and as a result cover instances with let, ask, help and others 51 (e.g. Want to [ask him to find other one of these] or [Somebody's helping him hold it] because the mine's real deep). Furthermore, we find various sentence length in relation to the periphrastic causatives. Some examples, especially as the children grow older, include elaborate sentences, where the periphrastic causative constructions 51 See the subsequent analysis for more details. 119 constitute a minor sub-clausal piece (e.g. I would like to sleep over her at her house every day because she [lets me stay up late] about ten o'clock or twelve thirty). Some sentences even contain multiple periphrastic causatives (e.g. And [I command Mark to come in] and hyp [I command Mark to be hypnotized by me and come in and be sliced]), although this may be due to inadequate inter-clausal separation, where the two clauses actually represent distinct utterances rather than a complex sentence. Table 15. Examples of (candidate) periphrastic causative constructions found in the CHILDES corpora Corpus Age (y; m; d) Produced causatives MALAKOFF 1;11;17 Let me kick. WELLS 1;11;30 I let you mend it. HIGGINSON 2;10 Be right back [I'm get my woolie blanket to keep me warm].* MANCHESTER 2;10;29 Can't get the dolly to work. DEMETRASWORKING 2;2;24 Have this get out.* VALIAN 2;5 Want to [ask him to find other one of these]. DEMETRASWORKING 2;7;4 [Let me play] and you can sing the grand old flag. WEIST 2;9;2 I gotta go get a chair to talk to you.* FLETCHERTHREE YEAR OLDS 3;0;24 He's got a beaker to drink out.** GARVEY 3;1 Let me drive the car. WEIST 3;4;18 [Let I drive it here] with my new car with Emma. WEIST 4;3;14 Because [they have water and duckbills swim].** KUCZAJ 4;3;15 [Somebody's helping him hold it] because the mine's real deep. GLEASON 4;5 You told [you told you to play that] at dinner time Daddy. WARREN 5;10 […] and it got um it made the water go up to the bricks and [made it overflow]. GLEASON 5;2;10 If I need help you help me.** CARTER 6 I would like to sleep over her at her house every day because [she lets me stay up late] about ten o'clock or twelve thirty. MACWHINNEY 6;0;27 And [I command Mark to come in] and hyp [I command Mark to be hypnotized by me and come in and be sliced]. *dubious but retained in the analysis **removed from the final dataset Table 15 also includes examples which have been flagged as periphrastic causatives initially, but later removed after the second revision of data. For example, the sentence He's got a beaker to drink out was initially kept, but after further scrutiny, it becomes clear that the two verb phrases are not related in the same way as in periphrastic causative construction (forming 120 a complex predicate), i.e. the noun ‘beaker’ is not actually a causee, but rather a theme which undergoes the action (the non-finite clause underlying the purpose can be paraphrased as in order to drink out). 52 The sentence Because they have water and duckbills swim was initially overlooked and included as a periphrastic causative because the water and duckbills was wrongly identified as a single entity forced to swim, but is actually a complex sentence with two independent clauses and predicates. In a similar fashion, the example If I need help you help me was removed from the final dataset; the sentence was initially registered as a periphrastic realization help you help me, although in reality represents a conditional sentence where it is necessary to put comma in between help and you. There are several examples where the causee is an inanimate noun (I'm get my woolie blanket to keep me warm or I gotta go get a chair to talk to you), where the thematic roles of the instrument and the causee often overlap, but these cases remain the minority of examples retained in the final dataset (see Section 4.2.2 for further analysis). In the first sentence, the ‘woolie blanket’ is arguably the instrument but it also performs an action, whereas in the latter case, the ‘chair’ can either be interpreted as one that is capable of speech, or simply an instrument which facilitates the action performed by the speaker. There are also examples which have been problematic to interpret, but were still retained due to shared syntax with the periphrastic causatives. Aside from I gotta go get a chair to talk to you, the example Have this get out was also retained in the final dataset, despite the fact that the construction has a twofold possibility interpretation-wise. In other words, it is possible that the sentence is in fact, one with two independent clauses which should have been separated by a full-stop, or that the speaker truly intended to produce a form of periphrastic causative, whereby (in a more loose interpretation) someone was asked to take out something. 52 The beaker can also be interpreted as the instrument which facilitates the process of drinking, thus resembling the conventional thematic relations captured by the periphrastic causative. Still, the example is different from other periphrastic causatives where the inanimate causee was ultimately kept. In the sentence Be right back I'm get my woolie blanket to keep me warm, there is a cohesion between the thematic roles of an instrument and the causee, but the example is retained in the final dataset because the sentence clearly denotes the patient (me), as opposed to the He's got a beaker to drink out, where the two may overlap. 121 4.2.2. Structural complexity and similarity When the overall distribution of periphrastic causatives is observed across age groups, the let construction is by far the most frequent one (see Table 16; for the corresponding examples, see Table 17). Part of the reason is that it includes the phrasal use of the verb, which semantically may barely qualify for the periphrastic causative, but pragmatically does not really imply causation in any degree (for instance, let me explain or let me introduce myself). This also relates to the fact that the most common realization in all age groups – let-ProPers-V – is considerably more frequent than let-N/MOD-N-V. 53 The second most frequent pattern across all age groups is make-ProPers-V (e.g. make him do it), followed by the third – help-ProPersV (e.g. help him do it). Other relevant patterns, in terms of sheer frequency, include various causee-based alternations with the aforementioned verbs, as well as less common patterns with get. The least common patterns are those with demonstrative pronouns in causee positions, as well as lexically more specific infinitival patterns with verbs that need not necessarily belong to traditional interpretations of periphrastic causatives. For instance, when it comes to TALKNP-into-V, only six instances were observed in the adult group (making it insignificant when normalised per million tokens as can be seen in Table 16 (freq.<0.5). Among the more interesting observations is the fact that periphrastic causative constructions are equally represented in the 4-6 age group and adult child-directed speech. This is not a tendency that was observed on ditransitive constructions, thus inviting questions as to why this might be the case. While the cross-examination of usage-frequencies regarding constructions is addressed in Section 6.1, it is worth noting that the overall frequencies of this particular construction use are basically the same, the chances for which must also be statistically low. So, the question is, what exactly can we infer from the sheer total frequency of the causative constructions? Does this mean that the language of children in 4-6 age group comes closely to that of adults regarding the overall use of relatively complex syntactic patterns? Does it mean that adults’ speech, when simplified for the purposes of conversing with children, starts to resemble the speech of 4-6 age groups? Naturally, there are some discrepancies in terms of individual patterns and verbs heading the periphrastic causatives 53 Note that the bare 'V' label does not necessarily imply that the construction, or the sentence, ends with the verb. Instead, the V label is used to highlight the fact that the pattern excludes instances which take the full infinitive form (which could be misinterpreted had the VP label been used). Full infinitives are designated as 'to-V'. 122 which might shed some light on the aforementioned questions, but there is no doubt that the causative construction production-related obstacles have been overcome by children 4-6 age group to a surprising degree. There is obviously a lot of similarity between 4-6 age group and adults, even when the individual patterns are taken into account, i.e. the most pronounced differences seem to be in some variations with the verbs get, let and make. In terms of percentages, the greatest jump in the use of a particular construction can be observed with the pattern get-ProPers-to-V, which more than doubles in the group of adults. This is also the only construction where the early groups prefer to use the nominal in the position of the causee as opposed to person pronouns, which they elsewhere almost always prefer. This is not the case with adult speech, but even in adult speech, the frequency of use between the get-ProPers-to-V and get-N/MOD-N-to-V comes closer than in other patterns in terms of pronoun-nominal competition. Although it is not completely clear why this might be the case, it surely corroborates the claim that sufficient input frequency is required to ‘tip the scales’ when it comes to competing patterns in children’s language. 54 The distribution of frequencies when it comes to causative constructions shares some main tendencies observed (albeit not all of them) in both ditransitive and relative constructions across different age groups. The frequency of the overall use of periphrastic causative as in the child-oriented adult speech is almost 2 times greater than in the 0-3 age group. As expected, the frequency at which they produce constructions requiring a greater level of syntactic complexity was lowest in the 0-3 age group, whereas when it comes to 4-6 and adults, the frequencies are approximately the same; the data does indicate that children indeed produce the periphrastic causatives by the age of four, but the discrepancies in frequencies suggest that the learning curve might not be complete. However, the main thing to be observed here is that the overall distribution of different periphrastic causatives demonstrates a striking resemblance between all three studied age groups. The patterns that seem to be most frequent in the adult input are also the ones that are most frequent in early speech, both in 0-3 and 4-6 age groups. The same goes for patterns where the input frequency was at its lowest, that is, the children in any of the two younger age groups did not demonstrate any need for a particular pattern that was underrepresented in the input. 54 I have also observed the nominal bias towards the infinitival complement in adult speech on the spoken transcripts from the BNC corpora with the causative help construction, which might suggest that the lexical differences are for some reason triggered by the syntactic environment on its own (Proroković, in preparation). 123 Table 16. Distribution of causative constructions across age groups (relative frequencies normalised per million tokens and rounded to the closest unit) 0-3 4-6 18 + get-ProDem-to-V 0 2 1 get-N/MOD-N-to-V 8 11 15 get-ProPers-to-V 5 7 18 have-ProDem-V 1 0 0 have-N/MOD-N-V 1 7 9 have-ProPers-V 1 6 12 help-N/MOD-N-V 5 13 25 help-N/MOD-N-to-V 0 0 2 help-ProPers-V 71 71 96 help-ProPers-to-V 1 5 4 let-ProDem-V 0 0 0 let-N/MOD-N-V 23 45 75 let-ProPers-V 374 633 550 make-ProDem-V 1 0 1 make-N/MOD-N-V 16 50 35 make-ProPers-V 75 132 135 TELL-NP-to-V55 12 82 82 CAUSE-NP-to-V56 0 0 2 ALLOW-NP-to-V57 0 0 1 TALK-NP-into-V58 0 0 0 TOTAL 594 1065 1065 Table 17. Examples of the produced periphrastic causative constructions in the CHILDES corpora (patterns from Table 16 indicated in square brackets) Construction pattern Example Corpus Age get-ProDem-to-V Oh you want to [get some of those to come]? WARREN 6;2 get-N/MOD-N-to-V Um xxx I''m gonna [get the cow to drink some milk]. BLOOM 2;8;12 get-ProPers-to-V And my can [get it to go for towing the blue towing in the the blue car]. NELSON 2;0;1 have-ProDem-V Yeah now we need to lock it we can''t [have that happen]. GARVEY 3;6 55 TELL category included the following verbs (as search queries): order, command, appoint, direct, require, request, demand, dictate, ordain, convince, motivate, urge, ask, beg, bid, tell, instruct, authorize, empower, entitle, endorse, sanction, support, insist, talk. 56 CAUSE category included the following verbs (as search queries): cause, force, compel, coerce, lead, oblige, drag, drive, press, pressure, blackmail, strong-arm. 57 ALLOW category included the following verbs (as search queries): allow, permit, enable, grant, license, approve. 58 TALK category included the following verbs (as search queries): order, command, appoint, direct, require, request, demand, dictate, ordain, convince, motivate, urge, ask, beg, bid, tell, instruct, authorize, empower, entitle, endorse, sanction, support, insist, talk. 124 have-N/MOD-N-V And then he drove away and then the car [had its wheels blow up] and then he got he flew back and then he xxx. WEIST 3;3;9 have-ProPers-V Have you wrap me that way. WEIST 3;3;9 help-N/MOD-N-V Help baby take coat off. MCCUNE 1;8 help-N/MOD-N-to-V You were [helping Granddad to water his courgettes]. THOMAS 18+ help-ProPers-V Do you want to [help me do my Noddy jigsaw]? BELFAST 3;7;13 help-ProPers-to-V Uh [help us to imagine]! HSLLD 4;9;23 let-ProDem-V Peter couldn''t [let that happen]. HSLLD 18+ let-N/MOD-N-V Slide [let those other people come up]. VANHOUTE N 3;3;28 let-ProPers-V Huh [let me kiss it]. DEMETRAS 2;7;2 make-ProDem-V We [make them all stand up]. SUPPES 2;10;1 3 make-N/MOD-N-V Can can you [make mummy get some breakfast]? LARA 2;10;1 5 make-ProPers-V For [making him catch the dog]. WOLFHEMP 6;5 TELL-NP-to-V Mommy I [might ask you to paint a tulip] because that''s one thing I''m not good at. ERVINTRIPP 5;8;11 CAUSE-NP-to-V And [forcing someone to eat it] tulips. ERVINTRIPP 18+ ALLOW-NP-to-V (…) they weren''t letting you they weren''t [allowing you to treat yourself] by relaxing the muscles. SNOW 18+ TALK-NP-into-V You''re not [talking me into letting you have that kitty here]? HSLLD 18+ Furthermore, the Spearman's rank correlation indicates high values when it comes to frequencies of the observed patterns within the causative construction (N=12). 59 The correlation is signficant on the 0.01 level and highest among the 0-3 age group and the adult group, and nearly the same value is found between 0-4 age group and the adults (see Table 18). Again, there is little to no difference between the frequency-based ranking of particular periphrastic causatives, which leads us to assume the same degree of correspondence in terms of preferred patterns across all three age groups. Table 18. Spearman’s rank correlation of the observed causative patterns across age groups 0-3 4-6 18+ 0-3 1 0.923** 0.958** 4-6 1 0.951** 18+ 1 *ρ=0,591 p=0.05 **ρ=0,777 p=0.01 59 Note that the Spearman’s rank correlation was conducted for the first 12 patterns because some were simply not found in the corpus for the youngest age groups, meaning that they could not be ranked amongst themselves. 125 Although there is a quite a bit of similarity across age groups in terms of distribution and frequency of usage, there are some interesting discrepancies can be more easily observed when ratios are taken into account. For instance, help-ProPers-V is one of the more frequent patterns found in adult speech, so it is not surprising that the same goes for younger age groups. However, the observed volume of production in younger age groups comes very close to adult speech (which means that, comparatively speaking, children produce it more than other patterns when contrasted with adults). This could be foreseen in 4-6 age groups, especially considering that the overall production of periphrastic causatives appears to be the same between them and adults. If anything, given the same overall volume of periphrastic causatives in the two groups, one could argue that the constructions somewhat drops in terms of expected frequency in 4-6 age group. However, given the rest of tendencies between 0-3 age groups and 4-6 age groups, which typically reflect a rise in the use of various causatives by a significant margin, it may be interesting to consider why the same rise is not exhibited in help-ProPers-V, i.e. why the same construction is so frequent from the onset. When compared to other causative patterns, this is the only one which does not rise considerably once the children turn 4 years of age. Well, one of the reasons that may explain this anomaly stems from the fact that the youngest age groups simply need to use the pattern at question more often. In other words, the situational context that the youngest children find themselves in, or the requirements they experience that necessarily need to be fulfilled by their interlocutors, lead them to use the construction in question. For example, it is certainly expected that children constantly need some sort of help with various tasks at hand, at which point they resort to constructions such as are you going to help me make this (BELFAST, 2;7;26) or help me move it out of the way (BLOOM, 2;10;19) etc. To further support this, the analysis revealed that out of the total 301 instances observed in the corpora (of the patterns help-ProPers-V and help-ProPers-to-V), 240 of them contained me as the causee pronoun (79.7%). Although the construction was not as pronouncedly represented in the adult speech, it was represented sufficiently for children to make early generalisations and inferences about the flexible use of that particular pattern (out of the total 1874, there were 936 instances with me in the causee position – or 49.9% in total). The second pattern where the tendencies differ across age groups more than they do with some other patterns is TELL-NP-to-V, which included a variety of verbs that have a causative implication to them. The analysis of the said pattern, along with the other two non- 126 prototypical periphrastic causatives considered (CAUSE-NP-to-V and ALLOW-NP-to-V), located only three verbs in the 0-3 and 4-6 age groups (Table 19). In other words, despite the fact that a large variety of verbs with causative implications were included in the analysis (43 of them; see Footnote 55, Footnote 56, and Footnote 57), only three types of instantiations were observed in 0-3 and 4-6 age groups. In the 0-3 and 4-6 age groups, we find the verbs tell, ask and allow, whereas in the adult group we find other verbs such as force, cause, order, press, beg, led, convince, pressed, drive, urge, talk, enable, drag, get, demand, request and command. Not only did the extraction method yield just three different verbs in the causative pattern in the younger age groups, but the verbs allow (in the 0-3 age group) and command (in the 4-6 age group) were observed only once. Table 19. Observed frequencies of particular verbs in the pattern V NP to V (excluding get and help) 0-3 4-6 18+ tell 36 72% tell 111 82.84% tell 931 58.33% ask 13 26% ask 22 16.42% ask 578 36.22% allow 1 2% command 1 0.75% allow 24 1.5% Total 100% 100% 96.05% The data in Table 19 clearly shows another similarity across age groups within the VNP-to-V pattern. Regardless of the fact that some instances of other verbs were found in the adult speech, it is clear that two verbs represent the majority of all periphrastic causatives when it comes to child-directed speech. But what seems to be even more interesting is that the same two are the most frequent ones in early speech. The results can be interpreted in the usage-based framework, as they also seem to favour the item-based tendencies in early speech. The most frequent verbs that the children hear are also the ones which completely rely on in their early speech. This is not to say that children do not necessarily understand or use the rest, especially as they grow older and reach 5 or 6 years of age, but it is indicative of the way in which they express themselves early on. Moreover, aside from the fact that tell and ask are most frequently found in the input, it is important to point out that they are also the ones which are semantically most flexible. We could easily look at the two as ‘umbrella terms’ for many others that have been included in the analysis, which tend to be more restricted in their meaning, entailing much more nuance when in terms of goals, participants in the action, manner etc. The fact that they are semantically flexible reflects on the speech of children, as well as that of adults; however, 133 one open to a greater number of schemas with different caused event verbs (e.g. make the car run vs. make the stone run). The fact that all age groups produce it so frequently in the causative construction tells us that children start to grasp this concept very early, both the intent expressed in the causative construction and potentially the difference between cars and other inanimate nouns (e.g. stone). Other frequent animate nouns in 0-3 age group include animals which usually refer to toys or animals in storybooks or other materials used to stimulate/elicit speech. In other words, the frequent causees in the 0-3 age group tend to be concrete objects, those that are observable or accessible to children. Most of these are not found in the input among the most frequent ones, which indicates that children utilize the function of causatives for expressing activities and notions relevant to them in immediate surroundings (my car in make my car go; BOHANNON, 3;0), whereas adults apply causatives more frequently on situationally less specific concepts (e.g. eyes in sometimes onions make your eyes cry; FLUSBERG; 18+). Moreover, with the progression of age, in the 4-6 age group the most frequent ones become more abstract (for example, light, water, anybody, anything and the like). This indicates both a linguistic and a cognitive development; (1) linguistic in the sense of the ability to generalise on the use of more abstract concepts in causative constructions and to apply them productively and flexibly, and (2) cognitive in the sense of understanding more complex phenomena, but more importantly, understanding which conditions they ought to satisfy to fill the ‘causee slot’ in the causative construction. When looking at the secondary verbs (verbs heading the caused event) in the periphrastic causative constructions, there again seems to be a considerable overlap in the most frequent choices for all age groups. Despite the fact that the secondary verb slot in the periphrastic causative construction is arguably constraint free semantics-wise (it is more or less possible to fill the slot with any verb), the cross age comparison indicates a resemblance in different age group preferences in using particular verbs as opposed to others. 65 For example, the verb go is the most frequent choice in all age groups, and the verb do also ranks highly for all three, while the verbs get, come, feel, disappear appear among the most frequent choices in 65 It is true that some verbs make better “candidates” for filling the slot, as they tend to depict an action which can indeed be triggered by the actions of another person (causer). Action verbs may be preferred as opposed to stative verbs (make someone go vs. make someone be), though not exclusively (for example, consider the frequently used make someone think or make someone feel). 134 all age groups. Interestingly, the verb get is among the top three choices for the 0-3 and 4-6 age groups, whereas in the adult group, though frequent, ranks considerably lower. This may be anchored in situation-specific or participant-specific factors which drive one item to be produced as opposed to others. For example, children may use get more frequently because they tend to be in need of something more often than adults (e.g. can you make mummy get some breakfast; LARA, 2;10;15), whereas feel may rank more highly in adult speech due to their appeal to emotions in children (e.g. does it make all your muscles feel better; FLUSBERG, 18+). It is difficult to assess just how much of an impact the input bears in relation to other pragmatic aspects of language production (and reproduction), but the overproduction of the same items that most frequently occupied certain slots in the input indicates item-based patterns that go beyond items heading certain phrases (such as ditransitive or causative verbs). Table 23. Secondary (caused event) verb in the V-NP-V pattern frequency ranking across age groups Rank 0-3 4-6 18+ 1 go go go 2 do get do 3 get do feel 4 come grow sit 5 stand play look 6 see feel come 7 feel come stand 8 sit fly get 9 eat clean put 10 write cool make 11 grow disappear fall 12 say fall disappear 13 disappear stand eat Of course, there are other important mechanisms that might be at play besides the pragmatic element of their speech and the relevance of the input. The ranking of certain verbs can be affected by the vocabulary deficiencies in the youngest age group. One of the things that we may notice is the high frequency of the verb see in the 0-3 age group, and one of the reasons why this might be the case is their inability to retrieve a more straightforward lexical item to express the intended meaning as opposed to using the periphrastic variant. In other words, construction such as make him see is another way of saying show him, and a similar relationship 135 (to a certain degree) is captured with help (both have been observed concomitantly with see in the 0-3 age group). The compensation of lexical inaptitude with syntactic prowess is a reflection of the fact that the development of one’s vocabulary is a lifetime-long process, i.e. the younger the children are, the higher the chances of them resorting to periphrastic alternatives if the more straightforward lexical ones are not retrievable at the given moment. A more detailed insight into the correspondence of verbs within the periphrastic causative construction is better illustrated in the level of overlap between the input (adult language) and the two age groups of children. The exact degree to which there is a similarity in using certain lexical items in the specific slots within the causative construction, is more clearly visible in the cross-group differences in the proportions of items that were also found in the input (see Figure 8). When it comes to both caused event verbs and causee nouns, the proportion of items that were also observed in the input is greater for the 0-3 age group than for the 4-6 age group. Figure 8. The proportion of lexical items observed in the input in the periphrastic V-NP-V pattern 20.0% 30.0% 40.0% 50.0% 60.0% 70.0% 80.0% Causee noun Caused event verbs Causee noun Caused event verbs 4-6 items observed in the input 50.4% 67.2% 0-3 items observed in the input 53.9% 76.5% Caused event lexical items overlap in the V-NP-V pattern 136 It is not particularly surprising that both age groups tend to overlap to a considerably greater extent when it comes to the verb slot, than when it comes to the noun slot. The primary reason for this is probably vocabulary-related while the other may be of pragmatic origin. In other words, the range of lexical items is naturally smaller for verbs than for nouns, while the second reason is probably the difference between children and adults when it comes to their choice of the referent; i.e. the production of periphrastic causative constructions entails that the speaker choses a particular causee which is to be forced into doing something, and the focus of children and their interlocutors is often oriented on different actors or entities within the experimental setting. Nevertheless, the results clearly indicate that there are more lexical items that have been observed in the adult speech produced by the 0-3 age group than by the 4-6 age group, which suggest a more restrictive character of the earliest speech as predicted by the first hypothesis (see hypothesis 1.3). 4.2.4. Animacy As far as the role of causee is concerned in the V-NP-V pattern, a look at the animacy may provide a new perspective on the language production in relation to age differences and the character of the nouns used in the causee slot. This element takes the central position in the causative construction and is crucial for the interpretation of the entire construction. For example, it may take both inanimate and animate forms, although it should arguably favour the animate ones given that they are expected to carry out an additional action, i.e. given that the causee tends to be the agent of the secondary action, it should more likely take on the animate form. Although this is a bold claim, a look at the most frequent caused event verbs reveals a potentially verb-induced constraints on the causee slot (see Table 23); i.e. the vast majority of them (across all age groups) are verbs that tend to semantically entail a relationship where the subject takes on the sentient form (e.g. feel, sit, look, come, stand, get, make, eat, write, say etc.). In other words, if we assume the greater likelihood of the causee taking on the role of the agent (or a sentient patient), we might also assume a greater likelihood of said causee to take on the animate form. This type of comparative analysis may shed additional information on the perception of animacy in terms of construction processing on the part of children; we may partly draw inferences on the competitiveness between the mechanisms of understanding/generalising (e.g. the semantic implication of the most frequent caused event verbs observed in the 137 periphrastic causatives should drive children to employ animate agents in the position of the causee) and input frequency (e.g. the most dominant pattern in the input should affect the children’s production regardless of its suitability). The results on the three categories of animates (inherent animates, contextual animates and inanimates) reveal where the shift in animacy actually occurs across the age groups in the V-NP-V pattern. It has already been shown that the use of inanimate nouns in the position of the causee rises with the progression of age; it rises considerably between the 0-3 age group and the 4-6 age group, and it continues to slightly rise in adult language. Naturally, the rise of inanimates comes at the expense of other categories of animacy and the Figure 9 illustrates the trends across age groups. Figure 9. Animacy of the causee in periphrastic causatives The rising trend in the use of inanimates that comes with the progression of age is accompanied with the declining trend in the use of inherent animates in the position of the causee. It is the highest in the 0-3 age group (44% of all uses), but drops in the 4-6 age group (41% of all uses) and continues to drop in the adult group (35% of all uses). The portion of different contextual animates employed in the 0-3 age group amounts up to 21% of all uses, 0-3 4-6 18+ Inherent Animates 44% 41% 35% Contextual Animates 21% 15% 18% Inanimates 35% 44% 47% 0% 5% 10% 15% 20% 25% 30% 35% 40% 45% 50% Proportion of use Animacy (type-based) 138 dropping to 15% in the 4-6 age group, and then rising again in adult language. If we go to Table 24, we can clearly see that the drop between the 0-3 age group and the 4-6 age group is in the use of animate toys in the position of the causee. This is not completely surprising, given that we can expect the youngest category to be oriented towards toys more than slightly older children. With the rise in cognitive power and the ability to handle more complex phenomena in terms of conceptualization, the children in the 4-6 age group may employ causatives anchored to their immediate surroundings less frequently (e.g. cause this makes the light come when I’m talking; HALL, 4;6). On the other hand, when it comes to adults, the examined speech was mostly child-directed which is why they may have resorted to animate toys more frequently when requesting of their young interlocutors to perform a certain action (e.g. make the fish bite the truck; SACHS, 18+). Table 24. Animacy of the causee 0-3 4-6 18+ INAN 35% 44% 47% ANc 23% 23% 16% ANp 22% 18% 19% ANt 14% 8% 13% ANb 7% 7% 6% When the animate causees are split into 4 categories based on the degree and context, inanimate things appear to be the most frequent ones across all age groups in the position of the causee. Naturally, if the rest of the categories of animacy were lumped together into one group, they would be more frequent than inanimates, but considering that these do not constitute a coherent and homogenous group, it seems best to treat them separately in this analysis. Although the overall patterns in terms of (in)animacy of the causee largely coincide across age groups, there seem to be some smaller discrepancies which may be telling for the differences in the use of causatives. The first thing to observe is that the adult group employs inanimate items more frequently than the younger age groups, but especially when compared to the 0-3 age group. Indeed, inanimates are fairly frequent in all age groups, but the saturation around inanimates tends to be greater in the adult group. One possibility is that the rise of lexical richness in the adult group results in the greater portion of inanimate items—if we presume that there is a greater variety of nouns that are inanimate, that is, that with the greater number of diverse nouns there is also a greater potential for the production of inanimates. However, the 139 amount of inanimate items in the position of the causee may suggest an additional aspect of the causative construction that becomes more prominent as the children grow older. An inanimate item in the position of the causee may suggest that the construction moves away from the prototypical role of the agent in the position of the causee (e.g. have them go this way; KUCZAJ, 3;4;8) towards the role of the patient (e.g. make my milk disappear; BLOOM, 2;7;13). To be fair, many of the nouns labelled as inanimate adopt animate characteristics and perform the roles of the agent in some sense (e.g. make the wheel go into the barrel; BLOOM, 2;8;12), but it seems legitimate to assume that they require a greater degree of abstraction and conceptualization than the typical animate agents. Whether the prototypical periphrastic causative construction employs the causee as agent rather than the patient of the secondary action is less of an issue here; in the context of this research, it is more important to observe the fluid function of the causative construction and how it affects its development and acquisition. In other words, the input children receive permits the use of inanimate items in the position of the causee to a surprisingly high degree, which should in turn result in similar numbers when it comes to the early language output. Indeed, the portion of inanimate nouns in the position of the causee is fairly frequent in the 0-3 age group (35% of all uses), but still considerably less than in 4-6 age group (44% of all uses) or the adult group (47% of all uses). This may indicate that in the competition between the cognitive mechanisms related to understanding (how children initially interpret the causative constructions and their function) and input frequency (what children hear the most), the limits of early cognition may somewhat skew the results, resulting in a preference towards animate beings in the position of the causee (which represents less of a conceptual challenge for children). Distributional learning plays a role in the sense that children learn to expect animate beings in the position of the agent across other transitive constructions, which may change and become more flexible in the causative construction (e.g. make the bed stand up and put the little person on it; BATES, 18+). When confronted with first causatives, for the sake of easier processing, children may focus on the ones which take animate beings in the position of the cause as this would not contradict the expectations they developed through the mechanisms of distributional learning. Finally, this interpretation of the results agrees well with Tomasello’s hypothesis about imitative learning which incorporates various cognitive mechanisms such as distributional learning, i.e. both input frequency and children’s ability to interpret the input and 140 formulate generalisations have to be taken into consideration for the complete account of early language development. 4.2.5. Lexical richness In order to get a better grasp on the length and complexity of sentences containing periphrastic causatives, we may look at the data pertaining to MLU (Table 25). 66 The difference in MLUs actually suggests that the most complex (or the longest) sentences containing causatives are found in the 4-6 age group, and not the one of adults. We can now observe that the data differs from the one on ditransitive constructions, at least in terms of which age group produces the most elaborate sentences that contain the construction in question. While the inconsistency in cross-constructional MLU data is described in more detail later (see Section 6.1), we can note that there is a large gap between the 0-3 and the 4-6 age groups, i.e. despite a lot of “noise” and nonce words being recorded in the youngest age groups, the data suggests that the 4-6 age group tends to produce sentences longer by at least 2 words on average when compared to the 0-3 age group. Table 25. General overview of the sentences containing make-NP-V across age groups 0-3 4-6 18+ Token count 3129 3163 29756 Word type count 457 536 1909 Character count (without spaces) 11591 11866 115342 Utterance count 386 298 3204 Open class words 1629 1543 14789 Closed class words 1454 1547 14382 MLU (morpheme-based) 8,59 11,25 9,82 When it comes to lexical density in utterances that centre around the periphrastic causative make-NP-V constructions, the numbers are very similar across age groups (see Table 26). Similarly to sentences with ditransitive constructions, the variation in lexical density in those containing causative constructions seems non-existent. The LD values for these 66 Again, it is important to note that these measures capture the sentences in their entirety, rather than just being restricted to the construction themselves, which means that they might encompass some additional sentence material that does not fall into the causative domain. 141 utterances revolve around 0.5, meaning that in utterances containing make-NP-V causative patterns for each content word there is also one grammatical word on average. The numbers are slightly higher than in utterances containing ditransitive constructions (discussed in the subsequent chapters in more detail), but the point remains that neither of the age groups produces utterances denser in information when compared to the two others. Considering that it centres around two verbs and (at least) one noun, the periphrastic causative construction in itself favours lexical words as opposed to function ones (for example, we might expect higher averages than 0.5). Again, we could expect that lexical density would increase in the younger age groups if they use more isolated forms of causative constructions instead of embedding them in more elaborate utterances, but this does not appear to be the case. Table 26. Lexical richness for the causative make-NP-V Lexical density and sophistication Lexical diversity: range and variation LD LS1 CVS1 NDW NDW-ER50 MSTTR-50 RTTR SVV1 0-3 0.51 0.14 0.56 321 34.8 0.63 7.62 10.17 4-6 0.5 0.14 0.8 417 32.9 0.66 8.91 21.48 18+ 0.52 0.19 0.86 427 35 0.68 9.8 20.96 When it comes to lexical sophistication (LS1), an increase is only visible between the 4-6 age group and adults. Interestingly, the speakers of all three age groups use words outside of the most frequent ones on less than 2 occasions per 10 words in utterances with causative constructions. While there is an increase in the child-directed language of adults, the question is how much of that has been intentionally simplified in terms of vocabulary. In this context, verb-related sophistication may be more interesting, particularly because the second verb in the construction is open to unconstrained insertion of all sorts of verbs. This index reveals a considerable spike between the 0-3 and the 4-6 age groups already, suggesting that children older than 3 increase the use of less frequent verbs. Additionally, the same trend can be observed even in terms of verb variation (SVV1), which shows a clear spike between the 0-3 and the 46 age groups, with little difference between the 4-6 age group and adults. In a sense, the measure of sophistication can be interpreted as an indicator of how much the child relies on the input. As stated before, sophistication is not really a measure of diversity (although they tend to correlate), but a measure of how many ‘rare’ words a speaker uses. If we assume that the input has a key role in early language development, going beyond the most frequent words in the input suggests a step away from its impact. Sophistication, especially verbal (since verbs head 142 phrases and are seen as integral in language development) can serve to illuminate a dynamic relationship between the role of input, word frequency and the children’s productive language use. If the production of less frequent items can consistently be tracked in its stepwise growth, the argument can serve to further bolster the claims of most frequent items being acquired first, and later guiding the acquisition of more complex (or in this case rarer) words and phrases. The measures of lexical diversity across age groups also indicate a mild rise in sentences containing causative make-NP-V constructions. In terms of MSTTR-50, there is an increase between the 0-3 and 4-6 age groups, as well as between the 4-6 age group and that of adults. The same goes for RTTR, with both values suggesting a slightly greater difference between the two young age groups than between the 4-6 age group and adults. When interpreted together with stable lexical density, we begin to understand that the increase in diversity, regardless of its size, reflects an expansion of utterances with lexical and grammatical items in proportionate numbers, or it shows an increase in the number of word forms without the increase in tokens. For example, if the utterances grew in the number of different lexical items solely, the diversity would increase, but so would density. It is true that the increase in diversity is relatively mild and it is questionable to which degree it match the hypothesized progression in LD. Nevertheless, given that we know that the utterances containing causative constructions grew in size between the 0-3 and 4-6 age groups, but decreased in size between the 4-6 and the adult group, we can speculate that the measures of diversity truly reflect the number of different word forms averaged per the number of tokens. 245 Waters, G. S., & Caplan, D. (1996). The capacity theory of sentence comprehension: Critique of Just and Carpenter (1992). Psychological Review, 103(4), 761–772. Wellman, H. M., Harris, P. L., Banerjee, M., & Sinclair, A. (1995). Early understanding of emotion: Evidence from natural language. Cognition and Emotion, 9, 117–149. Wieczorek, R. (2010). Using MLU to study early language development in English. Psychology of Language and Communication, 14(2), 59. Wolfe-Quintero, K., Inagaki, S., & Kim, H. Y. (1998). Second language development in writing: Measures of fluency, accuracy, & complexity (No. 17). University of Hawaii Press. Wolfe-Quintero, K., Inagaki, S., & Kim, H. Y. (1998). Second language development in writing : Measures of fluency, accuracy, and complexity (Report No. 17). Honolulu: University of Hawai’i, Second Language Teaching and Curriculum Center Wolff, P., & Song, G. (2003). Models of causation and the semantics of causal verbs. Cognitive psychology, 47(3), 276-332. Wonnacott, E. (2011). Balancing generalization and lexical conservatism: An artificial language study with child learners. Journal of Memory and Language, 65(1), 1-14. Yoder, P. J. (2006). Predicting lexical density growth rate in young children with autism spectrum disorders. American Journal of Speech-Language Pathology, 15, 378–388. Yoder, P., & Stone, W. L. (2006). A randomized comparison of the effect of two prelinguistic communication interventions on the acquisition of spoken communication in preschoolers with ASD. Journal of Speech, Language, and Hearing Research, 49, 698–711 Yoder, P. J., Warren, S. F., & McCathren, R. B. (1998). Determining spoken language prognosis in children with developmental disabilities. American Journal of Speech Language Pathology, 7(4), 77–87. Zenker, F., & Kyle, K. (2021). Investigating minimum text lengths for lexical diversity indices. Assessing Writing, 47, 100505. Ziegeler, D. (2004). Grammaticalisation through constructions: The story of causative have in English. Annual Review of Cognitive Linguistics, 2(1), 159-195. Zora, S., & Johns-Lewis, C. (1989). Lexical density in interview and conversation. York Papers in Linguistics, 14, 89-100. 246 8.2. CHILDES corpora included in the research: Following is the list of CHILDES corpora sources used in this research. Though the corpora was accessed via the Sketch Engine platform, any use of it must be accompanied by referencing the work(s) which the authors of the individual subcorpora published, as indicated on the TalkBank pages (see English and American corpora under https://childes.talkbank.org/access/). 103 The access to both English and American corpora according to the author and the name is also available on the following webpage: https://sla.talkbank.org/TBB/childes. Bates, E., Bretherton, I., & Snyder, L. (1988). From first words to grammar: Individual differences and dissociable mechanisms. Cambridge, MA: Cambridge University Press. Bernstein, N. (1982). Acoustic study of mothers’ speech to language-learning children: An analysis of vowel articulatory characteristics. Unpublished doctoral dissertation. Bos-ton University. Bloom, L. (1970). Language development: Form and function in emerging grammars. Cambridge, MA: MIT Press Bohannon, J. N., & Marquis, A. L. (1977). Children’s control of adult speech. Child Development, 1002–1008. Brent, M. R. & Siskind, J. M. (2001). The role of exposure to isolated words in early vocabulary development. Cognition, 81(2), 31-44. Brown, R. (1973). A first language: The early stages. Cambridge, MA: Harvard University Press. Clark, E. V. (1978a). Awareness of language: Some evidence from what children say and do. In R. J. A. Sinclair & W. Levelt (Eds.), The child’s conception of language. Berlin: Springer Verlag. Cruttenden, A. (1978). Assimilation in child language and elsewhere. Journal of Child Language, 5, 373–378. 103 Note that some corpora do not carry the names of the researchers authoring the references listed (e.g. THOMAS, MIAMI, WOLFHEMP etc.). For example, the longitudinal data compiled under the so-called NEW ENGLAND corpus was primarily studied by Catherine Snow and Barbara Pan, with a number of other authors working on the corpusrelated publications that are listed as citation information for the use of the corpus on the TalkBank pages. 247 Davis, Barbara L. & Peter F. MacNeilage (1995). The articulatory basis of babbling. Journal of Speech and Hearing Research, 38, 1199-1211. Demetras, M. (1989a). Changes in parents’ conversational responses: A function of grammatical development. Paper presented at ASHA, St. Louis, MO. Demetras, M. (1989b). Working parents’ conversational responses to their two-year-old sons. University of Arizona. Demetras, M., Post, K., & Snow, C. (1986). Feedback to first-language learners. Journal of Child Language, 13, 275–292. Dickinson, D. K., & Tabors, P. O. (Eds.) (2001). Beginning literacy with language: Young children learning at home and school. Baltimore: Paul Brookes Publishing. Ervin-Tripp, S. (1978). Some features of early child–adult dialogues. Language in Society, 7(3), 357-373. Evans, M. A., & Schmidt, F. (1991). Repeated maternal book reading with two children: Language-normal and language-impaired. First Language, 11(32), 269-286. Fletcher, P., & Garman, M. (1988). Normal language development and language impairment: Syntax and beyond. Clinical Linguistics and Phonetics, 2, 97–114. Forrester, M. (2002). Appropriating cultural conceptions of childhood: Participation in conversation. Childhood, 9, 255-278. Garvey, C. (1979). An approach to the study of children’s role play. The Quarterly News-letter of the Laboratory of Comparative Human Cognition, 12. Gathercole, Virginia C. (1986). The acquisition of the present perfect: Explaining differences in the speech of Scottish and American children. Journal of Child Language 13, 537-560. Gleason, J. B. (1980). The acquisition of social speech and politeness formulae. In H. Giles, W. P. Robinson, & S. M. P. (Eds.), Language: Social psychological perspectives. Oxford: UK: Pergamon. Gopnik, M. (1989). Reflections on challenges raised and questions asked. In P. R. Zelazo & R. G. Barr (Eds.), Challenges to developmental paradigms. Hillsdale, NJ: Lawrence Erlbaum Associates. Haggerty, L. (1929). What a two-and-one-half-year-old child said in one day. Journal of Genetic Psychology, 38, 75–100. Hall, W. S., Nagy, W. E., & Linn, R. (1984). Spoken words: Effects of situation and social group on oral word usage and frequency. Hillsdale, NJ: Erlbaum. 248 Hayes, D. P. (1986a). The Cornell Corpus. Technical Report Series 86-1. Ithaca, NY: Department of Sociology, Cornell University. Hayes, D. P., & Ahrens, M. G. (1988). Vocabulary simplification for children: A special case of ‘motherese’?. Journal of child language, 15(2), 395-410. Henry, A. (1995). Belfast English and Standard English: Dialect variation and parameter setting. New York: Oxford University Press. Hicks, D. (1990). Kinds of texts: Narrative genre skills among children from two communities. In A. McCabe (Ed.), Developing narrative structure. Hillsdale, NJ: Erlbaum. Higginson, R. P. (1985). Fixing-assimilation in language acquisition. Unpublished doctoral dissertation. Washington State University. Howe, C. (1981). Acquiring language in a conversational context. New York: Academic Press. Inkelas, Sharon & Yvan Rose. 2007. Positional Neutralization: A Case Study from Child Language. Language 83, 707-736. Jones, G., & Rowland, C. F. (2017). Diversity not quantity in caregiver speech: Using computational modeling to isolate the effects of the quantity and the diversity of the input on vocabulary growth. Cognitive Psychology, 98, 1-21. Korman, M. (1984). Adaptive aspects of maternal vocalizations in differing contexts at ten weeks. First Language, 5:44-45. Kuczaj, S. A., & Maratsos, M. P. (1975). What children can say before they will. Merrill-Palmer Quarterly, 21, 89–111. Lieven, E., Salomo, D. & Tomasello, M. (2009). Two-year-old children’s production of multiword utterances: A usage-based analysis. Cognitive Linguistics,20, 3, 481-508. Lieven, E., Salomo, D. & Tomasello, M. (2009). Two-year-old children’s production of multiword utterances: A usage-based analysis. Cognitive Linguistics, 20, 3, 481-508. Lust, B., & Blume, M. (2016). Research methods in language acquisition: Principles, procedures, and practices. Walter de Gruyter GmbH & Co KG. MacWhinney, B. (1991). The CHILDES project: Tools for analyzing talk. Hillsdale, NJ: Erlbaum. MacWhinney, B., & Bates, E. (1978). Sentential devices for conveying givenness and newness: A cross-cultural developmental study. Journal of Verbal Learning and Verbal Behavior, 17, 539–558. 249 MacWhinney, B., & Snow, C. (1990). The Child Language Data Exchange System: An update. Journal of Child Language, 17, 457-472. McCune, L. (1995). A normative study of representational play at the transition to language. Developmental Psychology 31(2), 198-206. Miranda, E., Camp, L., Hemphill, L., & Wolf, D. (1992). Developmental changes in children’s use of tense in narrative. Paper presented at the Boston University Conference on Language Development, Boston. Morisset, C. E., Barnard, K. E., & Booth, C. L. (1995). Toddlers' language development: Sex differences within social risk. Developmental Psychology, 31(5), 851-865. Nelson, K. (Ed.) (1989). Narratives from the crib. Cambridge, MA: Harvard University Press. Ninio, A., Snow, C., Pan, B., & Rollins, P. (1994). Classifying communicative acts in children’s interactions. Journal of Communications Disorders, 27, 157-188. Pearson, Barbara Z. (2002). Narrative competence among monolingual and bilingual school children in Miami. In D. K. Oller and R. E. Eilers (Eds.), Language and literacy in bilingual children (pp. 135-174). Clevedon, UK: Multilingual Matters. Peters, A. (1987). The role of imitation in the developing syntax of a blind child. Text, 7, 289– 311. R. A. Berman & D. I. Slobin (1994). Relating events in narrative: A crosslinguistic developmental study. Hillsdale, NJ: Lawrence Erlbaum Associates. Rollins, P. R., (2003). Caregiver contingent comments and subsequent vocabulary Comprehension. Applied Psycholinguistics. 24, 221-234 Sachs, J. (1983). Talking about the there and then: The emergence of displaced reference in parent–child discourse. In K. E. Nelson (Ed.), Children’s language, Vol. 4, Hillsdale, NJ: Lawrence Erlbaum Associates. Soderstrom, M., Blossom, M., Foygel, R., & Morgan, J. L. (2008). Acoustical cues and grammatical units in speech to two preverbal infants. Journal of Child Language, 35, 869-902. Suppes, P. (1974). The semantics of children’s language. American Psychologist, 29, 103–114. Tager-Flusberg, H., Calkins, S., Nolin, T., Bamberger, T., Anderson, M., & Chandwick-Dias, A. (1990). A longitudinal study of language acquisition in autistic and Down syndrome children. Journal of Autism and Developmental Disorders, 20, 1–21. Tardif, T., Gelman, S. A., & Xu, F. (1999). Putting the “noun bias” in context: A comparison of English and Mandarin. Child development, 70(3), 620-635. 250 Theakston, A. L., Lieven, E. V. M., Pine, J. M., & Rowland, C. F. (2001). The role of performance limitations in the acquisition of verb-argument structure: an alternative account. Journal of Child Language, 28, 127-152. Valian, V. (1991). Syntactic subjects in the early speech of American and Italian children. Cognition, 40, 21–81. Van Houten, L. (1986). Role of maternal input in the acquisition process: The communica-tive strategies of adolescent and older mothers with their language learning children. Paper presented at the Boston University Conference on Language Development, Boston. Van Kleeck, A., & Beckley-McCall, A. (2002). A Comparison of Mothers' Individual and Simultaneous Book Sharing With Preschool Siblings. American Journal of Speech-Language Pathology 11 (2). Warren-Leubecker, A. (1982). Sex differences in speech to children. Unpublished doctoral dissertation. Georgia Institute of Technology. Weist, R. M. & Zevenbergen, A. (2008). Autobiographical memory and past time reference. Language Learning and Development, 4(4), 291 – 308. Weismer, S. E. (2017). Typical talkers, late talkers, and children with specific language impairment: A language endowment spectrum?. In Language disorders from a developmental perspective (pp. 83-101). Psychology Press. Wells, C. G. (1981). Learning through interaction: The study of language development. Cambridge, UK: Cambridge University Press. 251 9. ABSTRACT 9.1. Abstract (in English) This doctoral dissertation investigates various developmental parameters in the context of first language acquisition through a comparative triconstructional analysis (ditransitive, causative, and relative constructions) in children whose first language is English. The research is based on corpus analysis, during which the mentioned constructions were extracted and analysed from the English part of the CHILDES corpora. For a comprehensive interpretation and understanding of first language acquisition, intra-linguistic evidence (i.e., language phenomena related to syntax and lexicon) is observed and compared across three age-defined groups (0-3, 4-6, and 18+). Given the specificities of the three observed constructions, the analyses are tailored to their characteristics (such as frequency, complexity, and variability of certain lexemes and structural patterns appearing in these constructions). Simultaneously, they are conducted to allow not only inter-age group comparisons but also inter-constructional comparisons (such as the degree of similarity in the representation and frequency rankings of certain lexical units, as well as measures like structural and lexical complexity, including the mean length of utterance, lexical density, and type-token ratios). The main contribution of this research is the methodologically unique way of monitoring the parallel development of children's speech at the lexical, morphological, and syntactic levels, which analytically and interpretatively unifies the mentioned constructions through the lens of cognitive theories of language acquisition that are primarily based on the ‘usage-based’ model. Among other things, the research results show what can be described as a linguistic version of Pareto's principle, where all three age groups in the production of the observed constructions largely rely on a few lexical units in certain construction slots, and where the majority of their linguistic expression when it comes to said constructions is occupied by a relatively small number of structural patterns. Significant lexical and structural similarities among the same constructions were confirmed at several levels across the observed age groups, along with a more “conservative” language use and a more pronounced concentration of certain items and syntactic patterns with the declining of age. Furthermore, the ‘item-based’ claims about “simpler” constructions being more lexically restricted production-wise than more “complex” constructions was not confirmed; rather, an examination of certain lexical items used in specific construction slots showed that item-based tendencies are present in all constructions across all three age groups. 252 In conclusion, the data on frequency and lexical and syntactic complexity obtained through comparative analysis of child and adult speech contribute to existing body of research on the nature of language acquisition and development. Keywords: language acquisition, corpus-based research, lexical and structural complexity, ditransitive constructions, periphrastic causative constructions, relative constructions 253 9.2. Abstract (in Croatian) Naslov disertacije: Razvoj leksika i strukturne složenosti kod usvajanja engleskog kao prvog jezika: dvoprijelazne, kauzativne i relativne konstrukcije Sažetak: U ovom doktorskom radu se istražuju različiti jezično-razvojni parametri u kontekstu usvajanja prvog jezika u vidu usporedne trokonstrukcijske analize (dvoprijelazne, kauzativne i relativne konstrukcije) kod djece kojoj je prvi jezik engleski. Istraživanje se temelji na korpusnom istraživanju prilikom koje su se ekstrahirale i analizirale spomenute konstrukcije iz engleskog dijela korpusa CHILDES. U svrhu cjelovitog tumačenja i razumijevanja usvajanja prvog jezika promatraju se i uspoređuju unutarjezični dokazi (odnosno jezične pojave vezane za sintaksu i leksik) između tri dobno određene grupe (0-3, 4-6 i 18+). S obzirom na posebnosti triju promatranih konstrukcija, analize su prilagođene njihovim osobitostima (poput učestalosti, složenosti i varijabilnosti određenih leksema i strukturnih obrazaca koji se pojavljuju u spomenutim konstrukcijama), ali istovremeno provedene na način da, osim međudobnih usporedbi, dopuštaju i međukonstrukcijske usporedbe (poput stupnja različitosti u zastupljenosti i hijerarhiji učestalosti određenih leksičkih jedinica, te mjera poput strukturne i leksičke složenosti, uključujući prosječan broj morfema na sto rečenica, leksičku gustoću, te omjer različnica i pojavnica). Glavni doprinos ovog istraživanja je metodološki jedinstveno praćenje paralelnog razvoja dječjeg govora na leksičkoj, morfološkoj i sintaktičkoj razini, koje analitički i interpretativno objedinjuje navedene konstrukcije kroz prizmu kognitivističkih teorija o usvajanju jezika temeljenima prvenstveno na uporabi (eng. usage-based model). Između ostalog, rezultati istraživanja pokazuju ono što se može opisati kao lingvistička inačica Paretovog pravila, gdje se sve tri dobne skupine u proizvodnji promatranih konstrukcija uglavnom oslanjaju na nekolicinu leksičkih jedinica u određenim konstrukcijskim položajima, te da najveći dio njihovog jezičnog izričaja kod proizvodnje istih zauzima relativno mali broj strukturnih obrazaca. Na nekoliko razina je potvrđena značajna leksička i strukturna sličnost kod istih konstrukcija među dobno različitim skupinama, ali i „konzervativnija“ upotreba jezika te izraženija koncentracija određenih elemenata i sintaktičkih uzoraka s padom dobi. Osim toga, 254 hipoteza da će „jednostavnije“ konstrukcije biti orijentirane na manji broj leksičkih elemenata (eng. item-based) od složenijih konstrukcija nije potvrđena; odnosno, pregled određenih leksičkih elemenata korištenih na određenim konstrukcijskim položajima pokazao je da su item-based tendencije prisutne u svim konstrukcijama i dobno različitim skupinama. U konačnici, podatci o učestalosti i leksičko-sintaktičkoj složenosti dobiveni komparativnom analizom govora djece i odraslih doprinose postojećim raspravama o prirodi jezičnog usvajanja i razvoja. Ključne riječi: jezično usvajanje, korpusno istraživanje, leksička i jezična složenost, dvoprijelazne konstrukcije, kauzativne konstrukcije, relativne konstrukcije 261 [tag="DT.*"] [lemma="who|whom"] [tag="J.* "| tag="DT.*"| tag="PPZ"| tag="POS"]{0,2}[tag="PP:?"|tag="DT.*"|tag="N.*"] [tag="V.*"] (e.g. Those whom X knew…) NP-whose-NP-VP [tag="N.*"] [lemma="whose"] [tag="J.* "| tag="DT.*"| tag="PPZ"| tag="POS"]{0,2}[tag="PP:?"|tag="DT.*"|tag="N.*"] [tag="V.*"] (e.g. Man whose X knew…) [tag="PP.*"] [lemma="whose"] [tag="J.* "| tag="DT.*"| tag="PPZ"| tag="POS"]{0,2}[tag="PP:?"|tag="DT.*"|tag="N.*"] [tag="V.*"] (e.g. He whose X knew…) [tag="DT.*"] [lemma="whose"] [tag="J.* "| tag="DT.*"| tag="PPZ"| tag="POS"]{0,2}[tag="PP:?"|tag="DT.*"|tag="N.*"] [tag="V.*"] (e.g. Those whose X knew…) NP-where-NP-VP [tag="N.*"] [lemma="where"] [tag="J.* "| tag="DT.*"| tag="PPZ"| tag="POS"]{0,2}[tag="PP:?"|tag="DT.*"|tag="N.*"] [tag="V.*"] (e.g. The place where X happens…) [tag="PP.*"] [lemma="where"] [tag="J.* "| tag="DT.*"| tag="PPZ"| tag="POS"]{0,2}[tag="PP:?"|tag="DT.*"|tag="N.*"] [tag="V.*"] (e.g. It where X happens…) [tag="DT.*"] [lemma="where"] [tag="J.* "| tag="DT.*"| tag="PPZ"| tag="POS"]{0,2}[tag="PP:?"|tag="DT.*"|tag="N.*"] [tag="V.*"] (e.g. There where X happens…) NP-when-NP-VP [tag="N.*"] [lemma="when"] [tag="J.* "| tag="DT.*"| tag="PPZ"| tag="POS"]{0,2}[tag="PP:?"|tag="DT.*"|tag="N.*"] [tag="V.*"] (e.g. The time when X happens…) [tag="PP.*"] [lemma="when"] [tag="J.* "| tag="DT.*"| tag="PPZ"| tag="POS"]{0,2}[tag="PP:?"|tag="DT.*"|tag="N.*"] [tag="V.*"] (e.g. It when X happens…) [tag="DT.*"] [lemma="when"] [tag="J.* "| tag="DT.*"| tag="PPZ"| tag="POS"]{0,2}[tag="PP:?"|tag="DT.*"|tag="N.*"] [tag="V.*"] (e.g. Then when X happens…) 262 10.2. Additional data on constructions For a detailed explanation of subsequent measures calculated with the ‘eval’ function in CLAN (Table 40, Table 41, Table 42, Table 43) see MacWhinney (2023a, pp.132-133). Table 40. MLUs and other data on constructions; ‘eval’ function in CLAN Constructio n Age grou p Total _Utts MLU_ Utts MLU_ Words MLU_ Morph emes FREQ_ types FREQ_ tokens FREQ_ TTR Verbs_ Utt ditransitive give-NP-NP 0-3 1014 978 6.56 7.16 682 6975 0.10 1.39 4-6 798 768 8.32 9.21 799 6960 0.12 1.63 18+ 11205 11052 8.47 9.30 3283 99654 0.03 1.66 causative make-NP-V 0-3 386 375 8.06 8.59 457 3129 0.15 2.35 4-6 298 280 10.43 11.25 536 3163 0.17 2.62 18+ 3204 3146 9.15 9.82 1909 29756 0.06 2.43 relative Nthat-NP-VP 0-3 206 203 10.23 11.29 431 2132 0.20 1.75 4-6 497 475 13.06 14.09 855 6513 0.13 2.22 18+ 5251 5144 13.59 14.78 3626 71902 0.05 2.23 random sample 0-3 20739 12222 3.89 4.16 3176 69791 0.05 0.50 4-6 16377 10748 5.24 5.61 3564 77242 0.05 0.72 18+ 18290 12088 4.45 4.75 3853 73818 0.05 0.61 Table 41. Part-of-Speech proportions in constructions; ‘eval’ function in CLAN Construction Age group %_No uns %_Pl urals %_Ve rbs %_pr ep %_ad j %_ad v %_co nj %_de t %_pr o ditransitive give-NP-NP 0-3 18% 9% 21% 2% 2% 2% 1% 8% 29% 4-6 18% 11% 19% 2% 2% 3% 1% 10% 26% 18+ 18% 8% 19% 3% 3% 4% 1% 10% 23% causative make-NP-V 0-3 11% 13% 31% 3% 2% 7% 1% 5% 23% 4-6 14% 18% 26% 4% 2% 7% 2% 5% 22% 18+ 12% 15% 28% 4% 3% 6% 2% 5% 22% relative Nthat-NP-VP 0-3 17% 27% 19% 4% 3% 5% 1% 10% 24% 4-6 18% 19% 18% 5% 3% 5% 1% 9% 23% 18+ 18% 19% 18% 7% 3% 5% 1% 10% 20% random sample 0-3 19% 10% 16% 4% 4% 6% 1% 6% 16% 4-6 16% 13% 16% 5% 4% 6% 1% 7% 17% 18+ 17% 12% 16% 5% 4% 7% 1% 7% 16% 263 Table 42. Tenses and forms in constructions; ‘eval’ function in CLAN Construction Age grou p %_Aux %_Mod %_3 S %_13S %_PAST %_PAST P %_PRES P ditransitive give-NP-NP 0-3 2% 3% 5% 0% 13% 1% 11% 4-6 1% 3% 11% 1% 17% 1% 10% 18+ 2% 4% 7% 1% 14% 3% 10% causative make-NP-V 0-3 0% 3% 6% 1% 6% 0% 7% 4-6 1% 2% 9% 1% 13% 1% 7% 18+ 1% 3% 8% 1% 10% 1% 7% relative Nthat-NP-VP 0-3 1% 1% 14% 2% 24% 5% 12% 4-6 1% 1% 12% 5% 25% 3% 9% 18+ 1% 1% 14% 5% 32% 5% 10% random sample 0-3 1% 2% 14% 1% 12% 4% 13% 4-6 1% 2% 13% 3% 18% 4% 11% 18+ 1% 2% 13% 2% 13% 4% 10% Table 43. Open class and closed class words ratios; ‘eval’ function in CLAN Construction Age group noun_verb open_closed #open-class #closed-class ditransitive give-NP-NP 0-3 0.8 0.8 3008 3860 4-6 0.9 0.8 2963 3868 18+ 0.9 0.8 43926 53936 causative make-NP-V 0-3 0.4 1.1 1629 1454 4-6 0.5 1.0 1543 1547 18+ 0.4 1.0 14789 14382 relative Nthat-NP-VP 0-3 0.9 0.8 935 1167 4-6 1.0 0.8 2816 3571 18+ 1.0 0.8 31498 39437 random sample 0-3 1.1 1.0 32199 33465 4-6 1.0 0.8 33658 39693 18+ 1.1 0.9 33516 36206 264 Table 44. Token-based animacy of the nouns in the observed construction slots Construction Constructio n slot Age group INAN ANc ANp ANt ANb ditransitive give-NP-NP 1st (indirect) object 0-3 2% 94% 3% 1% 0% 4-6 2% 95% 2% 1% 0% 18+ 6% 86% 6% 2% 0% ditransitive give-NP-NP 2nd (direct) object 0-3 94% 3% 1% 1% 1% 4-6 95% 1% 1% 2% 1% 18+ 96% 1% 0% 1% 2% causative make-NP-V causee 0-3 28% 39% 19% 10% 4% 4-6 45% 32% 13% 5% 5% 18+ 37% 37% 9% 9% 9% relative N-that-NP-VP relativized noun 0-3 89% 9% 0% 1% 0% 4-6 87% 8% 0% 4% 1% 18+ 90% 6% 1% 2% 1% For a detailed explanation of indices on lexical variation outlined in the subsequent tables (Table 45, Table 46, Table 47) see Lu (2012). Table 45. Measures of lexical sophistication Construction Age group LD LS1 LS2 VS1 VS2 CVS1 ditransitive give-NP-NP 0-3 0.44 0.25 0.40 0.00 0.00 0.04 4-6 0.45 0.23 0.40 0.02 0.12 0.24 18+ 0.45 0.26 0.41 0.03 0.29 0.38 causative make-NP-V 0-3 0.51 0.14 0.31 0.03 0.62 0.56 4-6 0.50 0.14 0.30 0.05 1.30 0.80 18+ 0.52 0.19 0.38 0.05 1.46 0.86 relative N-thatNP-VP 0-3 0.42 0.24 0.36 0.05 0.94 0.69 4-6 0.40 0.20 0.37 0.07 2.01 1.00 18+ 0.41 0.21 0.37 0.05 1.18 0.77 265 Table 46. Measures of lexical variation Construction Age grou p NDW NDW Z-50 NDW - ER50 NDW -ES50 TTR MST TR50 CTT R RTT R LOG TTR UBE R ditransitive give-NP-NP 0-3 257 28 29.70 28.70 0.19 0.56 4.96 7.02 0.77 13.63 4-6 360 27 31.10 30.60 0.21 0.61 6.20 8.77 0.79 15.53 18+ 379 31 33.80 30.70 0.21 0.64 6.29 8.89 0.79 15.60 causative make-NP-V 0-3 321 30 34.80 30.40 0.18 0.63 5.39 7.62 0.77 14.21 4-6 417 33 32.90 33.10 0.19 0.66 6.30 8.91 0.78 15.49 18+ 427 31 35.00 34.80 0.22 0.68 6.93 9.80 0.80 16.58 relative Nthat-NP-VP 0-3 412 30 35.40 32.90 0.19 0.64 6.26 8.85 0.78 15.43 4-6 470 36 36.00 32.60 0.18 0.66 6.46 9.14 0.78 15.61 18+ 604 34 36.10 33.80 0.20 0.70 7.77 10.99 0.80 17.32 Table 47. Measures of lexical variation Construction Age group VV1 SVV1 CVV1 LV VV2 NV ADJV ADVV MODV ditransitive give-NP-NP 0-3 0.09 2.12 1.03 0.34 0.04 0.63 0.02 0.03 0.06 4-6 0.15 6.62 1.82 0.38 0.06 0.61 0.04 0.04 0.07 18+ 0.16 8.24 2.03 0.39 0.07 0.62 0.06 0.03 0.09 causative make-NP-V 0-3 0.14 10.17 2.25 0.29 0.08 0.70 0.03 0.04 0.06 4-6 0.19 21.48 3.28 0.31 0.10 0.63 0.03 0.04 0.07 18+ 0.20 20.96 3.24 0.36 0.11 0.74 0.04 0.06 0.10 relative N-thatNP-VP 0-3 0.28 26.79 3.66 0.39 0.11 0.51 0.03 0.04 0.07 4-6 0.27 28.74 3.79 0.37 0.10 0.52 0.04 0.04 0.08 18+ 0.26 30.94 3.93 0.42 0.10 0.61 0.04 0.05 0.09