scieee AI-readable full text Open interactive document viewer

Graph Theory and Universal Grammar

Franco, Ludovico

Abstract

In the last few years, Noam Chomsky (1994; 1995; 2000; 2001) has gone quite far in the direction of simplifying syntax, including eliminating X-bar theory and the levels of D-structure and S-structure entirely, as well as reducing movement rules to a combination of the more primitive operations of Copy and Merge. What remain in the Minimalist Program are the operations Merge and Agree and the levels of LF (Logical Form) and PF (Phonological form). My doctoral thesis attempts to offer an economical theory of syntactic structure from a graph-theoretic point of view (cf. Diestel, 2005), with special emphases on the elimination of category and projection labels and the Inclusiveness Condition (Chomsky 1994). The major influences for the development of such a theory have been Chris Collins’ (2002) seminal paper “Eliminating labels”, John Bowers (2001) unpublished manuscript “Syntactic Relations” and the Cartographic Paradigm (see Belletti, Cinque and Rizzi’s volumes on OUP for a starting point regarding this paradigm). A syntactic structure will be regarded here as a graph consisting of the set of lexical items, the set of relations among them and nothing more.

Full text

Graph Theory and Universal Grammar Tesi di Dottorato in Linguistica Università of Firenze - Facoltà di Lettere e Filosofia/ Dipartimento di Linguistica In associazione con CISCL, Centro Interdipartimentale di Studi Cognitivi sul Linguaggio, Siena Ciclo XX – AA 2007/08 Settore disciplinare L-Lin/01 Candidato: Ludovico Franco Supervisore: Coordinatore del dottorato: Prof. Luigi Rizzi Prof. Leonardo Maria Savoia 2 3 Ringraziamenti Desidero ringraziare anzitutto il Prof. Luigi Rizzi, mio advisor a Siena e la Prof. Rita Manzini, mia tutor fiorentina. Ritengo che sia un privilegio enorme per un linguista quello di poter cogliere stimoli da persone così piene di senso ed entusiasmo. E che tanto hanno dato alla linguistica contemporanea. Un ringraziamento particolare va alla Prof. Adriana Belletti. Sono certo che senza la sua capacità di coinvolgere e appassionare nelle lezioni per il corso di Morfosintassi a Siena dell’A.A. 2003/2004 non mi sarei innamorato di questa disciplina (almeno non fino al punto da spendere le mie energie in un dottorato senza la borsa). E ai i colleghi del CISCL, in particolare Enzo, Giuliano e Ida. Altra citazione d’obbligo per la mia famiglia e per l’aiuto morale e, soprattutto, materiale nel corso di mesi difficili (a cominciare da Tempo Reale in cattive acque e le conseguenti, inevitabili difficoltà economiche occorse). Tuttavia, non la ringrazio come spesso vedo si usa fare negli Acknowledgments “per aver reso possibile la stesura della tesi”, bensì per aver permesso che non mi indebolisse tutto quanto è passato intorno, in particolare la mia assurda querelle con l’Università degli Studi di Firenze con annessa sospensione dal corso di laurea in Medicina. Una cosa che mi ha fatto male, che ancora stento a capire e che ha inevitabilmente condizionato il mio lavoro per il dottorato. Ringrazio poi Gigapedia.org, la mia Biblioteca d’Alessandria configurabile e le proxy lists in giro per la rete. Scuse sentite al Prof. Leonardo Savoia per aver agito d’istinto e per non aver fatto riferimento a lui quando sarebbe stato doveroso. Infine un ringraziamento al Prof. Andrea Moro, anche se non lo conosco personalmente e l’ho incontrato solo ad un paio di seminari, per una frase che ho letto in un suo bel libro divulgativo recente in merito al rapporto tra neuroscienze e linguistica teorica: “…Forse non tutto il male viene per nuocere, una visione parziale non è di per sé negativa, basta che sia esplicita e non abbia la pretesa di esaurire l’argomento […] Senza poi escludere che siano i mestieri di neuroscienziato e di linguista a essere vecchi: non è detto che non ne debba nascere uno nuovo”. Un bello stimolo. Una bella sfida possibile. Da ultimo una speranza: mi auguro che il mio pseudo-inglese, da qui in avanti, assomigli il più possibile ad una lingua naturale. Questa tesi è dedicata a mia moglie Salomé, straordinaria ogni giorno di più: senza di lei avrei perso il fil-o. 4 ABSTRACT In the last few years, Noam Chomsky (1994; 1995; 2000; 2001) has gone quite far in the direction of simplifying syntax, including eliminating X-bar theory and the levels of D-structure and S-structure entirely, as well as reducing movement rules to a combination of the more primitive operations of Copy and Merge. What remain in the Minimalist Program are the operations Merge and Agree and the levels of LF (Logical Form) and PF (Phonological form). My doctoral thesis attempts to offer an economical theory of syntactic structure from a graph-theoretic point of view (cf. Diestel, 2005), with special emphases on the elimination of category and projection labels and the Inclusiveness Condition (Chomsky 1994). The major influences for the development of such a theory have been Chris Collins’ (2002) seminal paper “Eliminating labels”, John Bowers (2001) unpublished manuscript “Syntactic Relations” and the Cartographic Paradigm (see Belletti, Cinque and Rizzi’s volumes on OUP for a starting point regarding this paradigm). A syntactic structure will be regarded here as a graph consisting of the set of lexical items, the set of relations among them and nothing more. 5 Quotes “Any inference rule can be non-trivially revised so that one fails to accept it where one once accepted it, or vice versa.” W.V. Quine "Graphics reveal data." E. Tufte 6 7 Preface / Guidelines This work consists of six chapters. In the first chapter I introduce how topologic networks could be useful for a wide range of cognitive matters, from a theoretic perspective, given some basic principles. I introduce as well the fascinating hypothesis of a topological and viral nature of language (spreading linguistic variation) made up by Piattelli Palmarini & Uriagereka in 2004. Then, I review the idea of David B. Searls who claims that many techniques used in bioinformatics and genetics, even if developed independently, may be seen to be grounded in generative linguistics and, symmetrically, the discover of a “language” gene: FOXP2. These discoveries (or beliefs) are reported here to show how fruitful could be a crossdisciplinary attitude in the search for the innermost nature of Language. That’s also the main reason for which I have tried to apply some basic principles of Graph Theory to Minimalist principles. The last section of the chapter is a rough historical survey (mainly based on Wildgen, 1994, 2000) of a “topological” way of thinking among human sciences, from Lullus, Bruno and Leibniz to Fillmore, Minsky and Langacker. Note that the first chapter is intended to be quite light and does not require a narrow competence in the field of generative linguistics. In the second chapter “Graphs and Transformations” I give a survey of some essentials of Graph Theory from a mathematical point of view, with major emphases on graph transformations including a translation of Chomsky grammars into graph grammars, showing the computational completeness of graph transformation. Unfortunately, graphs are quite generic structures that can be encountered in many variants in the literature, and there are also many ways to apply rules to graphs. One cannot deal with all possibilities in a basic survey and it would also go beyond the scope of my work. I focus on directed graphs (syntactic trees are a special kind of directed, label-edged graphs), because the directed graphs can be specialized into many other types of graphs, and I define graph grammars as a language-generating device with the more general notion of a transformation unit that models binary relations on graphs. Chapter three “From bare phrase structure to syntactic graphs” is the core of the work with the graph theoretic (re)definition of internal and external MERGE and the explanation of the way lexical graphs could capture 8 constituency, C-command and other important syntactic relations expressed in standard tree representations (assuming for example that different PF interpretations could be elegantly derived via syntactic graphs using traversal algorithms developed by Yasui (2004a,b)). I also introduce in this section other theoretical issues regarding a labelfree syntax, such as the ones developed by Chris Collins (2001; 2002), John Bowers (2001), and Joan Chen-Main (2006). Finally, I outline some light similarities of Brody’s Mirror Theory (1997) with my proposal. Chapter four “Topics of Persian Syntax and a graph based analysis of Persian Ezafe” introduces some challenging issues of Persian grammar and develops a graph based account of the “Ezafe puzzle” in Western Indoiranian languages, in which NP modifiers standardly occur postnominally and “link” to the noun head via an Ezafe particle (Ez), which may be invariant (Persian, Sorani), or agree with N in φ-features (Kurmanji, Zazaki). We will use in our research the precious cross-linguistical data collected by Larson and Yamakido (2005; 2006) and Samvelian (2006). After a basic (but articulated) sketch of Persian syntax, with major emphases on some syntactic aspects that I have already considered in previous works, such as word order and split-headness, Inverse Case Attraction of Persian relative clauses and Persian light verb constructions which seems to be an ouvert instance of Hale and Keyser (1993; 2002) compositional syntactic analysis (without theta-roles) of Argument Structure (see also Harley, Folli, Karimi, 2003), I give a detailed review of previous analyses of the Ezafe morpheme in the generative framework (cfr. Samiiam, 1983; 1994; Ghomeshi, 1997; Kahnemuyipour, 2000; Franco, 2004; Larson and Yamakido, 2005; Samvelian; 2006 among others) and I develop here a graph analysis of the Ezafe phenomenon based on Den Dikken and Singhapreecha (2004), where the authors give a cross-linguistic account (the point of departure for them was the comparison between French and Thai) of the noun phrases in which linkers occur, in terms of DP-internal Predicate Inversion (see also Moro, 1997; 2000). This approach prompts an analysis of relative-clause constructions that recognizes relative clauses as predicates of DP-internal small clauses, combining the attractions of the traditional approach and the Vergnaud/Kayne raising approach by assigning relative clauses an internal structure similar to the traditional one while giving it the external distribution of a predicate by treating it as the predicate of a noun phrase-internal small clause. In other words, we could say that a sort of Generalized Predicate Inversion in the Persian complex noun phrase is marked by the presence of the Ezafe mopheme. 9 In a graph theoretic perspective the assumption that a syntactic element could be interpreted as a linker, implies the theoretical necessity that certain linguistic items could be selected by the Lexicon as Edges (E), instead of vertexes (V) and this is a stimulating fact. The fifth chapter is a ground for some (possible and extreme) theoretical consequences of a graph theoretic analysis of language faculty. The Conclusion follows. This work also has three appendixes. The first involve the field of logics and is a rough review of Charles Sanders Peirce’s existential graphs (based mainly on Proni, 1992); the second is far more near to the questions and answers raised in this work, and has a straight linguistic topic, showing some of the most interesting dynamics of the Relational Grammar Paradigm established in the seventies by Perlmutter and Postal. Finally the third appendix is a brief application of a Graph theoretic derivation to an Austronesian Language (Tagalog). 16 arise in the individual, and, ideally, how they arose in the evolutionary history of our species. A few years ago, Marc Hauser, Noam Chomsky and Tecumseh Fitch published an influential paper, titled “The faculty of language” (Hauser et al., 2002). They ask which components of our linguistic abilities are uniquely human, and which components, or close analogues of those, can be found in other species (in a broad sense). In a talk during the Basque Country encounter with Noam Chomsky, held in 2006 in San Sebastian, Mark Hauser defines language: “as a mind-internal computational system designed for thought and often externalized in communication. That is, language evolved for internal thought and planning and only later was co-opted for communication, so this sets up a dissociation between what we do with the internal computation as opposed to what the internal computation actually evolved for. We [Hauser, Chomsky and Fitch] defined the faculty of language in the broad sense (FLB) as including all the mental processes that are both necessary and sufficient to support language. The reason why we want to do it in that way is because there are numerous things internal to the mind that will be involved in language processing, but that need not be specific to language. For example, memory is involved in language processing, but it is not specific to language. So it’s important to distinguish those features that are involved in the process of language computation from those that are specific to it. That’s why we developed the idea of the faculty of language in the narrow sense (FLN), a faculty with two key components: 1) those mental processes that are unique to language, and 2) those that are unique to humans. Therefore, it sets out a comparative phylogenetic agenda in that we are looking both for what aspects are unique to humans, but also what aspects are unique to language as a faculty8”. Therefore, a very good summing-up of the observations made in that seminal article is the one made by Dennis Ott (2007): “The picture that Hauser, Chomsky and Fitch offer is that of ‘blind’ recursive generation, coupled with interface components that transform the output of narrow syntax into representations encoding sound and meaning, and a range of modalities of performance, oral speech being one of them, that lie outside the narrow-syntactic system. suggesting that these performative components—conceptual-intentional and sensorimotor systems—are present in other species, such as higher primates—a conjecture that goes back to Descartes” (Ott, 2007: 78). 8 From M. Piattelli-Palmarini, J. Uriagereka and P. Salaburu Eds. (in press) Of Minds and Language: The Basque Country Encounter with Noam Chomsky, Oxford University Press. 17 I think that if the function of language (the narrow one) evolved only in humans, this suggests that its evolutionary history is incredibly short in biological terms. Indeed, there is evidence that language was born in humans about 100,000 to 40,000 years ago (Bickerton, 1990; Lass, 1997). The recent paradigm shift in evolutionary biology, commonly labelled the Evo-Devo paradigm, has given wide credibility to claims of saltational emergence of a trait, triggered by some minimal mutation, or as a by-product of a functionally unrelated change, such as, for example, brain growth. For an introduction, one of the many works of divulgation written by S. J. Gould could be an interesting reading. According to Anderson and Lightfoot (2003), which interestingly titled their book The Language Organ – Language as Cognitive Physiology, the theory of (universal) grammar must hold universally such that any person’s grammar can be attained on the basis of naturally available trigger experiences. “The mature grammar must define an infinite number of expressions as wellformed, and for each of these it must specify at least the sound and the meaning. A description always involves these three items and they are closely related; changing a claim about one of the items usually involves changing claims about the other two. The grammar is one subcomponent of the mind, a mental organ which interacts with other cognitive capacities or organs. Like the grammar, each of the other organs is likely to develop in time and to have distinct initial and mature states” (Anderson and Lightfoot (2003: 45). Modern physiology has discovered that the visual system recognizes triangles, circles or squares through the structure of the circuits that filter and recompose the retinal image (see the topographical model of Hubel and Wiesel, 1962). Certain nerve cells respond only to a straight line sloping downward from left to right within a specific, narrow range of orientations; other nerve cells to lines sloped in different directions. The range of angles that an individual neuron can register is set by a genetic program, but experience is needed to fix the precise orientation specificity (cf. the Nobel prize R. Sperry, 1961). In Anderson and Lightfoot (2003) is depicted this interesting experiment: “In the mid-1960s David Hubel, Torsten Wiesel, and their colleagues devised an ingenious technique to identify how individual neurons in an animal’s visual system react to specific patterns in the visual field (including horizontal and vertical lines, moving spots, and sharp angles). They found that particular nerve cells were set within a few hours of birth to react only to certain visual stimuli, and, furthermore, 18 that if a nerve cell is not stimulated within a few hours, it becomes totally inert in later life. In several experiments on newborn kittens, it was shown that if a kitten spent its first few days in a deprived optical environment (a tall cylinder painted only with vertical stripes), only the neurons stimulated by that environment remained active; all other optical neurons became inactive because the relevant synapses degenerated, and the kitten never learned to see horizontal lines or moving spots in the normal way” (Anderson and Lightfoot, 2003: 51. Cf. their work for other interesting interdisciplinary data). We commonly see the process of language learning as a similarly selective process: parameters are provided by the genetic equipment, and relevant experience fixes those parameters (Piattelli-Palmarini and Uriagereka, 2004). Notice that a certain mature cognitive structure, in the sense of Anderson and Lightfoot, emerges at the expense of other possible structures, which are lost irretrievably as the inactive synapses degenerate. Following Chomsky (2004), to a good approximation, the internal structure of the language organ is determined by four factors. The first factor is the genetic endowment, expressed in the initial state of the Language Faculty. The second factor is external linguistic data, which shapes the development of the language organ within narrow boundaries (cf. Guasti, 2003). The third factor comprises what is called “developmental constraints” in theoretical biology and general principles of biophysics. The fourth factor concerns the embedding of the Language Faculty within the mind, that is, the way it interfaces with other components (Ott, 2007). Anyway, it is interesting to report here that Steven Pinker and Ray Jakendoff (2005) disagree in many respects (dichotomies) with the article by Hauser, Fitch and Chomsky. These include: i) The Narrow/Broad dichotomy, which makes space only for completely novel capacities and for capacities taken intact from non-linguistic and non-human capacities, omitting capacities that may have been substantially modified in the course of human evolution; ii) The utility/original-function dichotomy, which conceals the possibility of capacities that are adaptations for current use; iii) The human/non-human dichotomy, which fails to distinguish similarity due to independently evolved analogous functions from similarity due to inheritance from a recent common 19 ancestor; iv) the core/non-core and syntax/lexicon dichotomies, which omit the vast set of productive linguistic phenomena that cannot be analyzed in terms of narrow syntax, and which thus incorrectly isolate recursion as the only unique development in the evolution of language. My personal opinion is that we acquire a productive system, a grammar, in accordance with the requirements of a genotype. We must assume that the linguistic genotype yields finite grammars, because they are represented in the finite space of the brain, but that they range over an infinity of possible sentences. This (virtual or real) genotype may act as a “parser” which interacts with the grammar to assign structures and meanings to incoming speech signals, and captures our capacity to understand spoken language in real time. I believe that the linguistic module contains abstract structures which are compositional (consisting of units made up of smaller units) and which fit a narrow range of possibilities. At this point, speculations concerning language evolution become highly relevant for linguistic theory and I suppose that the proposal of Piattelli Palmarini and Uriagereka (2004), for which language has a viral origin, could be seen as more than a simple provocation. Piattelli Palmarini and Uriagereka (henceforth: PP&U) argue that such a rapid transition (I mean, from non-language to language) cannot be easily accounted for in customary “adaptive” evolutionary terms and they propose that only a brain reorganization “of a drastic and sudden sort” could have given raise to such a state of affairs. This could be realistic simply assuming the Evo-devo paradigm cited above, but PP&U consider that two main facts suggest that this reorganization of human cognitive functions may have been “epidemic” in origin: i) Recent evolutionary accounts for the emergence of broad systems in organisms point in the direction of “horizontal” transmission of nucleic material, often of viral origin9 (from viruses to parasites to transposable elements: the example par excellence is 9 In situations of this sort, evolutionary rapidity is a consequence of boosting relevant numbers by having not individuals, but entire populations, carry relevantly mutated genetic material. Cf. Microbiology and Immunology On-Line Textbook: USC School of Medicine. 20 the origin of the Adaptive Immune System); ii) When the formal properties behind context-sensitivity in grammar (for example as displayed in syntactic displacement) are studied in Minimalist terms, a surprising parallelism surfaces with the workings of the Adaptive Immunity. The point of departure for PP&U is the consideration of the importance of “horizontal” transmission of mobile DNA sequences - called transposable elements - that are pervasive in the genomes of bacteria, plants and animals. These elements replicate fast and efficiently and it is common to find hundreds of thousands of copies of such elements in one single genome. Indeed, initial sequencing of the human genome revealed that as much as 45% of the total is constituted of DNA that originated from transposable elements10. Transposable elements can sometimes move “laterally” between species, a phenomenon known as horizontal transfer. Once these horizontal transfers of genetic material have successfully taken place, then ordinary “vertical” transmission perpetuates the new genome: for example, some researchers have suggested that the immune system of higher vertebrates is the product of the activity of a transposable element that was “domesticated” following horizontal transfer from a bacterium millions of years ago (cf. Hiom et al., 1998 cited in PP&U, 2004; and for a “broad spectrum” analysis the article of 1987, “Mitochondrial DNA and human evolution” by Cann et al.). Then, the authors try to justify the assumption operated by Chomsky of the structural perfection of the “language organ”, also considering the bizarre fact that creationists have used Chomsky’s ideas as “evidence” against evolutionary theory. PP&U search for biological counterparts that could be considered al least quasi-perfect and introduce the term “modular improvement”. One example given of modular improvement is the cardiovascular system of vertebrates considered as a fractal space filling network of branching tubes, under the assumption that the energy dissipated by this transportation system is minimized. It is an inspired account, in my opinion the one who consider the evolution of an entire mechanism (specifically chomskian Narrow Syntax) which 10 Stable insertion of transposons, that evolve new coding and/or regulatory functions, sometimes occurs, with dramatic evolutionary consequences (Karp, 2004). 21 establishes one or more interfaces to be most likely epigenetic in nature: viral interactions provide the right level of complexity11. Furthermore, complex co-evolutions between viruses and hosts are known to have happened, with structural changes in the host, which in addition get transmitted to its offspring. For example the expression of captured genes encoding immuno-regulatory proteins is supposed to be one of the mechanisms used by viruses in their interaction with the host’s immune system (cf. Hsu et al. 1990 cited in PP&U, 2004). Along the lines of this biological substrate, PP&U proposal develops: “Suppose that, at some point, humans only had some primitive formal, perhaps a form of proto-language in the sense of Bickerton (1990) or maybe even a system unrelated to symbolic communication. Narrow Syntax in the sense that concerns most syntacticians would not have arisen yet. Then a major mind/brain reorganization would have taken place, which one hopes the detection of the morphological virus may be related to. The technical question is: supposing we have an organized elementary syntactic structure, and furthermore an alien element which in some sense does not belong, what can the host do in order to eliminate it? First of all, it must detect the intruder. This is no trivial task in a set of mechanisms which, by all accounts, has virtually no holistic characteristics. One possibility is for the host to detect the intruder on the basis of not being able to integrate it semantically. Next, there has to be some sort of “immune response”, whereby the intruder is somehow eliminated. The issue here is “who” eliminates the virus, and “how”. One must bear in mind that all of this has to be done with systemic resources. One of the few simple ways that a set of mechanisms of the assumed complexity would have of proceeding with the immunization task would be to match the virus element in categorial type”. This is a kind of presupposed structure (nonterminal symbols, phrasal nodes) in phrase-structure grammars. It is as if a morphological “antigen” were detected and eliminated by a syntactic “antibody”. As to how the elimination proceeds, one has to allow the set of mechanisms the ability to delete the virus matched by the antibody, under a strong version of the match: full categorial identity. In turn, if the host behaves as immune systems do, it should keep a memory of the process (after a single exposure to a virus, immune cells memorize the intruder and provide resistance for life)” (from Piattelli Palmarini and Urigereka 2004: 361-362). 11 Viruses are exquisitely species and tissue specific, they code for structural proteins and can infect an entire population, and importantly for our purposes, unlike bacteria or other parasites, they can integrate into a genome. Unlike maliciously built computational viruses, biological viruses don’t have a purpose, thus may a priori result in a variety of consequences for an organism. Granted, the normal result is a more or less major disruption of the organism’s functions or structure, due to the rapid multiplication of the infecting virus at the expenses of the host’s own machinery, but this is not inevitable, and in principle a virus may sometimes be integrated stably, and inheritably, into the genome of its host (cf. Karp, 2004). 22 Presumably, then - if I understood correctly the assumptions made by the two authors - in the presence of a detected virus v of category X, the host will systematically respond with matching antibody category X, and the elimination of v under complete featural identity with the particular categorial values that X happens to exhibit or, otherwise the relevant host (derivation) would die (terminate). I report below an example of a viral linguistic derivation taken by PP&U (2004: 363): (1-1) a. [ _ [ T-agr seem [ [Jack] [to be …]]]]12 target source b. Virus () detection: [ T-[agr] [seem [ [Jack] [to be …]]]] c. Search for categorical match [ T-[agrD] [seem [ [DPJack] [to be …]]]] with antibody () d. Eliminate virus under [ T-[agrD] [seem [ [DPJack] [to be …]]]] categorical identity D value = DP values e. Systematize the sequence < b; c; d> as typical of the language As a result of the transformational process, there is a demonstrable sense in which the formal object is more complex than it was prior to the “immunization”, in that the basic “tree” relations are warped. This can be illustrated as in (2-1), taken from PP&U (2004: 363): (2-1) 12 In this instance the crucial feature in the target (of movement) are agreement features in Tense (T), and the source of the movement is Jack, which can appropriately check those uninterpretable features in terms of its own interpretable ones. In the process, the source element becomes accessible to the computation by way of Case valuation, which the target renders (Chomsky, 1995). 23 So, it is possible to argue that the warped object resulting from associating the antibody DP to the T with a viral feature - which happens to be of the Determiner sort (for example agreement in person/number) - creates new local relations (Collins, 1997). In particular, the viral antigen-antibody relation establishes a “chain”. For instance, TP1 in (2-1) (the mother of the Jack node) establishes the context for the lower link in the chain, while the T’2 in the example (the mother of the T2 node hosting the antigen) establishes the context for the higher link in the chain. The chain linking the two relevant sites for the immunization is {{Jack, T’2 }, {Jack, TP1}}, or {T’2, TP1} factoring out Jack. For PP&U a syntactic chain is analogous to secondary structuring in nucleic acids, that is, the establishment (through something like pseudo-knots, which are pairs of stem-loop elements in which part of one stem resides within the loop of the other) of relations between bases further apart in the linear sequence: relations other than the most elementary pairings which primary structure yields. Just as RNA secondary structures have numerous consequences (through the ability of information sharing of a sort which, without the pseudo-knot, would be too long-distance to be viable) so too chains have consequences (see paragraph 1.3 and fig. 1-1). Thus, in the PP&U view, the complex object in (2-1) is simply a rearrangement of more elementary lexical features, nothing more holistic than that. The above result is definitely topological: “After the immunization takes place, a (new) linguistic topology emerges, and the result lends itself to otherwise impossible interpretations” (Piattelli Palmarini and Uriagereka, 2004: 365). I realize that PP&U conjecture is essentially metaphorical. Nonetheless, I think this metaphor is productive and worth pursuing to its several interesting consequences. Indeed, there are reasons (hints) to believe that it may be more than just a metaphor. The next two paragraphs should be explicatory. 1.3 The language of genes (David B. Searls, 2002) In the 1980s, several researchers began to follow various threads of Chomsky’s legacy in applying linguistic methods to molecular biology. The 24 main results included the fundamental observation that formal representations could be applied to biological sequences — the extension of linguistic formalisms in new, biologically inspired directions — and the demonstration of “the utility of grammars in capturing not only informational but also structural aspects of macromolecules” (cf. Searls, 1992). David B. Searls in an interesting article published on Nature in 2002 demonstrated in which way is possible to apply linguistics to the structure of nucleic acids. In brief, Searls showed that a folded RNA secondary structure entails pairing between nucleotide bases that are at a distance from each other in the primary sequence, establishing (in some way) those relationships that in linguistics are called dependencies. The most basic secondary-structure element is the stem-loop13, in which the stem creates a succession of nested dependencies that can be captured in an idealized form by the following context-free base-pairing grammar, as reported in (Searls 2002: 211) (see Chapter two for a survey on Chomsky’s hierarchy). (3-1) [SgSc; ScSg; SaSu; SuSa; Sε ] (Where the ε in the last rule indicates that an S is simply erased.) For Searls this grammar affords any and every derivation of “hairpin” sequences of a form such as the following: (4-1) SgScgaSucgauSaucgaucgaSucgaucgaucgaucgauc Derivations from this grammar grow outward from the central S, creating the nested dependencies of the stem (Fig. 1-1a), analogous to such phenomena as nested relative clauses in natural language. We have to say that in a realistic stem-loop, the derivation would terminate in an unpaired loop of at least several bases and might also contain, for example, non-Watson–Crick base pairs. But such features are easily added to the grammar without affecting the fundamental result that any language consisting of RNA sequences that fold into these basic structures requires context-free expression. 13 The structure is also known as a hairpin or hairpin loop. It occurs when two regions of the same molecule, usually palindromic in nucleotide sequence, base-pair to form a double helix that ends in an unpaired loop. The resulting lollipop-shaped structure is a key building block of many RNA secondary structures (cf. Watson JD, Baker TA, Bell SP, Gann A, Levine M, Losick R. (2004). Molecular Biology of the Gene. 5th ed. CSHL Press) 25 Then Searls note that, in addition to stem-loop structures, arbitrarily branched folded structures may be captured by simply adding to the grammar above a rule SSS, whose application creates bifurcations in the derivation tree (Fig. 1-1b). The base-pairing dependencies remain non-crossing, although more complicated. The resulting grammar is formally ambiguous, meaning that there are guaranteed to be sequences in the language for which more than one derivation tree is possible. Thus, the string gaucgaucgauc can be derived as a single hairpin or as a branched structure (Fig. 1-1a,b). This linguistic property of ambiguity, reflected in natural languages in sentences that can be syntactically parsed in more than one way (for example, “She saw the man with the telescope”), directly models the biological phenomenon of alternative secondary structure. Furthermore, finding that the language of RNA is at least context-free has mathematical and computational consequences, for example, for the nature and inherent performance bounds of any algorithm dealing with secondary structure14 (Knudsen and Hein, 1999). These consequences show the importance of characterizing linguistic domains in the common terminology and methodology of formal language theory, so as to connect them immediately to the wealth of tools and understanding already available15. In the light of these practical consequences of linguistic complexity16 (cf. Shieber, 1985), a significant finding is that there exist phenomena in RNA that in fact raise the language even beyond context-free. The most obvious of these are so-called non-orthodox secondary structures, such as pseudoknots (Fig. 1-1c). This configuration induces crossserial dependencies in the resulting base pairings, requiring context-sensitive expression Searls argue that, given this further promotion in the Chomsky hierarchy, 14 For instance, the fast, regular expression search tools used commonly in bioinformatics (such as those in the popular Perl scripting language) are ruled out, as in their standard form they specify only regular languages (Knudsen and Hein, 1999). 15 For this reason, from early nineties bioinformatics textbooks have devoted whole chapters to the relationship of biological sequences to the Chomsky’s hierarchy. 16 Natural languages seem to be beyond context-free as well, based on linguistic phenomena entailing cross-serial dependencies, although in both domains such phenomena seem to be less common than nested dependencies. Thus, by one measure at least, nucleic acids may be said to be at about the same level of linguistic complexity as natural human languages. For some observations on this issue see Shieber, S. 1985. “Evidence against the context-freeness of natural language”. Linguist. Phil. 8, 333–343, 32 On top of all the more specific conceptual fields (arrays of nine concepts), stands a universal field, which contains those qualities of God that are at the origin of all further entities and their concepts. This semantic system has an ontological and metaphysical foundation in the tradition of Aristotelian and medieval logic (figure 2-1). Fig. 2-1. Lullus’ Ars Magna The idea that concepts/words form linear arrays, that the extremes may be linked together, and that a hierarchy of such arrays exists, is a sort of first realization of “field-semantics” (cf. Wildgen, 1998). Anyway, the interesting thing here, is that Lullus did not stop at the static idea of a (circular) field of concepts: he proposed a combinatory mechanism which may have been motivated by the “machinery” of medieval syllogistics, but which contained a new mathematical impulse which allowed the later development of computing machines by Leibniz, Pascal, and others (cf. Wildgen, 1994). Leibniz gave to the Lullus’ idea the name Ars Combinatoria, by which it is now often known. It’s interesting to notice that some computer scientists have adopted Lullus as a sort of founding father, claiming that his system of logic was the beginning of information science. This method was an early attempt to use logical means to produce knowledge. Lullus hoped to show that Christian doctrines could be obtained artificially from a fixed set of preliminary ideas. For example, one of the tables listed the attributes of God: goodness, greatness, eternity, power, wisdom, 33 will, virtue, truth and glory. Lullus knew that all believers in the monotheistic religions - whether Jews, Muslims or Christians - would agree with these attributes, giving him a firm platform from which to argue. A hierarchy of linear (and circular) fields, and a combinatorial dynamics on it, already constitute a powerful theoretical instrument for organizing the universe of concepts, and represent a sort of gate to the lexicon24. In the late sixteenth century, Giordano Bruno (1548-1600) began to elaborate the Lullian system, approaching to a new system of conceptual organization based on the analogy between the universe and the mind. He replaced Lullus’s closed linear field with a regular, bidimensional pattern extending to infinity. Bruno’s view of the fundamental workings of language based on his studies of ancient languages, approaches that of contemporary semiotics: “Images do not receive their names from the explanations of the things they signify, but rather from the condition of those things that do the signifying25”. A view which was repeated by Fernand de Saussure nearly 300 years later in the 1890’s with his idea that a word is composed of two parts: the “signified” and the “signifier”, and is at the foundation of contemporary linguistics. From a mathematical point of view, Lullus’s field is a circular segment divided into nine sub-segments. Wolgang Wildgen notes that: “If instead of linear segments the basic unit is a regular surface, we may consider either the filling of an (infinite) area by circular surfaces (spheres), or its filling by regular surfaces (polygons) or bodies (polyhedra). The corresponding mathematical problem is that of an (optimal) package of circles/spheres or polygons/polyhedra” (Wildgen 1998: 216). Bruno decided exactly that the filling of a surface by squares is most adequate: his semantic universe is constructed on the basis of a square grid. Note that Bruno’s system as a sort of internal dynamic that concerns what now we call “the fillers of the memory pattern26”. Every word in a pattern may be replaced by its metaphor or its metonym. Thus, for Bruno, a text 24 Studies upon Lullus of every variety appear now throughout the world, the majority of which are reviewed in the bibliographical bulletin of the journal Studia Lulliana, available online. 25 Giordano Bruno, On the Composition of Images, Signs & Ideas, (p. 31), ed. by Higgins, Dick New York: Willis, Locker & Owens, on-line available. 26 Cf. McElree, Brian, Foraker Stephani and Dyer Lisbeth. 2003. “Memory structures that subserve sentence comprehension” in Journal of Memory and Language 48, 1 : 67-91. 34 which was first generated along the lines of basic meanings associated with the terms filled into the system can produce a totally different text or text interpretation by using a set of metaphorical and metonymical processes. Some authors think that the consequences of Bruno’s parallel work on cosmology and artificial memory are a new model of semantic fields which was so radical in its time that the first followers (although ignorant of this tradition) are the Von-Neumann automata (cf. Sporns and Alexander, 2002) and the neural net systems of the late 1980s (cf. Wildgen 1998: 237). The tradition of Lullus and Bruno was still alive when Leibniz (1646-1716) designed his “De Synthesi et Analysi universali seu Arte inveniendi et judicandi”. Leibniz’ solution is an arithmetic one and can be interpreted as a precursor of feature-semantics27. It associates cognitively primitive (i.e. non-definable) concepts with prime numbers. All definable concepts correspond to nonprime numbers, which can be decomposed into prime numbers. Leibniz eliminates (as Lullus did) the basic distinction between subject and predicate, and practically considers only two levels: primitive and (by definition) composite concepts. Then, Leibniz sketches a constructive device, which generalizes the methods of Euclid and applies them to conceptual systems28. The transition from the arithmetical to the geometrical characteristic corresponds to the transition between possible (conceivable) worlds, as pure intention, to the real world, to the spatialization and temporization of intentional concepts. Central notions are geometrical congruence and the intersection of geometrical figures29. 27 A rough presentation: a semantic feature is a notational method which can be used to express the existence or non-existence of semantic properties by using plus and minus signs. (i) Man is [+HUMAN], [+MALE], [+ADULT] Woman is [+HUMAN], [-MALE], [+ADULT] Boy is [+HUMAN], [+MALE], [-ADULT] Girl is [+HUMAN], [-MALE], [-ADULT] Intersecting semantic classes share the same features. Some features need not be specifically mentioned as their presence or absence is obvious from another feature: This is a redundancy rule. 28 For a detailed explanation see Rutherford, D., 1998. Leibniz and the Rational Order of Nature. Cambridge University Press. 29 Note that Leibniz was the first to use the term “analysis situs” later used in the 19th century to refer to what is now known as topology. There are two takes on this situation. On the one hand, Mates (1986: 240), citing a 1954 paper in German by Freudenthal, argues: "Although for [Leibniz] the situs of a sequence of points is completely determined by the distance between them and is altered if those distances are altered, his admirer Euler, in the 35 In brief, Leibniz demonstrates only how the simplest notions like space, point, line, plane, circle, position in space are constructed. This type of conceptual characteristic has the merit that all entities defined can be constructed and Leibniz imagines how his system if elaborated can be used to describe plants and animals and to invent machines. The geometrical characteristic would allow man to do this with symbolic techniques in his imagination without the help of concrete figures and models30. 1.5.2 Gestalt theory and the topological psychology of Kurt Lewin At the origin of Gestalt theory stands philosophy and psychology, which were not yet institutionally separated in Germany in early 1900s. In the various schools of Gestalt psychology (Berlin, Graz, Leipzig) different aspects were foregrounded: psychophysiological aspects in Berlin (e.g., Wertheimer, Koffka, Köhler, Lewin), intellectual forces as Gestalt-foundation in Graz (e.g., famous 1736 paper solving the Königsberg Bridge Problem and its generalizations, used the term geometria situs in such a sense that the situs remains unchanged under topological deformations. He mistakenly credits Leibniz with originating this concept. ...it is sometimes not realized that Leibniz used the term in an entirely different sense and hence can hardly be considered the founder of that part of mathematics." But Hirano (1997) argues differently, quoting Mandelbrot (1977: 419): "...To sample Leibniz’ scientific works is a sobering experience. Next to calculus, and to other thoughts that have been carried out to completion, the number and variety of premonitory thrusts is overwhelming. We saw examples in 'packing,'. A “topological” interest for Leibniz is further reinforced by finding that for one moment its hero attached importance to geometric scaling. In "Euclidis Prota"..., which is an attempt to tighten Euclid's axioms, he states,...: 'I have diverse definitions for the straight line. The straight line is a curve, any part of which is similar to the whole, and it alone has this property, not only among curves but among sets.' This claim can be proved today." Thus the fractal geometry promoted by Mandelbrot drew on Leibniz's notions of selfsimilarity and the principle of continuity: natura non facit saltus. We also see that when Leibniz wrote, in a metaphysical vein, that "the straight line is a curve, any part of which is similar to the whole..." he was anticipating topology by more than two centuries (cf. wikipedia.org-> Leibniz). 30 Leibniz' critique of image-like models can be generalized to all too specific and ad hoc pictorial descriptions. In an interesting paper that trace the linguistic influences of Leibniz, Wolfgang Wildgen, who says: “A strategy already condammed by Leibniz 300 years ago is systematically tried by cognitive semantics which work with ad hoc figures and with pictures which have no theoretical status. Semantics of this type will soon accumulate a chaotic universe of ad hoc figures and will loose the capacity to find general and stable regularities which is a central aim of any scientific enterprise. Thus Leibniz geometrical characteristic is a kind of deconstruction of cognitive semantics in the style of Lakoff”. (Widgen 1998; p. 224) 36 Meinong, Benussi), emotional and symbolic aspects in Leipzig (e.g., Cornelius, Bühler) (cf. Casati and Varzi, 1999). Gestalt theory, in the most general terms, is a theory of mind and brain that proposes that the operational principle of the brain is holistic, parallel, and analog, with self-organizing tendencies; or, that the whole is greater than the sum of its parts. The classic Gestalt example is a soap bubble, whose spherical shape is not defined by a rigid template, or a mathematical formula, but rather it emerges spontaneously by the parallel action of surface tension acting at all points in the surface simultaneously. A good summing up is the one made by Casati and Varzi: “Gestalt theorists emphasise the figure-ground articulation of perceived configurations. Some portions of a perceived scene, due to their intrinsic wholeness, are perceived as figure, and are delineated against a background which is perceived as completing itself behind the figure. Now, the visual boundary that separates figure from ground is "oriented": it belongs to the figure and not to the ground31”. For linguistics issues, the most prominent figure of Gestalt is probably Kurt Lewin (1890-1947). As early as 1912 Kurt Lewin foresaw that a scientific psychology would have to make use of “topology” and of “the dynamics which could be conceived in a topological structure” (cf. Lewin 1969: 9). Many of Lewin’s ideas recall principles of “force dynamics” worked out in greater sophistication in the linguistic sphere by Talmy (1988), and Talmy, along with Petitot (1989) and others, tried to demonstrate the importance of topology for the understanding of a variety of different sorts of linguistic structuring (cf. paragraph 1.5.4). As Talmy notes, the conceptual structuring effected by language is illustrated most easily in the case of prepositions. A preposition such as' in is magnitude neutral (in a thimble, in a volcano), shape neutral (in a well, in a trench), closure-neutral (in a bowl, in a ball); it is not however discontinuity neutral (in a bell-jar, in a bird cage). Work on verb-aspect and the mass-count distinction, too, has profited from a topological orientation in the paradigm of Cognitive Linguistics (cf. paragraph 1.5.4). Lewin’s central idea was that of a “psychological life-space” (psychologischer Lebensraum): life-space is constituted by the individual and a situation relevant for the individual at a given moment. The life-space of an individual has two aspects: every partial domain of an individual's life-space 31 From Casati, Roberto, and Varzi, Achille., 1999, Parts and Places: The Structures of Spatial Representation. Cambridge, Mass.: MIT Press. 37 corresponds to a “psycho-domain”, containing the person, structures of the life-space specifically relevant for the person (individual situations) and structures of the life-space which are constituted independently from the person (standard situations)32 and psychic “locomotion”, such as paths in the life-space with preferred routes, barriers and obstacles. So, language has a major function in the Lewin framework: it transfers internal states of the person to his/her environment. Language enables a type of indirect locomotion in psychic space: for example, surrogate locomotion as in the speech act of ordering or in social coordination via language, change of social and personal relations and cognitive influence, which creates new possibilities of psychic locomotion via learning. From a narrow psychological point of view, Lewin begins with the opposition thing (intuitively: a closed connected unity) and region (intuitively: a space within which things are free to move). As Lewin points out, what is a thing from one psychological perspective may be a region from another: “A hut in the mountain has the character of a thing as long as one is trying to reach it from a distance. As soon as one goes in, it serves as a region in which one can move about.” (Lewin, 1969: 116) He defines the notion of a boundary zone z between two disconnected but proximate regions m and n, as the region, foreign to m and n, which has to be crossed in passing from one to the other. The whole m + n + z is then connected in the topological sense. (Lewin 1936: 121) The concept of a barrier he defines as a boundary zone which offers resistance to passage of things between one region and another (Smith, 1994). 32 For Lewin, the development of a child or an adult may be described as a change in lifespace. The life-space of a child alters as soon as it learns to grasp, to control, to walk, to speak, and so on. A prisoner has a dramatically reduced life-space and some situations may contain attractors (cf. emotional attractors, sympathy, love, etc.) or repellers (situations of frustration, anger), which provoke reactions of escape. Thus far we have considered the person to be an integral component of a life-space. But persons may themselves be regarded as a topological field with an inner area (intrapersonal domain), a periphery of this domain, and a sensor-motorical domain, which lies between the person and his/her context (the situation). The topological psychology of Lewin was later elaborated by Fritz Heider in his “attributional psychology”. Heider strengthened the relation between 'life span' categories and semantic categories, for example, perceptual, experimental, affecting, causing, evaluation, part-whole relations (possessive), can, trying, wanting, etc. (cf. Heider, 1958). Thus, the psychic and the symbolic world have the structure of a field, and topological and dynamic (vectorial) notions from contemporary mathematics are used in Heider (1958) to specify these fields. 38 Such resistance may be asymmetric; thus it may be greater in one dimension than in the opposite direction. Barriers effect the degree of communication between one region and another, or in other words the degree of influence of the state of one region on that of another region. Hence the notion of degree of influence, too, need not be symmetric (Smith, 1994). For example, the fact that A is in a certain degree of communication with B does not imply that B is in equally close communication with A. In Lewin’s paradigm two regions A and B are said to be parts of a dynamically connected region if a change of state of A results in a change of state of B. The notion of dynamic connectedness, too, is by what was said earlier a matter of degree. In fact we can distinguish a hierarchy of degrees of inter-linkage between regions, and here Lewin echoes discussions in the Gestalt-theoretical literature of the notions of “strong” and “weak” Gestalten (1969: 173). A strong Gestalt may be defined as a complex with a high degree of dynamic connectedness between its parts (examples given: an organism, an electromagnetic field). A weak Gestalt, for example a chess-club, has a lesser but still non-zero dynamic connectedness between its parts, while a purely summative whole (an Und-Verbindung in Gestalt terminology) is such that its separate units manifest a zero degree of dynamic connectedness33 (Smith, 1994). We have introduced the basic concepts of Lewin’s topological psychology in a rather general way, abstaining from any specific applications to psychological matters. It will extend too much the scope of our discussion; it is interesting to notice that Lewin’s critics, however, rightly drew attention to a certain crucial shortfall in his use of mathematical notions in his writings. As was correctly pointed out by his critics, Lewin rarely makes the mathematical theory of notions such as connectedness, boundary, separateness, and so on, do any substantial work within the framework of his investigations34. Anyway certain aspects of Lewin’s generalizations of standard topology have since shown themselves to be highly fruitful if applied to a linguistic field. In my opinion, these generalizations include: 33 Interestingly, in the light of our discussions, the notions central to Gestalt theory can be defined not merely on the basis of the notion of dynamic connectedness, but also in terms of structure-preserving transformations. 34 This criticism was put forward in an influential article by London (1944), an article which did much to thwart the further development of topological psychology (or of a topologically founded cognitive science) as Lewin had conceived it. 39 i) the recognition that it is possible to construct topology on a non-atomistic, mereological (cf. Casati and Varzi, 1999) basis which works in terms of wholes (regions) as well as parts; ii) the systematic employment of the notion of asymmetric boundary, a notion which turns out to be crucial in many cognitive spheres (cf. Smith, 1994); iii) the employment of topological ideas and methods also in relation to finite domains of objects. 1.5.3 Husserl’s topology Another precursor of the idea of using topology as a foundation for human sciences (and, in nuce, cognitive sciences) is Edmund Husserl (1959-1938). Husserl’s Logical Investigations (1900-01) contain a formal theory of part, whole and dependences that is used by the author to provide a framework for the analysis of mind and language of just the sort that is presupposed in the idea of a topological foundation for cognitive science. The title of the third of Husserl’s Logical Investigations is “On the Theory of Wholes and Parts” and it divides into two chapters: “The Difference between Independent and Dependent Objects” and “Thoughts Towards a Theory of the Pure Forms of Wholes and Parts”. Husserl’s theory is concerned also with the horizontal relations between the different parts within a single whole, relations which serve to give unity or integrity to the wholes in question (cf. Smith, 1994). In other words, in Husserl’s theory some parts of a whole exist merely side by side, they can be destroyed or removed from the whole without detriment to the residue. A whole all of whose parts manifest exclusively such “side-bysideness” relations with each other is called a heap or aggregate or, more technically, a purely summative whole (Und-Verbindung). In many wholes, however, and one might say in all wholes manifesting any kind of unity, certain parts stand to each other in relations of what Husserl called necessary dependence (which is sometimes, but not always, necessary interdependence). Such parts, for example the individual instances of hue, saturation and brightness involved in a given instance of colour, cannot, as a matter of necessity, exist, except in association with their complementary parts in a whole of the given type (cf. Smith, 1994). There is a huge variety of such lateral dependence relations giving rise to 40 correspondingly huge variety of different types of whole which more standard approaches of “extensional mereology” (cf. Casati and Varzi, 1999) are unable to distinguish. The connection between part and whole on the one hand and dependence on the other may be seen in the fact that every whole can be regarded as being dependent on its own constituent parts. It is one not inconsiderable advantage of Husserl’s theory that it allows a precise formulation of these and a range of related theses within a single framework. Roman Jakobson (1966) applied Husserl’s ideas on parts, wholes and categories from the Logical Investigations in different branches of linguistics, in the early development of categorial grammar and of phonology, respectively. Thus, Jakobson’s account of distinctive features is as he himself admits an application of Husserl’s idea35. 1.5.4 Fillmore’s frames and the “nouvelle vague” of topological semantics Lullus’s relational concept is elaborated by the concept of valence. Charles Fillmore would later (well, nine centuries later) call these schemata for the organization of concepts into a unitary macro-concept: frames (cf. Fillmore, 1977). Fillmore has been extremely influential in the areas of syntax and lexical semantics; he was one of the founders of cognitive linguistics, and developed the theories of Case Grammar (1968), and Frame Semantics (1976). In all of his research he has illuminated the importance of semantics, and its role in motivating syntactic and morphological phenomena. The framework of Case Grammar is a system of linguistic analysis, focusing on the link between the valence of a verb and the grammatical context it requires, created by Fillmore in the context of early Transformational Grammar. This theory analyzes the surface syntactic structure of sentences by studying the combination of deep cases (i.e. thematic roles) -- Agent, Patient, Benefactor, Location or Instrument -- which are 35 The topological background of Husserl's work makes itself felt already in his theory of dependence. It comes to the fore above all however in his treatment of the notion of phenomenal fusion (a sort of non-discrete Merge?): the relation which holds between two adjacent parts of an extended totality when there is no qualitative discontinuity between the two. For example, adjacent squares on a chess-board array are not fused together in this sense; but if we imagine “a band of colour that is subject to a gradual transition from red through orange to yellow”, then each region of this band is fused with its immediately adjacent regions (Smith, 1994). 41 required by a specific verb. For instance, the verb "give" in English requires an Agent (A) and Patient (P), and a Beneficiary (B); e.g. “Gianni (A) gives money (P) to Maria (B)”. Obviously, the heritage of Fillmore is still alive in contemporary generative grammar researches. According to Fillmore, each verb selects a certain number of deep cases which form its case frame. Thus, a case frame describes important aspects of semantic valency, of verbs, adjectives and nouns. Case frames are subject to certain constraints, such as that a deep case can occur only once per sentence. Some of the cases are obligatory and others are optional. Obligatory cases may not be deleted, at the risk of producing ungrammatical sentences. A fundamental hypothesis of case grammar is that grammatical functions, such as subject or object, are determined by the deep, semantic valence of the verb, which finds its syntactic correlate in such grammatical categories as Subject and Object, and in grammatical cases such as Nominative, Accusative, etc. Fillmore (1968) puts forwards the following hierarchy for an universal subject selection rule: (5-1) [Agent < Instrumental < Objective] That means that if the case frame of a verb contains an agent, this one is realized as the subject of an active sentence; otherwise, the deep case following the agent in the hierarchy (e.g. Instrumental) is promoted to subject. Frame semantics is a topological theory that relates linguistic semantics to encyclopaedic knowledge developed by Fillmore, and is, in many respects, a further development of his Case grammar (cf. also Eco, 1975). The basic idea is that one cannot understand the meaning of a single word without access to all the essential knowledge that relates to that word. For example, one would not be able to understand the word “sell” without knowing anything about the situation of commercial transfer, which also involves, among other things, a seller, a buyer, goods, money, the relation between the money and the goods, the relations between the seller and the goods and the money, the relation between the buyer and the goods and the money and so on. Thus, a word activates, or evokes, a frame of semantic knowledge relating to the specific concept it refers to (or highlights, in frame semantic terminology). A semantic frame is defined as a coherent structure of related concepts that are related such that without knowledge of all of them, one does not have complete knowledge of one of the either, and are in that sense 48 2.2.3 Connected and disconnected graphs A graph G is said to be connected if there is a path between any two distinct vertexes of G, else it is said to be disconnected. A multigraph is one in which there exists multiple edges between the same pair of vertexes. The multiplicity of the edge between the pair of vertexes is given by the number of multiple edges between the pair of vertexes under consideration. 2.2.4 Subgraphs A subgraph G1= (V1, E1) is derived from a directed or undirected graph G = (V, E) such that V1 and E1 are subsets of V and E and neither are null sets39. If V1=V then the subgraph is called a spanning subgraph. An induced subgraph has a set of vertexes (U) which is a subset of the set of vertexes (V) of the graph (G) and contains all the edges present in G that are incident with the chosen subset of vertexes (cf. fig 2-5). The constraint is that U ≠ 0. Such a subgraph of G is said to be induced by U (Gibbons, 2002). fig. 2-5 subgraph derivation 2.2.5 Complete graphs, isomorphism and degrees A complete graph (Kn) on V, a set of ‘n’ vertexes, is a loop-free undirected graph where for all vertexes x, y ∈ V, and x ≠ y, there is an edge. For a chosen value of ’n’, there exists only one complete graph. 39 It’s interesting to notice that instead of removing nodes and edges, one may add some nodes and edges to extend a graph such that the given graph is a subgraph of the extension. The addition of nodes causes no problem at all, whereas the addition of edges requires the specification of their labels, sources, and targets, where the latter two may be given or new nodes (cf. Kreowski et al. 2006). . 49 Fig. 2-6 complete graph When we have two undirected graphs G1= (V1, E1) and G2= (V2, E2), a function ƒ: V1 → V2 is called a graph isomorphism if ƒ is one-to-one and onto and for all vertexes x, y ∈ V1 and edges {x, y} ∈ E1 if and only if {ƒ (x), ƒ (y)} ∈ E2. So, G1 and G2 are called isomorphic graphs. fig. 2-7 these graphs are isomorphic, Indeed, the required isomorphism is given by v1 -> 1; v2->2; v3->3; v4->4; v5->5 The degree of any vertex ‘v’, deg (v), is the number of edges in the graph G that are incident with v. A loop at any vertex is counted as two edges. A degree of 1 for any vertex makes it a pendant vertex. An undirected graph (or multigraph40) with the same degree for each vertex is called a regular graph. The incoming, or in, degree of v is the number of edges in G that are incident into v. It is denoted by id (v). The outing, or out, degree of v is the number of edges in G that are incident from v. It is denoted by od (v). A loop at a vertex contributes a count of one to both in-degree and out-degree. 2.2.6 Planar and bipartite graphs The graph G is called a planar graph if G can be drawn in the plane of a paper with its edges intersecting only at vertexes of G; else it is called a 40 So, the Euler circuit in an undirected graph (or multigraph), G = (V, E), with no isolated vertexes, is one that traverses every edge of the graph exactly once. Since the Euler circuit traverses every edge of the undirected graph exactly once, it is not possible to have two distinct Euler circuits for the same graph. An open trail that traverses each edge only once is called an Euler trail (cf. Gibbons, 2002) 50 nonplanar graph. A graph G = (V, E) is said to be bipartite, if V = V1 U V2 with V1∩V2 = 0, and every edge of G is of the form {x, y} with x ∈ V1 and y ∈ V2. If each vertex in V1 is connected to every vertex in V2, then it is a complete bipartite graph. If |V1| =m and |V2|=n then the graph is denoted by Km, n. fig. 2-8 A complete bipartite graph 2.2.7 Elementary subdivisions In G = (V, E) a loop-free undirected graph, the operation of removing an edge e = {a, c} and adding the edges {a, b}, {b, c} to G-e, where b ∈ V is called elementary subdivision. The loop-free undirected graphs G1= (V1, E1) and G2= (V2, E2) are called homeomorphic if they are isomorphic or if both of them can be obtained from the same loop-free undirected graph by a sequence of elementary subdivisions41. 2.2.8 Hamilton Graphs The graph G = (V, E) with |V| ≥ 3, has a Hamilton cycle if there is a cycle in G that contains every vertex in V. A Hamilton path is a path in G that contains each vertex42. 41 The Kurotowski’s theorem states that a graph is nonplanar if and only if it contains a subgraph that is homeomorphic to either K5 or K3, 3 (Cf. Grimaldi, 2005). 42 In a famous topographic problem, the travelling salesman problem a person is supposed to visit each town in his district, and this he should do in such a way that saves time and money. Obviously, he should plan the travel so as to visit each town once, and so that the overall flight time is as short as possible. In terms of graphs, he is looking for a minimum weighted Hamilton cycle of a graph, the vertices of which are the towns and the weights on the edges are the flight times. it is widely believed that no practical algorithm exists (cf. Diestel, 2005). 51 fig. 2-9 An Hamilton cycle 2.2.9 Colouring a graph A proper colouring of an undirected graph G= (V, E) occurs when each of the vertexes of G are coloured so that if {x, y} is an edge in g then x and y are coloured with different colours, that is adjacent vertexes have different colours. The minimum number of colours needed to properly colour G is called the chromatic number of G (cf. footnote 37). 2.3 Trees Let’s now - given some basic principles and a terminology - try to define a (syntactic) tree into graph theoretical terms. In general terms, a tree is useful as a pictorial representation of structure. As a device for representing structure, it is applicable to any situation where a hierarchy of choices is made. A derivation is a perfect example of a hierarchy of choices, and as such a tree is ideal as a visual representation of a derivation. To construct a derivation tree, we start with a tree containing only the root node. For each step of the derivation, the tree is correspondingly extended. That is, every time a production is used to replace a non-terminal in the current sentential form by the string on the right hand side of the production, lines are drawn from the corresponding non-terminal in the tree to each symbol in the replacement string (cf. Harju, 2005). At each stage in the construction of the tree, reading from left to right, the leaf nodes will be the current sentential form. In formal language theory, and specifically in generative linguistics, tree diagrams are much used, and a tree diagram is an acyclic, connected graph. Here’s a definition: (3-2) A loop-free, undirected, connected graph G= (V, E) which has no cycles is called a tree and is denoted by T= (V, E). 52 fig. 2-10 A tree of operations for the arithmetic formula x(y+z)+y The spanning subgraph, for a connected graph, which is also a tree is the spanning tree43. It provides minimal connectivity between all vertexes of the graph. There are two algorithms, according to Diestel (2005) to search for the spanning tree in a graph: i) Depth First Search algorithm44 ii) Breadth First Search algorithm45. If G is a directed graph and the undirected graph associated with G is a tree, then G is a directed tree. It is a rooted tree if there is a unique vertex ‘r’ (root) such that the in-degree of r, id(r) =0 and for all other vertexes (v) the indegree, id (v) =1. The vertex with out-degree, out (v) =0 is the leaf or terminal vertex. All other vertexes are branch nodes or internal vertexes. The subtree at 43 Spanning trees are often optimal solutions to problems, where cost is the criterion. We may also wish to construct graphs that are as simple as possible, but where two vertices are always connected by at least two independent paths. These problems occur especially in different aspects of fault tolerance and reliability of networks, where one has to make sure that a breakdownof one connection does not affect the functionality of the network. Similarly, in a reliable network we require that a break-down of a node (computer) should not result in the inactivity of the whole network (cf. Harju, 2006) 44 Intuitively, in a deep-first search one starts at the root (selecting some node as the root in the graph case) and explores as far as possible along each branch before backtracking. Formally, a deep-first search DFS is an uninformed search that progresses by expanding the first child node of the search tree that appears and thus going deeper and deeper until a goal node is found, or until it hits a node that has no children. Then the search backtracks, returning to the most recent node it hadn't finished exploring. Cf T.H. Cormen, C.E. Leiserson And R.L. Rivest, Introduction to Algorithms, MIT Press, 1993. 45 Breadth-first search algorithm begins at the root node and explores all the neighbouring nodes. Then for each of those nearest nodes, it explores their unexplored neighbour nodes, and so on, until it finds the goal. Cf T.H. Cormen, C.E. Leiserson And R.L. Rivest, Introduction to Algorithms, MIT Press, 1993. 53 any vertex is the subgraph induced by that vertex as root and all of its decedents. If the edges or branches of the rooted tree are ordered then it is called an ordered rooted tree. Given m ∈ Z+, a rooted tree T = (V, E) is an m-ary tree if for all vertexes the od (v) ≤ m. When m =2 it is a binary tree. If od (v) = 0 or m for all v ∈ V then T is called a complete m-ary tree. For m=2, it is a complete binary tree and can be used to represent binary operations. In a complete mary tree each internal vertex has exactly m-children. If T is a complete binary tree of height h and all the leaves in T are at level h then, T is called a full binary tree. The level number is the number of paths from the root vertex to the vertex under consideration. If h is the largest level number achieved by a leaf of T= (V, E), then, T is said to have a height h. A rooted tree of height h is balanced if the level number of every leaf in T is (h-1) or h. A vertex v in a loop-free undirected graph G = (V, E) is called an articulation point if the subgraph G-v has more components than the given graph G. The removal of the articulation points disconnects the graph. In terms of communications systems and networks the articulation points indicate locations where the system is most vulnerable. A loop-free connected undirected graph without any articulation points is called biconnected and is said to be a “nonseparable graph”. A graph with weights assigned to its edges is called a weighted subgraph. Weights are positive real numbers attached to the edges. The might signify parameters such as cost, time, length etc. on each edge considered as a link. If x, y ∈ V, but (x, y) ∉ E then wt(x, y) = ∞. A set P of binary sequences (representing a set of symbols) is a prefix code if no sequence in P is the prefix of any other sequence in P. When a sequence of binary characters (0 and 1) are used to represent symbols, then, care should be taken that the binary code for one symbol does not form a prefix to the code of another symbol, else decoding will be difficult and might be incorrect. 2.3.1 The Shortest path and the Minimal spanning tree problems In our discussion, I think that is relevant to introduce this two problems and their solutions. The shortest path problem arises whenever there is a need to determine the shortest, cheapest, or most reliable path between one or many pairs of nodes in a network. The minimal spanning tree problem occurs when we need to design the simplest network (a spanning tree) that will connect topologically dispersed 54 system components so that they can communicate with each other and the total construction cost is minimized. The shortest path problem differs from the minimum spanning tree problem in that, the former is used to find the cheapest path between some nodes, while the latter is to find the cheapest tree that connects every node (cf. Diestel, 2005). Many of the most salient core ingredients of network techniques are captured by the shortest path problem. It is often encountered in the transportation and communication networks. 2.3.1.1 Dijkstra’ Shortest Path algorithm In a given weighted directed graph G= (V, E) the shortest path between any two vertexes is that path in which the sum of weights of all the constituent edges is the minimum. The main idea of the algorithm is to change the temporary labels associated with vertexes into permanent ones. The permanent label of a vertex denotes the shortest path weight from the source vertex to the current vertex. The source vertex is given a permanent label (0,-) and each of the other vertexes is given a temporary label (∞,-). Let P and P’ (= V-P) be the sets containing vertexes with permanent and temporary labels respectively. At each step, the algorithm chooses the vertex x ∈ P’ with the minimum temporary label, and makes it permanent, records its predecessor’s index and updates the temporary values of all the vertexes. It repeats this procedure till all nodes get permanent labels (cf. Harju, 2005). 2.3.1.2 Kruskal’s Minimal Spanning tree algorithm The main idea of this algorithm is to add at each step the cheapest edge (lowest weight) from the set of remaining edges. The subgraph developed at each step should not contain any cycles. Initially the counter i is set to 1 and an edge e1 in G= (V, E) is selected with the lowest weight. If edges e1, e2…ei have been selected for 1≤ i ≤n-2 where n = |V|, then, ei+1 is selected such that it has the lowest weight among the remaining edges and the subgraph being formed does not contain any cycles. Subsequently, i is replaced by i+1.The procedure is repeated till i< n-1.If i = n -1 then the subgraph of G determined bys the edges e1, e2…en-1 is connected with n vertexes and n-1 edges and is the minimal spanning tree for G (cf. Harju, 2006). 55 2.3.1.3 Prim’s Algorithm This gives the “optimal tree” for a graph. Initially the counter i is set to 1 and an vertex v1 ∈ V is placed in set P. Now N and T are defined as N = V-{v1} and T =0. For 1≤ i ≤n-2 where n = |V|, P is the set of vertexes and T is the set of edges and N = V-P. The edge of minimal weight in G that connects a vertex x in P to a vertex y (= vi +1) in N is added to the set T. Now, y is placed in set P and deleted from N. The counter is incremented by 1. The process is repeated till i=n. If i = n then the subgraph of G determined by the edges e1, e2…en-1 is connected with n vertexes and n-1 edges and is the optimal spanning tree for G. (cf. Diestel, 2005). 2.4 Graph transformations Graph transformation is a rule-based method that performs local changes on graphs (Diestel, 2005). With graph transformation rules it is possible to specify formally and visually for instance the semantics of rule-based systems (like the semantics of functional languages), specific graph languages, graph algorithms (like the search of all Eulerian cycles in a graph), and many more. The idea of a graph transformation rule is to express which part of a graph is to be replaced by another graph. Unlike strings, a subgraph to be replaced can be linked in many ways (i.e., by many edges) with the surrounding graph. Consequently, a rule also has to specify which kind of links are allowed; this is done with the help of a third graph that is common to the replaced and the replacing graph and requires that the surrounding graph may be linked to the replaced graph only with edges incident to this third graph (cf. Kreowski et al. 2006). Here is a formal definition of a rule of graph transformation, taken from Kreowski et al. (2006: 6): (4-2) A rule r = (L ⊇ K ⊇ R) consists of three graphs L; K; R such that K is a subgraph of L and R. The components L, K, and R of r are called left-hand side, gluing graph, and right-hand side, respectively. Here I show, following Kreowski et al. (2006), a simple example, concerning shortest paths. The figure below shows the two essential rules for the computation of shortest paths in distance graphs, that is graphs labeled with non-negative integers. The first rule adds (radd) a direct connection 56 between each two nodes that are connected by a path of length 2 and sums the distances up. Using this rule, one can compute the transitive closure of the given distance graph. If one applies the second rule (rselect), which chooses the shortest connection of two direct connections as long as possible, one ends up with shortest connection between each two nodes. Fig. 2-11. Graph transformation rules for the computation of shortest paths, taken from Kreowski et al. (2006; p. 7) 2.5 Graph Languages Analogously to Chomsky grammars (see paragraph 2.7) in formal language theory, graph transformation can be used to generate graph languages. A graph grammar consists of a set of rules, a start graph, and a terminal expression fixing the set of terminal graphs (cf. Ehrig et al. 1999). Such a terminal expression may consist of a set Δ ⊆ ∑ of terminal labels admitting all graphs that are labeled over Δ. Here is formal definition of a graph grammar: (5-2) A graph grammar is a system [GG] = (S; P; Δ) where S is the initial graph of GG, P is a finite set of graph transformation rules, and Δ ⊆ ∑ is a set of terminal symbols. The generated language of GG consists of all graphs G that are labeled over Δ and that are derivable from the initial graph S via successive application of the rules in P (Kreowski et al. 2006: 14). I give an example of a grammar generating connected graphs (see paragraph 2.2.). Consider connected = {°; P; *} where the start graph consists of a single node 57 and the terminal expression allows all graphs labeled only with * . Note that the symbol * denotes a special label in ∑ standing for unlabeled and being invisible in displayed graphs (cf. Gibbons, 2002). The rules in P = {p1; p2; p3} are depicted in Figure 2-12. The rule p1 adds a node v and an edge e such that v is the target of e, and takes as source of e an already existing node. The rule p2 is similar, the only difference being that the direction of the new edge e is inverted. The third rule p3 generates a new edge between two existing nodes. The new edge can also be a loop if the two nodes in the left-hand side of p3 are identified, for example, if they are one and the same node in the match of the left-hand side. It can be shown that the generated language of connected, L(connected), consists of all non-empty connected unlabeled graphs. Fig. 2-12. Rules for the generation of a (connected) graph 2.6 Formal grammars and linguistic theory I give in this section a rough presentation of Chomsky Grammars, before applying a translation of Chomsky Grammars into Graph Grammar (see above), in the final paragraph of this chapter concerning graph theory’s formalisms. In linguistics a formal grammar is a precise description of a formal language — that is, of a set of strings over some alphabet. The two main categories of formal grammars are generative grammars, which describe how to write strings that belong to a given language (generate), and analytic grammars, which describe how to recognize when strings are members in the language 64 diagram has been assumed (cf. Haegeman, 1996) to have properties such as: (a) There is one vertex, called the root vertex, that is dominated by no vertexes and from which there is a path to every vertex, where a path is any linear subset of a tree (cf. Chapter 2). (b) Every vertex other than the root has exactly one vertex that immediately dominates it. (c) The vertexes that each singular vertex immediately dominates are ordered from the left (or better, temporally); as a corollary of (c): (c1) says that a well-formed sentence needs to constitute a (single) connected graph with one special vertex as its root. (c2) need not (or should not be retained within the minimalist program), where linear order is assumed to play no significant syntactic role. The postulates (c1) and (c2) has been adopted in virtually every theory of phrase structure, and it specifically excludes a diagram (cf. Yasui, 2004a) with a closed route such as (3-3): (3-3) 1 2 3 4 5 6 The problematic node is clearly 6, which is dominated by two nodes, 4 and 5; thus, the structure given in (3-3) violates what assumed in (c2). Whether (c2) should be assumed or not depends merely on other assumptions on phrase structure. I try to show in the present contribution that a graph-based analysis of some of the fundamental assumptions in the minimalist program leads to the rejection of (c2) (or the growth of a linear model). In the syntactic graphs, to be proposed below, in fact the output of external MERGE is a tree with the property (c2) but that of internal MERGE (or movement) is a graph that has a “closed route”47. 47 This distinction could offer a natural explanation for the parametric difference in whmovement; the PF requirement of linearizing lexical items forces a graph with a closed route to be changed into a tree in either of the two possible ways, which correspond to ouvert and covert movement. 65 3.2 Graphs and Bare phrase structure One important simplification of phrase structure proposed since Chomsky (1994, 1995: Chapter 5) is the elimination of category and projection labels by the extensive use of lexical items themselves, which is motivated by the Inclusiveness Condition48. To meet the Inclusiveness Condition, (4-3) is to be assumed instead of (1-3) (=(2-3) in our revision from a graph theoretic point of view). (4-3) will it will will be be snowing V = {it, will, will, will, be, be, snowing} E={<will, it>, <will, will>, <will, will>, <will, be>, <be, be>, <be, snowing>} The structure given in (4-3) contains nodes with the same labels will and be. If nodes with the same label are to be identified as one (cf. Yasui, 2004a), the set V in (4) is non-distinct from {it, will, be, snowing}. Then, (4-3) is forced to be replaced by (5-3): (5-3) will it be snowing V={it, will, be, snowing} E={<will, it>, <will, will>, <will, will>, <will, be>, <be, be>, <be, snowing>} 48 The Inclusiveness Condition says that the output of a system does not contain anything beyond its input. It was first proposed in Chomsky (1995: 225) as a condition met by the computational system of human language, and taken to imply that the interface levels contain nothing more than arrangements of lexical features. In other words, a language which meets the inclusiveness condition cannot contain traces or indices left after movement. 66 The structure in (5-3) contains three loops (cf. paragraph 2.2), <will, will>, <will, will> and <be, be>, which express nothing other than intermediate projections (so they can be rejected; erased). Therefore, (6-3) is to be finally assumed here: (6-3) will it be snowing V={it, will, be, snowing} E={<will, it>, <will, be>, <be, snowing>} The structure in (6-3) might not look like a syntactic tree, but the set of nodes V is essentially the enumeration in the sense of Chomsky (1995: 225227), and E in (6-3) seems to meet the Inclusiveness Condition just because no projection and category labels are added (a similar conclusion is found in Yasui, 2004a). Furthermore, the order pair <α, β> is generally defined as {{α}, {α, β}}, and it looks quite close to Chomsky's (1995:244-245) definition of the object formed from α and β of the type α: {α, {α,β}}. If {α, β} is adopted instead of {α, {α,β}} as the definition of the object formed from α and β, the discrepancy between the formal definition and its graphical representation can disappear, which seems to be an interesting result (cf. Collins, 2002). In this graph theory based account for syntactic structure, it is evident that my primary influence (my debt) is the paper of Collins (2002), that is a sort of quest for a label free syntax. 3.3 External and Internal Merge First, let’s consider external MERGE, which is applied to two substructures α and β and produces a larger structure only if some syntactic relation holds between α and β (cf. Starke, 2001). To paraphrase it in the terms of a linguistic graph theory, MERGE is a kind of “linking operation” over two graphs49, with 49 According to the standard conception of syntactic structure, merging lexical items, α and β, is represented by the tree in (i), with no order assumed. (i) 67 an ordered pair added to the linking process. It is formally defined as (7-3): (7-3) MERGE definition in topological network terms: Given two graphs G1 =(V1, E1) and G2=(V2, E2), MERGE (G1, G2) give a graph G=(V, E) such that (i) V = V1 ∪ V2 and (ii) For some v1∈ V1, v2 ∈ V2, E = E1 ∪ E2 ∪ {<v1, v2>} or E= E1 ∪ E2 ∪ {<v2, v1>} The resulting set of vertexes is simply the union of the E sets of the two input graphs (this fact is crucial for the driftage of Chapter, 5). Some vertex in one graph enters into a local syntactic relation with some vertex in the other graph, whereby the two graphs are combined50. Tentatively, a derivation starts with a set of minimal graphs, each of which consists of a single lexical item, chosen from the lexicon before the connection among the items take place. For instance, the example in (4-3) could starts with (8-3) (8-3) . it . will . be . snowing G1=(V1, V2): V1={it}, E1=∅ G2=(V2, E2): V2={will}, E2=∅ G3=(V3, E3): V3={be}, E3=∅ G4=(V4, E4): V4={snowing}, E4=∅ Since the auxiliary be selects the progressive form of a verb (snowing), the ordered pair <be, snowing> is introduced as in (9-3): (9-3) . it . will . be . snowing The new node γ is introduced along with two directed edges connecting it with α and β. α and β are pronounced but γ has no phonetic value. If γ is the same type as α, α is the head of the structure labeled as γ, and α selects or agrees with β. The syntactic relation is reversed if γ is the same type as β (cf. Yasui, 2004b). 50 I assume that is the first member of each ordered pair to select the second member or to agree with it with its EPP feature (cf. Boskovic, 2002). 68 MERGE(G3, G4) = G5=(V5, E5) Where V5=V3∪V4={be, snowing} And E5=E3∪E4 {<be, snowing>}={<be, snowing>} Then, the item Will selects51 verbal and also case-checks for nominative; thus, the recursive application of MERGE will convert (9-3) into (10-3) and then (11-3) (=(6-3)): (10-3) . it . will . be . snowing MERGE (G2, G5)=G6=(V6, E6) V6=V2∪V5={will, be, snowing} E6=E2∪E5∪ {<will, be>}={<will, be>, <be, snowing>} (11-3) . it . will . be . snowing MERGE (G1, G6)=G7=(V7, E7) V7=V1∪V6={it, will, be, snowing} E7=E1∪E6∪ {<will, it>}={<will, it>, <will, be>, <be, snowing>} It does not matter which of the ordered pairs in E7 is added first. For example, it is easy to verify that adding the pair <will, it> before <will, be> or <be, snowing> will make the same result. It is important to notice that MERGE as defined in (7-3) does not prevent a graph from being combined with itself. This is, in fact, an instance of an 51 It is possible to find some similarities with some assumpion made in Bowers (2001). Cf. paragraph 2. for a discussion of his proposal. 69 internal MERGE operation (or movement). To illustrate this point, consider (12-3), an example involving wh-movement: (12-3) [I wonder] what Gianni will say. The bare phrase structure analysis of (12-3) probably could be the structure represented in (13-3)52. (13-3) [WH] what [WH] [WH] will will Gianni will say say what With a graph theoretic representation, (12-3) would have a structure like (14-3) before the internal MERGE of the item what: (14-3) . [WH] Gianni . . will . say . what V={[WH], will, Gianni, say, what} 52 Note that the vP structure and the movement (internal merge) of the subject/object are ignored here. 70 E={<say, what>,<will, say>,<will, Gianni>,<[WH], will>} MERGE {G, G} will produce (15-3), where the wh-checking relation of <[WH], what> is inserted: (15-3) MERGE (G, G)=G’=(V’, E’): a. V’=V ∪ V=V={[WH], will, Gianni, say, what} E’=E ∪ E ∪ {<[WH], what>}={<say, what>,<will, say>, <will, Gianni>, <[WH], will>, <[WH], what>} b. . [WH] Gianni . . will . say . what The representation given in (15-3) clearly violates one of the defining properties of tree from a classic perspective, (see c2): what is immediately dominated by say and [WH] (cf. also Chen-Main (2006) proposal, reviewed below in paragraph 3.9.3). More generally, internal MERGE on a tree always introduces one closed route, and the result is not a tree by definition (this fact has been noted for first by Yasui, 2004a). The example (15-3) might look too outrageous, given the widely accepted view that a sentence has a tree structure. Nothing, however, seems to be wrong with this kind of representation as a syntactic structure. We may ask if every lexical item has a “vertex” nature. We will address this question in Chapter 4 and 5. In brief, I have argued that internal MERGE is not different from external MERGE in adding one ordered pair to E. The difference is that in external MERGE, one member of a pair is a new lexical item introduced from the lexicon, while two of the lexical items already introduced form a pair in internal MERGE. 71 3.4 Constituency I believe that lexical graphs can capture constituency and other important syntactic relations expressed in standard tree representations. Take again (13) for example. Its traditional and minimalist representations are showed in (16-3 a,b), while its lexical graph is (16-3c). (16-3) The non-branching nodes in (16-3a) are eliminated in (16-3b), and the remaining projection nodes are represented by lexical items as labels, which are marked to distinguish them from those that have a phonetic realization. 72 The node be* expresses important syntactic relations: be and snowing are sisters; they form a constituent; and the constituent they form is the same type as be rather than snowing. The lower will* has a similar function. The upper will* captures the fact that it forms a larger structure with the constituent that is of the same type as will, and that the resultant structure is also the same type as will. The marked vertexes are removed in (5c), where the relations of selection and agreement are represented by directed edges. The four nodes in (16-3c) correspond to the four minimal projections in (16-3b). The upper will* and be* in (16-3b) correspond to the subgraphs rooted by will and be in (16-3c), respectively. More generally, a constituent of the type α can be defined as α itself or a sub-graph that consists of all the nodes α dominates. One constituent that falls out of this definition is the lower will* in (16-3b), which corresponds to the intermediate projection I’ in (16-3a). This is a welcome result, since an intermediate projection is syntactically and semantically invisible as Chomsky (1994:10) claims. Head-movement, which is strictly local, affects two nodes connected by a single edge. In a lexical graph, the subject is adjacent to the Infl just like the object is adjacent to the verb. Assuming the standard tree representation like (16-3a), Kayne (1984) and Haegeman (1996: 485) argue that cliticization of the object to the verb is upward and syntactic, but cliticization of the subject to the Infl is downward due to the intermediate node I’ and hence must be analyzed as a PFoperation. If a lexical graph is adopted, this conclusion can be circumvented since the problematic intermediate node I’ is absent, as discussed above (cf. also Yasui, 2004a). The configuration of a lexical graph alone does not distinguish a specifier from a complement. If syntactic structure is built first by satisfying selectional requirements, followed by agreement or formal feature checking, a specifier and a complement can be distinguished based on their derivational histories. Alternatively, a specifier, if it is not an expletive, results from Move or internal Merge; it is connected into the category selecting it and also to the one inducing agreement with it. I will assume that a specifier can be identified by its derivational history or its double connectedness in a lexical graph. I think that it is possible to give a basic condition for derivational steps: A derivation graph for an expression is a record of one possible sequence of steps taken to derive the expression in question from lexical items. 73 3.5 Graph theory and C-command Next, consider how c-command can be defined in a lexical graph. The example in (17-3a; adapted from Yasui 2004a, who discusses the same argument in a slightly different manner) shows a typical contrast in reflexive binding, and its bare phrase structure and lexical graph are (17-3b,c), respectively. (17-3) a. the mother of the boy talked about herfelf/*himself c. [past] the talk mother about of herself/*himself the boy In (17-3b), the upper the* is immediately dominated by the root, which dominates the reflexive, but the lower the* is immediately dominated by of*, which does not dominate the reflexive. Applying the same definition of ccommand to (17-3c) can account for the contrast; the upper the c-commands 80 In my opinion, the very interesting thing in Yasui account59 is the demonstration that one way to derive multiple PF-interpretations is to choose at a branching node either of its child nodes rather than always giving priority to its left child, while another possibility is to start a post-order traversal from any leaf node and proceed towards the root generally in the ascending order of the nodes. This is reasonable since all the leaf nodes are pronounced before the internal nodes in post-order. The root of a graph is unique in dominating all the other vertexes and that leaf vertexes are also unique in the opposite sense; they dominate no other nodes. If a post-order traversal starts from a leaf node, there should be as many head-final pronunciations as the number of leaf nodes. 3.9 Other leading hypothesis for a label-free syntax I introduce in this paragraph other influential theoretical issues regarding a label-free syntax, such as the ones developed by Chris Collins (2001; 2002), John Bowers (2001), and Joan Chen-Main (2006). I have to admit that the major influence for the development of my graph-based label(and level)-free syntax has been Chris Collins’ (2002) seminal paper “Eliminating labels”- 3.9.1 Collins (2001-2002): eliminating labels and projections The main point of Collins’ analysis is to suggest that it may be possible to eliminate labels in the minimalist framework. In other words, the operation Merge(V, X) yelds (b) rather than (a): (25-3) a. Merge (V, X) = {V, {V,X}} b. Merge (V, X) = {V, X} This is also the basic assumption of my work. Collins argues that given a principle of lexical access (Chomsky 2000) that he calls “The Locus Principle”, the labels in the theory of Bare Phrase Structure can be eliminated entirely, leaving bare Merge as the only operation of syntax apart from Agree. Here is Collins’ Locus Principle: (26-3) 59 It is interesting to notice that she also developed a C++ program for the derivation of lexical graphs (cf. Yasui 2003). 81 “Let X be a lexical item that has one or more probe selectors. Suppose X is chosen froma lexical array and introduced into the derivation. Then the probe/selectior of X must be satisfied before any new unsaturated lexical items are chosen from the lexical array. Let us call X the locus of the derivation”. (Collins 2002: 46) Following Collins’ terminology, a lexical item is said to be saturated if all of its selectors have been satisfied and unsaturated if one or more of them is still unsatisfied. When all the selectors of every lexical item in the array have been satisfied, the lexical array is saturated and the process of forming relations is complete. If any selector of any lexical item has not been satisfied, then the lexical array is unsaturated and the process of forming relations is incomplete. It is possibly to assume that when a selector of a lexical item has been satisfied, it is deleted from the lexical entry. A saturated lexical item thus contains no selectors, while an unsaturated lexical item contains at least one unsatisfied selector. The Locus Principle ensures that no relation can be formed between a lexical item and another unsaturated lexical item. Then, Collins attempts to extend the Minimal Link Condition to subcategorization in a label-free theory of Merge. One straightforward way to state the Minimal Link Condition is as follows (see Chomsky, 2002 and Rizzi, 1990): (27-3) Let P be a probe. Then, the goal G is the closest feature that can enter into an agreement relation with P. Collins proposes to account for the fact that subcategorization/selection conditions are severely constrained to apply to the nearest c-commanded category of the appropriate type by treating the subcategorization/selection feature as a kind of probe, hence subject to the Minimal Link Condition (MLC); in order to explain the fact that the functional projection in an example such as the following doesn’t block subcategorization: (28-3) John looks too happy to leave. Collins stipulates that the MLC applies to subcategorization in such a way that it is blocked he stipulates that the MLC applies to sub-categorization in such a way that it is blocked just in case there is an intervening lexical category ([+/-V, +/-N]). The problem is why a prenominal adjective doesn’t 82 block selection of N by a D element such as the: (29-3) a. [the [smart [student]]] b. [the [very smart [student]]] Considering that this is not a problem in the case of a branching AP such as very smart in (29-3b), Collins speculates (following Rubin, 1996) that perhaps prenominal adjectives are always branching categories, though he doesn’t really argue very strongly for such an approach. In any case, it is clear that Rubin’s theory, while introducing a new functional category Mod which selects AdvP and AP complements, is just another means of getting around the fact that X’-theory doesn’t provide a natural way of distinguishing modification from relations such as subcategorization and selection (cf. also Bowers, 2001). 3.9.2 A basic operation for syntax: Form Rel. (Bowers, 2001) In an influential yet still unpublished paper (book?) from 2001 John Bowers tries to show that the only operation needed in the syntax is “Form Relation” (FormRel), which combines pairs of lexical items, or features of lexical item, and forms ordered pairs in accordance with specific properties of those lexical items. Thus, Bowers argues that there is a very small set of ordered pairs that constitute the fundamental relations (in the mathematical sense) of natural language syntax. Bowers, following Collins (2001; 2002) assumes in his work just four basic linguistic relations: selection, subcategorization, modification, and agreement. In addition, the author argues that each time an ordered pair is formed, there is an immediate reflex in both at the phonological interface and the semantic interface (with principles called “Immediate Spell-Out” and “Immediate Interpretation,” respectively). Given these assumptions, Bowers tries to show that the notions of constituent structure (and so the tree adjoining model) and movement are simply artefacts of the fundamental legibility conditions that hold at the semantic and phonological levels, together with a small number of computational principles that either limit the search space of FormRel or limit the possible outputs of Spell-Out and Interpretation60. 60 Neither the idea that the primitives of syntactic theory should be relations rather than constituents nor the idea that Spell-Out and Interpretation should be immediate are totally new and unique to Bowers’ theory. Various similar proposals have been shared in 83 Bowers assumes, as mentioned above, that there is just one basic operation in narrow syntax, Form Relation (FormRel). FormRel is a binary operation that applies to lexical items a and b and forms an ordered pair (a, b). Given this operation, a network of syntactic relations is built up in the following way. First, an array A of lexical items is chosen from the Lexicon. Second, FormRel applies successively to pairs of lexical items, selected from A and from previously formed ordered pairs, continuing until all the selection and subcategorization (cf. Collins, 2002) features of every lexical items are satisfied and none are left unsatisfied (cf. Bowers, 2001: 16). The derivation in Bowers’ paper is regulated by the following computational principle, a slightly modified version of Collins’ (2003) Locus Principle (see above (26-3)): (30-3) The Locus Principle according to Bowers: “Suppose a lexical item l, called the Locus, containing unsatisfied selection and subcategorization features, is selected from a lexical array. Then all the subcategorization conditions and selectional requirements of l must be satisfied before a new lexical item can be selected as the Locus” (Bowers, 2001: 18). Bowers agrees with Collins, considering the fact that for him a lexical item all of whose sub-categorization conditions and selectors have been satisfied is said to be saturated; if any of them have not been satisfied, it is said to be unsaturated. Indeed, the Locus Principle rules out the possibility of a lexical item A forming a relation with an unsaturated lexical item B61. I give an example below to make clearer Bowers’ view. Consider the phrase read the books. The Locus Principle requires that the relation selection RSel (the, books) be established before the relation subcategorization RSub (read, the). If the latter was formed first, the Locus Principle would be violated, since the item the would be unsaturated (cf. Collins, 2002) at that point62. frameworks such as Perlmutter and Postal Relational Grammar (see Appendix B). 61 As shown in Collins and Ura 2001, this imposes an inherent order on the process of forming a network of relations between lexical items. 62 It is important to note that there are no constituents in a theory of this sort. In the example just discussed, there is no constituent [the book] in narrow syntax, nor is there one of the form [read [the books]] (with or without labels). Instead, there are simply two relations (read, the) and (the, books). 84 From a theoretical point of view, it’s interesting to say that, though there is a superficial similarity between an account based purely on relations and one that incorporates the operation Merge or its equivalent, due to the fact that both involve the construction of sets, it must be said the operation Merge goes far beyond what is involved in a relational theory. In the sentence above, the output of the (canonical) Merge operation would be a new syntactic object of the form: {read, {the, books}}. Despite the fact that the outputs of successive applications of Merge are only unordered sets, each operation results in a new syntactic object which incorporates the results of all the preceding operations: it is clearly a theory that incorporates a notion of constituent structure. In the relational theory of John Bowers, on the other hand, no new syntactic objects of this sort are produced. Instead, still considering the sentence above, there are just the two ordered pairs (read, the) and (the, books). Bowers’ account would be a revolutionary one, but I suppose that the notion of (a simplified) Merge is an axiom of syntactic theory63, and I have cited Bowers’ paper in my work mainly because is another important attempt to eliminate labels and to simplify the whole syntactic derivation process. Indeed, Bowers gives a simple and economical account of the possible linearization process in syntax (Bowers, 2001; p.27): (31-3) “Suppose a and b are lexical items and the ordered pair (a, b) is a member of the selection relation RSel. The linearization function (FL) operates on RSel(a,b) as follows: FL (RSel(a, b)) = a-b”. Thus Bowers’ FL is a very simple and general function which ensures that the phonetic form of the “first coordinate” of a sub-categorization relation precedes the phonetic form of the “second coordinate”. Here is an example to see how FL works. Let’s start by choosing from the lexicon the items read, the, and book. Assuming that the selects nouns and read subcategorizes determiners, the two relations (the, books) and (read, the) can be 63 The syntactic operation Move could dissolve into the simpler, more economic, and indispensable syntactic operation Merge in the simplified way shown in paragraph 3.3, and X-bar theory could be simplified from a graph theoretical point of view; in my idea of syntax, Merge is a necessary “linking operation” over two graphs, with an ordered pair added to the linking process and, in many respects I agree with Starke (2001) over the unnecessarily of operation Move by the fact that Move requires Merge, but not vice versa. 85 formed. The Locus Principle, in the form of Bowers, requires that they are formed in that order. The relation (read, the) couldn’t be formed first, because the would be unsaturated at that point. Here is a schemata of the process, which is a simplified version of the one in Bowers (2001: 29) (32-3) Bowers’ linearization process i) a. Select the, books from A [Locus] > the (unsaturated); books is saturated. b. FormRel(the, books)=(the, books) : FL((the, books))= the-books ii) a. Select read from A; select the from (the, books) formed at step ib) [Locus] > read (unsaturated); the is saturated. b. FormRel(read, the)=(read, the) : FL((read, the)) =the-books-read 3.9.3 Multi-dominance and Lexical graphs (Chen-Main, 2006) Finally, concerning linearization, I discuss briefly some of the things put in evidence by Chen-Main PhD dissertation (2006), that explores formal and linguistic consequences in a “multi-dominance” system that result from taking linearizability to be a property of well-formed syntactic structures. Her work is another intentional return to the notion that syntactic structures should be represented at the most basic level with nodes (vertexes) and edges, and an invitation to import useful ideas and results from the study of graphs. For Chen-Main, the point of departure is the consideration of how some common constructions such as wh-questions and coordinated constructions seem to allow lexical items to play multiple grammatical roles typically associated with distinct positions. As a prototypical example, in sentences like “What did Gianni eat?”, what is usually assumed to function as the object of eat, even though it appears sentence initially rather than in the canonical object position. In another Chen-Main example “Joe bakes and Sam sells cookies”, a single noun phrase, cookies, satisfies both verbs’ need for an object. Traditionally, this apparent “multiple-linking” is attributed to co-indexing distinct elements filling multiple positions, only one of which is pronounced. Alternatively, Chen-Main argue that a multiply-linked element can be conceptualized as an element immediately dominated by multiple parent 86 nodes. Under such a multi-dominance approach, trees no longer suffice for representing the immediate dominance relation. Rather, the set of syntactic structures is expanded to include non-tree graphs in a shape similar to those examples I have given in sections 3.2 and 3.3. The interesting fact about Chen-Main thesis is to examine how such multidominance structures64 are generated and how their terminals are linearized. That question is answered in her work by adopting the node-contraction operation, originally introduced into the Tree Adjoining Grammar formalism to allow for coordination (cf. also Citko, 2005). Chen-Main considers nodecontraction to be a general mechanism in the Tree grammar system. Here is my schematic resume of her proposal: node-contraction is involved not only in generating coordination, but also for cases traditionally dealt with via movement. The existence of island effects that prohibit movement from certain domains indicates that one must specify when multi-dominance cannot occur (by placing certain locality restrictions on node contraction at the derivational level, Chen-Main argues that a number of these island effects can be derived). The linearization quest is a matter of real interest (cf. Obsviously Kayne, 1994 and Chapter 5) and, as I have already mentioned, shares some similarities with my approach65. This proposal too leaves behind the oneparent-per-node restriction that characterizes trees (nodes are allowed to be immediately dominated by multiple parent nodes), but it differs from my view in the fact that the possibility of the elimination of labels is not considered there. Below, in (33-3) I show Chen-Main definition of a lexical graph: 64 To summarize the matter, multi-dominance has been explored by a number of researchers: Peters and Ritchie’s (1981) “phrase linking grammar” is a variety of multidominance syntax in which two types of immediate dominance are possible and a node is allowed to have a parent in both relations. The structures used by Goodall (1987) to analyze coordination allow a single lexical item to be part of multiple conjuncts. Later, Gärtner’s (1997) close examination of the widely followed Minimalist Program (Chomsky 1995) led to a proposal to replace the operations Merge, which combines two syntactic objects and forms a single combined object, and Move, which duplicates part of a syntactic object and merges it with the original object, with a single hybrid operation whose application allowed multidominated structures. In 2001, both Starke and Chomsky recast Move as a special case of Merge. Starke (2001) argued that Move could be reduced to a special case of Merge applied to non-adjacent nodes, and Chomsky (2001) introduced the terms External Merge, which merges nodes that have not been merged before, and Internal Merge, which re-merges a node that has previously been merged, resulting in multiply dominated nodes. (see also Citko, 2005). 65 However, the proposed process for deriving ordering information does not guarantee a linearization of terminals for every graph. A graph may be unlinearizable due to either lack of ordering information or conflicting ordering information (for a detailed account cf. Chen-Main, 2006; p. 56-62) 87 (33-3) Definition of a syntactic graph (Chen-Main, 2006; p. 53) A syntactic graph is a five-tuple <N, Q, ID, SP, L>, where N is a finite set, the set of nodes, Q is a finite set, the set of labels, ID is an irreflexive, intransitive, asymmetric relation in N × N, the immediate dominance relation SP is an irreflexive, intransitive, asymmetric relation in N × N, the sister precedence relation L is a function from N into Q, the labelling function, and such that the following conditions hold: a. Single Root Condition ∃ X ∈ N such that ∀ Y ∈ N, (X, Y) ∈ ID* b. Non-Overlapping Condition ∀ X, Y ∈ N, i. if (X, Y) ∈ ID, then (X, Y) ∉ SP, and ii. if (X, Y) ∈ SP, then (X, Y) ∉ ID. c. Acyclicity Condition66 ∀ X, Y∈ N, if (X, Y) ∈ ID+, then (Y, X) ∉ ID+. A corollary of this definition given is that syntactic trees could be considered as special cases of graphs. They are subject to an additional condition that avoid from not linearized structures. (34-3) Single Parent Condition for syntactic trees (based on Chen Main 66 Note that Chen-Main acyclity condition is a crucial point of departure from my proposal (cf. Paragraph 3.2; 3.3). 88 definition of a syntactic graph see above (33-3)) ∀X, Y, Z ∈ N, if (X ≠ Y) then ¬ (((X, Z) ∈ ID) and ((Y, Z) ∈ ID)) Resuming all these argument in just two words we may say that ChenMain (2006) is another interesting proposal that considers syntactic structures as directed graphs that meet certain well-formedness conditions, and that these conditions allow some non-tree syntactic structures. 3.10 Another way to simplify things: Mirror Theory (Brody, 1997) Mirror Theory is a syntactic framework developed in (Brody, 1997), where it is offered as a consequence of eliminating purported redundancies in Chomsky’s minimalism (Chomsky, 1995). A fundamental feature of Mirror Theory is its requirement that the syntactic head-complement relation mirror certain morphological relations (such as constituency). This requirement constrains the types of syntactic structures that can express a given phrase; the morphological constituency of the phrase determines part of the syntactic constituency, thereby ruling out other, weakly equivalent, alternatives. Another fundamental feature of Brody (1997) is the elimination of phrasal projection. Thus the X-bar structure on the left becomes the mirror theoretic structure on the right: (35-3) XP X ------> YP X’ Y Z X ZP (Brody, 1997) calls this systematic collapse of X, X’ and XP nodes “telescope”. Every node may now have phonetic content, and children are identified as specifiers or complements depending on their direction of branching; left-daughters are specifiers and right-daughters are complements (previously, as we know specifiers were children of XP, and complements were children of X’). Furthermore, the complement relation is a “word- 89 forming” relation, where according to the “mirroring” relation, the phonetic content of each head follows the phonetic content of its complement. For example, mirror theory can generate trees like the following, which given the “mirror” relation between morphology and syntax, is pronounced John sleep - s: (36-3) -s Johni Sleep ti Kobele et al. 2002 in a work based on Joshi (1987)67 give a formal representation of a Mirror grammar. A mirror theoretic tree can be viewed as a standard binary branching tree together with two functions; one, a function f from branches to a two element set {right; left}, the other, a function g from nodes to a two element set {strong; weak}. If a is the parent of a’, then a’ is a specifier (or left child) of a if f ((a; a’)) = left and a complement (or right child) of a otherwise. In basic terms, a mirror theoretic expression is defined to be a mirror theoretic tree along with a labelling function from the nodes of the tree to a set of labels. A label consists of a phonetic part (which is opaque to the syntax) and a finite sequence of syntactic features. A mirror theoretic grammar consists of a finite lexicon of ‘basic’ expressions, together with two structure building operations, merge and move, which build expressions from others either by adjoining structures, or by displacing sub-parts of structures. Each operation in Mirror Theory is feature driven, and ‘checks’ features (and thus a derived expression will have fewer features than the sum total of the features of the expressions (tokens) used to derive it). The expressions generated by the grammar are those in the closure of the lexicon under the structure building functions. A complete expression is one all of whose features have been checked, save for the category feature of the root, and the string language at a particular category is simply the yields of the complete expressions of that category. 67 I frankly suggest to read Joshi, Aravind K. 1987. An Introduction to Tree Adjoining Grammars. In A. Manaster-Ramer, editor, Mathematics of Language. John Benjamins, Amsterdam. 96 describes phases as self-contained components of derivation, and asserts that internal elements of a given phase must be on its phase’s edge, before moving out to another phase. Within a traditional paradigm, it is a must to describe Persian CP as head-initial, because this is an easy way to account for many facts concerning CPs in Persian: a) complement clauses follow matrix clauses; b) relative clauses follow matrix clauses; c) the interrogative particles aya/magar (if) are the “topmost” items in interrogative sentences; d) although wh-words do not necessarily move, when they move, it is to the beginning of an utterance. See the examples below: (10-4) a. ne-mi-dun-e [CP ke farda mi-yam] NEG-DUR-know-3S that tomorrow DUR-come-1S ‘He doesn’t know I’m coming tomorrow.’ (Mahootian 1997: 90) b. un mard-o [CP ke ruzname mi-xund] peyda kard that man-OM that newspaper DUR-read visible did ‘He found the man who was reading the newspaper.’ (Mahootian 1997: 34) c. aya in gorbe-ye-shoma-st? INTER this cat-Ez-you-is ‘Is this your cat?’ (Mahootian 1997:9) d. cera ma saket be-man-im? Why we quiet SBJN-remain-1P ‘Why do we remain quiet?’ (Dehdari 2006: 45) Concerning split headedness an implied statement was made by Karimi (2005) assuming the following clause structure for Persian: (Karimi, 2005: 7) (11-4) 97 I propose the following table (adapting and revising the one proposed in Dehdari, 2006) to summarize the evidences given by the empirical data: (Tab. 4-1) Relationship Head initial Head final Perf Aux - VP √ V – manner adverb √ V – Obj √ Verbal Copula – Pred √ V – PP √ Passive Aux – V √ Preposition – N √ Det – N √ Num – N √ N – Relative enclitics √ N – Gen (Ez) √ Adj – superlative (es. tarin) √ Complementizer – S √ Interrogative – S √ Tense (es. mi particle) – VP √ Adverbial Subordinator - S √ 98 It would be logic to assume that human parsers are quicker when the branching is consistent in one direction. Mixed-branching seems hard to process. However, Persian data seems to reveal a mixed model. Kayne (1994) challenges the above view. According to Kayne's theory, linear ordering is mapped from asymmetric c-command relations that hold between non-terminal nodes; thus, distinct word orderings should not reflect distinct hierarchical structures (cf. Chapter 5). Kayne (1994) argues that head final languages have a head-initial underlying structure with abstract functional categories, to which movement operations apply so as to derive the appareant head-final order. It is interesting to notice that Chomsky (1995) points out some weaknesses of Kayne’s theory such as its crucial reliance on non-branching nodes to deduce surface order. Another type of word order variation that is not strongly linked to clear syntactic (or semantic) differences has been referred to scrambling, the widespread phenomenon of Persian syntax. See the following example from Japanese, which is acceptable, without a heavy stress on the initial constituent or a pause after it: (12-4) Mary-o John-ga mi-ta Mary-Acc -Nom go-Past 'Mary, John saw.' One dominant approach represented by Saito (1985) and subsequent work is to regard (12-4) as derived by the syntactic operation of scrambling (cf. also Karimi, 1999 for Persian data). Scrambling, in a traditional framework, produces an adjoined structure and the moved constituent leaves a trace in its original position. So we have to assume hierarchical properties that differ from the non-scramblel (base-generated) instance. As we have seen in chapter 3 a syntactic graph theory allows to formulate algorithms that can deduce multiple PF interpretations from the shared syntactic structure. The flexibility in word order or the multiplicity of PF interpretation appears to be attested in head-final languages (Fukui (1993); Yasui (2004)). But also head initial languages71 have variable degrees of word order freedom. 71 For instance, English, which is head-initial, allows a certain amount of word order freedom by shifting a heavy NP rightward as shown in (ia,b), but a light constituent like the pronoun it cannot be shifted rightward, as shown in (id): 99 It can be said that head-final languages allow wider variation in word order than head-initial languages, and the variation in the former is always leftward; rightward word flexibility is highly limited. Another importantasymmetry between head-initial and head-final languages was pointed out by Bresnan (1972). Bresnan (1972: 42) states that only languages with clause-initial Complementizer permit a Complementizer-attraction transformation72. Fukui (1993) proposes a stimulating theory of the correlation between a value of the head-parameter and word order flexibility based on a grammatical operation (Move alpha) that creates a structure that is inconsistent with the value of a given parameter in a language is costly in the language, whereas one that produces a structure consistent with the parameter value is costless (Minimalist economy is well interpreted in this way). According to this idea, scrambling in head final languages is of no cost since it moves a constituent leftward and does not destroy the head-finality. Here is Fukui's account (1993: 400): a. A language has a costless optional movement (or shows flexible word order) only if it is head-final, and the operation produces a structure consistent with the head-final value (i.e., it is leftward). b. A language has a costly obligatory movement only if it is head-initial, and the operation produces a structure inconsistent with the head-initial value (i.e., it is leftward). (i) a. They brought the beautiful dress into my room, b. They brought into my room the beautiful dress. (Fukui (1993: 410)) c. They brought it into my room. d. *They brought into my room it. On the other hand, for example, Japanese scrambling moves a constituent leftward, whether it is light or heavy, as shown in (iia,b): (ii) a. sono utukusii doresu-o karera-wa watasi-no heya ni mottekita. that beautiful dress-Acc they-Top I-no room to brought (Fukui (1993: 410)) b. sore-o karera-wa watasi-no heya ni mottekita. it-Acc they-Top I-no room to brought 72 Logically possible but non-existent would be languages with a clause-final Complementizer that attracts a wh-phrase rightward. 100 Crucially, it is possible to observe that in Persian long distance scrambling occur when involve head-final elements73, while, interestingly, it is blocked by any intervening NP with the same case as the NP being long distance scrambled74 (cf. Karimi, 1999: 174176; Richards, 2002: 240). _________________ | | (13-4)a. *Sasan Kimea goft [ke _ ketab-â-ro az Sepide kharide] Sasan Kimea said that book-Pl-Spec.Acc from Sepide bought “Sasan, Kimea said [that _ bought the books from Sepide]” ________________________________ | | b. *be Ali Sasan be Kimea goft [ke ketab-ro _ dâde] to Ali Sasan to Kimea said that bookSpec.Acc gave “To Ali, Sasan said to Kimea [that he gave the book _ ] _________________________________________________ | | c. *dokhtar-â-ro Kimea pesar-â-ro tashvig kard [ke _ bebusand] girl-Pl-Spec.Acc Kimea boy Pl-Spec.Acc encouragement did that kiss “the girls, Kimea encouraged the boys [to kiss _] 73 Following Yasui (2004a), I think that an elegant account of Persian word order may be given if we simply admit the possibility of an algorithm shift in the course of the derivation (cf. Chapter 3). Indeed, the PF-interpretation of a standard syntactic tree is obtained by ignoring its nonterminal nodes and pronouncing its terminal nodes from left to right, and the ordering of terminal nodes comes from a value of the head parameter set for the language in question, while, obviously, a lexical graph requires a different (shiftable) PF-interpretive algorithm, since all its nodes need to be pronounced. Does a lexical graph possess the linearity required for its PF interpretation? I claim that it originates in its overall configuration. Lexical items are introduced into a syntactic derivation one by one, producing a larger structure at every step. Frampton and Gutmann (1999, 2002) make this point clear in their theory of crash-proof computation: lexical items are introduced automatically in the right order, and no crash caused by incorrect order of selection is possible. Then, it is natural to expect the right order of structure-building to be reflected in the PF word order. As I have already pointed out in the previous chapter, following Yasui (2004), tree traversal algorithms can be classified into two major categories: depth-priority and width-priority traversals. What seems to be relevant to traversal of natural languages is the former: starting from the root, we go as deep as possible until reaching some leaf node, typically the leftmost one (for spatial-temporal reasons); we move back to its mother node and visit the other children if any; the remaining nodes are traversed in the same manner. 74 Notice that this facts seems to hold in Japanese as well (cf. Saito, 1985:185, cited in Richards, 2002), where scrambling of a subject past another subjest is impossible: ________________ | | (i) *Sono Okasi-ga John-ga [ _ oisii-to] omotteiru this candy-Nom John-Nom tasty that thinks “This candy, John thinks [ _ is tasty] 101 4.1.3 Some observations on Persian relatives: a Case Attraction phenomenon Persian relative clauses are usually introduced by the complementizer ke (that), which is used regardless of the animacy, gender or function of the head noun (Karimi, 2001). In non-restrictive relative clauses, the head noun often carries an enclitic morpheme (Encl) which links the noun to the following relative clause (14-4). If the relativized noun is the object of the main sentence, then it may appear with the object marker râ as illustrated in (15-4). That’s an interesting empirical observation. The following examples are from Megerdoomian (2001). (14-4) zan-i ke injâ neshaste ast hamsar-e Nâder ast woman-Encl that here sit-Part is spouse-Ez Nader is `The woman that is sitting here is Nader’s wife” (15-4) ketâb-i-râ ke diruz kharide budam emruzsobh tamâm kard-am book-Encl-Obj that yesterday bought was-1sg today-morning finish did-1sg “This morning, I finished the book that I had bought yesterday.” The relative clause may be separated from the head noun by the main verb as illustrated below (see Megerdoomian, 2001 and Franco, 2004 for a more articulated discussion). In addition, several relative clauses could follow a head noun. The following example in taken from Ghomeshi (2002). (16-4) mâ pesar-ân-i-râ entekhâb mi-kon-im ke dar jang sherkat na-karde-and we boy-Plur-Encl-Obj choosing Imp-do-1pl that in war participation neg-done-3pl “We choose (the) boys that have not participated in the war.” If the head noun is the subject or direct object of the relative clause, it is often left as a gap as was shown in the examples in (14-4) and (15-4). However, even in such cases, the relativized noun may be replaced by a resumptive pronoun in the clause it originated from. Thus, in (17-4), an example taken from Megerdoomian doctoral thesis (2001), the head noun 102 plâk-e kuchak (small plaque) is the subject of the relative clause; it is substituted by the resumptive pronoun ân (it). The use of the resumptive pronoun usually occurs when the head noun is separated from the relative clause by an intervening verb (cf. McCloskey, 1992). In this example, the verb pey borde-and (have found) precedes the relative clause. (17-4) dâneshmand-ân be plâk-e kuchak-i dar maqz pey-borde-and ke ân niz tâkonun nâshenâxte mânde bud. scientist-Plur to plaque-Ez small-Encl in brain found-3pl that it also until now unknown remained was “Scientists have found a small plaque in the brain that until now had remained undiscovered”. Thus, when the head noun is the indirect object or is extracted from a Prepositional Phrase adjunct in the clause, a resumptive pronoun is used. In other words, the position from which the head noun originates is substituted by a pronoun that agrees with the head noun. This is exemplified in the sentences below: (18-4) in bache-hâ ke az ânhâ âdres mi-porsid-i... this kid-Plur that from them address Imp-ask-2sg “These kids from whom you asked for the address...” (19-4) shahr-i ke dar ân tazâhorât shode bud ... city-Encl that in it demonstrations become was “The city in which demonstrations took place...” (20-4) zan-i ke barây-ash ketâb kharid-i ... woman-Encl that for-Clitic(3sg) book buy-Past-2sg “The woman for whom you bought a book...” As shown in Franco (2006) and as already mentioned above (15-4), the morpheme /raa/; /ro/ in spoken form) - as the specific marker for accusative case - can accompany, at least in spoken language, the head noun of the relative clause that is the subject of the main clause and the object of the relative clause (21-4) in Persian. (21-4) Zan-i-ro [ke did-i] inja-st 103 woman-Acc that saw-2sg here-is “The woman whom you saw is here” While this phenomenon concerning the Persian language, known in the literature as Case Attraction, resembles to the “Inverse Attraction” discussed in Bianchi (1999) for Latin and Ancient Greek, it has its own peculiar characteristics: a) it is quite optional b) it blocks extraposition, as shown in (22-4), c) it is always the nominative case that is attracted to the accusative case (23-4a,b,c). (22-4). *Zan-i-ro inja-st [ke did-i] (23-4) a. pesar-i-ro [ke … ] boy-Encl-obj that NOM ⇔ ACC b. *be pesar-i-ro [ke … ] to boy-encl-obj that DAT ⇔ACC c. * az pesar-i-ro [ke … ] from boy-encl -obj that ABL ⇔ACC Then, this attraction only applies to the head noun of the restrictive relative clause. Since the head noun of the non-restrictive clause lacks the restrictive morpheme /-i/, it cannot attract the marker for the accusative case (24-4). (24-4) * an mard-e mosen-ro [ke diruz did-am] emruz raft that man-EZ old-obj that yesterday saw-I today went-3sg “That old man, whom I saw yesterday, went today”. 104 Karimi (2001), instead, argues that /raa/,/ro/ as the specificity marker for accusative case in Persian, cannot be generated with the relative head in the relativized position (subject position of the relative clause) (25-4 a;b). Hence, rejecting Kayne’s (1994) raising analysis, she suggests that the head noun is base-generated in the Spec of the larger DP and then it moves to the spec of KsP (the suggested projection that has the marker in its head) (26-4). (25-4). a. Kimea un pesar-i-ro [ke inja neshaste bud] be man mosarrefi kard. Kimea that boy-Acc that here sitting was-3sg to me Introduction do-past-3sg “Kimea introduced to me the boy who was sitting here” b. un… [ CP [ C’ ke [-i pesar-ro] inja neshaste bud] ] tha that encl boy-Acc here sitting was-3sg (26-4) [ksP [un-pesar-i] i [Ks’ –ro] [DP ti [ D’ ] [CP ] ] ] (Karimi, 2001) Given the possibility of the occurrence of the accusative case marker /raa/ with the subject of the main clause (when is the object of the relative clause), I propose that, at some point in the computational process of derivation, the case marker /raa/; /ro/ was present inside the relative clause. Adopting a raising analysis - as suggested by Kayne (1994) and discussed in Bianchi (1999) and Bhatt (2002) among others - for relative clauses in Persian we will come up with a structure as in (27-4) for sentences like the one in the example (21-4). The case marker for the accusative together with the head noun moves to the position of Spec of CP and then the head noun further moves to spec of DP. (27-4) [ CP [ DP [ zan ] K –i tK -ro] i [C’ ke ti did-i] inja ast If my proposal is right, so the empirical facts that I cited above are evidences for the possibility of a raising analysis to explain the syntax of relative clauses in Persian. The impossibility of generating the accusative marker with the head of the relative clause in a relativized subject position in Karimi (2001) appear as an inadequate argument for rejecting a raising 105 analysis75: as we have shown, the appearance of the same case marker with the subject of the main clause is an important empirical observation that supports a raising analysis. 4.1.4 Persian Noun Phrases [a rough guide to the Ezafe domain] The head of a noun phrase could be a noun or an infinitival verb. Pronouns and proper names may also head noun phrases and they function as possessors in forming complex noun phrases (such as possessive constructions: ketâb-e Saloomeh (Salome’s book)). Persian head noun is preceded (at the surface) by the determiners, the numeral constructions and the quantifiers, and it is followed by the modifiers, which usually consist of an adjectival phrase (AP). Superlative adjectives, however, do not appear in the AP; instead, they precede the head 75 In this note I show the basic syntactic interpretations for the two major competing analyses of relative clauses: the raising analysis and the matching analyses. The head raising analysis was originally proposed by Brame (1968), Schachter (1973), and Vergnaud (1974). Recent versions include the relevant one of Kayne (1994), among others. Under the head raising analysis that we are adopting, the head NP originates inside the relative clause CP, as shown in (i). (i) the [book]j [CP [which tj]i John likes ti] The matching analysis was originally proposed in Lees (1960) and Chomsky (1965) and has been discussed and extended in Sauerland (1998). The matching analysis postulates that corresponding to the external head there is an internal head which is phonologically deleted under identity with the external head. However, the internal head and the external head are not part of a movement chain. In fact Sauerland argues that in certain cases, the lexical material of the internal head does not need to be the same as the lexical material of the external head. It just needs to be similar enough. (ii) the [book] [CP [which book]i John likes ti] A more articulated discussion is impossible here, considering the aim of this work. Anyway, I want to show the consequences of an interesting observation made by Richard Larson (1985). Larson observed that headed relative clauses containing a trace in adjunct position, but neither a relative adverb or a stranded preposition, are grammatical only if the external head of the relative clause is a bare-NP adverb. (iii) a. the way [Opi that you talk ti] (Larson, 1985, from Bhatt, 2002) b.*the manner/fashion [Opi that you talk ti] c. You talk that way. d.*You talk that manner/fashion. The well-formedness of the operator-variable chain in (iiia) depends upon what the head NP is. Information about the head NP is required internal to the relative clause. Under a head raising or a matching analysis, the ill-formedness of (iiib) directly follows from the ungrammaticality of (8d). This explanation is not directly available under the head external analysis, and Larson, who is assuming the head external analysis, has to introduce a not economical feature transmission mechanism which makes the relevant information about the head NP available internal to the relative clause. 112 three main kinds of non-verbal constituent: bare noun heads, adjectival small clauses, and prepositional small clauses. Hale and Keyser analysis draws its primary inspiration from English (but includes an incredible set of data from native American languages), where the categorial status of adjectival and nominal verb roots is very clear. I give a representation of Hele and Keyser’s underlying structures for denominal (unergative and location/locatum) and deadjectival verbs: (40-4) This approach makes the difference between unergative and unaccusative verbs depend on more than the X-bar notation. It explains the “semanticmorphological” properties of verbs of these classes. In many languages, the verbalizing part of the structure is visibly morphologically realized as an affix or light verb, as in these examples from Hale and Keyser (2002). (41-4) a. negar egin “to cry” (Basque) 113 cry do jolas egin “to play” …and many more play do b. di-yin “to breath” (Navajo) do breath di-zheeh “to spit” … and many more do spit On such an approach, the thematic properties of a particular verb are dependent on the syntactic and semantic properties of the verbalizing functional element and of the non-verbal constituent which make it up. Changing the properties of the verbalizing element — the light verb — results in a change in Agent selection: the light verb is responsible for the presence or absence of an external argument. Similarly, the causative/ inchoative alternation in pairs like John opened the door/The door opened is also the result of varying the light verb, although the morphological consequences of this variation are invisible in languages such as English or Italian. Each of Hale and Keyser proposed underlying structures for English verbs, above, have natural non-incorporated counterparts in Persian complex predicate constructions, where the light verb and non-verbal element are realized separately. Furthermore, the agentivity of a particular complex predicate is dependent on the light verb involved, and the telicity of the complex predicate is dependent on the non-verbal element involved, in a very transparent way (Harley, Folli, Karimi, 2005). Persian is a language in which the complex syntactic nature of verbs is very easily discerned, and in which Hale and Keyser’s proposals concerning the structure of the verb phrase find wide confirmation. 4.1.5.1 Deriving Persian argument structures (Harley, Folli, Karimi 2003) We argued that unergatives are formed when a nominal element is incorporated into a light verb which selects for an external argument. Similarly, inchoativesresult when an adjectival element is incorporated into a light verb which does not select for an external argument. These structures translate naturally to Persian complex predicates (Haely, Folli, Karimi, 2003). Consider the representation of a complex predicates like gerye kardan, “weeping doing” that translates as a typical unergative like cry: (42-4) 114 Similarly, consider the syntax of a Complex predicate that translates as a typical inchoative, like bidâr shodan “awake becoming”: (43-4) Just as hypothesized by Hale and Keyser for the English causative/inchoative alternation, the alternation between the inchoative and the causative of awake in Persian is accomplished by changing the light verb from the equivalent of 'become' (shodan) to the causative 'make' (kardan). Probably, the Persian case constitutes the strongest possible evidence for the syntactic nature of l-syntax as proposed by Hale and Kaiser. 4.2 The Ezafe puzzle Now, I give a detailed review of previous analyses of the ezafe phenomenon in the generative framework (cfr. Samiiam, 1983; 1994; Ghomeshi, 1997; Kahnemuyipour, 2000; Franco, 2004; Larson and Yamakido, 2005; Samvelian; 2006 among others) and I develop a graph analysis of the Ezafe based on Den Dikken and Singhapreecha (2004), where the authors give a cross-linguistic account (the point of departure was the comparison between Franch and Thai) of the noun phrases in which linkers occur, in terms of DP-internal Predicate Inversion (see also Moro, 1997 and Den Dikkken 2006). 115 4.2.1 An introduction to Ezafe A number of West Iranian languages - Persian, different Kurdish dialects, Hawrami, Zazaki (Larson & Yamakido, 2005) - share several aspects in their noun phrase structure: i) The surface word order pattern is strongly head-initial. Adjectival modifiers, the possessor NP, prepositional phrases and the relative clause follow the head noun, which may only be preceded by some determiners (in example demonstratives, cardinals and quantifiers), and in very few cases by an adjective (cf. paragraph 4.1.3). ii) Possession is expressed by means of a bare NP (DP) which follows adjectival and some prepositional modifiers80. iii) Elements occurring between the head noun and the possessor NP are linked to the head and to one another by the Ezafe, realized as an enclitic morpheme. iv) Prepositional complements appear outside the Ezafe domain and follow the possessor NP (DP)81 The following Persian NP exemplifies these points: (44-4) in lebâs-e sefid-e bi âstin-e Maryam this dress-EZ white-EZ without sleeve-EZ Maryam “this Maryam’s sleeveless white dress” (Samvelian, 2006) 80 The expression of possession by means of an NP in close construction with the head noun, and, to some extent, the Ezafe construction is reminiscent of the Semitic construct state construction (Borer, 1988). Despite the fact that the possessor NP is not constrained to be strictly adjacent to the head noun, as it is the case in construct state nominals, the constituents occurring between the head noun and the possessor NP have been nevertheless assumed to be subject to some significant constraints, leading to the analysis of the Ezafe domain as a domain of bare heads (X°s) adjunction in Persian (Ghomeshi, 1997). This is reminiscent of the word-like properties of construct state nominals (Borer, 1988). 81 It’s interesting to notice that the word order pattern is identical to that of the Celtic noun phrase. However, unlike Celtic languages, the word order in the noun phrase does not parallel the one in the clause structure. Although Persian is verb final, the reversed order within the noun phrase has provided motivation for the application of the head movement analysis to the nominal domain in some works (Kahnemuipour 2000; Franco, 2005). 116 The Ezafe construction has been a particular focus of interest in different recent studies (Samiian, 1983; 1994; Ghomeshi 1997; Kahnemuyipour, 2000; Larson and Yamakido, 2005 among others). Actually this construction raises several issues in syntax (and also morphology, see Samvelian, 2006) mainly the status of the Ezafe itself. The Ezafe has generally been assumed by Persian grammars (see Lazard 1992) to be semantically vacuous. Furthermore, it can be iterated within the NP, occurring as many times as there are modifiers. But, at least in Persian, it is not the expression of a concord between the head noun and its dependants (this is not a case of some Kurdish dialects; see below). On the basis of these observations, Samiian (1983) and Ghomeshi (1997) propose not to view the Ezafe as a morpheme at all, but rather as an element inserted in Phonological Form (see Chomsky, 1981). For Ghomeshi (1997) - maybe the leading analysis of this phenomenon - the need for the Ezafe vowel results from the fact that nouns being non-projecting in Persian, a “phonological linker”, in example the Ezafe, must be present in order to indicate phrasing within the nominal constituent. This view of the Ezafe has been rejected in subsequent studies and various alternative analyses have been suggested. Ezafe has been seen: i) as a Case-marker (Samiian 1994, Larson and Yamakido 2005; 2006). ii) as a marker associated with the syntactic movement of the noun and realizing a “strong feature” (Kahnemuyipour, 2000). iii) as a the morphological ouvert item of a functional head in the domain of AP, in a “cartographic approach” (Franco, 2005, following Cinque, 1994). iv) as a suffix attaching to the head and to some of its intermediate projections, and marking them as awaiting a modifier or a complement. (Samvelian, 2006, following Nichols, 1986). In the following sections work, after an historical excursus and the review of the most interesting analysis of the Ezafe phenomenon in the generative paradigm, I will made a new graph based proposal following some observations developed by Den Dikken and Singhapreecha and Den Dikken (2006) for Chinese – Mandarin, which may lead us to consider Ezafe as a linker indicating subject predicate inversion. This view of the Ezafe is quite the opposite of the Case-marker analysis suggested by Samiian (1994) and Larson and Yamakido (2005; 2006), 117 according to which the Ezafe is rather a dependent marking device. 4.2.2 The Ezafe construction: a rough historical perspective and a basic overview of Ezafe domain The Ezafe construction is a specificity of those languages that display a head-initial word-order pattern within their DP (e.g. Persian, Kurdish dialects, Hawrami, Zazaki, Kermanian dialects, etc). The correlation between the head-initial word order pattern and the availability of the Ezafe may be accounted for on historical grounds82 (cf. Samvelian, 2006). The enclitic Ezafe has probably its origin in a demonstrative-relative morpheme in Old Iranian. In Persian, it can be related to hya (tya), a demonstrative, linking the head noun to adjectival modifiers, to the possessor NP and also to a relative clause in Old Persian: (45-4) Kāra hya manā [Darmesteter (1883) from Samvelian (2006)] ‘my army; the army which is mine’ (46-4) kāsaka hya kapautaka [Meillet (1931) from Samvelian (2006)] ‘the blue stone’ (47-4) vivānam jatā utā avam kāram hya dārayavahauš xšāyahiyhyā. [Meillet (1931) from Samvelian (2006)] ‘Beat Vivâna and this army which declares itself as a proponent of the king Darius.’ It’s interesting to notice that hya (tya) is not a simple linker, but that it further has a demonstrative value. The demonstrative hya (tya) can function as a head by itself: (48-4) ima tya adam akunavam [Meillet (1931) from Samvelian (2006)] ‘This is what I did’ 82 However, there is no necessary correlation between the head initial word order pattern within the NP and the Ezafe. Some South-western Iranian languages, which are very close to Persian, dispense with the Ezafe. In different Kermanian dialects, for instance, the Ezafe, although available, due to a close and longstanding contact with Persian, is hardly ever used or is optional (Rebuschi, 2002). Similar facts are also observed in some North Western Iranian languages, such as Tâti dialects (Lecoq 1989) or in Southern Kurdish dialects (Fattah 2000). 118 Hya (tya-) becomes –i in Middle Persian and progressively looses its demonstrative value to end up as a simple linker. Contrary to Persian, Kurdish and Zazaki (cf. Larson & Yamakido, 2006) have still a so-called “Demonstrative Ezafe”, different from the affixal Ezafe, which functions as a demonstrative pronoun heading nominal phrases, as shown in the following examples. (49-4) yê dwê … yê sêyê EZ second EZ third “the second one, the third one” (Kurmanji Kurdish, from Larson & Yamakido, 2006) (50-4) kitêb-î min o hî to book-EZ my and EZ-your ‘my book and yours (Sorani Kurdish from Samvelian, 2006) Nowadays, when available, in Western Iranian languages the Ezafe generally links the head noun to its adjectival modifiers and to the possessor NP, as illustrated by the following examples for Persian taken from Kahnemuyipour (2000) and for Sorani Kurdish taken from Samvelian (2006): (51-4) lebâs-e sefid-e donya (Persian) dressEZ whiteEZ Donya “Donya white dress” (52-4) kirasêk-î hin-î Narmîn (Sorani Kurdish) dress-EZ blue-EZ Narmîn “a blue dress of Narmin’s” Furthermore, the complement of an adjective can also be introduced by the Ezafe: (53-4) Saloomeh mašqul-e kâr ast (Persian) Salomé occupied-EZ work be-3sg-pres ‘Salomé is busy working’ 119 The fact that adjectives can take the Ezafe is not surprising. Indeed, adjectives behave like nouns in many respects, so that in Persian, for instance, it is quite impossible to establish a distinct class of adjectives, and several items are indistinctly used as nouns or adjectives, depending upon the context (Lazard, 1992). A more interesting empirical observation is the use of the Ezafe with some prepositional heads. This occurs in Persian, for instance, as illustrated by the following examples: (54-4) a. barâ-ye Saloomeh for-EZ Salomé b. aleyh-e Saloomeh against-EZ Salomè “for Salomé” ; “against Salomé” It has been argued by Ghomeshi (1997) that prepositions occurring with the Ezafe are not in fact prepositions, but nouns. Real prepositions such as bâ “with”, az “from”etc., by contrast, never occur with the Ezafe (cf. also Chapter 5). This assumption is arguably appropriate for locative prepositions, such as zir “under”, pošt “behind”, which are originally nouns and display a range of nominal properties (Samiian, 1994; Larson and Yamakido, 2005): the phrase they head can function as the subject or the direct object of a sentence. It faces however serious problems when applied to items such as barâ “for”, alâraqm “despite”, and aleyh “against”, which have none of the distributional and morphological properties of nouns83. Consequently, the analysis of these items as nouns is exclusively motivated by the fact that they can be marked by the Ezafe. Anyway, not only is it unclear what such an analysis could gain from it, but also far more mysterious is the way it could work! Contrary to Ghomeshi (1997), and according to Larson and Yamakido (2005), I think that is possible to consider these items as prepositions, which implies that, with regard to the Ezafe, prepositions have, simply, two distinct subtypes: those which take the Ezafe and those which do not take the Ezafe (this fact will have some consequences for some further discussions raised in the following Chapter). 83 For instance, the constituents they head can never be the subject or the direct object of a sentence (cf. Samvelian, 2006). 120 Apart from its typical uses to introduce adjectives and the possessor NP, the Ezafe morpheme can introduce prepositional phrases and adverbial phrases. The possibility for prepositional phrases to be introduced by the Ezafe in Persian is illustrated by the following example takan from the novel Yek ruz mânde be eyd-e pâk by Z. Pirzâd and cited in Samvelian (2006) (55-4) ne-mitavânest-am tasmim begiram sobh-hâ-ye [PP bâ mâdar-râ] bištar dust dâr-am yâ sobh-hâ-ye [PP bâ kabutar-hâ-râ]. NEG-can.Past-1sg decision take-1Sg morning-pl-EZ with mother-Acc more like.Pres-1sg or morning-Pl-EZ with pigeon-pl-Acc “I could not decide whether I loved better the mornings with mother or theones with the pigeons” Once again, contrary to what has been claimed in some previous studies (Ghomeshi 1997, Larson and Yamakido 2005), there is no ban on the presence of PPs within the Ezafe domain, whatever be the type of the head preposition. The situation is more contrasted with respect to relative clauses (cf. paragraph 4.1). In Persian, only reduced relatives can be linked to the head noun by the Ezafe (cfr. Samvelian 2006, from which are taken the following examples. I recommend to refer to her work for a cross-linguistic survey and for more details84): (56-4) in javân-e [az Suis bar gašte] this young-EZ from Switzerland return “this boy returned from Switzerland” 84 Samvelian also observes that Kurdish, by contrast, allows for all relative clauses to be introduced by the Ezafe: (i) mirov-ê ku min dît-î (Kurmanji) man-EZ.M.SG that I see.PAST ‘the man whom I saw’ (ii) aw şâr-a-y (ka) dît-mân (Sorani) that town-DEF-EZ (that) see.PAS-1.P ‘The town that we visited’ Furthermore in the Kurmanji dialect, even a non-relative subordinate clause which is a dependent of a noun can be introduced by the Ezafe, as illustrated by the following example: (iii) bi xayâl-â [ko aw ji bâjêr darkati bûn] at imagination-EZ that they from city out were ‘Imagining that they were outside the city...’ [taken from Bedir Khan and Lescot, 1970] 121 (57-4) * ketâb-i-e ke ru-ye miz ast book-rel-EZ that on-EZ table be.PRES “The book that is on the table” In the following section I will try to outline, from a diachronical perspective, the analyses of the Ezafe-morpheme in the generative literature of the last two decades. 4.2.3 Restriction on Ezafe (Samiian, 1983) Vida Samiian (1983) is the first detailed study on Ezafe in Persian within a modern syntactic framework, namely X-bar theory. The empirical facts mentioned by Samiian have been taken up in subsequent works (Ghomeshi 1997, Kahnemuipour 2000, Larson and Yamakido, 2005), although they have been accounted for in a radically different way. In this section, I will consider different restrictions on the Ezafe construction pointed out by Samiian (1983), then, in the following one I will then give a detailed account of Ghomeshi (1997). Samiian (1983) points out two major types of restrictions on the Ezafe construction: the constituents occurring within the Ezafe domain are strictly ordered and constrained with regard to their distribution. In depth, according to Samiian, different elements linked by the Ezafe to the head noun occur in a fixed order: (58-4) ketâb-e târiz-e sabz-e bijarzes-e Saloomeh book-EZ history-EZ green-EZ without value-EZ Salome “Salome’s green history book without any value” In the example (58-4) the head noun is followed by an attributive noun, an adjective modifier, a prepositional modifier and a possessor NP: (59-4) Head Noun – Attributive noun – Adjective modifier – Prepositional Modifier – Possessor NP Although Samiian (1983) considers this order to be a strict one, it must be noted that the only real constraint concerns the placement of the possessor NP, which must occur in the final position within the Ezafe domain, any other position being excluded for the possessor (cf. also Kahnemuipour, 128 account for the Ezafe vowel before attributive adjectives (as I have argued above it is not clear why attributive adjectives need case). Following the principles of Case Theory (cf. Haegeman, 1996) and Stowell’s Case Resistance Principle (1981), Samiian assumes that non-caseassigners need case, whereas Case-assigners do not. The trigger for the insertion of Ezafe is, then, the lack of Case-assigning properties. Hence, Ezafe appears only on categories that cannot assign case, such as [+N] categories, and doesn’t appear on case-assigning categories, such as [-N]. Given this, a question arises: why most prepositions in Persian take their complement via Ezafe? Since verbs and prepositions are both case-assigners by virtue of their [-N] feature, the latter are not expected to need some special device for taking a complement. However, the fact remains that the only phrasal category where Ezafe is not found is the verb (but see paragraph 4.3 for an interesting empirical fact). Samiian divides the prepositions in Persian in two groups — those which do not take Ezafe (Class 1; C1) and those which take Ezafe (Class 2, C2). C1 prepositions possess all the properties associated with prepositions, including the ability to directly assign case. C2 Prepositions, on the contrary, exhibit some nominal properties including the inability to assign case. In order to explain the syntactic behaviour of C2 Prepositions, Samiian adopts the Neutralization Hypothesis for German adjectives, for which she cites the work of van Riemsdijk (1991). Under this proposal, German adjectives are neutralized in their [+N] feature, that is, they are specified only for the [+V] feature, rather than fully specified [+V,+N] elements. Consequently, as a [+V] category they are nondistinct from the [+V,-N] category, from verbs. This provides an explanation for the fact that adjectives and [-N] categories in German share some properties: at least, they all can assign case. Following the same line of reasoning, Samiian suggests that C2 Prepositionss are neutralized with respect to their [-N] feature, thus they have only the feature specification [-V]. The full paradigm of Persian lexical categories according to Vida Samiian is given as follows in (71-4). (71-4) N: [-V,+N] V: [+V,-N] A: [+V,+N] C1 P: [-V,-N] 129 C2 P: [-V]86 (Samiian 1994: 38) Building on Samiian (1994) and following the Larsonian DP-shell structure87 (cf. Larson 1991-2006), Larson and Yamakido (2005) suggest too that Ezafe is a case marker. However, their account differs from the one suggested by Samiian. Larson and Yamakido make the interesting assumption that C2 prepositions are nouns and thus, by eliminating the distinction between most Persian prepositions and nouns, they are able to provide a unified syntactic account for all nominal modifiers. They propose that nominal modifiers are generated as arguments of D post-nominally, in the position of relative clauses. As [+N] elements they require Case, hence, in English, they move up to get caselicensed by the Determiner. Persian, however, has at its disposal the Ezafe marker, which according to them is a special “device” for making Case available in the base position. Thus, Ezafe allows the underlying postnominal position of nominal modifiers to emerge, since they are case-licensed in their base position. In other words, Larson and Yamakido propose an analysis in which (predicative) nominal modifiers originate postnominally as inner arguments of D, which only afterwards combines with NP and a subsequently merged light determiner (d) attracts the D head, deriving D–N word order. Furthermore, DP modifiers are considered to be lowest complements of the head. Under this account, nominal modifiers, in all languages, originate as arguments of D in a post nominal position. Those DP-modifiers that do not have Case features to be checked, PPs and CPs for instance, remain in situ. 86 It’s interesting to notice that C2 Ps are left with only the feature [-V] which makes them unable to assign case, since according to Samiian only categories specified for [-N] are caseassigners. Therefore, C2 Ps have to make use of a special device, more specifically the dummy case assigner Ezafe, in order to be able to take complements. 87 Here is a brief explanation of Richard Larson’s DP theory. In a paper from 1991 Larson discussed the syntactic projection of DP from the standpoint of generalized quantifier theory, and argued that, under the latter, the most appropriate analogy is not between DP and CP/TP (cf. Abney, 1987; Szabolsci, 1983), but rather between DP and VP. Specifically, Larson suggests that: (i) DP can be understood as projecting arguments according to a thematic hierarchy that is parallel to (but different in role-content from) that found in VP, (ii) That Determiners sort themselves into intransitive, transitive and ditransitive forms, much like Verbs, and (iii) that nominal modifiers, including relative clauses, project in the DP very much like adverbial elements in VP. 130 On the contrary, those modifiers that bear Case features (for instance APs) are required to move to a site where they can check Case. The hypotesis of Larson and Yamakido is fascinating. The prenominal position is assumed to be such a site and there are languages such as Persian however that have in their D-system an item that can be inserted to check Case on [+N] determiner complements, allowing the latter to remain in situ: Ezafe would be the Casemarking device. The analysis of Ezafe as a Case marker, anyway, faces several problems. The data on which it relies is not well grounded. The major argument invoked by both Samiian (1994) and Larson and Yamakido (2005) to support this analysis is the fact that constituents such as PPs and relative clauses, which no not require to be Case-marked, are excluded from the Ezafe domain, but other authors (cf. Samvelian, 2006 for another critical survey) have empirically demonstrated that in Persian PPs headed by P1s, when they are modifiers, as well as adverbial phrases occur within the Ezafe domain.88 Furthermore, since under Larson and Yamakido’s account, Ezafe is supposed to enable [+N] modifiers, which otherwise should move to a prenominal position, to remain in situ, it is unclear why in other languages, for instance, reduced relatives are allowed to remain in situ without being Casemarked, while in West Iranian languages such a Case-marking device is necessary in order to let them remain in a post-nominal position. The view of Ezafe as a Case-marker becomes even more problematic when other West Iranian languages are taken into account. First, as reported in Samvelian (2006), in almost all groups of Kurdish dialects restrictive relative clauses may be introduced by Ezafe; In addition, in Kurmanjî dialects, which have generally maintained morphological case marking, the Possessor NP linked to the head noun by Ezafe appears in oblique case, as shown by the following example, taken from Samvelian (2006): 88 Furthermore, though relative clauses cannot be linked to the head noun by Ezafe in Persian, reduced relatives, on the contrary, can be introduced by Ezafe. This is shown by the following attested examples: (i) ...nazm-e ostân tavassot-e in javân-e [az suis bar gashte]RRC bishtar hâsel xâhad âmad order-EZ province by means-EZ this young-EZ from Switzerland back turn.PP more gained AUX come.PAS ‘(...) the province’s order will be better established by this young man came back from Switzerland’ (ii) aks-e [châp shode dar ruznâme]RRC aks-e râvi-e dâstân ast17 photo-EZ publication become.PP in newspaper photo-EZ narrator-EZ story be.PRES ‘The photo published in the newspaper is the photo of the story’s narrator’ It is unclear why constituents such as PPs and reduced relatives would require to be Casemarked in Persian. 131 (72-4) mashl-ah narmîn-ê house-EZ.Fem Narmin-Obl.Fem “Narmin’s house” Viewing Ezafe as a Case marker would imply that the possessor NP Narmin is case marked twice for the same function (cf. also Kahnemuyipour, 2000) and this is problematic within any theory of Case marking. Given this, either the Persian Ezafe is a radically different item from its Kurdish counterpart, or Ezafe is not a Case-marker, neither in Persian nor in other Iranian languages. The fact that PPs and reduced relative clauses are linked to the head noun by Ezafe (at least in Persian) strongly supports the latter conclusion. 4.2.5 Ezafe and movement (Kahnemuyipour, 2000) Kahnemuyipour (2000) provides an explanation for Ezafe insertion based on syntactic movement. He notes that if Ezafe were a marker inserted only to identify constituent-hood, as proposed by Ghomeshi (1996), then the order of the modifier and the noun would be irrelevant. However, there are cases in Persian where the adjective precedes the noun and no Ezafe is inserted (cf. (73-4)). Moreover, the Ezafe vowel is ungrammatical in this context. (73-4) a. gol-âb flower-water “rose-water” vs. ?âb-e gol b. bozorg-mard big-man “great man” vs. mard-e bozorg 'big man' c. ketâb-xune book-house “library” vs. ?xune-ye ketâb (Kahnemuyipour 2000:3) Kahnemuyipour takes this fact to suggest that the Ezafe construction is associated with syntactic movement and that the Ezafe vowel is the realization of a strong [Mod] feature borne by modifiers. He assumes a leftbranching structure and a prenominal Merge position for all noun-modifying elements. Their postnominal surface position, then, is derived by movement of the noun. Referring to Cinque’s (1990) proposal about the base position of adjectives in the noun phrase, namely in the specifier of functional phrases above the NP, and according to my proposal, independently developed in Franco (2004), Kahnemuyipour suggests that modifiers in Persian too head 132 functional projections above the noun phrase and, furthermore, that they bear the strong feature [Mod]. The noun, which also bears the feature [Mod], moves up and head-adjoins to the modifier, thus checking its [Mod] feature against the [Mod] feature of the modifier. The [Mod] feature is then morphologically realized as Ezafe on the noun. Kahnemuyipour doesn’t tackle the issue of those prepositions that take their complement via Ezafe. It remains unclear whether they are considered to originate below the Ground complement (cf. Franco, 2004) and then move up to head-adjoin to it, which would mean that the preposition is modified by its complement. To resume in some points his proposal, Kahnemuyipour (2000) suggests that the Ezafe marker is associated with syntactic movement and suppose for Persian: i) Right branching structure. ii) Prenominal Merge position for adjectives, modifying nouns, possessors, etc. iii) Basing on Cinque (1990) and Rubin (1997), he assumes that adjectives are located in the heads of functional projections ModP above NP. iv) The adjectives (and all elements that modify nouns) bear a feature [Mod]. v) The feature [Mod] is strong and triggers overt movement of the noun. vi) The noun moves up, head-adjoins to the adjective and checks the strong feature [Mod] against the [Mod] feature of the Adjective89. vii) The strong feature [Mod] is morphologically realized on the noun element by Ezafe. viii) In this way, the postnominal surface position is derived by successive head adjunction to check the strong feature [Mod]. The representations of the derivational stages of this analysis are shown below: (74-4) sag-e siah-e gonde dog-Ez black-Ez big “big black dog” 89 Note that the order of adjectives in Persian doesn’t appear to be as strict as this proposal would predict (cf. Franco, 2004). 133 a. ModP Adj0 ModP gonde [Mod] Adj0 NP siâh-e [Mod] N0 (CP) sag-e b. c. ModP Adj0 ModP Adj0j Adj0 tj NP gonde Ni0 Adj0 [Mod] ti (CP) sang-e siâh-e 134 4.2.6 An inflectional affix? (Samvelian, 2006) Another view of the Ezafe phenomenon is the one of Samvelian (2006) already cited here for many interesting cross-linguistic examples, who argue that the Ezafe in Persian could be regarded as an inflectional affix, attaching to a head noun and to some of its intermediate projections and marking them as awaiting a modifier or a single NP complement. Samvelian also demonstrate that the affixal analysis of the Ezafe could be applied to Kurmanji, though the Ezafe construction does not display the same range of properties as in Persian. The analysis outlined in Samvelian’s paper entails that the Ezafe particle, whose origins in Modern Persian can be traced back to the Old Persian relative/demonstrative hya/tya (see above paragraph 4.2.1), has undergone a process of reanalysis-grammaticalization, being thus reinterpreted as a part of the nominal inflection. Given its enclitic status (at least in Modern Persian), it could be assumed that the conflict arising from the requirement for two opposite directions of attachment – morpho-syntactic attachment to the right and phonological attachment to the left – has been resolved by the reanalysis of the Ezafe as a nominal inflectional affix, thus aligning morpho-syntactic attachment with phonological attachment. Such a view assumes that Ezafe particle has ceased to function as a relative particle, specializing as a device for nominal attribution. Samvelian think that, on the basis of distributional, prosodic and morphological criteria, two major sets of inflectional affixes within the Persian NP will be established. The members of the first set, i.e. the definite suffix –(h)e and the plural suffix hâ, may be considered as word-level inflectional affixes: they attach to the head (i.e. the noun) within the NP and cannot be separated from it by any other inflectional affix. Furthermore, they bear lexical stress and cannot have wide scope over the coordination of two nouns. The members of the second set, i.e. the Ezafe, the determiner –i and personal enclitics, are argued to be phrasal affixes: they occur at the right edge of nominal non-maximal projections, and are located after Set (1) inflectional affixes. See the examples below, taken from Samvelian (2006; p. 13). (75-4) a. in bâr âhang hamân-i bud ke pesar-e-ye film-e hendi 135 barâ-ye doxtar-e mi-zad this time melody same-encl be.Pas that boy-Def-EZ film-EZ Indian forEZ girl-Def Imp-beat.Pas “This time the melody was the same as the one the boy in the Indian film was playing for the girl”. b. * ketâb-i-e Maryam book-Ind-EZ Maryam “a book of Maryam” c. * hamsâye-ye negarân-e bachecehâ-yash-e Maryam neighbour-EZ worried-EZ child-P-PAF.3.P-EZ Maryam “Maryam’s neighbour who is worried about his children” In brief, Samvelian’s work suggests a morphological account of these restrictions in terms of position class morphology (Stump, 2001) where collections of items compete for realization in a single position. In this perspective, the data in 75-4) receive a purely surface-based account, in terms of constraints on affix stacking, with no need for syntactic constraints. If this account is appropriate, it is expected that such examples would become grammatical in case of a reordering of the constituents within the NP so that affix stacking is avoided. This indeed seems to be the case, given the contrast between (76-4a) and (76-4b): (76-4) a. *qahremân-e [rânde shode az mihan-ash]-e in roman hero-EZ drive.Pas become.Pas from homeland-pron.3.S]- this novel “the hero of this novel, who is driven away from his homeland” b. qahremân-e [az mihan-ash rânde shode]-ye in roman hero-EZ from homeland-PAF.3.S drive.Pas become.Pas]-EZ this novel “the hero of this novel, who is driven away from his homeland” 136 4.3 An interesting fact: the tense Ezafe of the Behdînî-Kurdish Geoffrey Haig (2005) An interesting empirical fact is reported in a paper by Geoffrey Haig, who takes a look at the Ezafe particle in an another Western Iranian language, the Behdînî dialect of Kurdish (BK), spoken in North Iraq. In this language, one exponent of the Ezafe has undergone a sort of different development: it is arguably no longer part of an NP (or DP) at all, but is now a particle with a particular tense/aspect value (feature), presumably part of a Tense projection in the clause. Typical examples that Haig has taken from MacKenzie (1961), are the following: MacKenzie (1961) (77-4) xusk-a min ya çuy-î sîk-ê sister-Ez.F 1S.Obl Ez.F go:Pas-Ptcpl market-Oobl “My sister has gone to the market” (78-4) got-ê ku šah-ê wan yê mir-î say:Pas-to.him that king-Ez.M 3Pl.Obl Ez.M die:Pas-Ptcpl “(He) said to him that their King had died.” The paradigm of forms available to this particle is identical to the paradigm found for the NP-based Ezafe; MacKenzie (1961) had already pointed out that the two are etymologically identical. Haig (2005) refers to the latter as the Tense Ezafe. In MacKenzie’s data, the tense Ezafe was largely restricted to co-occurrence with state or locational predicates, and participial verb forms. But the more recent fieldwork of Haig shows that the Ezafe particle now regularly occurs with finite present tense verb forms. With finite present tense verb forms the Tense Ezafe contrasts with clauses lacking such particles, indicating that the particle is indeed now part of the system of tense/aspect distinctions in the language, adding a particular sense of immediacy or current relevance: (79-4) Ez yê xwarin-ê çê.di-k-im 1Sg Ez.M meal-Obl Imp-do-Pres-1Sg “I am making/preparing a meal (right now)” (Haig, 2005) 137 vs. (80-4) Ez xwarin-ê çêdi-k-im 1Sg meal-Obl imp-do-Pres-1Sg “I (generally) make Kurdish food” (Haig, 2005) Then, Haig tries to reconstruct the most plausible diachronic scenario that led to the current situation, noting a striking (and inspiring, for me) parallels to the developments of copula elements from nominal linkers / pronouns in Hebrew and Mandarin90 (cf. Li & Thompson 1977 and, also, Den Dikken, 2005) and suggesting that the pathway concerned, involved the reanalysis of an originally hybrid element, the Old Iranian ancestor of the Ezafe (cf. paragraph 4.2.1), which conflated a C and a D projection, with the latter ultimately leading to the Tense-projection position of the BK Tense-Ezafe. The parallels I have found in Haig (I mean the ones concerning the “evolution of the copula”, especially in Chinese Mandarin), reminded me the encounters with Den Dikken (1995; 2005) and Moro (1993; 1997; 2000) theories about copula and, in some ways, have been an unconscious trigger for the present analysis of Ezafe. 4.4. Ezafe: Linking the Complex (network of the) Noun Phrases The Ezafe morpheme in our analysis, partly based on Den Dikken and Singhapreecha (2004) is a linker which, in a traditional path, is introduced in the course of a Predicate Inversion operation that inverts a noun-phraseinternal predicate with its phrasal subject. The constituent preceding the Ezafe morpheme could be interpreted, in fact, as the subject of the inverted predicate, which tells us that the word-order effect of Predicate Inversion is undone later in the derivation via raising of the subject to the Specifier of a higher functional projection (in a vein of a Cartographic approach) whose head is empty in the base, but gets filled by a linker as it raises. This is the most plausible account in a traditional perspective. 90 To explain how a demonstrative pronoun (shi) developed into a copula, Li & Thompson (1977: 420) suggest that this change came from a topic mechanism: “the subject pronoun which is co-referential with the topic in the comment of a topic-comment construction is reanalyzed as a copula morpheme in a subject predicate construction. As Li and Thompson (1977) have pointed out, the copula of Mandarin shi, originated from the demonstrative prononun, which was reanalyzed as a copula when the topic grammaticalized into a subject, the topic-comment construction thus becoming a subject-predicate contruction (cf. also Baker, 2003) 144 could have been - given an “edge status” to some items - between vertex-items and edge-items (intuitively determiners, verbs and complementizers would be edges; referential items vertexes), avoiding even the need of a set of relations (the relations would be subsumed in the Lexicon in this way). I think that the second option is irrational, and no language (natural or artificial) can be computed in that way. On the contrary, the first possibility is somewhat stimulating and I will try to make an articulated discussion of the consequences of a proposal of this kind in the following chapter. 145 5. Some theoretical issues and implications In this chapter I will lay some possible theoretical foundations, based upon the representational investigations made in the third chapter of this work, and assumed for of the analysis of Ezafe in the fourth chapter. 5.1. On the possible “Edgeness” of functional items As I have postulated in the last paragraph of the previous chapter - mainly based on the syntactic analysis of Persian - some items of the lexicon may be selected as edges instead of vertexes. The presence of the Ezafe morpheme as a linker in Persian noun phrases has been our first empirical example (interpreted in the vein of Den Dikken and Singhapreecha (2004)). This observation leads to a strong similarity with the representation of chemical structures (going far beyond the metaphor of Baker (2001)), as we will see (lying) pervasively in the present section. However, the very interesting fact is that this kind of observation could be assumed as the ground for an empirical (weak) generalization that has several consequences for the study of syntax in a graph theoretic perspective: The only items that may be structurally (not pragmatically) deleted in syntax are functional items (in example complementizers, prepositions or postpositions, determiners, copulas, linkers). Functional items are represented as edges. Thus, while a vertex (a lexical item) has necessarily a phonological content, an edge may be pronounced or not. I assume that, in a graph theoretic perspective, this is one of the fundamental observations concerning the empirical parametric differences across natural languages. Many examples will follow, with the aim to contextualize this simple, yet hopefully not trivial deduction. An unconscious trigger for this issue has been represented by the powerful Lewis formulas (definitely, a sort of directed graphs), commonly used for the schemata of chemical elements and molecules93. As we will see in depth as a metaphor, but as we may intuitively introduce here, lexical items (such as nouns, verbs or adjectives) may be seen as protons inside a nucleus, while functional items as electrons that exist outside of the nucleus in areas of high probability (but without spatial substance) called orbits, or shells. Again, 93 Notice that the discover of the deeply important chemical phenomenon of isomerism has been anticipated by the graphical notations of Crum Brown in the 1850s. 146 (fuzzy) topology plays a crucial role. Our hypothesis represents an economical way to explain some of the parametric (or, say, typological) differences of many languages (in example the absence (vs. presence) of determiners (as in the Slavic languages or Chinese Mandarin) or the absence (vs. presence) of copula (as in Russian or Chinese Mandarin). 5.1.1. On the different status of lexical vs. functional items In some way, a preliminary theoretical X’—Theory ground for my assumptions could be found in the seminal work of Fukui and Speas (1986). They suggested that the projections of lexical heads L° are recursively iterable L’s94, while the projections of functional heads F°s are non iterable F’ (these projections are driven in example by the discharge of F°s unique subcategorization feature onto the complements of F°s) and closed F’’ (being this projections driven by the discharge of a unique agreement feature onto a maximal projection that moves into the (necessarily) unique Spec position). Therefore, given the relativized X-Theory of Fukui and Speas (1986) any functional head has a lonely complement (sister to F° and dominated by F’)) and (at most) one specifier (sister to F’ and dominated by F’’), who agrees with F° and closes off the F’’ projection. 5.1.2. The ontology of functional items Functional items are necessarily non-productive closed classes. Taking their innermost nature as an issue for the discussion, it is possible to say that they lack a descriptive content. As argued in the groundbreaking work of Abney (1987): “Their semantic contribution is second order, regulating or contributing to the interpretation of their complement. They mark grammatical or relational features, rather than picking out a class of objects”. (Abney 1987: 65). I hope this point could be useful for my proposal. Functional items are edges in a graph based syntactic model just because they are relational items, without a specific weigh (or, at some level, a referential nature)95. 94 In Fukui ans Speas (1986) reinterpretation of X’-Theory these projections are driven, for example, by the discharge of theta features or subcategorization features. 95 I think that categories are not unitary notions, but emerge at the interface of different 147 Take Persian Ezafe as an example. As a functional morpheme in what is canonically assumed to be the nominal “domain”, such an item is useful to specify a reference. In more general terms, a noun provides a predicate, while a determiner “picks out a particular member’s of the predicate’s extension” (Abney 1987: 76). Furthermore von Fintel (1995) argues that functional morphemes have a logical-relational (for example permutation-invariant) semantics: “Logicality means insensivity to specific facts about the world, suggest a purely mathematical relationship” (von Fintel, 1995: 179) 5.2. Case: from Government and Binding to Graph Theory Even though only pronouns show overt morphological case in languages such as Italian (i.e io vs. me) or English (I vs. me) - taking as a starting point the Government and Binding (GB) Paradigm (cf. Chomsky, 1981; Haegeman, 1996) - it is widely assumed that all NPs have Case (called abstract Case) that matches the morphological Case that shows up on pronouns. Appeal is made to other languages with much richer case systems than Italian or English to “back up” this claim. Reviewing the basis of GB Case Theory we may observe that (in phrase structure terms): a. Nominative Case is assigned to the NP specifier of I[+fin] such as in (5-1). (5-1) components of human cognitive and communicative capacities (cf. Muysken, 2005). Categories are far more ambiguous than a commonsensical view could assume. Crosslinguistically, there is an infinite series of examples concerning categories underspecification, overlapping phenomena (cf. Comrie et al. 2005) and there is also an Indonesian language claimed to be (almost) mono-categorial (cf. Gil, 2001). Note also that in recent years, within the generative framework, Marantz and other researchers coming from Distributed morphology paradigm take a strong stand (in many, still unpublished papers) and propose that basically all lexical categories (nouns, verb, adjectives) are formed in a similar manner from category-less roots. This claim predicts that we should find productive triplets (noun, adjective, verb) of every category-less root (cf. Marantz, 2001). 148 b. Accusative case is assigned to the NP sister of V or P. The C[for] which is homophonous with the preposition for acts like P for Case assignment. Note that the subject of a non-finite clause could not receive Case from I[-fin] since only I[+fin] assigns Nominative Case. (5-2) c. Genitive Case, finally, is simply assigned to the specifier of N. What is the same about these positions that receive Case and the positions that assign Case? Chomsky observed that every maximal projection (=XP) that dominates the NP that receives Case also dominates the head that assigns it (at least, if we do not count the IP that intervenes between the C[for] and the NP). A (rough, avoiding the notion of c-command) definition of government comes from this observation: a GOVERNS b iff 149 a. a is a head [±N,±V] or I[+fin] or C[for], and b. every XP that dominates a also dominates b, and c. every XP (other than IP) that dominates b also dominates a. In this definition, a and b stand for particular categories. Proposition (a) requires that a is one of the heads N, V, A, P, I[+fin] or C[for]. Almost always, b is an NP, since are nouns that need Case (which is assigned under the government relation). Proposition (b) determines how high up the tree a head may govern: if every maximal projection above the head must also dominate the NP in question, then the NP must be below the maximal projection of the head (e.g. VP for V, IP for I[+fin]). Proposition (c), finally, provides the lower limit of government by not allowing the head to govern down into another maximal projection other than IP. Together, Proposition (b) and (c) establish locality constraints on the government relation for each head (cf. Rizzi, 1990; Manzini, 1992). The Case assignment rules in terms of government are quite simply ad elegant. Referring to the structures represented above we may say that: a. I[+fin] assigns nominative case to the NP specifier that it governs. b. N assigns genitive case to the NP specifier that it governs. c. V, P, C[for] assign accusative case to the NP that they govern. GB requires that all NPs must have Case at S-structure by the Case Filter (Cf. Haegeman, 1996). CASE FILTER: *NP if it does not have Case at S-structure. With one further assumption, we will have the motivation for Amovement in passive, unaccusative, and raising constructions (Cf. Zamparelli, 1995-2000). The famous Burzio’s Generalization (Burzio, 1986) states that predicates which do not assign a semantic role to their external argument cannot assign Case to their complement(s). This provides the answer to why the passive object must (cross-linguistically, can) move. In GB, passive verbs do not assign a semantic role to their external argument position, so they have lost the ability to assign Case to their complement. Therefore, the NP object cannot remain in place at S-structure and must move to a position where it can get Case: the specifier of I[+fin] where nominative case is assigned. The same observation accounts for the 150 unaccusative and raising constructions: the NP which cannot receive Case in its (Deep structure) position is the one which must move; and it may only move to a position which does assign Case, the specifier of I[+fin]. Now, it is interesting to notice that, contrary to GB assumptions, more recent researches (Bittner and Hale, 1996; Chomsky, 1998) assume structural Case as a syntactic feature that is incapable of inducing movement. Thus, Case gets checked (erased) as a result of structural factors that exist independently of Case itself. In the Minimalist Program, Chomsky argues that Casechecking is “ancillary” to other feature-checking mechanisms. I assume that, cross-linguistically, Case morphemes, revealing a relation between items in what we may continue to call a clause (or in a phrase: see Case marked adjectives), are functional elements that may be represented as edges. A trivial proof is that natural languages make extensive use of prepositions (other available functional elements) where Case morphology is not a given syntactic possibility. Thus, morphological Case on nominals is a common device to express the syntactic (and semantic) relationships between clausal constituents. However, it is important to remark that the languages of the world that use this strategy vary greatly with respect to the number of Case categories represented in their inflectional system.96 A problem, concerning the GB paradigm, is that if we assume a fine structure of the inflectional phrase (cf. Pollock 1989; Belletti, 1990), and according to Chomsky and Belletti we refine the observation in (5-1) saying that “NP is nominative if governed by Agr”, many cross-linguistical data, such as Japanese and Icelandic quirky subjects (Ura, 2000; Sigurdsson, 2002) or split-ergativity patterns (i.e. agreement vs. Case), such the ones in Georgian aorist (cf. Harris, 1981; Marantz, 1984; Comrie, 1989) raise objections. See the examples below: 96 The minimal Case paradigm contains two members, since paradigmatic relationships between word forms are ultimately based on binary oppositions (minimal pairs). This implies that whenever a language has an overtly marked Case category expressing a specific function, a corresponding zero-marked base form is counted as a Case (the “default case” or “direct case”) even if it has non-specific function describable in positive terms. In such instances, the base form receives its case status only by virtue of contrasting with a functionally (and formally) marked Case category. An example, given in the World Atlas of Language Structures (edited by Comrie et al. 2005), is Mapudungun (Araucanian; Chile), which has only one overt case suffix -mew ~ -mu expressing diverse oblique functions such as place, cause, and instrument. On the other vertex of the continuum there are languages which show very large paradigms. The languages with the largest paradigms, again according to the World Atlas of Language Structures, are Hungarian with (under some analyses) 21 productive cases, followed by Kayardild (Tangkic; Queensland, Australia) with 20, and Lak (Nakh-Daghestanian; eastern Caucasus, extensively studies at Max Plank Institute Department of Evolutionary Anthropology) with 19 cases. (Iggesen, 2005). 151 Japanese (5-3) a. Taroo-ni hebi-ga kowa-i (Ura, 2000) Taroo-Dat snake-Nom fearful-Pres “Taroo is fearful of snakes” b. Taroo-ni eigo-ga dekir-u Taroo-Dat English-Nom understand-Pres “Taroo understands English” As we may observe easily, dative NP in the examples above in (5-3a,b) behave as subjects (in example, they can bind subject oriented anaphors). A similar quirkyness is shared by Icelandic: Icelandic (5-4) Mér var hjàlpaò (Sigurdsson, 2002) Me-dat was helped “I was helped” (5-5) a. Akur leiddist Aki-dat bored “Aki was bored” b. Akur virtist hafa leiddist Aki-dat seem to-have bored “Aki seemed to have been bored” c. Viò töldum Akur hafa leiddist We believe Aki-dat to-have bored “We believe Aki to have been bored” Icelandic quirky Subjects, summarizing a survey made up by Sigurdsson (2002), have been tested for reflexivisation, ECM infinitives, Raising, Control etc. Some examples have been given in (5-5). Furthermore, a so-called split-ergativity due to the tense/aspect specification of the clause is found somewhere. Basic facts concerning the relation between Case and agreement are illustrated in the following examples. Georgian is 152 well-known for its split-ergativity of this kind97. In Georgian, for example, the aorist tense system demands ergativity and the present tense system demands accusativity (data are from Comrie, 1978 and Lyle, 1997, cited in Ura, 2005; cf. also Harris, 1981). Georgian (5-6) a. PRESENT SUBJ in transitive: Nominative + subject agreement SUBJ in unergative: Nominative + subject agreement SUBJ in unaccusative: Nominative + subject agreement OBJ in transitive: Accusative + object agreement b. AORIST SUBJ in transitive: Ergative + subject agreement SUBJ in unergative: Ergative + subject agreement SUBJ in unaccusative: Absolutive + subject agreement OBJ in transitive: Absolutive + object agreement Geogian (5-7) a. PRESENT a’. Student-i midis. Intransitive student-NOM go(PRES) ‘The student goes.’ a’’. Student-i ceril-s cers Transitive student-NOM letter-ACC write(PRES) ‘The student writes the letter.’ 97 Besides Georgian, Hindi and many other Indo-Aryan languages, Burushaski, Tibetan, Nepali, Samoan, etc. show similar split-ergativity due to the tense/aspect specification of the clause (see Comrie 1978; Dixon, 1994, and Palmer, 1994 for a list of such languages). 153 b. AORIST b’. Student-i mivida. Intransitive student-ABS go(AOR) ‘The student went.’ b’’. Student-ma ceril-i dacera. Transitive student-ERG letter-ABS write(AOR) ‘The student wrote the letter.’ (from Ura, 2005) Although it is a well-known fact that agreement in Georgian is extraordinarily puzzling and highly resistant to a systematic explanation, the data in (5-6) and (5-7) reveal the weakness of GB assumptions about Case, showing a discrepancy between abstract Case and Case morphology / agreement. In brief, the set of data presented here means that Case morphology is not a systematic reflex of abstract Case. We may argue that Case has not to be viewed under licensing, or filtering, but it is instead Case that interprets syntax and expresses a relational feature (at PF level). Anyway the feature values are largely self-explanatory. In languages lacking morphological case98, grammatical relations are expressed by word order and/or morphologically and prosodically independent function words (generally, prepositions and postpositions), and partly also by morphological devices on the verb. Here are some cross-linguistic examples concerning what, following and partially revising Comrie (1989), we may call the Alignment of (Structural) Case Marking: i. No morphological case Thai (5-8) phom kin khaw I eat rice “I eat rice” 98 Remember that in these languages, the default form receives its Case status by pairing it with functionally and formally Case-marked categories.