Taking a Free Ride in Morphophonemic Learning
Abstract
As language learners begin to analyze morphologically complex words, they face the problem of projecting underlying representations from the morphophonemic alternations that they observe. Research on learnability in Optimality Theory has started to address this problem, and this article deals with one aspect of it. When alternation data tell the learner that some surface [B]s are derived from underlying /A/s, the learner will under certain conditions generalize by deriving all [B]s, even nonalternating ones, from /A/s. An adequate learning theory must therefore incorporate a procedure that allows nonalternating [B]s to take a «free ride» on the /A/ →[B] unfaithful map.
Full text
Abstract As language learners begin to analyze morphologically complex words, they face the problem of projecting underlying representations from the morphophonemic alternations that they observe. Research on learnability in Optimality Theory has started to address this problem, and this article deals with one aspect of it. When alternation data tell the learner that some surface [B]s are derived from underlying /A/s, the learner will under certain conditions generalize by deriving all [B]s, even nonalternating ones, from /A/s. An adequate learning theory must therefore incorporate a procedure that allows nonalternating [B]s to take a «free ride» on the /A/ →[B] unfaithful map. Key words: chain shift, learning, morphophonemics, opacity, Optimality Theory; Arabic, Choctaw, German, Japanese, Rotuman, Sanskrit. * Special thanks to participants in the UMass summer 2004 phonetics/phonology group (SPe/oG) for their comments and suggestions. For additional comments, I am very grateful to Shigeto Kawahara, John Kingston, Joe Pater, Alan Prince, Bruce Tesar, and an anonymous reviewer. Catalan Journal of Linguistics 4, 2005 19-55 Taking a Free Ride in Morphophonemic Learning* John J. McCarthy University of Massachusetts at Amherst. Department of Linguistics Amherst, MA 01003 USA [email protected] Come on and take a free ride (free ride) (Winter 1972) Table of Contents 1. Introduction 2. Exemplification and explanation of the issue 3. The Free-Ride Learning Algorithm 4. Limitations of and Extensions to the FRLA 5. Conclusion Appendix: Further Free-Ride Examples References Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 19
1. Introduction Language learning is a central problem of linguistic research, and it is important to show that learners can in principle acquire the grammar of the ambient language from the limited data available to them. Formal studies of language learnability are concerned with determining, for a particular linguistic theory, whether and under what conditions learning can succeed. Ideally, any proposed linguistic theory will be accompanied by an explicit learning model from which learnability results can be deduced. Optimality Theory (Prince and Smolensky 2004) has been coupled with a learning model almost from the beginning. There is by now a very substantial body of work discussing the constraint demotion learning algorithm of Tesar and Smolensky (1994), extensions, and alternatives. Although much of this literature deals with phonotactic learning, it has begun to address morphophonemics: how do learners use data from phonological alternations to work out a grammar and a lexicon of underlying representations? (See, e.g., Dinnsen and McGarrity 2004; Ota 2004; Pater 2004; Tesar et al. 2003; Tesar and Prince forthcoming; Tesar and Smolensky 2000: 77-84; Tessier 2004.) Morphophonemic learning presents problems that are not met with in phonotactic learning. This article describes one such problem and offers a solution to it. The OT learnability literature standardly assumes that phonotactic learners posit identity maps from the lexicon to observed surface forms:1the learner hears [A], takes /A/ to be its underlying source, and seeks a grammar that will perform the identity map /A/ →[A]. If the learner never hears *[B], however, and this gap is nonaccidental, then his (or her) grammar must not map anything to *[B]. In particular, the /B/ →[B] identity map must not be sanctioned by the grammar. Instead, the grammar must treat a hypothetical /B/ input unfaithfully, mapping it to zero or to something that is possible in the target language. Morphophonemic learning is different. As learners begin to analyze morphologically complex words, they discover morphophonemic alternations for which the identity map is insufficient.2For example, once a learner becomes aware of paradigmatic relationships like Egyptian Arabic fihim/fihmu ‘he/they understood’ and sets up underlying /fihim/, he is then committed to the unfaithful map /fihim-u/ → fihmu. But what happens with nonalternating forms? Is the identity map still favored? In this article, I argue that nonalternating forms are sometimes derived by unfaithful maps as well. As I will show, there are examples with the following properties. Data from alternations indicate that some surface [B]s derive from underlying /A/s. Other evidence, some of it distributional and some involving opacity, argues that all surface [B]s, even the nonalternating ones, derive from underlying /A/s. I will propose a learning principle according to which learners who have discovered the /A/ →[B] unfaithful map from alternations will attempt to generalize it, projecting /A/ inputs for all surface [B]s, whether they alternate or not. In other 20 CatJL 4, 2005 John J. McCarthy 1. I will not attempt to list the many works that assume the identity map in phonotactic learning, nor will I attempt to identify the originator of this idea. 2. Alderete and Tesar (2002) argue that the identity map is insufficient even in phonotactic learning. Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 20
words, the nonalternating [B]s attempt to take a «free ride» on the independently motivated /A/ →[B] map.3The learner accepts this tentative hypothesis about underlying representations if it yields a grammar that is consistent (in the sense of Tesar 1997) and that is more restrictive (in the sense of Prince and Tesar 2004) than the grammar would be without this hypothesis. These two requirements, consistency and greater restrictiveness, are arguably sufficient to rule out inappropriate generalization of unfaithful maps.4 Section 2 and the appendix of this paper document the phenomenon: there are languages where alternations show that some [B]s derive from /A/s and further evidence shows that all [B]s must derive from /A/s. Section 3 presents and applies the proposed learning algorithm. Section 4 describes a limitation of this algorithm and proposes a modification to address this limitation. 2. Exemplification and explanation of the issue This section describes a single example, coalescence in Sanskrit. To establish the generality of the problem, however, this article also includes an appendix with several more examples. The examples in the appendix are somewhat more complex than Sanskrit, but they also establish that the scope of the learning problem goes well beyond the details of the Sanskrit analysis. In Sanskrit, alternations like those in (1) show that some surface long mid vowels [e] and [o] are derived by coalescence from /ai/ and /au/ sequences, respectively. (1) Sanskrit coalescence (de Haas 1988; Gnanadesikan 1997; Schane 1987; Whitney 1889) /tava indra/tavendra ‘for you, Indra (voc.)’ /hita upadaiʃa/hitopadeʃa‘friendly advice’ Coalescence occurs in both internal and external sandhi, and it is abundantly supported by alternations. The result of coalescence is a compromise in the height of the two input vowels. The result is long because the moras associated with the input vowels are conserved in the output. These are typical properties of coalescence cross-linguistically (de Haas 1988). Since all surface mid vowels are long in Sanskrit, all are at least potentially the result of coalescence. For many words with mid vowels, there are alternations that support a coalescent source. But it is unlikely that relevant alternation data will exist for all mid vowels, especially in the limited experience of the language learner. When presented with an [e] and no relevant alternations, what does the learner do? Taking a Free Ride in Morphophonemic Learning CatJL 4, 2005 21 3. I have borrowed the term «free ride» from Zwicky (1970). The situation Zwicky has in mind is quite different from what I am talking about: an analyst, instead of positing a rule /A/ →[C], posits a rule /A/ →[B] that is ordered before an independently motivated rule /B/ →[C], letting the output of one rule take a free ride on the other. 4. Ricardo Bermúdez-Otero has drawn my attention to Bermúdez-Otero (2004), which discusses a similar problem from the perspective of stratal OT. Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 21
(To simplify the exposition, I will only talk about [e] until late in section 3, when I will look at [o] as well.) A persistent bias toward the identity map would favor positing underlying /e/ for nonalternating [e]. There is a good reason, however, to think that learners actually adopt a free-ride strategy, deriving nonalternating [e] from /ai/. The reason: by deriving all [e]s from /ai/s, we explain why Sanskrit has no short [e]s. If all mid vowels are derived by coalescence and if the product of vowel coalescence is always a long vowel, then there is no source for short [e]. The traditional approach to cases like this involves positing a morpheme structure constraint (MSC) that prohibits underlying mid vowels, short or long. Classic MSC’s are not durable; they cannot restrict the output of rules. The rule of coalescence is therefore free to create (long) mid vowels. Short mid vowels are not observed in surface forms because they are barred from the lexicon by the MSC and no rule creates them. On this view, even nonalternating [e] has to be derived by coalescence because underlying /e/ is ruled out by the MSC.5This type of coalescence is said to be non-structure-preserving (de Haas 1988), since its output is something that is not allowed in the lexicon. In OT, richness of the base (ROTB) forbids MSCs, but an abstractly similar analysis is possible. Gnanadesikan (1997: 139-153) analyzes Sanskrit with a chain shift: underlying /ai/ maps to [e], while underlying /e()/ maps to [i()] or [a()] —it does not matter which. Gnanadesikan assumes a constraint set similar to (2). (2) a. IDENT(Vowel Height) (hereafter ID(VH)) For every pair (V, V´), where V is an input vowel and V´is its output correspondent, if V and V´differ in height, assign a violation mark. b. IDENT-ADJ(Vowel Height) (hereafter ID-ADJ(VH)) For every pair (V, V´), where V is an input vowel and V´is its output correspondent, if V and V´differ by more than one degree of height, assign a violation mark. c. *MID Mid vowels are prohibited. d. *DIPH Diphthongs are prohibited. e. UNIFORMITY (hereafter UNIF) No output segment has multiple correspondents in the input (≈no coalescence). For present purposes, I will assume that this is the entire constraint set and that the only unfaithful maps that are possible are coalescence, which violates UNIF, and changes in vowel height, which violate the IDconstraints. 22 CatJL 4, 2005 John J. McCarthy 5. The Alternation Condition of Kiparsky (1973) was specifically formulated so as not to exclude such an analysis. The Alternation Condition says, approximately, that no neutralization rule can apply to all instances of a morpheme. But coalescence in Sanskrit is not a neutralization rule if all [e]s are derived from /ai/. Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 22
Because mid vowels are prohibited in general, *MID must dominate ID(VH), so any input /e/ or /e/ drawn from the rich base will map unfaithfully to a high or low vowel: (3) *MID » ID(VH) In this comparative tableau (Prince 2002), the winning candidate’s violation marks are displayed as integers in the first row below the constraints. Subsequent rows compare the winner with various losers, which are preceded by a tilde. The W indicates that *MID favors the winner over this loser; the L says that ID(VH) favors the loser over the winner. (The loser’s number of violation marks is also shown in small print.) This tableau is telling us that winner-favoring *MID dominates loser-favoring ID(VH). Coalescence of diphthongs requires a ranking where *DIPH dominates both UNIF and ID(VH): (4) *DIPH » UNIF, ID(VH) The winning candidate in (4) incurs two marks from ID(VH) because each of its input vowels stands in correspondence with an output vowel of a different height. For Gnanadesikan, the key to analyzing Sanskrit is the constraint ID-ADJ(VH), since it blocks the *MID-satisfying map of /ai/ to [i] or [a] that would otherwise be expected. The idea is that fusing /ai/ to create a vowel that is either high or low establishes correspondence between one of the input vowels and an output vowel that is two steps away from it in height. ID-ADJ(VH) specifically excludes this map even though it offers satisfaction of *MID: (5) ID-ADJ(VH) » *MID As the subscripts indicate, the candidates in (5) are products of segmental coalescence, not segmental deletion. Therefore they owe allegiance, IDENT-wise, to both of their input parents. Taking a Free Ride in Morphophonemic Learning CatJL 4, 2005 23 /be()/ *MID ID(VH) →bi() or ba()l ~ be()1WL /va1-i2/ID-ADJ(VH) *MID → ve1,2 l ~ vi1,2 or va1,2 1 WL /va1-i2/*DIPH UNIF ID(VH) → ve1,2 12 ~va 1i21WL L Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 23
This analysis is a typical chain shift.6Underlying /ai/ becomes [e], whereas underlying /e/ and /e/ become something else, either a high vowel or a low one (it does not matter which). The constraint ID-ADJ(VH) blocks the fell-swoop map in which /ai/ coalesces to form a high or low vowel. In any chain shift /A/ →[B] and /B/ →[C], there must be a faithfulness constraint that is violated by the forbidden /A/ →[C] map but not by either of the permitted maps (Moreton 2000). ID-ADJ(VH) is such a constraint. No matter how Sanskrit is analyzed —with a MSC against mid vowels or the ranking g*MID » ID(VH)kin (3)— it presents a learning problem. Learners who hear nonalternating mid vowels will assume that they are derived from underlying mid vowels. How are they able to override this natural bias toward the identity map and instead learn a grammar that excludes underlying mid vowels with a MSC or, what amounts to almost the same thing, maps all of them unfaithfully? To see this problem in more concrete terms, we will look at how learning of an OT grammar is accomplished (there being no extant proposals for the learning of MSCs). The learner is armed with the learning procedure known as biased multirecursive constraint demotion (BMCD) (Prince and Tesar 2004). In multirecursive constraint demotion (Tesar 1997), learning proceeds on the basis of a comparative tableau like those given above, though with multiple inputs considered simultaneously. (See Prince 2002 on comparative tableaux in constraint demotion.) Constraints are ranked one at a time starting from the top of the hierarchy. A constraint is rankable if it favors no losers among the candidates that have not yet been accounted for. (A losing candidate has been accounted for if it is dispreferred relative to the winner by some constraint that has already been ranked.) Once some constraint has been ranked, the tableau with this new ranking is then submitted to the same procedure, recursively. The procedure terminates when all constraints have been ranked or when there are irreducible conflicts (as in (16) below), in which case no grammar can be found. The bias in BMCD favors grammars that are more restrictive (Prince and Tesar 2004; see also Hayes 2004, and Itô and Mester 1999). The bias says which constraints to rank first (=higher) when there are two or more constraints that favor no losers. By preference, markedness constraints are ranked higher. If there are no rankable markedness constraints, the ranking preference perforce goes to a faithfulness constraint. When several faithfulness constraints are available for ranking, the algorithm chooses the one that, by accounting for certain candidates, allows a markedness constraint to be ranked on the next recursive pass. (This is an intentional simplification; see Prince and Tesar 2004 for further details that are not important in the current context.) At the earliest stages of learning, there is presumably little or no awareness of morphological structure and no awareness of alternations, so only the phonotactics are being learned (on this assumption, see Hayes 2004; Pater and Tessier 2003; Prince and Tesar 2004). The phonotactic learner seeks a grammar that performs identity maps from the inferred lexicon to the observed surface forms. In the case of Sanskrit, 24 CatJL 4, 2005 John J. McCarthy 6. On chain shifts in acquisition, see Dinnsen (2004); Dinnsen and Barlow (1998); Dinnsen, O’Connor, and Gierut (2001). Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 24
the ambient language contains words with the vowels [i], [i], [a], [a], and [e]. A tableau is given in (6); all constraints are assumed to be unranked, though the algorithm works equally well if they have some ranking imposed on them already. (6) Sanskrit at onset of phonotactic learning (hypothetical examples) I will refer to this as a «support tableau», after Tesar and Prince (forthcoming), since it provides the support for the rankings found by BMCD. In this tableau, the initially rankable constraints (i.e., the constraints that favor no losers) are *DIPH, ID(VH), ID-ADJ(VH), and UNIF. The bias in BMCD favors ranking markedness constraints first, and of these only *DIPH is a markedness constraint, so it goes in the top rank of the constraint hierarchy. But because *DIPH favors no winners either, ranking it does not account for any losing candidates. On the next recursive pass through BMCD, no rankable markedness constraints can be found; *MID is the only remaining markedness constraint, and it favors the loser from the input /be/. Therefore, BMCD must rank a faithfulness constraint, and it selects the faithfulness constraint ID(VH) (see (7)), because adding ID(VH) to the ranking accounts for the loser from input /be/, and this will allow a markedness constraint, *MID, to be ranked on the next pass. (7) Support tableau (6) after second ranking pass Taking a Free Ride in Morphophonemic Learning CatJL 4, 2005 25 lexicon cands. *MID *DIPH ID(VH) ID-ADJ(VH) UNIF /ba/→ba (similarly /ba/) ~ bi 1W1W ~ be 1W1W /bi/→bi (similarly /bi/) ~ ba 1W1W ~ be 1W1W /be/ →be 1 ~ bi or ba L1W lexicon cands. *DIPH ID(VH) *MID ID-ADJ(VH) UNIF /ba/→ba (similarly /ba/) ~ bi 1W1W ~ be 1W1W /bi/→bi (similarly /bi/) ~ ba 1W1W ~ be 1W1W /be/→be 1 ~ bi or ba 1WL Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 25
All of the loser rows have been shaded to indicate that these candidates have been accounted for by ranking ID(VH). Since no unaccounted-for candidates remain, *MID is obviously free to be ranked. Because of the bias in BMCD, the remaining faithfulness constraints go below *MID, yielding the result in (8). (8) Sanskrit ranking after phonotactic learning *DIPH » ID(VH) » *MID » ID-ADJ(VH), UNIF The grammar in (8) is consistent with the observed vowel system of Sanskrit, but it is also consistent with a proper superset of Sanskrit in which short mid vowels are permitted: *[be]. (We can infer this because, if this grammar is given the input /be/, it yields the output *[be].) This grammar, then, is insufficiently restrictive, allowing things that Sanskrit does not allow. This is an instance of the Subset Problem (Angluin 1980; Baker 1979): learners presented with only positive evidence cannot proceed from a less restrictive grammar to a more restrictive one. The bias in BMCD is intended to address this problem by ensuring that learners always favor the most restrictive grammar that is consistent with the primary data, but BMCD is insufficient in this case (see also Alderete and Tesar 2002). The goal of this article is to explain how learners find the more restrictive grammar in cases like Sanskrit. Another way to look at the problem with (8) is that Sanskrit is harmonically incomplete relative to the markedness constraints provided: it lacks short mid vowels, yet there is no markedness constraint against them. (On harmonic (in)completeness, see Prince and Smolensky 2004: 164, 219-220.) As Gnanadesikan’s analysis shows, Sanskrit’s harmonically incomplete system can be derived by judiciously deploying the faithfulness constraint ID-ADJ(VH), but the phonotactic learner has as yet no evidence from alternations that show mid vowels being derived from underlying diphthongs. Worse yet, even with that evidence, the morphophonemic learner also fails to acquire the correct grammar, as we will see shortly. When a system is unlearnable because it is harmonically incomplete relative to a set of markedness constraints, there is a straightforward solution: make it harmonically complete, and therefore learnable with BMCD, by enlarging the set of markedness constraints. If the universal constraint component CON includes a constraint against short mid vowels, then BMCD will without difficulty find a grammar that is no less restrictive than Sanskrit requires. Such a constraint is not implausible as a narrowly tailored solution for Sanskrit,7but it is no help in dealing with the larger problem of learning similar systems. In the appendix of this article (see also McCarthy 2004), I document additional cases of the same learning problem and show for each of them that they cannot be solved by the expedient of introducing an additional markedness constraint. Because it is simple and familiar, I will stick with the Sanskrit example to exemplify the proposal made here, but bear in mind that any putative alternative solutions, to be of any real interest, need to deal with the cases in the appendix as well. 26 CatJL 4, 2005 John J. McCarthy 7. I am grateful to John Kingston for raising this issue. Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 26
As language development proceeds, the Sanskrit learner begins to attend to morphological structure and this leads, by a process I will not be discussing (though see the references in section 1), to the discovery of unfaithful maps from input to output. Let us say, for instance, that the learner has realized that the isolation forms tava and indra are components of the phrase tave ndra. The learner adds this new information about /ai/ →[e] unfaithful maps to the support tableau, yielding (9). (9) Sanskrit support tableau after some morphological analysis Three comments on this tableau: (i) Recall that I am making the simplifying assumption that there is coalescence but no deletion (because MAX » UNIF). The correspondence relations in the /va1+i2/ rows presuppose this: the vowels in all candidates bear the indices of their input correspondents /a1/ and /i2/. ID(VH) assigns a mark for every pair of corresponding segments that differ in height. The candidate [ve1,2] has two such pairs, (a1, e1,2) and (i2, e1,2), whereas [vi1,2] has only the pair (a1, i1,2). Unlike [ve1,2], however, [vi1,2] violates ID-ADJ(VH) because the pair (a1, i1,2) contains corresponding vowels that differ by two degrees of height and not just one. (ii) This tableau contains the /be/ →[be] identity map because there are forms that never alternate and/or forms whose alternations have escaped the learner’s experience or attention. This lingering identity map is the crux of the learning problem: because it is there in the tableau, BMCD continues to find a grammar that is insufficiently restrictive. (iii)This tableau preserves the ranking (8) that was acquired during phonotactic learning. With the added information about /va1+i2/ →[ve1,2], this ranking is now incorrect —the bottom row has loser-favoring constraints dominating the highest-ranking winner-favoring constraint. Because of this discrepancy between the ranking and the result, a new round of BMCD is required. Taking a Free Ride in Morphophonemic Learning CatJL 4, 2005 27 lexicon cands. *DIPH ID(VH) *MID ID-ADJ(VH) UNIF /ba/→ba (similarly /ba/) ~ bi 1W1W ~ be 1W1W /bi/→bi (similarly /bi/) ~ ba 1W1W ~ be 1W1W /be/→be 1 ~ bi or ba 1WL /va1+i2/→ve1,2 21 1 ~ va1i21WL L L ~ vi1,2 or va1,2 1LL 1W1 Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 27
(16) Support tableau (15) after first ranking pass But now it is clear that *VCDOBST and ID(voice) place antithetical demands on ranking, since both favor some unaccounted-for losers. This conflict is irreducible, so no further ranking is possible. BMCD has failed to find a grammar for German under the specific assumptions about underlying representations embodied in (15). This failure is unsurprising: devoicing is a neutralization process in German, since the [d]/[t] contrast is maintained in onset position, so there is no hope of deriving all [t]s from /d/s. Because BMCD fails to find a grammar, this new support tableau is discarded, as is the hypothesis it embodies about the German lexicon. Instead, the original support tableau (14) is correctly retained by the learner.11, 12 Taking a free ride on a neutralization process is a bad choice. The FRLA requires learners to consider this choice, but through BMCD it also supplies a way of rapidly detecting the error and recovering from it. 34 CatJL 4, 2005 John J. McCarthy 11. If the learner of German has insufficient phonotactic data at the point when the FRLA is activated, then he is in danger of reaching the wrong conclusion about the lexicon. Consider, for example, the effect of removing /daŋk/ from the support tableau, so there are no words in the support with surface syllable-initial voiced obstruents. In that case, a grammar —the wrong one— will be found by BMCD. (I am grateful to Anne-Michelle Tessier for raising this point.) In reality, there is probably nothing to worry about. Morphophonemic learning requires prior morphological analysis, whereas phonotactic learning can precede all morphological analysis. Therefore, morphophonemic learning is unlikely to begin before phonotactic learning is already quite far advanced. It is therefore reasonable to assume that the morphophonemic learner has already accumulated data that are sufficiently representative of the language’s phonotactic possibilities. 12. The German example indicates the need to say more about how FRLA interfaces with other aspects of learning. Bruce Tesar points out that the two losers from input /tat/ in support tableau (14) contribute no information about ranking, since no constraint favors them; hence, the standard approach to identifying informative losers, error-driven learning (Tesar 1998), would not include them in the support. Since these losers are informative in tableau (15), Tesar suggests that FRLA needs to apply error-driven learning to τnto find any losers that were uninformative in τobut have become informative in τnas a result of the change in underlying representations. lexicon cands. *VCDOBSTCODA *VCDOBST ID(voice) /daŋk/→daŋk 1 ~ taŋk L1W /dad/ →tat 2 ~ dad 1W2WL ~ dat 1W1L /rad/→rat 1 ~ rad 1W1WL Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 34
(ii) A grammar is found, but the ranking is unchanged. In this situation, nothing is gained in restrictiveness by taking the free ride. I assume that the lingering bias toward the identity map disfavors this use of the free ride. (See footnote 13 on why this is only an assumption rather than a solid fact.) As an illustration of this possibility, imagine a language Sanskrit´that is identical to real Sanskrit except that the inventory includes short mid vowels. Suppose the learner of Sanskrit´is in the early stages of morphophonemic analysis and has just discovered the coalescence alternation, so = {(ai, e)}. The pre-free-ride support tableau τois given in (17). The FRLA produces the support tableau (18), which is Sanskrit´with all (e, e) identity maps replaced by (ai, e). (17) Sanskrit´without free ride – τo Taking a Free Ride in Morphophonemic Learning CatJL 4, 2005 35 lexicon cands. *MID *DIPH ID(VH) ID-ADJ(VH) UNIF /ba/→ba (similarly /ba/) ~ bi 1W1W ~ be 1W1W /bi/→bi (similarly /bi/) ~ ba 1W1W ~ be 1W1W /be/→be 1 ~ bi or ba L1W /be/→be 1 ~ bi or ba L1W /va1+i2/→ve1,2 12 1 ~ va1i2L1WL L ~ vi1,2 or va1,2 L1L1W1 Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 35
(18) Sanskrit´with (ai, e) free ride – τn These two support tableaux, old and new, get identical grammars from BMCD: g*DIPH » ID-ADJ(VH) » ID(VH) » *MID » UNIFk. Nothing is gained in restrictiveness by deriving nonalternating [e] from /ai/, since the lingering /e/ →[e] identity map in (18) will not allow *MID to dominate ID(VH). In situations like this, where restrictiveness is not at stake, we may surmise that the learner does not take the pointless free ride. Therefore, the original support tableau (17) is retained and (18) is discarded.13 In Sanskrit´, nonalternating [e ] is derived from /e /, not /ai/, because [e] exists and is derived from /e/. (iii) A more restrictive grammar is found. This is the situation in real Sanskrit. Assume that the morphophonemic learner has already arrived at the stage represented by the support tableau (19), ranked according to (12). 36 CatJL 4, 2005 John J. McCarthy 13. It is only an assumption rather than an established fact that learners choose (17) over (18), favoring the identity map when all else is equal. Ordinary linguistic data is unhelpful in such situations, though other evidence may be relevant. For related discussion, see e.g. Kiparsky (1973) or Dresher (1981) on the Alternation Condition. lexicon cands. *MID *DIPH ID(VH) ID-ADJ(VH) UNIF /ba/→ba (similarly /ba/) ~ bi 1W1W ~ be 1W1W /bi/→bi (similarly /bi/) ~ ba 1W1W ~ be 1W1W /be/→be 1 ~ bi or ba L1W /ba1i2/→be1,2 12 1 ~ ba1i2L1WL L ~ bi1,2 or ba1,2 L1L1W1 /va1+i2/→ve1,2 12 1 ~ va1i2L1WL L ~ vi1,2 or va1,2 L1L1W1 Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 36
(19) Sanskrit without free ride – τo At this point, = {(ai, e)}. A copy of (19) is made, substituting the (ai, e) unfaithful map for all (e, e) identity maps. The table with this free ride is shown in (20). (20) Sanskrit with (ai, e) free ride – τn BMCD is applied to the new support tableau (20), yielding the ranking in (21). Taking a Free Ride in Morphophonemic Learning CatJL 4, 2005 37 lexicon cands. *DIPH ID-ADJ(VH) ID(VH) *MID UNIF /ba/→ba (similarly /ba/) ~ bi 1W1W ~ be 1W1W /bi/→bi (similarly /bi/) ~ ba 1W1W ~ be 1W1W /be/→be 1 ~ bi or ba 1WL /va1+i2/→ ve1,2 211 ~ va1i21 WLLL ~ vi1,2 or va1,2 1 W1LL 1 lexicon cands. *DIPH ID-ADJ(VH) ID(VH) *MID UNIF /ba/→ba (similarly /ba/) ~ bi 1W1W ~ be 1W1W /bi/→bi (similarly /bi/) ~ ba 1W1W ~ be 1W1W /ba1i2/→be1,2 211 ~ ba1i21 WLLL ~ bi1,2 or ba1,2 1 W1LL 1 /va1+i2/→ve1,2 211 ~ va1i21 WLLL ~ vi1,2 or va1,2 1 W1LL 1 Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 37
(21) Ranking from τn(20) *DIPH » ID-ADJ(VH) » *MID »ID(VH), UNIF Since a grammar has been found, it needs to be compared for restrictiveness with the grammar derived from the original support tableau (19). We already saw that ranking in (12), and it is repeated in (22): (22) Ranking of τo(19) *DIPH » ID-ADJ(VH) » ID(VH) » *MID » UNIF These grammars differ in restrictiveness. The grammar in (21) for the new support tableau (20) is more restrictive because it contains the ranking g*MID »ID(VH)k, which effectively denies faithful treatment to input mid vowels. The grammar in (22) for the old support tableau (19) has the opposite ranking of these two constraints and thereby permits input mid vowels to emerge faithfully. This grammar describes a proper superset of Sanskrit. The rankings in (21) and (22) are competing descriptions of Sanskrit based on different assumptions about the lexicon. Both grammars are consistent with the primary data, but (21) is more restrictive. Formally, the FRLA favors (21) because (21) has a higher r-measure (see (13)): the r-measure of (21) is 5, while the r-measure of (22) is 4.14 Since (21) is more restrictive, the learner adopts (21)’s support tableau (20) and the attendant hypotheses about underlying representations. In the discussion thus far, I have pretended that a single unfaithful map, learned from alternations, is sufficient with the FRLA to lead learners to the more restrictive grammar (21). This is a simplification; Sanskrit learners cannot get to (21) until contains at least two unfaithful maps, (ai, e) and (au, o). The reason for this is that taking a free ride on only one of these unfaithful maps will leave the other identity map in the support tableau. For instance, if = {(ai, e)}, then the support tableau in which (e, e) has been eliminated will still contain some (o, o) identity maps. And the presence of (o, o) in a tableau is sufficient to force ID(VH) to dominate *MID, as in the less restrictive grammar (22). This matter is already addressed in the FRLA. The learner constructs a support tableau for each of the subsets of (including itself, of course). If the learner has not yet encountered or attended to alternation data that support the (au, o) unfaithful map, then will not yet contain that map and the less restrictive grammar (22) will persist. Eventually, when the learner does get a handle on this other unfaithful map, he will be presented with a support tableau that, by virtue of two free rides, contains no (e, e) or (o, o) identity maps, and BMCD will find the more restrictive grammar in (21). 38 CatJL 4, 2005 John J. McCarthy 14. The r-measure of (21), for instance, is the sum of the number of markedness constraints that dominate ID-ADJ(VH) (that is, 1) and the number of markedness constraints that dominate ID(VH) and UNIF (that is, 2 each). The maximum r-measure for this constraint set is 6. Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 38
The FRLA is defined in this way so that grammar learning itself, rather than some special extragrammatical mechanism for data reduction, is responsible for discovering the generalization that all surface mid vowels derive from diphthongs. By itself, an unfaithful map is the meagerest sort of generalization: (A, B) pairs the input segment (or string) /A/ with output [B], abstracting away only from the actual form(s) in which this pairing occurs (see footnote 9). Unfaithful maps do not include information about the context in which they occur, nor do they state higher-order linguistic generalizations —e.g., there is no procedure for generalizing the (ai, e) map to all mid vowels. Rather, the capacity for learning this or any other linguistic generalization is reserved to grammar learning by BMCD. In this way, we avoid an unwelcome redundancy, where the learning mechanism preprocesses the data in a way that is closely attuned to what the constraints are looking for. A final point. Once the learner has successfully adopted a free ride, he ought to use this knowledge to figure out the underlying representations of newly-encountered words. A free ride is possible only when unfaithful mappings are biunique (that is, invertible): whenever the learner hears a novel word with [e], he can safely infer underlying /ai/ without waiting to hear other members of the paradigm. As yet, however, I have no concrete proposal for how learners might exploit this undoubtedly useful information. This section concludes with a formal statement of the FRLA. After the definitions in (23), (24) gives the algorithm in Algol/C-inspired pseudocode. (In definition (23), the distinction between phonological strings and their concatenative decompositions has been elided for clarity. In (24), the /* */ annotations surround comments.) (23) Definitions a. (Support) tableau A (support) tableau τis a set of ordered 6-tuples (input, winner, ℜ(input, winner), loser, ℜ(input, loser), tableau-row), where «input» is an underlying representation, «winner» and «loser» are output candidates, ℜ(input, winner) and ℜ(input, loser) are their respective correspondence relations with the input, and «tableau-row» is an ordered n-tuple of elements chosen from {W, L, Ø}, with one element for each constraint in CON. b. Unfaithful map in τ An unfaithful map in the (support) tableau τis a member (αi, βi) of ℜ(input, winner) in some member of τ, where αiℜβiviolates some faithfulness constraint in CON. The set of all unfaithful maps in a tableau τis designated by τ. Taking a Free Ride in Morphophonemic Learning CatJL 4, 2005 39 Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 39
(24) Free-ride learning algorithm Given: an initial support tableau τand a set τof all unfaithful maps in τ. FRLA( τ , τ ) τ o:= τ /*initialize result*/ For every pi∈℘( τ )/*℘( τ )=powerset of τ */ τ n:= Subst( τ o, pi) /*substitute unfaithful maps in subset pifor identity maps in τ o*/ if BMCD( τ n) = undefined /*if BMCD fails, go to next pi*/ continue else if R-measure(BMCD( τ n)) > R-measure(BMCD( τ o)) τ o:= τ n/*if τ nhas a higher r-measure, then it is now the one to beat*/ endif endfor return( τ o) Remarks: °°FRLA(τ, ) accepts as input a support tableau with a list of unfaithful maps contained in it. It returns a support tableau that may or may not be different from its input. °°Subst(τ, ) accepts as input a support tableau and a list of unfaithful maps. For every unfaithful map (α, β) in , it locates any winners in τwith (β, β) identity maps (i.e., it checks whether (β, β) ∈ ℜ(input, winner)). The morphemes underlying these winners are altered by substituting αfor β. As in «surgery» (Tesar et al. 2003), these changes are carried over to all forms in τthat contain the affected morphemes, and constraint assessments are updated to reflect the altered inputs. °°BMCD(τ) is defined by Prince and Tesar (2004). It accepts a support tableau as input and returns a constraint hierarchy or undefined if no hierarchy can be found. °°R-measure( ) is defined by Prince and Tesar (2004) and in (13) above. It accepts as input a constraint hierarchy (with constraints tagged as markedness or faithfulness), and it returns a positive integer or zero. The complexity of this learning algorithm is, in the worst case, proportional to the cardinality of ℘( ) —that is, 2N, where N is the number of unfaithful maps in . Exponentially increasing complexity looks bad, but in this case it is not as bad as it seems. As the learning problem scales up from examples of single processes to whole languages, or as the learner processes more alternation data, N grows very slowly. Most segments are mapped faithfully most of the time; alternationsupported unfaithful maps are more eye-catching, but they are also very much the exception. There is no fixed upper limit on N, but in practice languages with more than, say, 20 distinct alternation-supported unfaithful maps are probably not common. There is, then, little danger of a combinatorial explosion in more realistic learning situations than have been contemplated here. 40 CatJL 4, 2005 John J. McCarthy Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 40
Furthermore, there may be ways of improving the FRLA’s efficiency in actual practice. The various members of ℘( ) are not equally likely to yield a grammar on their respective passes through the main loop of FRLA(τ, ). Because finding a grammar is roughly equivalent to finding a linguistically significant generalization, grammars are more likely to be found for free rides on those subsets of that (i) violate the same faithfulness constraint(s) and (ii) are maximal subject to (i). For example, the Sanskrit learner would be well advised to consider pi={(ai, e), (au, o)} on an earlier pass through the loop than pi={(ai, e)} or pi={(au, o)}, since (ai, e) and (au, o) violate the same faithfulness constraint, UNIF. Once the free ride has been found, FRLA continues in search of other free rides, but the complexity of the problem is reduced significantly: because there are no remaining (e,e) or (o, o) identity maps in the support tableau, it is no longer necessary to examine members of ℘( ) that include maps with [e] or [o] as output.15 4. Limitations of and Extensions to the FRLA The FRLA has a significant limitation: it works for across-the-board free rides, but it does not work for contextually restricted free rides. Colloquial Arabic word-final vowels are a case in point (for further analytic details, see McCarthy 2004). Word-final vowels are always short, though vowel length is contrastive in other positions. Alternations like [œjilqa]~[œjilqani] ‘he will find ~ he will find me’ are straightforwardly accounted for under the assumption that some word-final short vowels are derived from underlying long vowels: the mapping of unsuffixed /jilqa/ to [œjilqa] is a familiar effect of NONFINALITY (NONFIN) and WEIGHT-TO-STRESS (WSP), which, by dominating ID(long), rule out *[jilœqa] and *[œjilqa], respectively. The naive expectation is that Arabic would have a contrast between final long vowel stems like /jilqa/ and final short vowel stems like hypothetical /jitba/. The contrast would be neutralized word-finally —[œjilqa] and [œjitba]— and preserved before suffixes —[jilœqani] and [œjitbani]. In reality, though, there is no such contrast: no stems behave like hypothetical *[œjitba]~[œjitbani], with a stem-final vowel that is short before suffixes. Every word-final vowel, although short on the surface, must be derived from an underlying long vowel, since all word-final short vowels alternate with long vowels when suffixed. Underlying stem-final short vowels served up by the rich base are presumably deleted, given the language’s propensity for syncope. Deletion of underlying final short vowels is accomplished by the ranking gFINAL-C » MAX(V)k(see (25)). (FINAL-C requires every word to end in a consonant.) Final underlying long vowels are protected from deletion, however, by MAX(V) (see (26)), for which there is solid support elsewhere in the language.16 Taking a Free Ride in Morphophonemic Learning CatJL 4, 2005 41 15. Bruce Tesar raises a related issue: could there be effects of the order in which the unfaithful mappings in are tried out? Schematically, if = {(A, B), (C, D)}, is it possible for the discovery of a free ride on (A, B) to block the discovery of a free ride on (C, D)? I have been unable to construct an example with this property, but of course this is no guarantee that the problem never arises. 16. On why there are no stems like *[œjitb]~[œjitbani], see McCarthy (2004). Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 41
(25) FINAL-C » MAX(V) (hypothetical example) (26) MAX(V) » FINAL-C Colloquial Arabic therefore has a chain shift, but a chain shift that is limited to word-final position: /V#/ →[V#] and /V#/ →[Ø#]. Alternation data tell the learner that some [V#]s derive from /V#/s, and the FRLA ought to get the learner over the next hump to conclude that all [V#]s derive from /V#/s, even when they occur in nonalternating, never-suffixed stems like [œhuwwa] ‘he’. If the learner fails to get over that hump, then he will acquire a grammar that maps /huwwa/ faithfully to [œhuwwa]. This less restrictive grammar includes the ranking gMAX(V)»FINAL-Ckand wrongly allows stems like *[œjitba]~[œjitbani]. Unfortunately, the FRLA in its present form is not up to this task. It can deal with an across-the-board free ride like Sanskrit’s, but not with a contextually limited one like Arabic’s. Alternation data tell the learner that there is a (a, a) unfaithful map in the language. But replacing all (a, a) identity maps with (a, a) maps is a gross error: it would entail deriving [œkatab] ‘he wrote’ from /katab/, and that is impossible because vowel length is contrastive in nonfinal open syllables, as shown by examples like [œkatab] ‘wrote’ and [œkatib] ‘writing’. What is required is a way for learners to discover contextually limited free rides like the one in Colloquial Arabic: all (a, a) identity maps should be replaced by (a, a), but only word-finally. Identity maps in other positions must be left intact. There are two ways to fix this problem with the FRLA. One way is to greatly increase the power of the currently quite limited theory of unfaithful maps, so that they can be distinguished by the contexts in which the maps occur. For reasons already discussed (see the end of section 3), this seems like a poor idea: it requires a special auxiliary learning mechanism, independent of the grammar, that is more or less able to reproduce the interactional possibilities of NONFIN and WSP. The contexts in which unfaithful maps occur are defined by the grammar, and any special learning mechanism charged with discovering and recording those contexts would simply duplicate the effect of the grammar. 42 CatJL 4, 2005 John J. McCarthy /jitba/ FINAL-C MAX(V) →œjitb 1 ∼œjitba 1WL /jilqa/MAX(V)FINAL-C →œjilqa 1 ∼œjilq 1WL Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 42
The other way of improving the FRLA is to provide it with a wider range of hypotheses about the lexicon. One of those hypotheses is a lexicon where only word-final [a]s take a free ride on (a, a). That is, the learner must consider lexicons in which various subsets of the (a, a) identity maps in the support tableau have been replaced by (a, a). These possibilities are listed in (27) for a dataset that contains a form with underlying /a/ determined from alternations ([œjilqa]~[jilœqani]), two forms with underlying nonfinal /a/ ([kiœtab], [œkatib]), and two forms with nonalternating surface [a] that might be eligible for a free ride ([œhuwwa], [œkatab]). The input segments that take a free ride are underlined. (27) Hypothesized lexicons a. /jilqa/, /kitab/, /katib/, /huwwa/, /katab/ b. /jilqa/, /kitab/, /katib/, /huwwa/, /katab/ c. /jilqa/, /kitab/, /katib/, /huwwa/, /katab/ d. /jilqa/, /kitab/, /katib/, /huwwa/, /katab/ e. /jilqa/, /kitab/, /katib/, /huwwa/, /katab/ f. /jilqa/, /kitab/, /katib/, /huwwa/, /katab/ g. /jilqa/, /kitab/, /katib/, /huwwa/, /katab/ h. /jilqa/, /kitab/, /katib/, /huwwa/, /katab/ Actual surface forms [œjilqa], [kiœtab], [œkatib], [œhuwwa], [œkatab] Lexicon (27a) has no free rides. Lexicon (27h) has an across-the-board free ride: all [a]s everywhere are derived from /a/. Lexicon (27e) is the desired result: only word-final [a] takes a free ride on the (a, a) unfaithful map. Other lexicons represent other choices about which [a]s to grant a free ride to. When presented with this wider range of possible lexicons, the FRLA will select the correct one, (27e). The short explanation is that BMCD cannot find a grammar for (27b, c, d, f, g, h), and (27e) is preferred to (27a) because (27e) allows for a grammar with a higher r-measure. The long explanation is that no grammar can be found for (27b, c, d, f, g, h) because no constraint ranking can derive [œkatab] from /katab/, /katab/, or /katab/ while also preserving the long vowels of /kitab/ and /katib/. For example, (28) is a support tableau that presupposes the lexicon in (27h), where all [a]s are derived from /a/. Taking a Free Ride in Morphophonemic Learning CatJL 4, 2005 43 Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 43
surface [ŋ] to explain why there is no rendaku voicing in [sakatoŋe]. For example, one might assume in turbid fashion that the [+voice] feature of underlying // remains present in surface structure, though unpronounced, and its presence is sufficient to block rendaku. The problem with this analysis is that there is no guarantee that surface [ŋ] derives from //; it might simply derive from /ŋ/, in which case rendaku would not be blocked. In fact, the analysis wrongly predicts a contrast between medial [ŋ]s that do and do not block rendaku, the former derived from /g/ and the latter from /ŋ/. There is no such contrast, however. The MSCbased analysis and Itô and Mester’s stratal approach do not predict this contrast because all surface []s and [ŋ]s derive from earlier []s. Gnanadesikan’s analysis of Sanskrit suggests a different way of looking at Japanese, however. Suppose that all surface [ŋ]s are derived from underlying //s and the grammar maps input /ŋ/ to something other than [], such as [n]. This is a chain shift: // →[ŋ] (medially) and /ŋ/ →[n] (everywhere). Abstractly, the learning problem presented by this chain shift is the same as we saw with Sankrit. For concreteness, I will borrow several markedness constraints from Itô and Mester’s (2003) analysis. (32) a. *VV Intervocalic is prohibited (a lenition constraint). b. * Voiced velar stops are prohibited (cf. Ohala 1983: 196-197). c. *ŋ Velar nasals are prohibited. We will look at their interaction among themselves and with two faithfulness constraints, IDENT(Nasal) and IDENT(Place). For the phonotactic learner of Japanese, the support tableau includes observed initial []s and medial [ŋ]s, both of which are derived by the identity map. (33) Support tableau at phonotactic stage 50 CatJL 4, 2005 John J. McCarthy lexicon cands. *VV**ŋID(Nas) ID(Place) /a/ →a 1 ~ da L 1W ~ ŋaL 1W1W ~ na L 1W1W /aŋa/ →aŋa1 ~ aa1W1WL 1W ~ ana L 1W Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 50
On the first pass through BMCD, *VV is identified as the only rankable markedness constraint. On the next pass, we seek a faithfulness constraint that can be ranked so as to free up one of the remaining markedness constraints for ranking. ID(Place) frees up *ŋ, which accounts for the remaining forms. The residual constraints, *and ID(Nas) are ranked so as to conform to the markedness-first bias. The resulting ranking is given in (34). (34) Result of phonotactic learning for Japanese *VV » ID(Pl) » *ŋ» *» ID(Nas) As the learner grows aware of data from alternations, he learns that there are unfaithful maps // →[ŋ] that occur when stem-initial // becomes medial as a result of the morphology. This unfaithful map is added to the support tableau and BMCD is applied once again. (35) Support tableau after some morphophonemic learning BMCD gives this support tableau the same ranking, (34). The commitment to the identity map for nonalternating forms means that [aŋa] must derive from /aŋa/ —and the presence of the (ŋ, ŋ) identity map in the support tableau forces the ranking ID(Pl) » *ŋ. The proposed chain shift, where underlying /ŋ/ maps to [n], is apparently not learnable under these assumptions about the input. It is learnable with the free-ride algorithm, however, since its basic character is similar to Sanskrit. In discussing Sanskrit, I noted that a local solution to the learning problem was possible by adding another markedness constraint. No such move will solve the problem in Japanese, however. Surface [ŋ]s all act like they are derived from underlying //s with respect to blocking rendaku. Fiddling with the constraints will not somehow magically subvert the commitment to the identity map. Taking a Free Ride in Morphophonemic Learning CatJL 4, 2005 51 lexicon cands. *VV**ŋID(Nas) ID(Place) /a/ →a1 ~ da L 1W ~ ŋaL 1W1W ~ na L 1W1W /aŋa/ →aŋa1 ~ aa1W1WL 1W ~ ana L 1W /a-a/ →aŋa11 ~ aa1W1WLL ~ ana L 11 W Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 51
References Alderete, John; Tesar, Bruce (2002). «Learning Covert Phonological Interaction: An Analysis of the Problem Posed by the Interaction of Stress and Epenthesis». Report no. RuCCS-TR-72. New Brunswick, NJ: Rutgers University Center for Cognitive Science. [Available on Rutgers Optimality Archive #543, http://roa.rutgers.edu/.] Angluin, Dana (1980). «Inductive Inference of Formal Languages from Positive Data». Information and Control 45: 117-135. Baker, Carl Lee (1979). «Syntactic Theory and the Projection Problem». Linguistic Inquiry 10: 533-581. Beckman, Jill N. (1998). Positional Faithfulness. University of Massachusetts, Amherst, doctoral dissertation. [Available on Rutgers Optimality Archive #234, http://roa.rutgers.edu/. Published as Positional Faithfulness: An Optimality Theoretic Treatment of Phonological Asymmetries in 1999 by Garland.] Bermúdez-Otero, Ricardo (2004). «The Acquisition of Phonological Opacity». University of Newcastle upon Tyne, unpublished manuscript. [Available on Rutgers Optimality Archive #593, http://roa.rutgers.edu/.] Churchward, Clerk Maxwell (1940). Rotuman Grammar and Dictionary. Sydney: Australasia Medical Publishing Co. [Reprint 1978, AMS Press, New York.] de Haas, Wim (1988). A Formal Theory of Vowel Coalescence: A Case Study of Ancient Greek. Dordrecht: Foris. Dinnsen, Daniel A. (2004). «A Typology of Opacity Effects in Acquisition». Indiana University, Bloomington, IN, unpublished manuscript. Dinnsen, Daniel A.; Barlow, Jessica A. (1998). «On the Characterization of a Chain Shift in Normal and Delayed Phonological Acquisition». Journal of Child Language 25: 61-94. Dinnsen, Daniel A.; McGarrity, Laura W. (2004). «On the Nature of Alternations in Phonological Acquisition». Studies in Phonetics, Phonology, and Morphology 11: 23-42. [Available at http://www.indiana.edu/~sndlrng/DinnsenMcGarrity% 2004.pdf.] Dinnsen, Daniel A.; O’Connor, Kathleen; Gierut, Judith (2001). «An Optimality Theoretic Solution to the Puzzle-Puddle-Pickle Problem». Handout of paper presented at the 75th Annual Meeting of the Linguistic Society of America, Washington, DC. Dresher, B. Elan (1981). «On the Learnability of Abstract Phonology». In: Baker, Carl Lee; McCarthy, John (eds.). The Logical Problem of Language Acquisition. Cambridge, MA: MIT Press, pp. 188-210. Gnanadesikan, Amalia (1997). Phonology with Ternary Scales. University of Massachusetts, Amherst, doctoral dissertation. [Available on Rutgers Optimality Archive #195, http://roa.rutgers.edu/.] Goldrick, Matthew (2000). «Turbid Output Representations and the Unity of Opacity». In: Hirotani, Masako (ed.). Proceedings of the North East Linguistics Society 30. Amherst, MA: Graduate Linguistic Student Association, pp. 231-246. Goldrick, Matthew; Smolensky, Paul (1999). «Opacity, Turbid Representations, and Output-Based Explanation». Paper presented at the Workshop on the Lexicon in Phonetics and Phonology, University of Alberta, Edmonton. Guerssel, Mohammed (1977). «Constraints on Phonological Rules». Language 3: 267305. 52 CatJL 4, 2005 John J. McCarthy Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 52
Guerssel, Mohammed (1978). «A Condition on Assimilation Rules». Linguistic Analysis 4: 225-254. Hayes, Bruce (1986). «Inalterability in CV Phonology». Language 62: 321-351. Hayes, Bruce (1995). Metrical Stress Theory: Principles and Case Studies. Chicago: The University of Chicago Press. Hayes, Bruce (2004). «Phonological Acquisition in Optimality Theory: The Early Stages». In: Kager, René; Pater, Joe; Zonneveld, Wim (eds.). Fixing Priorities: Constraints in Phonological Acquisition. Cambridge: Cambridge University Press, pp. 158-203. [Available on Rutgers Optimality Archive #327, http://roa.rutgers.edu/.] Itô, Junko; Mester, Armin (1999). «The Phonological Lexicon». In: Tsujimura, Natsuko (ed.). The Handbook of Japanese Linguistics. Oxford: Blackwell, pp. 62-100. Itô, Junko; Mester, Armin (2001). «Structure Preservation and Stratal Opacity in German». In: Lombardi, Linda (ed.). Segmental Phonology in Optimality Theory. Cambridge: Cambridge University Press, pp. 261-295. Itô, Junko; Mester, Armin (2003). «Lexical and Postlexical Phonology in Optimality Theory: Evidence from Japanese». In: Fanselow, Gisbert; Féry, Caroline (eds.). Resolving Conflicts in Grammars: Optimality Theory in Syntax, Morphology, and Phonology. Linguistische Berichte, Sonderheft 11. Hamburg: Helmut Buske, pp. 183-207. [Available at http://people.ucsc.edu/~ito/PAPERS/lexpostlex.pdf.] Kager, René (1999). Optimality Theory. Cambridge: Cambridge University Press. Kiparsky, Paul (1973). «Phonological Representations». In: Fujimura, Osamu (ed.). Three Dimensions of Linguistic Theory. Tokyo: TEC, pp. 3-136. Lombardi, Linda; McCarthy, John (1991). «Prosodic Circumscription in Choctaw Morphology». Phonology 8: 37-71. McCarthy, John J. (1999). «Sympathy and Phonological Opacity». Phonology 16: 331-399. McCarthy, John J. (2000). «The Prosody of Phase in Rotuman». Natural Language and Linguistic Theory 18: 147-197. McCarthy, John J. (2003a). «Comparative Markedness». Theoretical Linguistics 29: 1-51. McCarthy, John J. (2003b). «Sympathy, Cumulativity, and the Duke-of-York Gambit». In: Féry, Caroline; van de Vijver, Ruben (eds.). The Syllable in Optimality Theory. Cambridge: Cambridge University Press, pp. 23-76. McCarthy, John J. (2004). «The Length of Stem-Final Vowels in Colloquial Arabic». University of Massachusetts, Amherst, unpublished manuscript. [Available on Rutgers Optimality Archive #616, http://roa.rutgers.edu/. To appear in Perspectives on Arabic Linguistics 18, ed. by Dilworth B. Parkinson, John Benjamins, Amsterdam & Philadelphia.] McCarthy, John J.; Prince, Alan (1993). «Generalized Alignment». In: Booij, Geert; Marle, Jaap van (eds.). Yearbook of Morphology. Dordrecht: Kluwer, pp. 79-153. [Available on Rutgers Optimality Archive #7, http://roa.rutgers.edu/.] McCarthy, John J.; Prince, Alan (1995). «Faithfulness and Reduplicative Identity». In: Beckman, Jill; Walsh Dickey, Laura; Urbanczyk, Suzanne (eds.). University of Massachusetts Occasional Papers in Linguistics 18: Papers in Optimality Theory. Amherst, MA: Graduate Linguistic Student Association, pp. 249-384. [Available on Rutgers Optimality Archive #60, http://roa.rutgers.edu/.] McCarthy, John J.; Prince, Alan (1999). «Faithfulness and Identity in Prosodic Morphology». In: Kager, René; van der Hulst, Harry; Zonneveld, Wim (eds.). The Taking a Free Ride in Morphophonemic Learning CatJL 4, 2005 53 Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 53
Prosody-Morphology Interface. Cambridge: Cambridge University Press, pp. 218309. [Available on Rutgers Optimality Archive #216, http://roa.rutgers.edu/.] McCarthy, John J.; Wolf, Matthew (2005). «Less Than Zero: Correspondence and the Null Output». University of Massachusetts, Amherst, unpublished manuscript. [Available on Rutgers Optimality Archive #722, http://roa.rutgers.edu/.] Moreton, Elliott (2000). «Faithfulness and Potential». University of Massachusetts, Amherst, unpublished manuscript. Nicklas, Thurston (1974). The Elements of Choctaw. University of Michigan, doctoral dissertation. Nicklas, Thurston (1975). «Choctaw Morphophonemics». In: Crawford, James M. (ed.). Studies in Southeastern Indian Languages. Athens, GA: University of Georgia Press, pp. 237-250. Ohala, John (1983). «The Origin of Sound Patterns in Vocal Tract Constraints». In: MacNeilage, Peter (ed.). The Production of Speech. New York: Springer-Verlag, pp. 189-216. Ota, Mitsuhiko (2004). «The Learnability of the Stratified Phonological Lexicon». Journal of Japanese Linguistics 20: 19-40. [Available on Rutgers Optimality Archive #668, http://roa.rutgers.edu/.] Pater, Joe (2004). «Course Handout». University of Massachusetts, Amherst, unpublished manuscript. Pater, Joe; Tessier, Anne-Michelle (2003). «Phonotactic Knowledge and the Acquisition of Alternations». In: Solé, Maria-Josep; Recasens, Daniel; Romero, Joaquín (eds.). Proceedings of the 15th International Congress of Phonetic Sciences. Barcelona: Universitat Autònoma de Barcelona, pp. 1177-1180. [Available at http://people.umass.edu/pater/pater-tessier.pdf.] Prince, Alan (2002). «Arguing Optimality». Rutgers University, New Brunswick, NJ, unpublished manuscript. [Available on Rutgers Optimality Archive #562, http://roa.rutgers.edu/.] Prince, Alan; Smolensky, Paul (2004). Optimality Theory: Constraint Interaction in Generative Grammar. Malden & Oxford: Blackwell. [Revision of 1993 technical report no. RuCCS-TR-2, Rutgers University Center for Cognitive Science, New Brunswick, NJ. Available on Rutgers Optimality Archive #537, http://roa. rutgers.edu/.] Prince, Alan; Tesar, Bruce (2004). «Learning Phonotactic Distributions». In: Kager, René; Pater, Joe; Zonneveld, Wim (eds.). Fixing Priorities: Constraints in Phonological Acquisition. Cambridge: Cambridge University Press, pp. 245-291. [Available on Rutgers Optimality Archive #353, http://roa.rutgers.edu/.] Schane, Sanford (1987). «The Resolution of Hiatus». In: Bosch, Anna; Need, Barbara; Schiller, Eric (eds.). Papers from the 23rd Annual Regional Meeting of the Chicago Linguistic Society, Vol. 2: Parasession on Autosegmental and Metrical Phonology. Chicago: Chicago Linguistic Society, pp. 279-290. Schein, Barry; Steriade, Donca (1986). «On Geminates». Linguistic Inquiry 17: 691-744. Tesar, Bruce (1997). «Multi-Recursive Constraint Demotion». Rutgers University, New Brunswick, NJ, unpublished manuscript. [Available on Rutgers Optimality Archive #197, http://roa.rutgers.edu/.] Tesar, Bruce (1998). «Error-Driven Learning in Optimality Theory Via the Efficient Computation of Optimal Forms». In: Barbosa, Pilar; Fox, Danny; Hagstrom, Paul; 54 CatJL 4, 2005 John J. McCarthy Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 54
McGinnis, Martha; Pesetsky, David (eds.). Is the Best Good Enough? Optimality and Competition in Syntax. Cambridge, MA: MIT Press, pp. 421-435. Tesar, Bruce; Alderete, John; Horwood, Graham; Merchant, Nazarré; Nishitani, Koichi; Prince, Alan (2003). «Surgery in Language Learning». In: Garding, Gina; Tsujimura, Mimu (eds.). Proceedings of the 22nd West Coast Conference on Formal Linguistics. Somerville, MA: Cascadilla Press, pp. 477-490. Tesar, Bruce; Prince, Alan (forthcoming). «Using Phonotactics to Learn Phonological Alternations». In: Cihlar, Jonathan E.; Franklin, Amy L.; Kaiser, David W.; Kimbara, Irene (eds.). Papers from the 39th Regional Meeting of the Chicago Linguistics Society, Vol. 2: The Panels. Chicago: Chicago Linguistic Society. [Available on Rutgers Optimality Archive #620, http://roa.rutgers.edu/.] Tesar, Bruce; Smolensky, Paul (1994). «The Learnability of Optimality Theory». In: Aranovich, Raul; Byrne, William; Preuss, Susanne; Senturia, Martha (eds.). Proceedings of the Thirteenth West Coast Conference on Formal Linguistics. Stanford, CA: CSLI Publications, pp. 122-137. Tesar, Bruce; Smolensky, Paul (2000). Learnability in Optimality Theory. Cambridge, MA: MIT Press. Tessier, Anne-Michelle (2004). «Looking for OO-Faithfulness in Phonological Acquisition». University of Massachusetts, Amherst, unpublished manuscript. Ulrich, Charles (1986). Choctaw Morphophonology. UCLA, doctoral dissertation. Whitney, William Dwight (1889). Sanskrit Grammar. Cambridge, MA: Harvard University Press. Winter, Edgar (1972). Free Ride. Epic Records. Zwicky, Arnold M. (1970). «The Free-Ride Principle and Two Rules of Complete Assimilation in English». In: Campbell, M. A. et al. (eds.). Papers from the Sixth Regional Meeting of the Chicago Linguistic Society. Chicago: Chicago Linguistic Society, pp. 579-588. Taking a Free Ride in Morphophonemic Learning CatJL 4, 2005 55 Cat.Jour.Ling. 4 001-252 7/2/06 11:43 Página 55