scieee AI-readable full text Open interactive document viewer

Estimation of dynastic life-cycle discrete choice models

Gayle, George-Levi,Golan, Limor,Soytas, Mehmet A.

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Gayle, George-Levi; Golan, Limor; Soytas, Mehmet A. Article Estimation of dynastic life-cycle discrete choice models Quantitative Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Gayle, George-Levi; Golan, Limor; Soytas, Mehmet A. (2018) : Estimation of dynastic life-cycle discrete choice models, Quantitative Economics, ISSN 1759-7331, The Econometric Society, New Haven, CT, Vol. 9, Iss. 3, pp. 1195-1241, https://doi.org/10.3982/QE771 This Version is available at: https://hdl.handle.net/10419/217126 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/ Quantitative Economics 9 (2018), 1195–1241 1759-7331/20181195 Estimation of dynastic life-cycle discrete choice models George-Levi Gayle Department of Economics, Washington University in St. Louis and Federal Reserve Bank of St. Louis Limor Golan Department of Economics, Washington University in St. Louis and Federal Reserve Bank of St. Louis Mehmet A. Soytas Graduate School of Business, Ozyegin University This paper explores the estimation of a class of life-cycle discrete choice dynastic models. It provides a new representation of the value function for these class of models. It compare a multistage conditional choice probability (CCP) estimator based on the new value function representation with a modified version of the full solution maximum likelihood estimator (MLE) in a Monte Carlo study. The modified CCP estimator performs comparably to the MLE in a finite sample but greatly reduces the computational cost. Using the proposed estimator, we estimate a dynastic model and use the estimated model to conduct counterfactual simulations to investigate the role Nature versus Nurture in intergenerational mobility. We find that Nature accounts for 20 percent of the observed intergenerational immobility at the bottom of income distribution. That means that 80 percent of mobility at the bottom of the income distribution is explained by economic decision and economic/institutional constraints. Keywords. Discrete choice models, dynastic models, intergenerational mobility, nature versus nurture. JEL classification. C13, J13, J22, J62. 1. Introduction The importance of parents’ altruism toward their children and children’s altruism toward their parents has long been recognized as an important factor underlying the economic behavior of individuals. Economic models that incorporate these intergenerational links are normally referred to as dynastic models. Many important economic George-Levi Gayle: [email protected] Limor Golan: [email protected] Mehmet A. Soytas: [email protected] We thank the participants of 3rd Annual All Istanbul Meeting, Sabanci University 2013; Society of Labor Economists Annual Meetings 2015; European Economic Association 30th Annual Congress 2015; Econometric Society (11th) World Congress 2015; Southern Economic Association 85th Annual Meetings 2015; 10th International Conference on Computational and Financial Econometrics 2016; and the seminar participants at Izmir University of Economics, TOBB University of Economics and Technology, and Marmara University. The views expressed are those of the individual authors and do not necessarily reflect official positions of the Federal Reserve Bank of St. Louis, the Federal Reserve System, or the Board of Governors. ©2018 The Authors. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at http://qeconomics.org.https://doi.org/10.3982/QE771 1196 Gayle, Golan, and Soytas Quantitative Economics 9 (2018) behaviors—and hence the welfare effect of many public policies—critically depend on whether these dynastic links are explicitly modeled. Several papers have documented that (i) the distribution of wealth is more concentrated than that of labor earnings and (ii) it is characterized by a smaller of fraction of households owning a larger fraction of total wealth over time. There are different models of dynastic transfers explaining the persistence in wealth and income across generations (e.g., the Loury (1981), model of transmission of human capital and the Laitner (1992), model of bequests); however, in these models fertility is exogenous. Barro and Becker (1988, 1989) develop dynastic models with endogenous fertility; however, in their models endogenizing fertility leads to a lack of persistence in earnings and wealth because wealthier households have more children and therefore dynastic transfers do not depend on wealth and income. The data clearly show persistence in income across generations. Subsequently, dynastic models with endogenous fertility that capture the dynastic persistence of income and wealth have been analyzed extensively, but such models have not been estimated mainly because of computational feasibility considerations. This paper develops an estimator for dynastic models of dynastic transfers and estimates a model quantifying the different factors generating the persistence of income. Alvarez (1999) combined the main features of the above-mentioned models by incorporating the fertility decision into the Laitner (1992)andLoury (1981)dynastictransfer models. While some models, as Laitner (1992), incorporated an elaborate finite life-cycle model for adults in each generation, in other models there is one period of childhood and one period of adulthood. The framework we study incorporates all these elements and develops a model in which altruistic parents make discrete choices of birth, labor supply, and discrete and continuous investment choices in children. In particular, in order to accommodate many models in the literature, parents choose time with children and a continuous monetary investment in their children every year over their life cycle. The model can also be extended to include bequests. The model is a partial equilibrium model, and as in most dynastic models and in the basic setup, there is one decision-maker in a household; however, we show that it can be easily extended to a unitary household.1 While the study of dynastic models has been widespread in the economic literature, these studies have been largely theoretical or quantitative theory. However, the estimation of these models and the use of these estimated models to conduct counterfactual policy analysis are nonexistent. There are two main reasons for this gap; the first is data limitation and the second is computational feasibility. Ideally, one would need data on the choices and characteristics of multiple generations linked across time to estimate these dynastic models. The number of generations needed for estimation can be reduced to two by analyzing the stationary equilibrium properties of these model. Recently, data on the choices and characteristics of at least two generations have become available in the National Longitudinal Survey of Youth (NLSY79), Panel Study of Income Dynamics (PSID), and a number of European administrative datasets. 1In a companion paper, we extend the current framework to incorporate nonunitary households (Gayle, Golan, and Soytas (2014)). Quantitative Economics 9 (2018) Dynastic life-cycle discrete choice models 1197 There are two main estimators used in the literature to estimate dynamic discrete choice models: full solution method using the “nested fixed point” algorithm (NFXP) (see Wolpin (1984), Miller (1984), Pakes (1986), and Rust (1987) for early examples) and “conditional choice probability” (CCP) (see Hotz and Miller (1993), Altug and Miller (1998), and Aguirregabiria (1999)) estimators that do not require the solution to the fixed points. More recently, Aguirregabiria and Mira (2002) showed that an appropriately formed CCP-based estimator, “nested pseudo likelihood” (NPL), is asymptotically equivalent to an NFXP estimator. The major limitation of the NFXP estimation procedure is that it suffers from the curse of dimensionality (i.e., as the number of states in the state space increases, the number of computations increases at a rate faster than linear). Dynastic models add an additional loop to this estimation procedure: a nested fixed-point squared. Therefore, this estimation procedure suffers from the curse of dimensionality squared. However, even with a CCP estimator or an NPL estimator, estimation of the dynastic model requires dealing with further complications that are not present in single-agent dynamic discrete choice models. The main difficulty is deriving the representation of the value functions of the problem. This difficulty is associated with the nonstandard nature of the problem. A dynastic model has finite number of periods in the life cycle in each generation and infinitely many generations are linked by the altruistic preferences. This framework does not fit into a finite horizon dynamic discrete choice model since in the last period, there is a continuation value associated with the next generation’s problem that is linked to the current generation by the transfers and the discount factor. Therefore, we need to find a representation for the next generation’s continuation value if we want to treat the problem as a standard finite-period problem and solve it by backward induction.2In this paper, we propose a new estimation procedure based on a representation of the period value functions in terms of period primitives. In particular, we show that an appropriately defined alternative representation of the continuation value enables us to apply a CCP estimator to dynastic models. The general principles used in the estimation technique are well known in the literature,3and hence the main contribution of this paper is showing how these principles can be combined to estimate dynastic models. In a Monte Carlo study, we demonstrate that a multistage CCP estimator based on the new value function representation have good small-sample properties that compare favorably to a full solution NFXP estimator. For this comparison, we use a pseudo maximum likelihood estimator (PML) so that our results would be more comparable to those of the NFXP maximum likelihood estimator. We use the GMM version of the estimator developed in this paper to estimate a dynastic model of intergenerational transmission of human capital with unitary households. The estimated model captures well the labor supply, time with children, and fertility decisions of households. We then demonstrate the usefulness of our framework for 2Obviously, we can always solve the problem by NFXP if we assume that the problem is stationary in the generations. In this case, the solution to the dynamic programming problem requires solving the fixedpoint problem for the period value functions. However, as one can easily anticipate, we encounter the same computational burden of full solution. Therefore, our specific interest is CCP-type estimators. 3See Hotz and Miller (1993), Hotz, Miller, Sanders, and Smith (1994), Altug and Miller (1998), and Aguirregabiiria and Miria (2002) for the seminal contributions from which these general principles are derived. 1198 Gayle, Golan, and Soytas Quantitative Economics 9 (2018) policy analysis. This is done by conducting counterfactual simulations to investigate the role of the automatic transmission of education across generation (Nature) on integenerational mobility at bottom of the income distribution. We find that without the Nature on the intergenerational education production function mobility at the bottom of the income distribution would have been 20 percent higher. That means that 80 percent of mobility at the bottom of the income distribution is explained by economic decision and economic/institutional constraints. Lastly, not accounting for the reoptimization of subsequent generations in the model, as is done in the approach outlined in this paper, will overstate the effect of Nature on mobility by between 20 and 90 percent. Dynastic models have been used to study numerous topics in economics. These topics include explaining the cross-sectional correlation between parental wages and fertility (see Jones, Schoonbroodt, and Tertilt (2010), for a detailed overview of this literature), the relationship between inequality and growth (see, e.g., De la Croix and Doepke (2003)), the relationship between human capital formation and social mobility (see Heckman and Mosso (2014), for a survey of this literature), the relation among bequests, saving, and the distribution of wealth and earnings (see De Nardi (2004); Cagetti and De Nardi (2008), among others),4and the optimality of different ways of funding social security. These models have been used to shed light on the effect of education, child care subsidies, child labor regulations, and wealth and income redistribution policies on individual welfare. Reviewing this vast and diverse literature is beyond the scope of this paper; however, a short review of two of the literature segments will suffice to illustrate the need to estimate these models, and hence the wide applicability of our estimation technique. The first segment explains the widespread negative cross-sectional correlation between parental wage and fertility. The basic dynastic model as formulated by Barro and Becker (1989) cannot explain this negative correlation because wealthier parents increase the number of offspring, keeping transfer levels the same as less wealthy parents. Attempts in the literature to account for this negative correlation range from appropriately calibrating the model parameters so that the substitution effects are larger than the income effects, introducing the quality of children as a choice variable with an appropriate assumption about the cost of child-rearing (Becker and Lewis (1973), Becker and Tomes (1976), Moav (2005)),5to introducing nonhomotheticity in preferences (see, e.g., Galor and Weil (2000), Greenwood and Seshadri (2002), or Fernandez, Guner, and Knowles (2005)). As summarized in Alvarez (1999), depending on the functional form assumptions of the primitives and values of the structural parameters, dynastic models could generate the negative correlation between parental wages and fertility.6Therefore, 4For example, the De Nardi (2004) model explicitly focused on the transmission of physical and human capital from parents to children and intergenerational links. She shows that such a model can can induce savings behavior that generates a distribution of wealth that (i) is much more concentrated than that of labor earnings and (ii) also makes the rich keep large amounts of assets in old age to leave bequests to their descendants. 5See Jones, Schconbroodt, and Tertilt (2010, Section 5.2). 6Recently, Mookherjie, Prina, and Ray (2012) demonstrated that incorporating dynamic analysis of return to human capital can help explain the negative cross-sectional correlation between parental wages and fertility. Quantitative Economics 9 (2018) Dynastic life-cycle discrete choice models 1199 whether the basic dynastic model can explain this negative cross-sectional correlation between parental wages and fertility is an empirical question requiring careful exploration of the source of identification and estimation (see Gayle, Golan, and Soytas (2014, 2015), for examples of these types of analysis). The effects of the social security system on both capital accumulation and wealth distribution have been of great interest to economists and policy-makers for decades (see, for instance, Kotlikoff and Summers (1981), Caballé and Fuster (2003), among others). However, the optimal form of funding social security may depend on whether or not these intergenerational links are explicitly modeled. For example, Fuster, Imrohoroglu, and Imrohoroglu (2007) argued that when households insure members in the same family line, privatizing social security without compensation is favored by 52% of the population. If social security participants are fully compensated for their contributions and the transition to privatization is financed by a combination of debt and a consumption tax, 58% experience a welfare gain. These gains and the resulting public support for social security reform depend critically on a flexible labor market. If the elasticity of the labor supply is low, then support for privatization disappears. Therefore, it is important to estimate these models because policy implications often depend on the value of key structural parameters. In Fuster, Imrohoroglu, and Imrohoroglu (2007), the key structural parameter was the elasticity of labor supply, but in other models it may be the altruism parameters themselves. The rest of the paper is organized as follows. Section 2presents the basic gender-less life-cycle dynastic model with only discrete choices. Section 3presents the generic estimator of the life-cycle model and presents the Monte Carlo study. Section 4extends the framework to include continuous choices and transfers, intra-household behaviors, and gender. Section 5presents the basic framework of our empirical application. Section 6 presents our empirical results. Section 7concludes, all proofs are provided in an appendix, and additional tables are provided in the Supplementary Material (Gayle, Golan, and Soytas (2018)). 2. Theoretical framework The theoretical framework is developed to allow for estimation of a rich group of dynastic models and allows us to address many relevant policy questions. This section develops a model of altruistic parents who make transfers to their children. The transfers are discrete and can allow for (i) discrete time investment in children and (ii) monetary investment with discrete levels. Section 4extends this basic framework to allow for continuous choices and transfers. This allows us to use the framework to analyze bequests or any continuous monetary transfers by parents to their children. We incorporate two important aspects of the problem. First, fertility is endogenous. Endogenous fertility has important implications for intergenerational transfers and the quantity-quality tradeoffs made by parents when they choose transfers as the well as number of offspring. Second, we include a life cycle for each generation. The life cycle is important to understanding fertility behavior, spacing of children, and the timing of different types of 1200 Gayle, Golan, and Soytas Quantitative Economics 9 (2018) investments. This section analyzes a model with one gender-less decision-maker. We later extend this framework to a unitary household.7 We build on previous dynastic models that analyze transfers and intergenerational transmission of human capital. In some models, such as Loury (1981)andBecker and Tomes (1986), fertility is exogenous, whereas in others, such as Becker and Barro (1988) and Barro and Becker (1989), fertility is endogenous. The Barro–Becker framework is extended in our model by incorporating a life-cycle behavior model, based on previous work, such as Heckman, Hotz, and Walker (1985)andHotz and Miller (1988), into an infinite-horizon model of dynasties. Our life-cycle model includes individuals choices about time allocation decisions, investments in children, and fertility. We formulate a partial equilibrium discrete choice model that incorporates life-cycle considerations of individuals from each generation into the larger framework. Adults in each generation derive utility from their own consumption, leisure, and the utility of their adult offspring. The utility of adult offspring is determined probabilistically by the educational outcome of childhood, which in turn is determined by parental time and monetary inputs during early childhood, parental characteristics (such as education), and luck. Parents make decisions in each period about fertility, labor supply, time spent with children, and monetary transfers. For simplicity, the only intergenerational transfers are transfers of human capital, as in Loury (1981). However, the framework can include any other choice of transfer that is discrete. We assume no borrowing or savings for simplicity. The model assumes that the educational outcome of children is revealed at the last period of parent’s life cycle regardless of the birth date of the children. This assumption is similar to the Barro–Becker assumptions. In the parents’ life cycle, adult children’s behavior and choices do not affect the choices of parents. As in Barro–Becker, the choices can only be made by the children in their own life cycle which starts immediately after the parents’ life cycle ends.8 In the model, adults live for Tperiods. Each adult from generation g∈{0∞} makes discrete choices about labor supply (ht), time spent with children (dt),andbirth (bt),ineveryperiodt=1T. For labor time, individuals choose no work, part-time, or full-time (ht∈(012)); for time spent with children individuals choose none, low, or high (dt∈(012)). The birth decision is binary (bt∈(01)). The individual does not make any choices during childhood, when t=0. All the discrete choices can be combined into one set of mutually exclusive discrete choice, represented as k, such that k∈(0117).LetIkt be an indicator for a particular choice kat age t;Ikt takes the value 1if the kth choice is chosen at age tand 0otherwise. These indicators are defined 7Treatment of households, with two decision-makers (with separate utility functions), marriage, and divorce, is involved and is beyond the scope of this paper. See Gayle, Golan, and Soytas (2014) for more details on one such model. 8In a model where adult children’s behavior and choices do affect investment in children and fertility of the parents, solutions to the problems are significantly more complicated and it is not clear whether a solution exists. Quantitative Economics 9 (2018) Dynastic life-cycle discrete choice models 1201 as follows: I0t=I{ht=0}I{dt=0}I{bt=0} I1t=I{ht=0}I{dt=0}I{bt=1}  I16t=I{ht=1}I{dt=2}I{bt=1} I17t=I{ht=2}I{dt=2}I{bt=1} (1) Since these indicators are mutually exclusive, then 17 k=0Ikt =1.Wedefineavector, x, to include the time-invariant characteristics of the individual’s education, skill, and race. Incorporating this vector, we further define the vector zto include all past discrete choices as well as time-invariant characteristics, such that zt=({Ik1}17 k=0{Ikt−1}17 k=0 x). We assume the utility function is the same for adults in all generations. An individual receives utility from discrete choice and from consumption of a composite good, ct. The utility from consumption and leisure is assumed to be additively separable because the discrete choice, Ikt , is a proxy for leisure and is additively separable from consumption. The utility from Ikt is further decomposed into two additive components: a systematic component, denoted by u1kt(zt), and an idiosyncratic component, denoted by εkt. The systematic component associated with each discrete choice krepresents an individual’s net instantaneous utility associated with the disutility from market work, the disutility/utility from parental time investment, and the disutility/utility from birth. The idiosyncratic component represents a preference shock associated with each discrete choice kthat is transitory in nature. To capture this feature of εkt , we assume that the vector (ε0tε17t)is independent and identically distributed across the population and time and is drawn from a population with a common distribution function, Fε(ε0tε17t). The distribution function is assumed to be absolutely continuous with respect to the Lebesgue measure and has a continuously differentiable density. Per-period utility from the composite consumption good is denoted u2t(ctzt).We assume that u2t(ctzt)is concave in c;thatis,∂u2t(ctzt)/∂ct>0and ∂2u2t(ctzt)/∂c2 t<0. Implicit in this specification is the inter-temporally separable utility from the consumption good, but not necessarily for the discrete choices, since u2tis a function of zt,which is itself a function of past discrete choices but is not a function of the lagged values of ct. Altruistic preferences are introduced under the same assumption as the Barro– Becker model: Parents obtain utility from their adult offspring’s expected lifetime utility. Two separable discount factors capture the altruistic component of the model. The first, β, is the standard rate of time preference parameter, and the second, λN−ν,isthe intergenerational discount factor, where Nis the number of offspring an individual has over her lifetime. Here, λ(0<λ<1) should be understood as the individual’s weighting of her offsprings’ utility relative to her own utility. The individual discounts the utility of each additional child by a factor of −ν,where0<ν<1. We let earnings (wt) be given by the earnings function wt(ztht),whichdependson the individual’s time-invariant characteristics, choices that affect human capital accumulated with work experience, and the current level of labor supply (ht). The choices 1202 Gayle, Golan, and Soytas Quantitative Economics 9 (2018) and characteristics of parents are mapped onto their offspring’s characteristics (x)via a stochastic production function of several variables. The offspring’s characteristics are affected by their parents’ time-invariant characteristics, their parents’ monetary and time investments, and the presence and timing of siblings. These variables are mapped into the child’s skill and educational outcome by the function M(x|zT+1)where zT+1includes all parental choices and characteristics and contains information on the choices of time inputs and monetary inputs. Because zT+1also contains information on all birth decisions, it captures the number of siblings and their ages. We assume there are four mutually exclusive educational outcomes for offspring: less than high school (LH), high school (HS), some college (SC), and college (Coll). Therefore, M(x|zT+1)is a mapping of parental inputs and characteristics into a probability distribution over these four outcomes. We normalize the price of consumption to 1. Raising children requires parental time (dt)and market expenditure. The per-period cost of raising children is denoted pcnt . Therefore, the per-period budget constraint is given by wt≥ct+pcnt(2) The sequence of optimal choice for both discrete choice and consumption is denoted as Io kt and co t, respectively. We can thus denote the expected lifetime utility at time t=0of a person with characteristics xin generation g, excluding the dynastic component, as UgT (x) =E0T  t=0 βt17  k=0 Io ktu1kt(zt)+εkt+u2tco tztx(3) The total discounted expected lifetime utility of an adult in generation gincluding the dynastic component is Ug(x) =UgT (x) +βTλE0N−ν N  n=1 Ug+1nx nx(4) where Ug+1n(x n)is the expected utility of child n(n=1N) with characteristics x n.9 In this model, individuals are altruistic and derive utility from their offspring’s utility, subjecttodiscountfactorsβand λN−ν.10 This formulation is similar to the one in Barro– Becker, but it is extended to allow for differences in gender and “types.” To simplify presentation of the model, we assume that pcnt is proportional to an individual’s current earnings and the number of children, but we allow this proportion to depend on the state variables. This assumption allows us to capture the differential expenditures on children made by individuals with different incomes and characteristics. 9Note that this formulation can be written as an infinite discounted sum (over generations) of per-period utilities as in the Barro–Becker formulation. 10Note that since we add life-cycle, the regularity condition that implies that the discount factor of the children’s utilities, βTλN−νis between zero and one is satisfied for any N,asβis also between zero and one. Quantitative Economics 9 (2018) Dynastic life-cycle discrete choice models 1209 is invertible. The representation is obtained by combining known results16 from discrete choice estimation of stationary infinite-horizon problems with the finite horizon properties of the dynastic life-cycle model. 3.2 Estimation We parameterized the period utility by a vector θ2,ukt(ztθ2); the period transition on the observed states is parameterized by a vector θ3,F(zt|zt−1IkT =1θ3); the intergenerational transitions on permanent characteristics is parameterized by a vector θ4, Mn(x|zT+1θ4); and the earnings function is characterized by a vector θ5wt(xhtθ5). Therefore, the conditional value functions, decision rules, and choice probabilities now also depend on θ≡(θ2θ3θ4θ5βλν). Standard estimates of dynamic discrete choice models involve forming the likelihood functions from the CCPs derived in equation (16). This involves solving the value function for each iteration of the likelihood function. The method used to solve the value function depends on the nature of the optimization problems and normally falls into one of two cases: (i) Finite-horizon problems: The problem has an end date (as in a standard life-cycle problem); hence future value function is obtained by backwards induction. (ii) Stationary infinite-horizon problem: The valuation is obtained by a contraction mapping. A dynastic discrete choice model in unusual because it involves both a finite-horizon problem and an infinite-horizon problem. Solving both problems for each iteration of the likelihood function is computationally infeasible for all but the simplest of models. We avoid solving the stationary infinite-horizon problem in estimation by replacing the terminal value in the life-cycle problem with equation (20). This converts the problem into a finite-horizon problem that can be solved by backward recursion, with the flow utility function given by υk(zT)=ukT (zT)+λNT(zT)−ν x VxNT  n=1 Mn kx|zT(21) The per-period utility in the terminal period, ukT (zT), is parameterized by θ2.The intergenerational transition function, Mn k(x|zT), can be treated as known since it can be estimated from the data. Given Fε(ε0tε17t)and calculating V(x )via equation (20),17 we can calculate the ex ante value function at Tusing V(z T)= 17 k=0I0 kI(zTεT)[υk(zT)+εkT ]fε(εT)dεT. The conditional value function for T−1 is given by υk(zT−1)=ukT−1(zT−1)+βzTV(z T)F(zT|zT−1IkT =1). This is continued backward given υk(zT−1)to form value function at T−2,andsoon. 16See Aguirregabira and Mira (2002) and Pesendorfer and Schmidt-Dengler (2008) for the use and derivation of this inversion in the context of stationary infinite horizon problems. 17This manipulation is possible because the alternative value function in equation (20) is a function of only the parameters of the model and the CCPs. The CCPs can be estimated directly from the data then backward recursion becomes possible because the decision in the last period, T, is similar to a static problem when the value of children is replaced with equation (20). 1210 Gayle, Golan, and Soytas Quantitative Economics 9 (2018) The backward induction procedure outlined above shows that only Mn k(x|zT)in equations (21)and(20) depends on the next generation’s outcome. Thus, we can estimate the intergenerational problem with only two generations of data, as is the case in the standard stationary discrete choice models (see for example Rust (1987)). To estimate the intergenerational problem, we let Idtg,zdtg,andεdtg, respectively, indicate the choice, observed state, and unobserved state at age tin the generation gof dynasty d. Forming the CCPs for each individual in the first observed generation of dynasty dat all ages tyields the components necessary for estimation. Estimation proceeds in two steps. Step 1: In the first step, we estimate the CCP, transition, and earnings functions necessary to compute the inversion in equation (20). The expectation of observed choices conditional on the observed state variables gives an empirical analog to the CCPs at the true parameter values of the problem, θo 1, allowing us to estimate the CCPs; we denote this estimate by  pk(zdt1).Wealsoestimateθ3,θ4,andθ5, which parameterize the transition and earnings functions F(zt|zt−1IkT =1θ3),Mn(x|zT+1θ4),andwt(xhtθ5), respectively, in this step. Step 2: The second step can be estimated two ways, the first is a PML (as used in Aguirregabira and Mira (2002)) and the second is a GMM (as used in the original Hotz and Miller (1993)). We can use a PML method and not a pure maximum likelihood estimator because part of the likelihood function is concentrated out using the data. With D dynasties, the PML estimates of θ0=(θ2βλν)are obtained via θ0PML =argmax θ0D  dt1=1 T  t=0 17  k Idt1lnpk(zdt1;θ0 θ3 θ4 θ5)(22) where pk(zdt1;θ0 θ3 θ4 θ5)is the CCP defined in equation (16) with the conditional value function replaced with υk(zdt1θ0 θ3 θ4 θ5), which is calculated by backward recursion using the estimated choice probabilities and the transition functions outlined in Step 1. An alternative second-step GMM estimator is formed using the inversion found in Hotz and Miller (1993). Under the assumption that εis distributed independently and identically as type I extreme values, then the Hotz and Miller inversion implies that logpk(zdt1;θ0 θ3 θ4 θ5)/pK(zdt1;θ0 θ3 θ4 θ53) =υk(zdt1θ0 θ3 θ4 θ5)−υK(zdt1θ0 θ3 θ4 θ5) (23) for any normalized choice K.Wecanuse  pk(zdt1), estimated from Step 1, to form an empirical counterpart to equation (23) and estimate the parameters of our model. The moment conditions can be obtained from the difference in the conditional valuation functions calculated for choice kand a base choice(say K=0). The following moment conditions are produced for an individual at age t∈{1755}: ξjdt(θ0)≡υk(zdt1θ0 θ3 θ4 θ5)−υ0(zdt1θ0 θ3 θ4 θ5)−ln pk(zdt1)/  p0(zdt1)(24) Quantitative Economics 9 (2018) Dynastic life-cycle discrete choice models 1211 Therefore, there are 17 orthogonality conditions and thus j=117. Letting ξdt(θ0) be the vector of moment conditions at t, these vectors are defined as ξdt(θ0)= (ξ1dt(θ0) ξ2dt(θ0)ξ17dt(θ0)). Therefore, E[ξdt(θo 0)|zdt]converges to 0for every consistent estimator of true CCPs, pk(zdt1;θ0 θ3 θ4 θ5),fort∈{1755},andwhereθo 0is the true parameter of the model. Define ξd(θ0)≡(ξd1(θ0)ξdT (θ0))as the vector of moment restrictions for a given individual over time and define a weight matrix as (θ0)≡Et[ξd(θ0)ξd(θ0)]. Then the GMM estimate of θ0is obtained via θ02SGMM =argmin θ01/D D  d=1 ξd(θ0) 1/D D  d=1 ξd(θ0)(25) where is a consistent estimator of (θo). 3.3 Monte Carlo study We present a numerical example of a model with human capital investments and intergenerational transfers. We use the example to examine the performance of the proposed estimation technique. Using simulated data from the numerical example, we estimate the parameters of the model using the NFXP and PML estimators. The estimation is done for varying sample sizes (i.e., for 1000,10,000,20,000,and40,000). NFXP estimation of life-cycle dynastic models is possible only in the simplest dynastic structure. Hence for the Monte Carlo study we choose a simple model which can be estimated by both NFXP and PML. To the best of our knowledge, there is no empirical application of life-cycle dynastic model which is estimated by NFXP. Instead, all the empirical application of life-cycle dynastic model specify the terminal value at the end of an individual’s life cycle as a reduced-form function of the state variables. Dynastic models estimated in this fashion are not suitable for conducting counterfactual policy analysis. For illustration purposes, we start with the model in which the per-period utility function, uk(zt), has a linear form. In each period, t∈{01}, the individual chooses whether to invest or not, Ik∈{01}. We assume that individuals can have at most one child, N≤1. The utilities associated with each choice are given by uk(zt)=ztif k=0 (1−θ)ztif k=1 where Fε(εt)is the distribution of the choice-specific, unobservable part of the utility; it is assumed to be independently distributed type 1extreme value. In the environment in this example, the individual begins the life cycle with a particular set of character traits denoted by zt∈(0506070809).Notethatatt=0the individual has not made any choices yet, so the vector z0depends fully on initial characteristics x. The value of z1is given by the transformation function Fk(zt|zt−1)that given 1212 Gayle, Golan, and Soytas Quantitative Economics 9 (2018) by the transition matrix: F0(zt|zt−1)=⎛ ⎜ ⎜ ⎜ ⎜ ⎜ ⎝ 085 013 002 0 0 004 085 009 002 0 001 004 085 009 001 0001 005 085 009 00001 ⎞ ⎟ ⎟ ⎟ ⎟ ⎟ ⎠ and F1(zt|zt−1)=⎛ ⎜ ⎜ ⎜ ⎜ ⎜ ⎝ 10000 01090 00 013 027 0600 001 011 028 060 0004 013 023 06 ⎞ ⎟ ⎟ ⎟ ⎟ ⎟ ⎠  The individual’s traits in the next period are determined by the probabilities in the corresponding row, where each row corresponds to one of the initial values z0∈ (0506070809), and each column represents character traits in the next period, z1∈(0506070809). The transition is such that an individual with character traits z0=05who chooses not to have a child, such that the choice vector I0=0,willhave characteristics z1=05with a probability of 085. In this simplified model, the next generation’s initial characteristics z 0depend only on the sum of the financial investment decisions in the life cycle. The educational outcome of the offspring is determined by the intergenerational transition function: Mz 0|zT+1=⎛ ⎜ ⎝ 10 0 0 0 001040401 00004 006 09⎞ ⎟ ⎠ where zT+1can take values in {012}. The next generation’s starting character traits are determined by the probabilities given in the row, where each row corresponds to one of the values of zT+1∈(012)and the first row represents investment level zT+1=0.Ifthe individual invests nothing, then the next generation will have the lowest consumption value with complete certainty. The transition is such that an individual who opts to invest twice in the life cycle has a probability of 09that the next generation will start his life cycle with the characteristics z 0=09. We simulated the model for the parameter values, (θ2βλ)=(02508095),where θis the structural parameter of interest that gives the marginal cost of investment, and λ and βare the generational and time discount factors, respectively. We solve the dynamic problem for datasets of 1000,10,000,20,000,40,000 individual dynasties and repeat the simulation 100 times. For the CCP estimation, the initial consistent estimates are esti- Quantitative Economics 9 (2018) Dynastic life-cycle discrete choice models 1213 Table 1. Simplified discrete choice Monte Carlo simulation results. Pseudo Maximum Likelihood Nested Fixed Point (ML) Sample Size Sample Size 1000 10,000 20,000 40,000 1000 10,000 20,000 40,000 θ=025 Mean 024473 024935 024886 024881 022714 024571 023320 024477 Std. Dev. 004991 001328 000915 000668 004884 001354 002135 001019 Bias −000527 −000065 −000114 −000119 −002286 −000429 −001680 −000523 MSE 000249 000017 000008 000005 000288 000020 000073 000013 λ=08 Mean 080425 079745 079797 079673 077538 078966 076934 078855 Std. Dev. 011241 003175 002157 001587 009211 003244 003656 002063 Bias 000425 −000255 −000203 −000327 −002462 −001034 −003066 −001145 MSE 001253 000100 000046 000026 000901 000115 000226 000055 β=095 Mean 094208 095245 095037 095136 093441 095227 094603 095027 Std. Dev. 006276 001893 001301 000934 005322 001983 001820 001236 Bias −000792 000245 000037 000136 −001559 000227 −000397 000027 MSE 000396 000036 000017 000009 000305 000039 000034 000015 Avg. Comp. time 065 288 606 1260 3476 3764 4675 5098 Note: The pseudo maximum likelihood corresponds to the estimation conducted by the new estimator using PML and maximum likelihood (ML) estimation is by the nested fixed point (NFXP). All simulations were conducted using the programming language GAUSS on a 2-CPU 1.66-GHz, 3-GB RAM laptop computer. The unit of time is seconds. The mean, empirical standard deviation, bias, and mean squared error (MSE) of each parameter estimate are reported in the respective column for each sample size. The bias and the MSE are calculated relative to the original data-generating value of the parameter. The data-generating value of the parameter is also reported at the center of the summary statistics block for that parameter. mated nonparametrically using the generated sample. Next, we estimate the model by NFXP and PML.18 Table 1presents the results of the estimation. We find that the finite-sample properties of the estimators improve monotonically with sample size. In the NFXP estimation, the mean square error (MSE) of θdrops quickly as the sample size increases. The results for the discount factors are similar: MSEs fall as the sample size increases. In the PML estimation, we observe a similar pattern for all estimators. We obtain similar results from the NFXP and PML parameters. For the sample size of 1000,thePMLestimate of the MSE of θ0is 000249 compared with 000288 from the NFXP. The PML estimate of the MSE of λis 001253 compared with 000901, and the PML estimate of the MSE of β is 000396 compared with 000305. For the sample sizes of 10,000,20,000,and40,000,the MSEs obtained from PML estimation is lower than the MSEs obtained from the NFXP, but the magnitudes are still very close. In terms of biases, the two estimation algorithms are also quite similar. The major difference between the two estimation algorithms is computational time, which varies greatly between the NFXP and PML even though we 18As illustrated in the estimation section, intergenerational models at the final step can be estimated either by the PML or GMM method. For this simulation study, we used the PML because it is more comparable to the full solution maximium likelihood. 1214 Gayle, Golan, and Soytas Quantitative Economics 9 (2018) simulate a very simple model. The average computational time for the NFXP for a sample of 1000 is 3476seconds, but it is only 065 seconds for the PML estimation, meaning the PML was 530 times faster. For the sample size of 40,000, computation times are 5098 and 126seconds for the NFXP and PML, respectively, a ratio of 404. 4. Extensions The dynastic framework developed so far in this paper has three major drawbacks. First, parts of the parental investment and transfers from parents to children are monetary in nature. Additionally, for exposition purpose we assume that there were not borrowing or saving. Monetary investment and/or parental transfers, such as paying for college or purchasing a house for their children, are most naturally characterized as a continuous choice. Also it is natural to introduce borrowing and saving as a continuous choice. Second, the framework assumes that gender does not matter. However, there are significant differences in the cost, choices, and opportunities over an individual’s lifetime that are gender specific. Third, which is related to gender but not specific to it, is that individuals normally form households and it takes a man and a woman to reproduce, and fertility is central to the model. In this section, we consider extensions to the basic framework that account for these three shortcomings. 4.1 Continuous choice and transfer For the estimation technique developed above to be applicable to a dynastic framework, two features must be present. First, all choices must be discrete, and second, all systematic state variables, at the initial stage and in every period during the life cycle, must have a discrete support. We replace these assumptions with two weaker assumptions. The first is that there must be at least one discrete choice variable. This requirement is easily satisfied as birth decision is naturally discrete. The second is that the initial systematic state variables (i.e., endowment that an individual starts adult life with) must belong to a finite set with discrete support. This is weaker than the original assumption and is a less restrictive requirement; it is satisfied in a nontrivial number of economic dynastic models—for example, in models where human capital is the major intergenerational transfer and even in models of bequests once the amount transferred is discretized. In practice, in most dynamic programming models, the state space is normally discretized. This requirement, however, relaxes the assumption that state space is discrete for the entire lifetime and that all choice variables are discrete. While bequests and initial wealth still must be discrete, the framework allows for any transfers and investments the parents make during their lifetime and map into discrete initial conditions of the child, such as education, houses, or other assets that are discrete in nature. If these assumptions are satisfied, then we can modify the representation and then use the estimation strategy for the mixed discrete and continuous choice model.19 19See Altug and Miller (1998), Gayle and Miller (2004), Gayle and Golan (2012), and Gayle (2015)for application of CCPs estimators with mixed discrete and continuous choices. Quantitative Economics 9 (2018) Dynastic life-cycle discrete choice models 1215 For illustration purposes, we extend our framework to include continuous choice of assets and bequests, assuming that we observed data on the per-period assets, At,which is continuous. We assume that individuals beginning their life as adults with asset level A0. This level is a bequest from the parents. The initial level of assets, j, is discrete with: A0=[A1 0AJ 0]. The budget constraint is given by At+1−(1+r)At=wt(x ht)−pcnt −ct(26) where ris the interest rate for borrowing and savings, and the right-hand side is the household income net of expenditures on children and consumption. A few remarks are in order. First, in order to map the assets at age T,AT, to a discrete bequest level that individuals start their life with, A 0, we define a transition function Pr(A 0=Aj 0|AT). Second, there are different ways to model markets completeness or incompleteness that will translate into different restrictions on savings and assets levels. For illustration purposes, we will assume the interior solution for all asset choices and will ignore such restrictions in this presentation. We do not restrict the initial and terminal asset levels to be nonnegative. However, the framework can be adjusted to include all these different types of constraints. Let us further assume that the parents’ asset levels can potentially affect educational outcomes of children: higher savings of parents increase the probability of a higher level of educational attainment of the child.20 We redefine the vector of state variables zt to capture these new assumptions, zt=({Ik1}17 k=0{Ikt−1}17 k=0A1At−1x) with x∈{A0x1x|X|}, a discrete set with finite support. Thus xincludes all the characteristics that a person is endowed with at the beginning of life. In this application, it included the initial (discrete) levels of assets inherited from the parents. As before, M(x|zT+1)is the intergenerational transition probability of xconditional on a parent’s endowment, x, and the parent’s choices over his/her lifetime. It includes the education, inherited assets, and potentially skills, for example (as well as traits such as gender and race). As before, it is derived from an education production function, M(x|zT), and is augmented to incorporate Pr(A 0=Aj 0|AT), the assets transition functions. Let Io kt and Ao tbe the sequence of optimal choice over the parent’s lifetime. Also, plugging the budget constraint in the utility from consumption, we redefine the systematic part of current utility in equation (8)as ukt(ztAt)=u1kt(zt)+utwt(xht)−pcnt −At+1+(1+r)Atzt(27) Then the lifetime expected utility excluding the dynastic component at the start of an adult’s life becomes UgT (x) =E0T  t=0 βt17  k=0 Io ktu1ktztAo t+εktx(28) 20Assets can be a proxy of the ability to pay for college, for example. However, we allow for assets to impact educational outcomes in order to illustrate the general nature of the extension. One can think of the continuous variable as expenditure on children if observed in the data. 1216 Gayle, Golan, and Soytas Quantitative Economics 9 (2018) The preference shock εkt is associated with the discrete choices in period tand not the continuous choice variables; therefore, it is still indexed with k.Asbefore,wecanwrite the value function of the problem, which represents the expected present discounted value of lifetime utility from following Ioand Ao t,givenztand εt,as V(z t+1εt+1)=max It+1At+1 EIA T  t=t+1 βt−t 17  k=0 Iktukt(ztAt)+εkt +βT−tλN(zT)−ν N  n=1 ETUg+1nx n|zTzt+1εt+1 (29) By Bellman’s principle of optimality, the value function can be defined recursively as V(z tεt)= 17  k=0'Io kt(ztεt)uktztAo t(zt)+εkt +βVzεfε(ε)dεdFkz|ztAt]( where fε(ε) is the continuously differentiable density of Fε(ε0tε17t),and Fk(z|ztAt)is a transition function for state variables that is conditional on choices Io kt =1and At=A0 t.NotethatIo kt(ztεt)is a function of ztand εt, while Ao t(zt)is a function of only zt. This is a consequence of the additive separability of the preferences shock, which will not affect the continuous choice as demonstrated below. The ex ante value function is then V(z t)= 17  k=0 pk(zt)uktztAo t(zt)+Eε[εkt|Ikt =1zt] +βVzdFkz|ztAt (30) In this form, V(z t)is now a function of the CCPs, the continuous choice decision rule, the expected value of the preference shock, the per-period utility, the transition function, and the ex ante continuation value. All components except the conditional probability, the continuous choice decision rule and the ex ante value function are primitives of the initial decision problem. By writing the CCPs and the continuous choice decision rule as a function of just the primitives and the ex ante value function, we can characterize the optimal solution of the problem (i.e., the ex ante value function) as implicitly dependent on the primitives of the original problem. Let us define the conditional value function, υk(ztAt),as υk(zt)=max Atukt(ztAt)+βVzdFkz|ztAt(31) Quantitative Economics 9 (2018) Dynastic life-cycle discrete choice models 1217 Therefore, the probability of observing choice k, conditional on zt,pk(zt),isstillgiven by pk(zt)= k=k 1υk(zt)−υk(zt)≥εkt−εtkfε(εt)dεt(32) However, the optimal continuous choice is found in two steps. First, find the optimal choice conditional on Ikt =1, which is defined as Atk(zt). This is characterized by the following Euler equation: ∂ukt(ztAt) ∂At =−β ∂VzdFkz|ztAt ∂At (33) Then substitute it into the conditional valuation function: υk(zt)=uktztAtk(zt)+βVzdFkz|ztAt(34) and find the optimal discrete choice: Io(ztεt)=argmax I 17  k=0 Iktυk(zt)+εkt Finally, we obtain the optimal continuous choice by setting Ao t(zt)=Atk(zt)if Io kt(zt εt)=1. We now can find an alternative value function that is a function of only pk(zt), Atk(zt), and the primitives of the model. We can now state a more general version of Proposition 1. Proposition 2. There exists an alternative representation for the ex ante conditional value function at time tthat is a function of only the primitives of the problem and the CCPs as follows: υk(zt)=uktztAtk(zt) + T  t=t+1 βt−t 17  s=0[ps(zt)ustztAtk(zt) +Eε(εst|Ist=1zt)dFo k(zt|zt)(35) +λβT−tNT(zT)−ν NT  n=1 x Vx × KT  s=0Mn kx|zTps(zT)dFo k(zT|zt) 1218 Gayle, Golan, and Soytas Quantitative Economics 9 (2018) where Fo k(zt|zt)is the t−tperiod-ahead optimal transition function,recursively defined as Fo k(zt|zt)= ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ Fzt|ztIkt =1Atk(zt) for t−t=1 17  r=0 zt−1 pr(zt−1)Fzt|zt−1Irt−1=1At−1k(zt−1)Fo k(zt−1|zt) for t−t>1 where NT(zT)is the number the children induced from zT,KTis the number of possible choice combinations available to the individual in the terminal period (in which birth is no longer feasible), and Mn k(x|zT)=M(x|zT)conditional on IkT =1for the nth child born in a parent’s life cycle. This representation is similar to the one in Proposition 1except for the inclusion of Atk(zt)and the replacement of an integral for a summation deal with the continuous state variables over the life cycle. The inversion—and hence the estimation—follows through as before except we now need a first-stage consistent estimate of Atk(zt)as well. This is obtained as Atk(zt)=E[At|ztIkt =1].21 4.2 Household and gender We extend the basic framework to include household decisions and gender. To the best of our knowledge, no other paper estimates dynastic models with household decisions. There are many models of household decisions; here, we show how to extend the model to incorporate a unitary decision-maker. The framework can be extended to deal with collective household decisions: see Gayle, Golan, and Soytas (2014) for an application of this estimation technique to a noncorporative collective model of household behavior. Let an individual’s gender, subscripted as σ, take the value of mfor a male and f for a female: σ={f m}. Gender is included in the vector of invariant characteristics xσ. Let Kdescribe the number of possible combinations of actions available to each household. Individuals get married at time zero, and for simplicity we assume there is no divorce (see Gayle, Golan, and Soytas (2014) for an application with marriage and divorce). Households are assumed to live for Tperiods and die together. Time zero is normalized to account for the normal age gap between married couples, which would imply that men have a longer childhood than women. All individual variables and earnings are indexed by the gender subscript σ. We omit the gender subscript when a variable refers to the household (both spouses). The state variables are extended to include the gender of the offspring. Let the vector ζtindicate the gender of a child born at age t,whereζt=1 if the child is a female and ζt=0otherwise. The vector of state variables is expanded to include the gender of the offspring is as follows: zt={Ik1}K k=0{Ikt−1}K k=0ζ0ζt−1xfxm 21We are assuming that there is no additional stochastic element in the determination of Atk(z). Quantitative Economics 9 (2018) Dynastic life-cycle discrete choice models 1225 assignment functions. All these functions are fundamental parameters of our model and are estimated outside the estimation of the preferences, discount factors, and the net costs of raising children. The first-stage estimates also include equilibrium objects such as the CCPs. Below we present estimates on the main earnings equation, the unobserved skills function, and the intergenerational education production function. The estimates of the marriage assignment functions and the CCPs are included in the Supplementary Material (Gayle, Golan, and Soytas (2018)). Table 3. Estimates of earnings equation: Dependent variable: Log of yearly earnings. Variable Estimate Variable Estimate Variable Estimate Demographic Variables Fixed Effect Age squared −40e–4Female ×Full-time work −0125 Black −0154 (10e–5)(0010)(0009) Age ×LHS 0037 Female ×Full-time work (t−1)0110 Female −0484 (0002)(0010)(0007) Age ×HS 0041 Female ×Full-time work (t−2)0025 HS 0136 (0001)(0010)(0005) Age ×SC 0050 Female ×Full-time work (t−3)0010 SC 0122 (0001)(0010)(0006) Age ×COL 0096 Female ×Full-time work (t−4)0013 COL 0044 (0001)(0010)(0006) Current and Lags of Participation Female ×Part-time work (t−1)0150 Black ×HS −0029 Full-time work 0938 (0010)(0010) (0010)Female ×Part-time work (t−2)0060 Black ×SC 0033 Full-time work (t−1)0160 (0010)(0008) (0009)Female ×Part-time work (t−3)0040 Black ×COL 0001 Full-time work (t−2)0044 (0010)(0011) (0010)Female ×Part-time work (t−4)−0002 Female ×HS −0054 Full-time work (t−3)0025 (0010)(0008) (0010)Individual specific effects Yes Female ×SC 0049 Full-time work (t−4)0040 (0006) (0010)Female ×COL 0038 Part-time work (t−1)−0087 (0007) (0010)Constant 0167 Part-time work (t−2)−0077 (0005) (0010) Part-time work (t−3)−0070 (0010) Part-time work (t−4)−0010 Hausman Statistics 2296 (0010)Hausman p-value 0000 No. of Observations 134,007 No. of Individuals 14,018 R2044 0278 Note: Standard errors are listed in parentheses. LHS indicates completed education of less than high school; HS indicates completed education of high school but not college; SC indicates completed education of some college but not a graduate; COL indicates completed education of at least a college degree. 1226 Gayle, Golan, and Soytas Quantitative Economics 9 (2018) Earnings equation and unobserved skills.Table3presents the estimates of the earnings equation and the function of unobserved (to the econometrician) individual skill (see also Gayle, Golan, and Soytas (2014)). The top panel of the first column shows that the age-earnings profile is significantly steeper for higher levels of completed education; the slope of the age-log-earnings profile for a college graduate is about three times that of an individual with less than a high school education. However, the largest gap is for college graduates; the age-log-earnings profile for a college graduate is about twice that of an individual with only some college. These results confirm that there are significant returns to parental time investment in children in terms of the labor market because parental investment significantly increases the likelihood of higher education outcomes, which significantly increases lifetime labor market earnings. The bottom panel of the first column and the second column of Table 3show that male full-time workers earn 26times more than part-time male workers and female full-time workers earn 23times more than females part-time workers (see also Gayle, Golan, and Soytas (2014)). It also shows that there are significant returns to past fulltime employment for both genders; however, females have higher returns to full-time labor market experience than males. The same is not true for part-time labor market experience; males’ earnings are lower if they worked part-time in the past, while there are positive returns to the most recent female part-time experience. However, part-time experiences 2and 3years in the past are associated with lower earnings for females; these rates of earnings reduction are, however, lower than those for males. These results are similar to those in Gayle and Golan (2012) and perhaps reflect statistical discrimination in the labor market in which past labor market history affects employers’ beliefs about workers’ labor market attachment in the presence of hiring costs.24 These results imply there are significant costs in the labor market in terms of the loss of human capital from spending time with children, if spending more time with children comes at the expense of working more in the labor market. These costs may be smaller for females than males because part-time work reduces compensation less for females than for males. If a female works part-time for 3years, for example, she loses significantly less human capital than a male working part-time for 3years instead of full-time. This difference may give rise to females specializing in child care; this specialization comes from the labor market and production function of a child’s outcome, as is the current wisdom. The unobserved skill (to the econometrician) is assumed to be a parametric function of the strictly exogenous time-invariant components of the individual variables. This assumption is used in other papers (e.g., MaCurdy (1981), Chamberlain (1986), Nijman and Verbeek (1992), Zabel (1992), Newey (1994), Altug and Miller (1998), and Gayle and Viauroux (2007)). It allows us to introduce unobserved heterogeneity to the model while still maintaining the assumption on the discreteness of the state space of the dynamic programming problem needed to estimate the structural parameters from the dynastic model. Also because the unobserved skill is estimated in the first step, and hence is data in the subsequence steps this does not introduce the standard initial condition problem. The Hausman test statistic shows that we cannot reject this correlated fixed effect 24These results are also consistent with part-time jobs differing from full-time jobs for males more than for females. Quantitative Economics 9 (2018) Dynastic life-cycle discrete choice models 1227 Table 4. Three-stage least squares estimation of the education production function. Variable High School Some College College High school father 0063 0003 −0002 (0032)(0052)(00435) Some college father 0055 0132 0055 (0023)(0038)(0031) College father −0044 0008 0120 (0032)(0051)(0042) High school mother 0089 0081 −0019 (0040)(0065)(0052) Some college mother 0007 −0041 0017 (0030)(0049)(0039) College mother 0083 0120 0040 (0036)(0057)(0047) Mother’s time −0014 0080 0069 (0021)(0034)(0027) Father’s time 0031 0100 0026 (0019)(0029)(0025) Mother’s labor income −0025 −0013 0005 (0009)(0014)(0011) Father’s Labor Income 0001 0001 0002 (0003)(0004)(0003) Female −0002 0135 0085 (0017)(0028)(0022) Black 0020 0082 0043 (0039)(0063)(0051) No. of siblings under age 3 −0014 −0107 −0043 (0017)(0027)(0022) No. of siblings between ages 3 and 6 −0029 −0047 −0012 (0019)(0030)(0025) Constant 0855 −0231 −0359 (0108)(0172)(0140) Observations 1335 1335 1335 Note: Standard errors are listed in parentheses; the excluded class is less than high school. Data are from the FamilyIndividual File of the Michigan Panel Study of Income Dynamics (PSID), and include individuals surveyed between 1968 and 1997. Instruments: Mother’s and father’s labor market hours over the child’s first 8 years of life, linear and quadratic terms of mother’s and father’s age when the child was 5 years old. specification. Column (3) of Table 4presents the estimate of skills as a function of unobserved characteristics; it shows that blacks and females have lower unobserved skills than whites and males. This could capture labor market discrimination. Education increases the level of skills but it increases at a decreasing rate with the level of completed education. The rates of increase for blacks and females with some college and a college degree are higher than those of their white and male counterparts. This pattern is reversed for blacks and females with a high school diploma. Notice that skills are another transmission mechanism through which parental time investment affects labor market earnings in addition to education. Intergenerational education production function. A well-known problem with the estimation of production functions is the simultaneity of the inputs (time spent with chil- 1228 Gayle, Golan, and Soytas Quantitative Economics 9 (2018) dren and income). As is clear from the structural model, the intergenerational education production function suffers from a similar problem. However, because the output of the intergenerational education production (i.e., completed education level) is determined across generations while the inputs, such as parental time investment, are determined over the life cycle of each generation, we can treat these inputs as predetermined and use instruments from within the system to estimate the production function. Table 4presents results of a three-stage least squares estimation of the system of individual educational outcomes; the estimates of the two other stages are available in the Supplementary Material (Gayle, Golan, and Soytas (2018)). The system includes the linear probabilities of the education outcomes, Pr(Ed σ|zT+1),aswellasthelaborsupply, income, and time spent with children equations. The estimation uses the mother’s and father’s labor market hours over the first 5years of the child’s life as well as linear and quadratic terms of the mother’s and father’s age on the child’s fifth birthday as instruments. The estimation results show that controlling for all inputs, a child whose mother has a college education has a higher probability of obtaining at least some college education and a significantly lower probability of not graduating from high school relative to a child with a less educated mother; while the probability of graduating from college is also larger, it is not statistically significant. If a child’s father, however, has some college or a college education, the child has a higher probability of graduating from college. This is consistent with the findings of Rios-Rull and Sanchez-Marcus (2002). We measure parental time investment as the sum of the parental time investment over the first 5years of the child’s life. The total time investment (i.e., the sum of the per-period investment of the first 5years of a child’s life) is a variable that ranges between 0and 10 because low yearly parental investment is coded as one and high yearly parental investment is code as two. The results in Table 5show that while a mother’s time investment significantly increases the probability of a child graduating from college or having some college, a father’s time investment significantly increases the probability of the child graduating from high school or having some college. These estimates suggest that while a mother’s time investment increases the probability of a high educational outcome, a father’s time investment truncates low educational outcome. However, the time investment of both parents is productive in terms of their children’s education outcomes. It is important to note that mothers’ and fathers’ hours spent with children are at different margins, with mothers spending significantly more hours than fathers. Thus, the magnitudes of the discrete levels of time investment of mothers and fathers are not directly comparable since low and high investment of time differs across genders. 6.2 Second-stage estimation This section presents estimates of the intergenerational and intertemporal discount factors, the preference parameters, and child care cost parameters. Table 6presents the discount factors. It shows that the intergenerational discount factor, λ,is0795.Thisimplies that in the second-to-last period of the parent’s life, a parental valuation of their child’s utility is 795% of their own utility. The estimated value is in the same range of values obtained in the literature calibrating dynastic model (Rios-Rull and SanchezMarcos (2002), Greenwood, Guner, and Knowles (2003)). However, these models do not Quantitative Economics 9 (2018) Dynastic life-cycle discrete choice models 1229 Table 5. Structural estimates of discount factors and utility parameter. Variable Estimates Variable Estimates Discount factors Disutility/Utility of Choices β0816 Wife Husband (0002)Labor supply λ0795 No work Part-time −0512 (0200)(0005) υ0248 No work Full-time 0207 (0168)(0009) Marginal Utility of Income Part-time No work −2023 Family labor income 0480 (0003) (0004)Part-time Part-time −1168 Children ×Family labor income −0466 (0009) (0066)Part-time Full-time −0605 Children ×HS ×Family labor income 1216 (0008) (0065)Full-time No work −0408 Children ×SC ×Family labor income 1279 (0007) (0066)Full-time Part-time −124532 Children ×COL ×Family labor income 1300 (0011) (0065)Full-time Full-time 0001 Children ×HS spouse ×Family labor income −1017 (0010) (0066)Time with children Children ×SC spouse ×Family labor income −0995 Low Medium 0502 (0066)(0014) Children ×COL spouse ×Family labor income −0992 Low High 0564 (0066)(0013) Children ×Black ×Family Labor Income −0108 Medium Low −0169 (0004)(0008) Medium Medium 0129 (0010) Medium High 0593 (0013) High Low −0364 (0007) High Medium 0353 (0011) High High −0140 (0012) Birth 0701 (0025) Note: Standard errors are listed in parentheses. LHS indicates completed education of less than high school; HS indicates completed education of high school but not college; SC indicates completed education of some college but not a graduate; COL indicates completed education of at least a college degree. The excluded choice is no work, no time with children, and no birth for both spouses. include the life cycle. The estimated discount factor, β,is081. The discount factor is smaller than typical calibrated values; however, the few papers that have estimated it 1230 Gayle, Golan, and Soytas Quantitative Economics 9 (2018) Table 6. Model fit and counterfactual. Labor Supply Time Investment Whites Wife Wife Data Model 1 NN1NN1Model 2 NN2NN2Data Model 1 NN1NN1Model 2 NN2NN2 No work 02634 02599 03918 03750 03649 03730 02896 Low 06363 08999 08485 08464 09530 06004 06816 Part-time 01596 01622 01416 01451 01790 02193 02219 Medium 02257 00531 00808 00820 00251 02260 01771 Full-time 05770 05779 04666 04800 04561 04076 04886 High 01380 00470 00707 00716 00220 01736 01413 Husband Husband Data Model 1 NN1NN1Model 2 NN2NN2Data Model 1 NN1NN1Model 2 NN2NN2 No work 00290 00250 00233 00233 00245 00145 00201 Low 08237 09592 09459 09457 09796 08688 08985 Part-time 00306 00361 00336 00329 00447 00295 00295 Medium 01008 00238 00310 00310 00113 00807 00608 Full-time 09404 09390 09432 09439 09307 09560 09505 High 00755 00170 00231 00232 00091 00505 00407 Blacks Wife Wife Data Model 1 NN1NN1Model 2 NN2NN2Data Model 1 NN1NN1Model 2 NN2NN2 No work 01998 01309 01258 01232 03045 03309 02467 Low 06837 09046 08678 08694 07840 06550 07460 Part-time 01002 02150 02145 02154 02449 02142 02257 Medium 02192 00497 00690 00687 01210 01964 01371 Full-time 07000 06541 06597 06614 04506 04550 05277 High 00971 00457 00632 00618 00951 01486 01169 Husband Husband Data Model 1 NN1NN1Model 2 NN2NN2Data Model 1 NN1NN1Model 2 NN2NN2 No work 00640 00596 00587 00575 00788 00473 00567 Low 08338 09729 09638 09641 09299 09124 09320 Part-time 00423 00555 00471 00468 00553 00397 00403 Medium 00744 00123 00158 00158 00326 00479 00336 Full-time 08937 08850 08942 08957 08659 09130 09030 High 00919 00148 00204 00201 00375 00397 00344 Birth Whites Data Model 1 NN1NN1Model 2 NN2NN2 No birth 09014 09551 09387 09383 09701 08148 08833 Birth 00986 00449 00613 00617 00299 01852 01167 Blacks Data Model 1 NN1NN1Model 2 NN2NN2 No birth 08955 09249 09129 09135 08847 08349 08784 Birth 01045 00751 00871 00865 01153 01651 01216 Note: Model 1 is the baseline model with all parameters estimated. Model 2 is the model with discount factors calibrated (β=090λ=095). NNii =12removes Nature from the intergenerational educational production with full reoptimization. NNii=12removes Nature from the intergenerational educational production but does not allow for reoptimization of subsequent generations. Quantitative Economics 9 (2018) Dynastic life-cycle discrete choice models 1231 find similar values (e.g., Arcidiacono, Sieg, and Sloan (2007), find it to be 0.8).25 Lastly, the discount factor associated with the number children, υ,is025 which implies that the marginal increase in value from the second child is 068 and from the third child is 060. Identification of the discounts factors are nontrivial in dynamic discrete models. Here, we have three discount factors to identify instead of one as in the standard discrete choice models. However, note that past home hours, when the children are young, affect only the transition functions and not the current utility, so we have the common exclusion restrictions used to identify dynamic discrete choice models (Magnac and Thesmar (2002), Fang and Wang (2015)). Then the identification of the intergenerational discounts factors follows by a direct application of the proof of Proposition 2 in Fang and Wang (2015) to our setting. Table 5also presents the marginal utility of income, which is positive and increasing with the number of children except for a household with a college graduate wife and a husband with at least a high school education. Also, a husband’s education decreases the marginal utility of income for families with children. The marginal utility of income for families with children is also lower for black families. The right panel of Table 5presents our estimates of the disutility/utility from various combination of household choices. As is usual in discrete choice models, these are estimated relative to an outside choice, which is both spouses not (i) working, (ii) giving birth, or (iii) spending any time with young children. We also use an additive specification in which the costs of birth, work, and time with children are additively separable. First, every labor supply choice of the household carries with it a disutility relative to the reference choice except for households in which both spouses work full-time (which statistically is no different from disutility/utility reference) and when the wife does not work and the husband works full-time. In the data, if both spouses spend low time with children and there is no birth, then both spouses are equally likely to be observed working full-time than not working; hence the equal utility for both sets of choices. Second, there are no distinct patterns to utility from time with children; these estimates are highly nonlinear, perhaps reflecting that it is a mixture of leisure and disutility. However, giving birth provides a positive utility. This implies that, although parents get utility from the quality of their children, they also get some instantaneous utility from a birth. 6.3 Model fit In this section, we first assess the ability of our model to reproduce the basic stylized facts by race, gender, and marital status. We assess how well our model predicts the choices of labor supply, home hours with young children, and birth. Our model is over identified and passes the standard over identifying restrictions J-test. In the estimation, the CCPs are targeted; in the model fit analysis, we simulate a sample of individuals and determine whether the individuals in our simulated sample behave like the individuals in our data. In some regards, this exercise is equivalent to a graphical summary of our model’s over identification test. 25We are not aware of dynastic models in which the time discount factor is estimated. 1232 Gayle, Golan, and Soytas Quantitative Economics 9 (2018) Table 6presents the model’s fit. The model matches the labor supply patterns between gender and across race well. While it also matches the variation across race and gender for parental time with children, the levels are not similar in all cases. In examining the birth decisions, the model produces the differences in birth rates across households of different race, but it underpredicts the fecundity of whites by about a half. This lower birth rate is partly rationalized by the lower time with children predicted by the model. Nevertheless, our empirical model specification is very parsimonious: We do not include race, education, or marital status in the preference parameters for the disutility/utility of the different choices. In addition, the only unobserved heterogeneity is estimated from the earnings equations. Still, the model performs well in replicating the data based primarily on the economic interactions embodied in it. 6.4 The effect of nature versus nurture on intergenerational mobility One of the major benefits of our approach is the ability to do full blown counterfactual analysis. An alternative approach which has been used in empirical structural models to incorporate intergenerational/altruistic concerns is to modify the standard dynamics structural estimation methods by introducing an approximation for the value parents place on their children’s adults outcome as a function of some state variables, normally the educational outcome or test scores (see, e.g., Bernal (2008), Brown and Flinn (2011), and Del Boca, Flinn, and Wiswall (2013) among others). The main advantage of this alternate approach is that the estimation is easier and standard techniques in the literature can be used. The major disadvantage is that welfare/counterfactual analysis is subject to Lucas’ critique. That is, in a counterfactual environment the value parents place on their children quality changes in two ways, the value of the state variables and the functional form of the mapping between the state variables and the utility derived from those state variables. The is made obvious by an examination of equations (18)and(20). This alternative approach does not allow the functional form of the mapping to change. From equations (18)and(20), one can see that our framework can nested this alternative approach, since the approximation used in the alternative approach is equivalent to conducting a welfare/counterfactual analysis holding fixed the CCPs and transition used in the calculation of the value of a child. To illustrate the bias induced by ignoring the fact that the children themselves will reoptimize when we change the economic environment, we asked the counterfactual question of how much of the mobility across generations is due to the automatic transmission education across generations as opposed to differences in parent investment between parents of different educational background. To do this, we eliminate the automatic transition of education in production function.26 We report the results from two models. Model (1) is the model estimated in above. Model (2) is a model in which the discount factors (βand λ)aresettovalues commonly used in the literature (β=090 and λ=095) and all other parameters estimated. 26Operationally, we set the effect on education in the intergenerational production function equal to that of a high school graduate irrespective of the mother’s and father’s education level. Quantitative Economics 9 (2018) Dynastic life-cycle discrete choice models 1233 Figure 1. Counterfactuals and mobility. Note: Model 1 is the baseline model with all parameters estimated. Model 2 is the model with discount factors calibrated (β=090λ =095). NN(i) i =12removes Nature from the intergenerational educational production with full re-optimization. NN(i) i =12removes Nature from the intergenerational educational production but Do Not allow for re-optimization of subsequent generations. Table 6presents the summary of labor supply, time investment, and birth rate by gender and race for the data, baseline models and counterfactual simulations. It shows that if we eliminate the portion of parental education that is transmitted automatically across generation then parents will reoptimize and change labor supply, time investment, and fertility behaviors. Therefore, a pure statistical decomposition would be inappropriate for answering the question of how much mobility would change if there were no automatic (Nature) transmission of education from parents to children. The columns NN1presents the counterfactual estimates of our model and NN1the estimation results of the approximated model; similarly, NN2presents the results of the counterfactual from our model with (β=090 and λ=095)andNN2the results of the counterfactual of the approximated model. The columns NN1and NN2show that not taking into account that the all subsequent generations will also reoptimize induces significant bias with the bias being greater the larger the discount factors. To obtain a number that summarize the impact on mobility, Figure 1presents the probability that a children born in a family in the bottom 20 percent of the family income distribution will end up in a family with family income above the median of the next generation family income distribution. It shows that in the baseline model, that is, model 1, that only 30 percent of children born in the bottom 20 percent will end up with families earning above the median. However, if the automatic transition of education was eliminated that probability would increase by about 20 percenttoabout40 percent. However, ignoring the re-optimization of subsequent generations, we overestimate the impact of eliminating the automatic transition of education across generations on mobility by about 25 percent. Model (2) shows the similar qualitative patterns but shows 1234 Gayle, Golan, and Soytas Quantitative Economics 9 (2018) that the overestimate of the impact of “Nature” on mobility could be as high as 90 percent, which illustrate the gain from using the approach outline in this paper. 7. Conclusion This paper provides a new representation of the value function for discrete choice dynastic models that partially overcomes the curse of dimensionality of dynastic models by exploiting properties of the stationary decision rules. The representation can be used in multistage CCP estimators to estimate a rich class of dynastic models including investment in children’s human capital, monetary transfers, unitary households, endogenous fertility, and a life cycle within each generation. Under certain conditions, we show that the framework can also accommodate continuous choice variables. The paper extends methods used in the literature for the estimation of single-agent nondynastic models to the dynastic setting. The paper compares the performance of a multistage CCP estimator based on the new value function representation with a modified version of the full solution MLE using simulations and finds that the estimates are close to the nested fixed-point estimates as the sample increases but the computation time is reduced substantially. The paper then provides an application of a unitary household model in which households choose labor supply, time with children, and fertility; human capital is transmitted across generations by monetary and time investments of the parents. We then used the estimate model to conduct counterfactual simulations, investigating the role of the automatic transmission of education across generations (Nature effect) in accounting for the intergenerational immobility at the bottom of the income distribution. We find that without the Nature effect in the intergenerational education production function mobility at the bottom of the income distribution would have been 20 percent higher. Finally, not accounting for the reoptimization of sequent generations in the model, as is done in the approach outlined in this paper, will overstate the effect of the Nature on mobility by between 20 and 90 percent. Appendix Proof of Proposition 1. Recall the conditional value function in equation (14): υk(zt)=ukt(zt)+β zt+1 V(z t+1)F(zt+1|ztIkt =1) (46) We begin by noting that V(z t+1)= 17  s=0 ps(zt+1)ust+1(zt+1) +Eε(εst+1|Ist+1=1zt+1)(47) +β zt+2 V(z t+2)F(zt+2|zt+1Ist+1=1) Quantitative Economics 9 (2018) Dynastic life-cycle discrete choice models 1241 Laitner, J. (1992), “Random earnings differences, lifetime liquidity constraints, and altruistic intergenerational transfers.” Journal of Economic Theory, 58 (2), 135–170. [1196] Loury, G. C. (1981), “Intergenerational transfers and the distribution of earnings.” Econometrica, 49 (4), 843–867. [1196,1200] MaCurdy, T. E. (1981), “An empirical model of labor supply in a life-cycle setting.” Journal of Political Economy, 89 (6), 1059–1085. [1226] Magnac, T. and D. Thesmar (2002), “Identifying dynamic discrete decision processes.” Econometrica, 70 (2), 801–816. [1231] Miller, R. A. (1984), “Job matching and occupational choice.” Journal of Political Economy, 92 (6), 1086–1120. [1197] Moav, O. (2005), “Cheap children and the persistence of poverty.” Economic Journal, 115 (500), 88–110. [1198] Mookherjee, D., S. Prina, and D. Ray (2012), “A theory of occupational choice with endogenous fertility.” American Economic Journal: Microeconomics, 4 (4), 1–34. [1198] Newey, W. K. (1994), “The asymptotic variance of semiparametric estimators.” Econometrica, 62 (6), 1349–1382. [1226] Nijman, T. and M. Verbeek (1992), “Nonresponse in panel data: The impact on estimates of a life cycle consumption function.” Journal of Applied Econometrics, 7 (3), 243–257. [1226] Pakes, A. (1986), “Patents as options: Some estimates of the value of holding European patent stocks.” Econometrica, 54 (4), 755–784. [1197] Rios-Rull, J.-V. and V. Sanchez-Marcos (2002), “College attainment of women.” Review of Economic Dynamics, 5 (4), 965–998. [1207,1228] Pesendorfer, M. and P. Schmidt-Dengler (2008), “Asymptotic Least Squares Estimators for Dynamic Games.” The Review of Economic Studies, 75 (3), 901–928. [1207,1209] Rust, J. (1987), “Optimal replacement of GMC bus engines: An empirical model of Harold Zurcher.” Econometrica, 55 (5), 999–1033. [1197,1210] Wolpin, K. I. (1984), “An estimable dynamic stochastic model of fertility and child mortality.” Journal of Political Economy, 92 (5), 852–874. [1197] Zabel, J. E. (1992), “Estimating fixed and random effects models with selectivity.” Economics Letters, 40 (3), 269–272. [1226] Co-editor Peter Arcidiacono handled this manuscript. Manuscript received 21 September, 2016; final version accepted 24 October, 2017; available online 24 January, 2018.