Inference in dynamic discrete choice problems under local misspecification
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Bugni, Federico A.; Ura, Takuya Article Inference in dynamic discrete choice problems under local misspecification Quantitative Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Bugni, Federico A.; Ura, Takuya (2019) : Inference in dynamic discrete choice problems under local misspecification, Quantitative Economics, ISSN 1759-7331, The Econometric Society, New Haven, CT, Vol. 10, Iss. 1, pp. 67-103, https://doi.org/10.3982/QE917 This Version is available at: https://hdl.handle.net/10419/217137 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/
Quantitative Economics 10 (2019), 67–103 1759-7331/20190067 Inference in dynamic discrete choice problems under local misspecification Federico A. Bugni Department of Economics, Duke University Takuya Ura Department of Economics, University of California, Davis Single-agent dynamic discrete choice models are typically estimated using heavily parametrized econometric frameworks, making them susceptible to model misspecification. This paper investigates how misspecification affects the results of inference in these models. Specifically, we consider a local misspecification framework in which specification errors are assumed to vanish at an arbitrary and unknown rate with the sample size. Relative to global misspecification, the local misspecification analysis has two important advantages. First, it yields tractable and general results. Second, it allows us to focus on parameters with structural interpretation, instead of “pseudo-true” parameters. We consider a general class of two-step estimators based on the K-stage sequential policy function iteration algorithm, where Kdenotes the number of iterations employed in the estimation. This class includes Hotz and Miller (1993)’s conditional choice probability estimator, Aguirregabiria and Mira (2002)’s pseudo-likelihood estimator, and Pesendorfer and Schmidt-Dengler (2008)’s asymptotic least squares estimator. We show that local misspecification can affect the asymptotic distribution and even the rate of convergence of these estimators. In principle, one might expect that the effect of the local misspecification could change with the number of iterations K. One of our main findings is that this is not the case, that is, the effect of local misspecification is invariant to K. In practice, this means that researchers cannot eliminate or even alleviate problems of model misspecification by choosing K. Keywords. Single-agent dynamic discrete choice models, estimation, inference, misspecification, local misspecification. JEL classification. C13, C61, C73. Federico A. Bugni: [email protected] Takuya Ura : [email protected] We thank the three anonymous referees for comments and suggestions that have significantly improved this paper. We are also grateful for helpful discussions with Victor Aguirregabiria, Peter Arcidiacono, Joe Hotz, Shakeeb Khan, Matt Masten, Arnaud Maurel, Jia Li, and the participants at the Duke Microeconometrics Reading Group and the Yale Econometrics Lunch Group. Any and all errors are our own. The research of the first author was supported by NIH Grant 40-4153-00-0-85-399 and NSF Grant SES-1729280. ©2019 The Authors. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at http://qeconomics.org.https://doi.org/10.3982/QE917
68 Bugni and Ura Quantitative Economics 10 (2019) 1. Introduction This paper investigates the effect of model misspecification on inference in single-agent dynamic discrete choice models. Our study is motivated by two observations regarding this literature.1First, typical econometric frameworks used in empirical studies are heavily parametrized and are therefore subject to misspecification. Second, several methods can be used to estimate these models, including Rust (1987,1988)’s nested fixed-point estimator, Hotz and Miller (1993)’s conditional choice probability estimator, Aguirregabiria and Mira (2002)’s pseudo-likelihood estimator, and Pesendorfer and Schmidt-Dengler (2008)’s asymptotic least squares estimator. While the literature has studied the behavior of these estimators under correct specification, their properties under misspecification have not been explored. To the best of our knowledge, our paper is among the first ones to investigate the effect of misspecification on inference in these types of models. In this paper, we propose a local misspecification approach in dynamic discrete choice models. By local misspecification, we mean that the econometric model is allowed to be misspecified, but the amount of misspecification vanishes as the sample size increases. Local misspecification is an asymptotic device that can provide concrete conclusions in the presence of misspecification while keeping the analysis tractable. As with any other asymptotic device, local misspecification is just an approximation to a finite sample situation (in this case, with misspecification) and should not be taken literally. Our local approach to misspecification in dynamic discrete choice models yields relevant conclusions in several dimensions. First, local misspecification can constitute a reasonable approximation to the asymptotic behavior when the mistakes in the specification of the model are small. Second, there are multiple available estimation methods, and their performance under misspecification is not well understood. We believe their relative performance under local misspecification is a relevant comparison criterion. Finally, while our approach to misspecification is admittedly “local” in that it is assumed to vanish, we allow the rate at which this occurs to be completely arbitrary. In particular, we allow the misspecification to disappear at a faster, equal, or even slower rate than the regular parametric convergence rate of √n. While this rate will affect the asymptotic properties of the estimators under consideration, it will not alter the main qualitative conclusions of our paper. We consider a class of two-step estimators based on the K-stage sequential policy function iteration algorithm along the lines of Aguirregabiria and Mira (2002), where K denotes the number of iterations employed in the estimation. By appropriate choice of the criterion function, this class captures the K-stage maximum likelihood type estimators (K-ML) and the K-stage minimum distance type estimators (K-MD). This class includes most of the previously mentioned estimators as special cases. Our main theoretical contribution is to characterize the asymptotic distribution of two-step K-MD and K-ML estimators under local misspecification. We show that local 1See the survey papers by Aguirregabiria and Mira (2010) and Arcidiacono and Ellickson (2011), and references therein.
Quantitative Economics 10 (2019) Inference in dynamic discrete problems 69 misspecification can affect the asymptotic distribution and even the rate of convergence of these estimators. We are particularly interested in the asymptotic behavior of these estimators as we vary the number of iterations K. We obtain three main results. Our first result is related to the asymptotic behavior of K-ML estimators. Under correct specification, Aguirregabiria and Mira (2002)proved that the asymptotic distribution of K-ML estimators is invariant to K. Under local misspecification, however, one might reasonably expect a different result. Intuitively, every stage of the policy function iteration algorithm brings the estimator closer to imposing the fixed-point/equilibrium conditions implied by the model. Given that the model is incorrectly specified, one might then conjecture that increasing the number of iterations would result in an estimator of inferior quality (e.g., more bias).2Our first main result is to show that this intuition is incorrect. We formally show that K-ML estimators are asymptotically equivalent for all K. Our second result is to show an analogous result for K-MD estimators, that is, given the choice of weight matrix, K-MD estimators are asymptotically equivalent for all K. If we combine these findings, we can conclude that the researcher cannot eliminate or even alleviate a problem of model misspecification by choosing the number of iterations K. Additional iterations are computationally costly and produce no change in asymptotic efficiency. Thus, from a practical viewpoint, we recommend using either the 1-ML or the 1-MD estimator. Finally, our third result is to compare K-MD and K-ML estimators in terms of asymptotic mean squared error. We show that an optimally-weighted K-MD estimator depends on the unknown asymptotic bias and is thus generally unfeasible. In turn, the feasible K-MD estimator with a weight matrix that minimizes asymptotic variance could have an asymptotic mean squared error that is higher or lower than that of the K-ML estimator or the K-MD estimator with identity weight matrix. In other words, given a particular choice of the number of iterations K(e.g., K=1), the presence of local misspecification implies that we cannot make clear-cut recommendations regarding the weight matrix for the K-MD estimator, and how this compares with the K-ML estimator. From a technical viewpoint, our analysis exploits a distinctive feature of single-agent dynamic discrete choice problems known as the “zero Jacobian property.” This property was used by Aguirregabiria and Mira (2002) under correct specification to obtain their results for K-ML estimators. One of our technical contributions is to use this property under local misspecification to derive analogous results for both K-ML and K-MD estimators. As we have explained, this paper uses a local approach to the problem of model misspecification. In practice, researchers typically specify econometric models that may contain nonvanishing errors, that is, global misspecification. Relative to the global misspecification analysis, our local misspecification approach has two important advantages. First, allowing for global misspecification in our dynamic discrete choice model typically makes the problem intractable, and generally valid results are thus hard to obtain. In contrast, the local misspecification yields concrete and general conclusions. Second, recall that the literature has produced several estimation methods to estimate 2We are grateful to an anonymous referee for suggesting this interpretation.
70 Bugni and Ura Quantitative Economics 10 (2019) the structural parameter of interest in a dynamic discrete choice problem. Under global misspecification, the different estimators typically converge in probability to different pseudo-true parameters which may or may not be related to the true structural parameter. This makes the results hard to interpret and compare. In contrast, under local misspecification, these different estimators are shown to consistently estimate the true structural parameter value. We can then compare their robustness to misspecification via their asymptotic distributions. This paper relates to a vast literature on inference under model misspecification. White (1982,1996) considered the problem of maximum likelihood estimation under global misspecification. Newey (1985a,b)andTauchen (1985) investigated the power properties of the model specification tests under local misspecification. More recently, Schorfheide (2005) considered a locally misspecified vector autoregression process and proposed an information criterion for the lag length in the autoregression model. Bugni, Canay, and Guggenberger (2012) compared inference methods in partially identified moment (in)equality models that are locally misspecified. Kitamura, Otsu, and Evdokimov (2013) considered a class of estimators that are robust to local misspecification in the context of moment condition models. None of the previously mentioned references consider inference in dynamic discrete choice problems, whose specific features are central to the results in this paper. Few references explore the issue of misspecification in dynamic discrete choice problems. For example, Norets and Takahashi (2013) considered surjective dynamic discrete choice models and show that these are necessarily correctly specified. Last, Chernozhukov et al. (2016) considered two-step estimators in a dynamic discrete choice model with a locally misspecified first step, and proposed estimators that are robust to this issue. In contrast, this paper allows both steps to be locally misspecified. The remainder of the paper is structured as follows. Section 2describes the dynamic discrete choice model and introduces the possibility of its local misspecification. Section 3develops a general result for two-step K-stage estimators under high-level conditions. Section 4applies the general result to K-ML estimators (Section 4.1)andK-MD estimators (Section 4.2). Section 5presents results of Monte Carlo simulation and Section 6concludes. The Appendix of the paper collects all the proofs and intermediate results. The following notation is used throughout the paper. For any s1s2∈N,0s1×s2and 1s1×s2denote a (s1×s2)-dimensional matrix composed of zeros and ones, respectively, and Is1×s2denotes (s1×s2)-dimensional matrix equal to the left upper block of the (max(s1s2)×max(s1s2))-dimensional identity matrix. We use ·to denote the Euclidean norm. For any s-dimensional column vector V,diag{V}is the (s ×s)- dimensional matrix with Vas its diagonal. For sets of finite indices S1={1|S1|} and S2={1|S2|},{M(s1s2)}(s1s2)∈S1×S2denotes the (|S1|×|S2|)-dimensional column vector equal to the vectorization of {{M(s1s2)}|S1| s1=1}|S2| s2=1. For any differentiable matrix function F(y):Ra×b→Rc×d,∂F(y)/∂y ∈Rac×bd denotes the usual matrix of derivatives. Finally, “w.p.a.1” abbreviates “with probability approaching one.”
Quantitative Economics 10 (2019) Inference in dynamic discrete problems 71 2. Setup Section 2.1 describes the dynamic discrete choice model assumed by the researcher. This paper allows this model to be incorrectly specified. Section 2.2 describes the nature of the model misspecification. 2.1 The econometric model An economic agent is assumed to behave according to the discrete Markov decision framework in Aguirregabiria and Mira (2002). In each period t=1T ≡∞,the agent is assumed to observe a vector of state variables stand to choose an action at∈A≡{1|A|} with the objective of maximizing the expected discounted utility. The vector of state variables st=(xtεt)is composed by two subvectors. The subvector xt∈X≡{1|X|}represents a scalar state variables observed by the agent and the researcher, whereas the subvector εt∈R|A|represents an action-specific state vector only observed by the agent. The agent’s future state variables (xt+1εt+1)are assumed to follow a Markov transition probability density dPr(xt+1εt+1|xtεtat)that satisfies: dPr(xt+1εt+1|xtεtat)=gθg(εt+1|xt+1)fθf(xt+1|xtat) where gθg(·)is the (conditional) distribution of the unobserved state variable with parameter θgand fθf(·)is the transition probability of the observed state variable with parameter θf. The utility is assumed to be time separable and the agent discounts future utility by a known discount factor β∈(01).3The current utility function of choosing action at under state variables (xtεt)is given by uθu(xtat)+εt(at) where uθu(·)is nonstochastic component of the current utility with parameter θu,and εt(at)denotes the atth coordinate of εt. The researcher’s goal is to estimate the unknown parameters in the model, θ≡ (θgθuθf)∈Θ,whereΘis the compact parameter space. Also, we denote θ=(αθf)∈ Θ≡Θα×Θfwith α≡(θuθg)∈Θα. Following Aguirregabiria and Mira (2002), we impose the following regularity conditions on the primitive elements of the econometric model. Assumption 1. For every θ∈Θ,assume that: (a) For every x∈X,gθg(ε|x) has finite first moments and is twice differentiable in ε, (b) ε={ε(a)}a∈Ahas full support, (c) gθg(ε|x),fθf(x|x a),and uθu(xa) are twice continuously differentiable with respect to θ. 3This follows Aguirregabiria and Mira (2002, footnote 12) and Magnac and Thesmar (2002).
72 Bugni and Ura Quantitative Economics 10 (2019) By Blackwell (1965)’s theorem and its generalization by Rust (1988), the optimal decision rule is stationary and Markovian, that is, the time subscript can be dropped. Furthermore, the optimal value function Vθis the unique solution of the following Bellman equation: Vθ(xε) =max a∈Auθu(x a) +ε(a) +β(xε) Vθxεgθgε|xfθfx|x adxε(2.1) By integrating out the unobserved error, we obtain the smoothed value function: Vθ(x) ≡ε Vθ(xε)gθg(ε|x)dε which is the unique solution of the smoothed Bellman equation: Vθ(x) =ε max a∈Auθu(xa) +ε(a) +β x∈X Vθxfθfx|xagθg(ε|x)dε (2.2) We now turn to the description of the conditional choice probability (CCP), denoted by Pθ(a|x), which is the model-implied probability that an agent chooses action awhen the observed state is x. Since the agent chooses an action in A,Pθ(|A||x) = 1−a∈˜ APθ(a|x) for all x∈X. Thus, the vector of model-implied conditional choice probabilities (CCPs) is completely characterized by Pθ≡{Pθ(a|x)}(ax)∈˜ A×Xwith ˜ A≡ {1|A|−1}. For the remainder of the paper, we use ΘP⊂[01]|˜ A×X|to denote the parameter space for the vector of CCPs. The vector of CCPs is a central equilibrium object in the model. Lemma 2.1 shows that the CCPs are the unique fixed point of the policy function mapping. By utility maximization, the vector of CCPs is determined by the following equation: Pθ(a|x) ≡ε 1a=argmax ˜ a∈Auθu(x ˜ a) +β x∈X Vθxfθfx|x ˜ a+ε(˜ a)dgθg(ε|x) which can be succinctly represented as follows: Pθ=ΛθVθ(x)x∈X(2.3) Also, notice that Equation (2.2) can be rewritten as Vθ(x) = a∈A Pθ(a|x)uθu(xa) +Eθε(a)|xa+β x∈X Vθxfθfx|xa(2.4) where Eθ[ε(a)|xa]denotes the expectation of the unobservable ε(a) conditional on the state being xand on the optimal action being a. Under our assumptions, Hotz and Miller (1993) show that there is a one-to-one mapping between the CCPs and the (normalized) smoothed value function. The inverse of this mapping allows us to re-express {Eθ[ε(a)|xa]}(ax)∈A×Xas a function of the vector of CCPs. By combining this with Equation (2.4), we can express {Vθ(x)}x∈Xas a function of Pθ. An explicit formula for
Quantitative Economics 10 (2019) Inference in dynamic discrete problems 73 such function is provided in Aguirregabiria and Mira (2002, Equation (8)), which we succinctly express as follows: Vθ(x)x∈X=ϕθ(Pθ) (2.5) By combining Equations (2.3)and(2.5), we obtain the following fixed-point representation of the vector of CCPs: Pθ=Ψθ(Pθ) (2.6) where Ψθ≡Λθ◦ϕθis the policy function mapping. As explained by Aguirregabiria and Mira (2002), this operator can be evaluated at any vector of CCPs, optimal or not. For any arbitrary P≡{P(a|x)}(ax)∈˜ A×X,Ψθ(P) provides the current optimal CCPs of an agent whose future behavior is according to P. Under the current assumptions, the policy function mapping has several properties that are central to the results of this paper. Lemma 2.1. Under Assumption 1,Ψθsatisfies the following properties: (a) Ψθhas a unique fixed-point Pθ, (b) The sequence PK=Ψθ(PK−1)for K≥1,converges to Pθfor any initial P0∈ΘP, (c) The Jacobian matrix of Ψθwith respect to Pis zero at Pθ. Following the literature on estimation of dynamic discrete choice models, the researcher estimates θ=(αθf)using a two-step procedure. In a first step, he uses fθfto estimate θf. In a second step, he uses Ψ(αθf)(P) and the first step to estimate α.The following assumption ensures that the model is identified. Assumption 2. The parameter θ=(αθf)∈Θis identified as follows: (a) θfis identified by fθf,that is,fθfa =fθfb implies θfa =θfb, (b) αis identified by the fixed point condition Ψ(αθf)(P) =Pfor any (θfP)∈Θf×ΘP, that is,∀θf∈Θf,Ψ(αaθf)(P) =Pand Ψ(αbθf)(P) =Pimplies αa=αb. Magnac and Thesmar (2002) provide sufficient conditions for Assumption 2.Also, Assumption 2implies the higher level condition used by Aguirregabiria and Mira (2002, conditions (e)–(f) in Proposition 4). Under these conditions, we can deduce certain important properties for the model-implied CCPs. Lemma 2.2. Under Assumptions 1–2, (a) Pθis continuously differentiable, (b) ∂Pθ/∂θ =∂Ψθ(Pθ)/∂θ, (c) αis identified by P(αθf)for any θf∈Θf,that is,∀θf∈Θf,P(αaθf)=P(αbθf)implies αa=αb.
74 Bugni and Ura Quantitative Economics 10 (2019) Lemmas 2.1 and 2.2 are well-known results under correct specification. At the risk of being repetitive, we include these in the paper for two reasons. First, we note that these properties belong to the econometric model, regardless of whether it is correctly specified or not. Second, the results in the paper will repeatedly make reference to these properties. Thus far, we have described how the model specifies two conditional distributions: the CCPs and the transition probabilities. The remaining element of the specification is the marginal distribution of the state variables, which is left completely unspecified. 2.2 Local misspecification We now describe the true data generating process (DGP), denoted by Π∗ n(axx),and explain its relationship to the econometric model in Section 2.1. Hereafter, a superscript with asterisk denotes true value. By definition, the DGP is the product of the transition probability, the CCPs, and the marginal distribution of the state variable, that is, for all (axx)∈A×X×X, Π∗ naxx=f∗ nx|ax×P∗ n(a|x) ×m∗ n(x) (2.7) where f∗ nx|ax≡Π∗ naxx ˜ x∈X Π∗ nax ˜ x P∗ n(a|x) ≡ ˜ x∈X Π∗ nax ˜ x (ˇ aˇ x)∈A×X Π∗ nˇ ax ˇ x(2.8) m∗ n(x) ≡ (ax)∈A×X Π∗ naxx For the same reason as before, P∗ n(|A||x) =1−a∈˜ AP∗ n(a|x) for all x∈X.Thus,the vector of true CCPs is completely characterized by P∗ n≡{P∗ n(a|x)}(ax)∈˜ A×X∈ΘP. Section 2.1 specifies Pθ(a|x) as the econometric model for P∗ n(a|x) and fθf(x|ax) as the econometric model for f∗ n(x|ax). This paper allows the econometric model to be misspecified, that is, inf (αθf)∈Θα×Θf P(αθf)−P∗ nfθf−f∗ n >0(2.9) but requires the misspecification to vanish asymptotically according to the following assumption. Assumption 3. The model is locally misspecified in the following sense:
Quantitative Economics 10 (2019) Inference in dynamic discrete problems 81 where BΠ∗and Π∗are as in Assumption 3,and ΥML and are the following matrices: ΥML ≡∂P θ∗ ∂α ∂Pθ∗ ∂α−1∂P θ∗ ∂α Σ−∂Pθ∗ ∂θ f∈Rdα×(|A×X|+dθf) ≡⎡ ⎣ I|A×X|×|A×X|I|A×X|×|A×X| I|A×X|×|A×X| ∂G1Π∗ ∂Π∗⎤ ⎦∈R(|A×X|+dθf)×|A×X×X| (4.2) where G1denotes the first component of Gin Assumption 8,that is,ˆ θfn ≡G1(ˆ Πn), ≡⎡ ⎢ ⎢ ⎢ ⎢ ⎢ ⎣ 10|˜ A|×| ˜ A| 0|˜ A|×| ˜ A| 0|˜ A|×| ˜ A|2 0|˜ A|×| ˜ A| 0|˜ A|×| ˜ A| 0|˜ A|×| ˜ A||X| ⎤ ⎥ ⎥ ⎥ ⎥ ⎥ ⎦ ∈R|˜ A×X|×| ˜ A×X| Σ≡⎡ ⎢ ⎢ ⎢ ⎢ ⎢ ⎣ Σ10|˜ A|×|A| 0|˜ A|×|A| 0|˜ A|×|A|Σ2 0|˜ A|×|A| 0|˜ A|×|A| 0|˜ A|×|A|Σ|X| ⎤ ⎥ ⎥ ⎥ ⎥ ⎥ ⎦ ∈R|˜ A×X|×|A×X| (4.3) and,finally,for all x∈X, x≡m∗(x)diag1/P∗(a|x)a∈˜ A+1|˜ A|×| ˜ A|/1− a∈˜ A P∗(a|x)∈R|˜ A|×| ˜ A| Σx≡I|˜ A|×|A|−P∗(a|x)a∈˜ A×11×|A|/m∗(x) ∈R|˜ A|×|A| m∗(x) ≡ (ax)∈A×X Π∗axx∈R As discussed in Theorem 3.1,Theorem4.1 shows that the K-ML estimator converges at a rate of nmin{1/2δ}and that changes in Khave a no effect on its asymptotic distribution. Theorem 4.1 also reveals that the asymptotic distribution is normal with asymptotic bias and variance given by ABML =ΥML ××BΠ∗×1[δ≤1/2] AVML =ΥML ××diagΠ∗−Π∗Π∗××Υ ML ×1[δ≥1/2] Inthecaseofδ>1/2, the local misspecification is irrelevant relative to sampling error and has no effect on the rate of convergence or the asymptotic distribution. A very different situation occurs when δ<1/2. In this case, the local misspecification is overwhelming relative to sampling error and dominates the asymptotic distribution. The rate of convergence of the estimator is nδand, at this rate, the asymptotic distribution is a pure bias term. Finally, we have a knife-edge case with δ=1/2. The rate of convergence
82 Bugni and Ura Quantitative Economics 10 (2019) is the usual parametric rate √n, but the asymptotic distribution can be biased. In consequence, an adequate characterization of the estimator’s precision is the asymptotic mean squared error: ΥML ××diagΠ∗−Π∗Π∗+BΠ∗B Π∗××Υ ML The K-ML estimator is a “partial” ML estimator in the sense that it plugs in the firststep estimator into the second step. In other words, it is not a “full” maximum likelihood estimator with respect to the entire parameter vector θ=(αθf). Because of this feature, the usual optimality results for maximum likelihood estimation need not apply. In fact, the next subsection will describe a K-MD estimator that can be more efficient than the K-ML estimator, even in the absence of local misspecification. Remark 4.1. Theorem 4.1 applies to any K∈Nbut does not extend to K→∞.Under some additional conditions, however, Aguirregabiria and Mira (2002, Proposition 3) showed that if the K-ML estimator converges as K→∞,itwilldosotoasolutionof Rust (1987)’s nested fixedpoint estimator. As noted in Aguirregabiria and Mira (2002, footnote 16), this result presumes the convergence of the K-ML estimator as K→∞, which has not been shown in the literature. 4.2 K-MD estimators To specialize the general result to K-MD estimators, we set the sample objective function Qnto QMD n(α θfP)≡−ˆ Pn−Ψ(αθf)(P)ˆ Wnˆ Pn−Ψ(αθf)(P)(4.4) where ˆ Pnis the sample frequency estimator of the vector of CCPs in Equation (4.1)and ˆ Wn∈R|˜ A×X|×| ˜ A×X|is the weight matrix. In the special case of K=1and ˆ P0 n=ˆ Pn,the K-MD estimator coincides with the estimators considered in Hotz and Miller (1993)and Pesendorfer and Schmidt-Dengler (2008).8We impose the following condition regarding the weight matrix. Assumption 10. ˆ Wn=W∗+opn(1),where W∗∈R|˜ A×X|×| ˜ A×X|is positive definite and symmetric. In principle, we could generalize Assumption 10 by allowing the weight matrix to be a function of the parameters of the problem. Similar results would then follow from longer arguments. Theorem 4.2 is a corollary of Theorem 3.1 and characterizes the asymptotic distribution of the K-MD estimator under local misspecification. 8To be precise, Pesendorfer and Schmidt-Dengler (2008, Equations (18)–(19)) consider a sample criterion function that allows ˆ Pnin Equation (4.4) to differ from the sample frequency estimator. We could also incorporate this feature in our setup at the expense of using longer arguments.
Quantitative Economics 10 (2019) Inference in dynamic discrete problems 83 Theorem 4.2 (K-MD). Suppose Assumptions 1–4and 7–10.Then,for any K ˜ K≥1, nmin{1/2δ}ˆαK-MD n−α∗ =nmin{1/2δ}ˆα˜ K-MD n−α∗+opn(1) d →ΥMDW∗××NBΠ∗×1[δ≤1/2]diagΠ∗−Π∗Π∗×1[δ≥1/2] where BΠ∗and Π∗are as in Assumption 3,ΥMD(W ∗)is the following matrix: ΥMDW∗≡∂P θ∗ ∂α W∗∂Pθ∗ ∂α−1∂P θ∗ ∂α W∗Σ−∂Pθ∗ ∂θ f∈Rdα×(|A×X|+dθf) with is as in Equation (4.2)and Σis as in Equation (4.3). Remark 4.2. The asymptotic distribution of the K-ML estimator is a special case of that of the K-MD estimator with W∗=. As discussed in previous theorems, Theorem 4.2 shows that the K-MD estimator converges at a rate of nmin{1/2δ}and that changes in Khave no effect on its asymptotic distribution. Theorem 4.2 also reveals that the asymptotic distribution is normal with asymptotic bias and variance given by ABMDW∗=ΥMDW∗××BΠ∗×1[δ≤1/2] AVMDW∗=ΥMDW∗××diagΠ∗−Π∗Π∗××ΥMDW∗×1[δ≥1/2] (4.5) In the knife-edge case with δ=1/2, the asymptotic variance and bias can coexist. In consequence, an adequate characterization of the estimator’s precision is the asymptotic mean squared error: ΥMDW∗××diagΠ∗−Π∗Π∗+BΠ∗B Π∗××ΥMDW∗(4.6) We now briefly discuss the optimality in the choice of W∗in K-MD estimation. First, consider the case in which local misspecification is asymptotically irrelevant, that is, δ>1/2.Inthiscase,theK-MD estimator presents no asymptotic bias and the asymptotic variance and mean squared error coincide. Provided that relevant matrices are nonsingular, standard arguments in GMM estimation imply that the minimum asymptotic variance and mean squared error among K-MD estimators are both equal to ⎛ ⎝ ∂Pθ∗ ∂α Σ−∂Pθ∗ ∂θ fdiagΠ∗−Π∗Π∗Σ−∂Pθ∗ ∂θ f−1∂Pθ∗ ∂α⎞ ⎠ −1 This minimum can be achieved by the following “feasible” choice of limiting weight matrix: W∗ AV ≡Σ−∂Pθ∗ ∂θ fdiagΠ∗−Π∗Π∗Σ−∂Pθ∗ ∂θ f−1 (4.7)
84 Bugni and Ura Quantitative Economics 10 (2019) We say that Equation (4.7) is a feasible choice because it can be consistently estimated. As pointed out in Remark 4.2,theK-ML estimator has the same asymptotic distribution as the K-MD estimator with W∗ ML =. Except under special conditions on the econometric model (e.g. ∂Pθ∗/∂θ f=0|˜ A×X|×dθf ), the K-ML estimator is not necessarily optimal among the K-MD estimators. Next, consider the knife-edge case in which local misspecification vanishes at the rate of sampling error, that is, δ=1/2. Once again, standard arguments in GMM estimation imply that the minimum asymptotic mean squared error among all K-MD estimators is ⎛ ⎝ ∂P θ∗ ∂α Σ−∂Pθ∗ ∂θ fdiagΠ∗−Π∗Π∗+BΠ∗B Π∗Σ−∂Pθ∗ ∂θ f−1∂Pθ∗ ∂α⎞ ⎠ −1 According to McFadden and Newey (1994, p. 2165), any limiting weight matrix W∗that minimizes the asymptotic mean squared error should satisfy the following condition: for some matrix C∈R|˜ A×X|×| ˜ A×X|, ∂Pθ∗ ∂αW∗=C∂Pθ∗ ∂αΣ−∂Pθ∗ ∂θ fdiagΠ∗−Π∗Π∗+BΠ∗B Π∗ ×Σ−∂Pθ∗ ∂θ f−1 (4.8) In other words, Equation (4.8) characterizes the class of optimal limiting weight matrices. A simple example of an optimal limiting weight matrix is W∗ AMSE ≡Σ−∂Pθ∗ ∂θ fdiagΠ∗−Π∗Π∗+BΠ∗B Π∗Σ−∂Pθ∗ ∂θ f−1 (4.9) By Equation (4.8), the consistent estimation of any optimal limiting weight matrix requires the consistent estimation of the asymptotic bias. Unfortunately, the researcher is not aware of this feature of the population distribution and, in this sense, estimating the optimal limiting weight matrix under local misspecification is infeasible in practice. In particular, note that the feasible limiting weight matrix W∗ AV that minimizes asymptotic variance could fail to be optimal (i.e., Equation (4.8) may not be satisfied when W∗=W∗ AV ).9 Finally, we could consider the case in which local misspecification is asymptotically overwhelming, that is, δ<1/2. In this case, the asymptotic distribution collapses to the pure bias term in Equation (4.5), and the asymptotic mean squared error coincides with the square of the asymptotic bias. As in the case with δ=1/2, minimizing asymptotic mean squared error is infeasible in the sense that it depends on the unknown asymptotic bias. Also, any feasible choice of limiting weight matrix such as W∗ AV could fail to be optimal. 9This is the case in some of our Monte Carlo simulations, in which using W∗ AV produces an asymptotic mean squared error that is larger than that the one obtained by using ˆ Wn=I|˜ A×X|×| ˜ A×X|.
Quantitative Economics 10 (2019) Inference in dynamic discrete problems 85 5. Monte Carlo simulations This section investigates the finite sample performance of the two-step estimators considered in previous sections under local misspecification. We simulate data using the classical bus engine replacement problem studied by Rust (1987). In each period t=1T ≡∞, the bus owner decides whether to replace the bus engine or not to minimize the discounted present value of his costs. In any representative period, his choice is denoted by a∈A={12},wherea=2represents replacing the engine, a=1represents not replacing the engine, and the current engine mileage is denoted by x∈X≡{120}. The researcher assumes that all individuals in his sample have the same deterministic part of the utility (profit) function, given by uθu(xa) =−θu1×1[a=2]−θu2×1[a=1]x (5.1) where θu≡(θu1θu2)∈Θu≡[−BB]2with B=10. In addition, the researcher also assumes that the errors are i.i.d. extreme value type I, independent of x,thatis, g(ε =e|x) = a∈A expe(a)exp−expe(a)(5.2) which does not have unknown parameters. Finally, the observed state is assumed to evolve according to the following Markov chain: fθfx|xa=(1−θf)×1a=1x=minx+1|X| +θf×1a=1x=x+1a=2x=1 (5.3) where θf∈Θf≡[01]. The researcher correctly assumes that β=09999. His goal is to estimate θ=(α θf)∈Θ=Θα×Θfwith α=θu∈Θα=Θu. The researcher correctly specified the error distribution and state transition probabilities, which satisfy Equation (5.3)withθf=025. Unfortunately, he does not correctly specify the utility function. The correct utility function is as follows: uθun(xa) =−θu1×1[a=2]−θu2×1[a=1]x+τn×1[a=1]x2(5.4) with θu1=1,θu2=005,andτn=−0025n−δwith δ∈{1/31/21}. Notice that this form of model misspecification is analogous to the first illustration discussed in Section 2.2.10 By the arguments in Section 2,thetrueCCPsP∗ nare determined by the true error distribution (Equation (5.2)), true state transition probabilities (Equation (5.3)with θf=025), and the true utility function (Equation (5.4)). Our choices of δinclude a case in which the local misspecification is asymptotically irrelevant (δ=1), one case in which local misspecification is the knife-edge case (δ=1/2), and one case in which the local misspecification is overwhelming (δ=1/3). In addition, we also consider a case in which the econometric model is correctly specified. 10Section 2.2 provides two other examples of local misspecification. We have also conducted Monte Carlo simulations based on these designs. For the sake of brevity, these are presented in the Supplemental Material (see Bugni anf Ura (2019)).
86 Bugni and Ura Quantitative Economics 10 (2019) Our simulation results will be the average of S=20,000 independent datasets of observations {(aixix i)}i≤nthat are i.i.d. distributed according to Π∗ n. We present simulation results for sample sizes of n∈{2005001000}. We generate marginal observations of the state variables according to the following distribution: m∗ n(x) ∝1+log(x)11 Together with previous elements, this determines the true joint DGP Π∗ naccording to Equation (2.7). Given any sample of observations {(aixix i)}i≤n, the researcher estimates the parameters of interest θ=(θu1θu2θf)using a two-step K-stage PI algorithm described in Sections 3–4. In the first step, the researcher estimates P∗and θ∗ fusing preliminary estimators ˆ P0 n=ˆ Pnand ˆ θfn = n i=1 1ai=1x i=xixi=|X| n i=1 1ai=1xi=|X| In the second step, the researcher estimates (θu1θu2)using the K-stage policy function iteration algorithm using criterion function Qnequal to (a) pseudo-likelihood function QML nin Section 4.1 and (b) the weighted minimum distance function QMD nin Section 4.2 with two limiting weight matrices: identity (i.e., W∗=I|˜ A×X|×| ˜ A×X|)and asymptotic variance minimizer (i.e., W∗=W∗ AV 12). We show results for number of stages K∈{12310}.13 We now describe simulation results for the estimator of θu2. We focus on θu2because we consider the linear coefficient of the utility function to be more interesting than the constant coefficient.14 Table 1describes results under correct specification. As expected, all estimators appear to converge to a distribution with zero mean and finite variance. Also as expected, the number of iterations Kdoes not seem to affect the bias or the variance of the estimators under consideration. Similar to Aguirregabiria and Mira (2002), we detect small differences between the results with K=1and those with K>1, especially for the smallest sample size. This effect tends to vanish as the sample size increases. These findings could be rationalized by the higher-order analysis in Kasahara and Shimotsu (2008). The K-ML estimator and the K-MD estimator with W∗=W∗ AV are similar and more efficient than the K-MD estimator with W∗=I|˜ A×X|×| ˜ A×X|. Table 2provides results under asymptotically irrelevant local misspecification, that is, δ=1. According to our theoretical results, the asymptotic behavior of all estimators 11Recall from Section 2that this aspect of the model is left unspecified by the researcher. 12This matrix is calculated using numerical derivates and Monte Carlo integration with a sample size that is significantly larger than those used in the actual Monte Carlo simulations. 13In accordance to our asymptotic theory, the simulation results with K∈{49}are almost identical to those with K∈{310}. These were eliminated from the paper for reasons of brevity and are available from the authors upon request. 14The results for θu1are qualitatively similar and are available from the authors upon request.
Quantitative Economics 10 (2019) Inference in dynamic discrete problems 87 Table 1. Simulation results under correct specification, that is, τn=0. K-MD(I|˜ A×X|×| ˜ A×X|)K-MD(W∗ AV )K-ML KStatistic n=200 n=500 n=1000 n=200 n=500 n=1000 n=200 n=500 n=1000 1√nBias 007 002 001 006 002 001 006 002 001 √nSD 025 025 024 023 023 022 022 022 022 nMSE 007 006 006 006 005 005 005 005 005 2√nBias 000 001 000 000 000 000 001 000 000 √nSD 024 025 024 023 023 022 022 022 022 nMSE 006 006 006 005 005 005 005 005 005 3√nBias 000 000 000 000 000 000 000 000 000 √nSD 024 025 024 023 023 022 022 022 022 nMSE 006 006 006 005 005 005 005 005 005 10 √nBias 000 000 000 000 000 000 000 000 000 √nSD 025 025 024 023 023 022 022 022 022 nMSE 006 006 006 005 005 005 005 005 005 Table 2. Simulation results under local misspecification with τn∝n−1. K-MD(I|˜ A×X|×| ˜ A×X|)K-MD(W∗ AV )K-ML KStatistic n=200 n=500 n=1000 n=200 n=500 n=1000 n=200 n=500 n=1000 1√nBias 011 004 003 010 004 003 009 004 003 √nSD 025 025 024 024 023 022 023 023 022 nMSE 007 006 006 007 006 005 006 005 005 2√nBias 004 003 002 004 003 002 004 003 002 √nSD 025 025 024 023 023 022 022 022 022 nMSE 006 006 006 005 005 005 005 005 005 3√nBias 004 003 002 004 003 002 004 003 002 √nSD 025 025 024 023 023 022 022 022 022 nMSE 006 006 006 005 005 005 005 005 005 10 √nBias 004 003 002 004 003 002 004 003 002 √nSD 025 025 024 023 023 022 022 022 022 nMSE 006 006 006 005 005 005 005 005 005 should be identical to the correctly specified model. This is confirmed in our simulations, as Tables 1and 2are virtually identical. Table 3provides results under local misspecification that vanishes at the knife-edge rate, that is, δ=1/2. According to our theoretical results, this should produce an asymptotic distribution that has nonzero bias and is not affected by the number of iterations K. By and large, these predictions are confirmed in our simulations. Given the presence of asymptotic bias, we now evaluate efficiency using the mean squared error. In this simulation design, the K-MD estimator with W∗=I|˜ A×X|×| ˜ A×X|is now also slightly more efficient than the K-ML estimator or the K-MD estimator with W∗=W∗ AV . This is also consistent with our theory: while K-MD estimator with W∗=I|˜ A×X|×| ˜ A×X|has more
88 Bugni and Ura Quantitative Economics 10 (2019) Table 3. Simulation results under local misspecification with τn∝n−1/2. K-MD(I|˜ A×X|×| ˜ A×X|)K-MD(W∗ AV )K-ML KStatistic n=200 n=500 n=1000 n=200 n=500 n=1000 n=200 n=500 n=1000 1√nBias 050 046 046 051 049 050 052 050 050 √nSD 028 028 027 026 026 024 025 025 024 nMSE 033 029 028 033 031 030 033 031 031 2√nBias 042 044 045 046 048 049 047 049 049 √nSD 029 028 026 026 025 024 024 024 024 nMSE 026 027 028 027 029 030 028 030 030 3√nBias 042 044 045 046 048 049 047 049 049 √nSD 028 028 026 026 025 024 024 024 024 nMSE 026 027 028 027 029 030 028 030 030 10 √nBias 042 044 045 046 048 049 047 049 049 √nSD 029 028 026 026 025 024 024 024 024 nMSE 026 027 028 027 029 030 028 030 030 Table 4. Simulation results under local misspecification with τn∝n−1/3and using the correct scaling. K-MD(I|˜ A×X|×| ˜ A×X|)K-MD(W∗ AV )K-ML KStatistic n=200 n=500 n=1000 n=200 n=500 n=1000 n=200 n=500 n=1000 1n1/3Bias 041 040 041 043 044 045 045 046 046 n1/3SD 014 012 010 013 010 009 012 010 009 n2/3MSE 018 018 018 020 020 021 022 022 022 2n1/3Bias 037 040 041 040 043 045 043 045 046 n1/3SD 014 012 010 012 010 009 011 010 009 n2/3MSE 016 017 018 018 020 021 020 021 022 3n1/3Bias 037 040 041 040 043 045 043 045 046 n1/3SD 014 012 010 012 010 009 011 010 009 n2/3MSE 016 017 018 018 020 021 020 021 022 10 n1/3Bias 037 040 041 040 043 045 043 045 046 n1/3SD 014 012 010 012 010 009 011 010 009 n2/3MSE 016 017 018 018 020 021 020 021 022 variance that the other two, it also appears to have less bias, resulting in less overall mean squared error. Table 4provides results under asymptotically overwhelming local misspecification, that is, δ=1/3. According to our theoretical results, the presence of this local misspecification changes dramatically the asymptotic distribution of all estimators. In particular, these no longer converge at the regular √n-rate, but rather at n1/3-rate. In fact, at √n-rate, the asymptotic bias is no longer bounded (see Table S1 in the Supplemental Material). Once we scale the estimators at the appropriate n1/3-rate, they converge to an asymptotic distribution dominated by the bias. Furthermore, our theoretical results in-
Quantitative Economics 10 (2019) Inference in dynamic discrete problems 89 dicate that the number of iterations Kdoes not affect this asymptotic distribution. These predictions are clearly depicted in Table 4. In line with the results in Table 3,theK-MD estimator with W∗=I|˜ A×X|×| ˜ A×X|is slightly more efficient than the K-ML estimator and the K-MD estimator with W∗=W∗ AV . 6. Conclusion Single-agent dynamic discrete choice models are typically estimated using heavily parametrized econometric frameworks, making them susceptible to model misspecification. This paper investigates how misspecification can affect inference results in these models. This paper considers a local misspecification framework, which is an asymptotic device in which the mistake in the specification vanishes as the sample size diverges. In this paper, we impose no restrictions on the rate at which these specification errors disappear. Relative to global misspecification, the local misspecification analysis has two important advantages. First, it yields tractable and general results. Second, it allows us to focus on parameters with structural interpretation, instead of “pseudo-true” parameters. We consider a general class of two-step estimators based on the K-stage sequential policy function iteration algorithm, where Kdenotes the number of iterations employed in the estimation. By appropriate choice of the criterion function, this class includes Hotz and Miller (1993)’s conditional choice probability estimator, Aguirregabiria and Mira (2002)’s pseudo-likelihood estimator, and Pesendorfer and Schmidt-Dengler (2008)’s asymptotic least squares estimator. We show that local misspecification can affect the asymptotic distribution and even the rate of convergence of these estimators. In principle, one might expect that the effect of the local misspecification could change with the number of iterations K.The main finding in the paper is that this is not the case, that is, the effect of local misspecification is invariant to K.Inparticular,(a)K-ML estimators are asymptotically equivalent and (b) given the choice of the weight matrix, K-MD estimators are asymptotically equivalent. In practice, this means that researchers cannot eliminate or even alleviate problems of model misspecification by choosing K. Additional iterations are computationally costly and produce no change in asymptotic efficiency. Under correct specification, the comparison between K-MD and K-ML estimators in terms of asymptotic mean squared error yields a clear-cut recommendation. Under local misspecification, this is no longer the case. In particular, local misspecification can introduce an unknown asymptotic bias which complicates this comparison. In the presence of asymptotic bias, the optimality of the estimator should be evaluated using the asymptotic mean squared error. We show that an optimally-weighted K-MD estimator depends on the unknown asymptotic bias and is thus generally unfeasible. In turn, the feasible K-MD estimator with a weight matrix that minimizes asymptotic variance could have an asymptotic mean squared error that is higher or lower than that of the K-ML estimator or the K-MD estimator with identity weight matrix.
90 Bugni and Ura Quantitative Economics 10 (2019) Appendix A.1 Additional notation Throughout this Appendix, “s.t.” abbreviates “such that,” and “RHS” and “LHS” abbreviate “right-hand side” and “left-hand side,” respectively. Furthermore, “LLN” refers to the strong law of large numbers, “CLT” refers to the central limit theorem, and “CMT” refers to the continuous mapping theorem. Given the true DGP Π∗ n,Equation(2.8) defined transition probabilities f∗ n,CCPsP∗ n, and marginal distribution of states m∗ n. The unconditional probability of (ax) ∈A×X is analogously defined by J∗ n(ax) ≡ ˜ x∈X Π∗ nax ˜ x and J∗ n≡{J∗ n(ax)}(ax)∈A×X. The limiting DGP f∗, transition probabilities f∗,CCPsP∗ were defined in Assumption 3. The other limiting objects are analogously defined by J∗≡limn→∞J∗ nand m∗≡limn→∞m∗ n. The sample analogue DGP ˆ Πn, transition probabilities ˆ fn,CCPs ˆ Pnwere defined in Equation (4.1). The sample analogue marginal distribution of states ˆ mnand unconditional probabilities ˆ Jnare analogously defined. For any (ax) ∈A×X, ˆ mn(x) ≡ (a˜ x)∈A×X ˆ Πnax ˜ x ˆ Jn(ax) ≡ ˜ x∈X ˆ Πnax ˜ x ˆ mn≡{ˆ mn(x)}x∈X,and ˆ Jn≡{ˆ Jn(ax)}(ax)∈A×X. A.2 Proofs of theorems Proof of Theorem 2.1. Since Θ=Θα×Θfis compact and (P(αθf)−P∗ n),(fθf−f∗ n) is a continuous function of (αθf), the arguments in Royden (1988, pp. 193–195) implies that ∃(α∗θ∗ f)∈Θthat minimizes (P(αθf)−P∗ n),(fθf−f∗ n). By Assumption 3(b), this minimum value is zero, that is, ∃(α∗θ∗ f)∈Θs.t. (P(α∗θ∗ f)−P∗)(fθ∗ f−f∗)=0or, equivalently, P(α∗θ∗ f)=P∗and fθ∗ f=f∗. Now suppose that this also occurs for (˜ θf˜α) ∈Θ. We now show that (θ∗ fα∗)= (˜ θf˜α). By triangle inequality fθ∗ f−f˜ θf≤fθ∗ f−f∗+f˜ θf−f∗and since θ∗ fand ˜ θf both satisfy fθf−f∗=0, we conclude that fθ∗ f−f˜ θf=0and so fθ∗ f=f˜ θf. By Assumption 2, this implies that θ∗ f=˜ θf. By repeating the previous argument with P(αθ∗ f)instead of fθf, we conclude that α∗=˜α. Proof of Theorem 3.1. Without loss of generality, we can consider the neighborhood Nto be “rectangular,” in the sense that N=Nα×Nθf×NP,whereNλdenotes the neighborhood of λ∗for any λ∈{α θfP}. This can always be achieved by replacing Nwith ˜ N⊆Nthat has the desired structure. This proof will make repeated reference to Nα.
Quantitative Economics 10 (2019) Inference in dynamic discrete problems 97 Condition (e). For any λ∈{α θfP}, direct computation shows that ∂2QMD ∞(α θfP) ∂λ∂α=2∂ ∂α ∂Ψθ(P) ∂λ W∗P∗−Ψθ(P)−∂Ψθ(P) ∂λ W∗∂Ψθ(P) ∂α(A.8) This function is continuous and if we evaluate at (α θfP)=(α∗θ∗ fP∗),weobtain ∂QMD ∞α∗θ∗ fP∗ ∂λ∂α=−2∂Ψθ∗P∗ ∂λ W∗∂Ψθ∗P∗ ∂α=−2∂P θ∗ ∂λ W∗∂Pθ∗ ∂α where the first line uses that P∗=Ψθ∗(P∗)and ∂Ψθ∗(P∗)/∂α =∂Pθ∗/∂α.Toverifytheresult, it suffices to consider the last expression with λ=α. By assumption, this expression is square, symmetric, and negative definite, and, consequently, it must be nonsingular. Condition (f). By Young’s theorem and Equation (A.8)withλ=Pand (α θfP)= (α∗θ∗ fP∗), ∂2QMD ∞α∗θ∗ fP∗ ∂P∂α=−2∂Ψθ∗P∗ ∂P W∗∂Ψθ∗P∗ ∂α=−2∂Ψθ∗(Pθ∗) ∂P W∗∂Ψθ∗(Pθ∗) ∂α =0|˜ A×X|×dα where the last equality uses that the Jacobian matrix of Ψθ∗with respect to Pis zero at Pθ∗=P∗. Part 2: Verify the conditions in Assumption 6. Assumption 6(b) holds as a corollary of Lemma A.2. To verify Assumption 6(a), consider the following argument. By direct computation, ∂QMD nα∗θ∗ fP∗ ∂α =2∂Ψθ∗P∗ ∂α ˆ Wnˆ Pn−Ψθ∗P∗=2∂P θ∗ ∂α W∗ˆ Pn−P∗+opn(1) where the last equality uses that Ψθ∗(P∗)=P∗,∂Ψθ∗(P∗)/∂α =∂P θ∗/∂α,ˆ Pn−P∗= opn(1),and ˆ Wn−W∗=opn(1). We then conclude that nmin{δ1/2}∂QMD nα∗θ∗ fP∗/∂α ˆ θfn −θ∗ f=⎡ ⎣ 2∂P θ∗ ∂α W∗0|A×X|×dθf 0dθf×|A×X|Idθf×dθf ⎤ ⎦nmin{δ1/2}ˆ Pn−P∗ ˆ θfn −θ∗ f +opn(1) From this and Lemma A.2, we conclude that the desired result holds with ζ∼⎡ ⎣ ∂P θ∗ ∂α W∗Σ0|A×X|×dθf 0dθf×|A×X|Idθf×dθf ⎤ ⎦ ×NBΠ∗×1[δ≤1/2]diagΠ∗−Π∗Π∗×1[δ≥1/2] (A.9) This completes the verification of Assumptions 5–6and so Theorem 3.1 applies. The specific formula for the asymptotic distribution relies on Equations (A.8)and(A.9).
98 Bugni and Ura Quantitative Economics 10 (2019) A.3 Proofs of lemmas Proof of Lemma 2.1. This proof follows from Aguirregabiria and Mira (2002, Propositions 1–2). Proof of Lemma 2.2. Parts (a)–(b) follow from Rust (1988, pp. 1015–6). Part (c) follows from combining Lemma 2.1 and Assumption 2. Lemma A.1. Suppose Assumptions 3–7.Then nmin{δ1/2}"ˆ Jn−J∗ ˆ θfn −θ∗ f# d →×NBΠ∗×1[δ≤1/2]diagΠ∗−Π∗Π∗×1[δ≥1/2] (A.10) with as in Equation (4.2). Proof. Under Assumption 7, the triangular array CLT (e.g., Davidson (1994, p. 369)) implies that √nˆ Πn−Π∗ nd →N0diagΠ∗−Π∗Π∗ If we combine this with Assumption 3, nmin{δ1/2}ˆ Πn−Π∗d →NBΠ∗×1[δ≤1/2]diagΠ∗−Π∗Π∗×1[δ≥1/2](A.11) Also, notice that nmin{δ1/2}"ˆ Jn−J∗ ˆ θfn −θ∗ f#=nmin{δ1/2}F( ˆ Πn)−FΠ∗ where F:R|A×X×X|→R|A×X|+dθfis defined as follows. For coordinates j≤|A×X| where jrepresents the corresponding coordinate (ax) ∈A×X,Fj(z) ≡˜ x∈Xz(ax˜ x), and for coordinates j>|A×X|,Fj(z) ≡G1j(z). By definition of G1,ˆ θfn =G1(ˆ Πn),and by Assumption 8,θ∗ f=G1(Π∗)and Fis continuously differentiable at Π∗.Bydirectcomputation, =∂F(Π∗)/∂Π. Then the result follows from the delta method and Equation (A.11). Lemma A.2. Suppose Assumptions 3–7.Then nmin{δ1/2}"ˆ Pn−P∗ ˆ θfn −θ∗ f# d →⎡ ⎣ Σ0|˜ A×X|×dθf 0dθf×| ˜ A×X|Idθf×dθf⎤ ⎦ ××NBΠ∗1[δ≤1/2]diagΠ∗−Π∗Π∗1[δ≥1/2] (A.12) with as in Equation (4.2)and Σas in Equation (4.3).
Quantitative Economics 10 (2019) Inference in dynamic discrete problems 99 Proof.LetF:R|A×X|+dθf→R|˜ A×X|+dθfbe defined as follows. For coordinates j≤|˜ A× X|with jrepresenting coordinate (ax) ∈˜ A×X,Fj(z) ≡z(ax)/a∈Az(˜ ax), and for j> |˜ A×X|,Fj(z) =zj.NoticethatF(( ˆ J nˆ θ fn))≡(ˆ P nˆ θ fn)and F((J∗θ∗ f))≡(P∗θ∗ f) by definition of F.IfweverifyFis continuously differentiable at z=(J∗θ∗ f)and ∂FJ∗θ∗ f ∂z=⎡ ⎣ Σ0|˜ A×X|×dθf 0dθf×| ˜ A×X|Idθf×dθf⎤ ⎦(A.13) then the result follows from the delta method and Lemma A.2.Wedothisnext. Consider (j ˇ j) ∈{1|˜ A×X|}×{1|A×X|} representing (ax) ∈˜ A×Xand (ˇ a ˇ x) ∈A×X.Foranyj>|˜ A×X|,Fjis continuously differentiable and ∂Fj(z)/∂zˇ j= 1[j=ˇ j].Foranyj≤|˜ A×X|, ∂Fj(z) ∂zˇ j=1[x=ˇ x]⎡ ⎢ ⎢ ⎢ ⎢ ⎢ ⎣ ⎛ ⎜ ⎜ ⎜ ⎜ ⎜ ⎝ ˜ a∈A z(˜ ax) −z(ˇ ax) ˜ a∈A z(˜ ax)2 ⎞ ⎟ ⎟ ⎟ ⎟ ⎟ ⎠ 1[a=ˇ a]+⎛ ⎜ ⎜ ⎜ ⎜ ⎜ ⎝ −z(ˇ ax) ˜ a∈A z(˜ ax)2 ⎞ ⎟ ⎟ ⎟ ⎟ ⎟ ⎠ 1[a= ˇ a]⎤ ⎥ ⎥ ⎥ ⎥ ⎥ ⎦ =1[x=ˇ x] ˜ a∈A z(˜ ax) ⎡ ⎢ ⎢ ⎢ ⎢ ⎣ 1[a=ˇ a]− z(ˇ ax) ˜ a∈A z(˜ ax) ⎤ ⎥ ⎥ ⎥ ⎥ ⎦ provided that ˜ a∈Az(˜ ax) >0. Since ˜ a∈AJ∗(˜ ax) > 0for all x∈X,Fis continuously differentiable at ((J∗θ∗ f)). By combining the formula for the derivatives from all coordinates, Equation (A.13) follows. Lemma A.3. For any λ ˜ λ∈{θfα},the following algebraic results hold: ∂lnPθ∗(a|x) (ax)∈A×X ∂λ =∂P θ∗ ∂λ Σ (ax)∈A×X J∗(ax)∂lnPθ∗(a|x) ∂λ ∂lnPθ∗(a|x) ∂˜ λ=∂P θ∗ ∂λ ∂Pθ∗ ∂˜ λ with and Σas in Equation (4.3). Proof. Before deriving the results, consider some preliminary observations. For any λ∈{α θfP},a∈APθ∗(a|x) =1and so ∂Pθ∗(|A||x)/∂λ =−a∈˜ A∂Pθ∗(a|x)/∂λ. Also, for any λ∈{α θfP}and (ax) ∈A×X,P∗(a|x) =Pθ∗(a|x) and so (∂Pθ∗(a|x)/∂λ)(1/P∗(a| x)) =∂lnPθ∗(a|x)/∂λ. For the first result, consider the following derivation: ∂P θ∗ ∂λ Σ =∂Pθ∗(a|x) (ax)∈˜ A×X ∂λ ×diag{xΣx}x∈X
100 Bugni and Ura Quantitative Economics 10 (2019) =∂Pθ∗(a|x) (ax)∈˜ A×X ∂λ ×diagdiag1/P∗(a|x)a∈˜ A−1/P∗|A||x1|˜ A|×1x∈X =∂lnP∗ θ(a|x)(ax)∈A×X ∂λ where the last equality uses the preliminary observations. For the second result, consider the following derivation: ∂P θ∗ ∂λ ∂Pθ∗ ∂˜ λ=∂Pθ∗(a|x) (ax)∈˜ A×X ∂λ ××∂Pθ∗(a|x)(ax)∈˜ A×X ∂˜ λ =∂Pθ∗(a|x) (ax)∈˜ A×X ∂λ diagm(x)diag1/P∗(a|x)a∈˜ Ax∈X ×∂Pθ∗(a|x)(ax)∈˜ A×X ∂˜ λ +∂Pθ∗(a|x) (ax)∈˜ A×X ∂λ diagm(x)1|˜ A|×| ˜ A|/1− a∈˜ A P∗(a|x)x∈X ×∂Pθ∗(a|x)(ax)∈˜ A×X ∂˜ λ = (ax)∈˜ A×X m(x)∂ln Pθ∗(a|x) ∂λ ∂Pθ∗(a|x) ∂˜ λ + x∈X m(x) P|A||x a∈˜ A ∂Pθ∗(a|x) ∂λ ˜ a∈˜ A ∂Pθ∗(˜ a|x) ∂λ = (ax)∈A×X J∗(ax)∂lnPθ∗(a|x) ∂λ ∂lnPθ∗(a|x) ∂˜ λ where the last equality uses the preliminary observations. A.4 Review of results on extremum estimators The purpose of this section is to state well-known results regarding the consistency and asymptotic normality of extremum estimators under certain regularity conditions. These results are referenced in our formal arguments. Relative to the standard versions in the literature (e.g., McFadden and Newey (1994)), our results allow for: (a) a rate of convergence that may differ from √nand (b) a sequence of DGPs that may change with sample size. Both of these features are important for our theoretical results. We omit the proofs for reasons of brevity but theses are available from the authors upon request. Theorem A.1. Assume the following:
Quantitative Economics 10 (2019) Inference in dynamic discrete problems 101 (a) Qn(θ) converges uniformly in probability to Q(θ) along {pn}n≥1. (b) Q(θ) is upper semicontinuous,that is,for any {θn}n≥1with θn→˜ θ,limsupQ(θn)≤ Q( ˜ θ). (c) Q(θ) is uniquely maximized at θ=θ∗. Then ˆ θn=argmaxθ∈ΘQn(θ) satisfies ˆ θn=θ∗+opn(1). Theorem A.2. Consider an estimator ˆ θnof a parameter θ∗s.t.ˆ θn=argmaxθ∈ΘQn(θ). Furthermore, (a) ˆ θn=θ∗+opn(1), (b) θ∗belongs to the interior of Θ, (c) Qnis twice continuously differentiable in a neighborhood Nof θ∗w.p.a.1, (d) For some δ>0,nδ∂Qn(θ∗)/∂θ d →Zfor some random variable Zalong {pn}n≥1, (e) supθ∈N∂2Qn(θ)/∂θ∂θ−H(θ)=opn(1)for some function H:N→Rk×kthat is continuous at θ∗, (f) H(θ∗)is nonsingular. Then nδ(ˆ θn−θ∗)=−H(θ∗)−1nδ∂Qn(θ∗)/∂θ +opn(1)d →−H(θ∗)−1Zalong {pn}n≥1. References Aguirregabiria, V. and P. Mira (2002), “Swapping the nested fixed point algorithm: A class of estimators for discrete Markov decision models.” Econometrica, 70, 1519–1543. [67, 68,69,71,73,77,80,82,86,89,98] Aguirregabiria, V. and P. Mira (2010), “Dynamic discrete choice structural models: A survey.” Journal of Econometrics, 156, 38–67. [68] Arcidiacono, P. and P. B. Ellickson (2011), “Practical methods for estimation of dynamic discrete choice models.” Annual Review of Economics, 3, 363–394. [68] Arcidiacono, P. and R. A. Miller (2011), “Conditional choice probability estimation of dynamic discrete choice models with unobserved heterogeneity.” Econometrica, 79, 1823– 1867. [75] Blackwell, D. (1965), “Discounted dynamic programming.” The Annals of Mathematical Statistics, 36, 226–235. [72] Bugni, F. A., I. A. Canay, and P. Guggenberger (2012), “Distortions of asymptotic confidence size in locally misspecified moment inequality models.” Econometrica, 80, 1741– 1768. [70,76] Bugni, F. A., and T. Ura (2019), “Supplement to ‘Inference in dynamic discrete choice problems under local misspecification’.” Quantitative Economics Supplemental Material, 10, https://doi.org/10.3982/QE917.[85]
102 Bugni and Ura Quantitative Economics 10 (2019) Chernozhukov, V., J. C. Escanciano, H. Ichimura, and W. K. Newey (2016), “Locally robust semiparametric estimation.” Working paper. [70] Davidson, J. (1994), Stochastic Limit Theory. Oxford University Press. [98] Gourieroux, C. and A. Monfort (1995), Statistics and Econometric Models,Vol.2.Cambridge University Press. [93,96] Hotz, J. V. and R. T. A. Miller (1993), “Conditional choice probabilities and the estimation of dynamic models.” Review of Economics Studies, 60, 497–529. [67,68,72,82,89] Hotz, J. V., R. T. A. Miller, S. Sanders, and J. Smith (1994), “A simulation estimator for dynamic models of discrete choice.” Review of Economics Studies, 61, 265–289. [80] Kasahara, H. and K. Shimotsu (2008), “Pseudo-likelihood estimation and bootstrap inference for structural discrete Markov decision models.” Journal of Econometrics, 146, 92–106. [86] Kitamura, Y., T. Otsu, and K. Evdokimov (2013), “Robustness, infinitesimal neighborhoods, and moment restrictions.” Econometrica, 81, 1185–1201. [70] Magnac, T. and D. Thesmar (2002), “Identifying dynamic discrete decision processes.” Econometrica, 70, 801–816. [71,73] McFadden, D. and W. K. Newey (1994), “Large sample estimation and hypothesis testing.” In Handbook of Econometrics (R. F. Engle and D. L. McFadden, eds.), Handbook of Econometrics, Vol. 4, 2111–2245, Elsevier. [84,100] Newey, W. K. (1985a), “Generalized method of moments specification testing.” Journal of Econometrics, 29, 229–256. [70,76] Newey, W. K. (1985b), “Maximum likelihood specification testing and conditional moment tests.” Econometrica, 5, 1047–1070. [70,76] Norets, A. and S. Takahashi (2013), “On the surjectivity of the mapping between utilities and choice probabilities.” Quantitative Economics, 4, 149–155. [70] Pesendorfer, M. and P. Schmidt-Dengler (2008), “Asymptotic least squares estimators for dynamic games.” Review of Economic Studies, 75, 901–928. [67,68,80,82,89] Rothenberg, T. J. (1971), “Identification in parametric models.” Econometrica, 39, 577– 591. [80] Royden, H. L. (1988), Real Analysis. Prentice-Hall. [90] Rust, J. (1987), “Optimal replacement of GMC bus engines: An empirical model of Harold Zurcher.” Econometrica, 55, 999–1033. [68,82,85] Rust, J. (1988), “Maximum likelihood estimation of discrete control processes.” SIAM J. Control and Optimization, 26, 1006–1024. [68,72,98] Schorfheide, F. (2005), “VAR forecasting under misspecification.” Journal of Econometrics, 128, 99–136. [70]
Quantitative Economics 10 (2019) Inference in dynamic discrete problems 103 Tauchen, G. (1985), “Diagnosing testing and evaluation of maximum likelihood models.” Journal of Econometrics, 30, 415–443. [70,76] White, H. (1982), “Maximum likelihood estimation of misspecified models.” Econometrica, 50, 681–700. [70] White, H. (1996), Estimation, Inference and Specification Analysis. Econometric Society Monographs, Vol. 22. Cambridge University Press. [70,93] Co-editor Christopher Taber handled this manuscript. Manuscript received 10 July, 2017; final version accepted 28 April, 2018; available online 9 May, 2018.