Polynomial regressions and nonsense inference
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Ventosa-Santaulària, Daniel; Rodríguez-Caballero, Carlos Vladimir Article Polynomial regressions and nonsense inference Econometrics Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Ventosa-Santaulària, Daniel; Rodríguez-Caballero, Carlos Vladimir (2013) : Polynomial regressions and nonsense inference, Econometrics, ISSN 2225-1146, MDPI, Basel, Vol. 1, Iss. 3, pp. 236-248, https://doi.org/10.3390/econometrics1030236 This Version is available at: https://hdl.handle.net/10419/103623 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/3.0/
Econometrics 2013,1, 236-248; doi:10.3390/econometrics1030236 OPEN ACCESS econometrics ISSN 2225-1146 www.mdpi.com/journal/econometrics Article Polynomial Regressions and Nonsense Inference Daniel Ventosa-Santaul` aria 1,* and Carlos Vladimir Rodr´ ıguez-Caballero 2 1Centro de Investigaci´ on y Docencia Econ´ omicas (CIDE), Divisi´ on de Econom´ ıa, Carretera M´ exico-Toluca 3655 Col. Lomas de Santa Fe, Delegaci´ on ´ Alvaro Obreg´ on, M´ exico 01210, Mexico 2Center for Research in Econometric Analysis of Time Series (CREATES) and Department of Economics and Business, Aarhus University, Fuglesangs All´ e 4, Building 2622 (203), Aarhus V 8210, Denmark; E-Mail: [email protected] *Author to whom correspondence should be addressed; E-Mail: daniel.v[email protected]; Tel.: +52-5727-9800 (ext. 2723). Received: 6 August 2013; in revised form: 28 October 2013 / Accepted: 7 November 2013 / Published: 18 November 2013 Abstract: Polynomial specifications are widely used, not only in applied economics, but also in epidemiology, physics, political analysis and psychology, just to mention a few examples. In many cases, the data employed to estimate such specifications are time series that may exhibit stochastic nonstationary behavior. We extend Phillips’ results (Phillips, P. Understanding spurious regressions in econometrics. J. Econom. 1986,33, 311–340.) by proving that an inference drawn from polynomial specifications, under stochastic nonstationarity, is misleading unless the variables cointegrate. We use a generalized polynomial specification as a vehicle to study its asymptotic and finite-sample properties. Our results, therefore, lead to a call to be cautious whenever practitioners estimate polynomial regressions. Keywords: polynomial regression; misleading inference; integrated processes Classification: MSC 62J05; 62M10; 62F05; 62F12; 91B84
Econometrics 2013,1237 1. Introduction There is some research on the effects of the nonstationarity of the variables on nonlinear relationships (spurious inference on linear regressions was uncovered by [1], and later explained by [2]). In [3], it is shown (both in finite samples and asymptotically) that six nonlinear tests (the Ramsey Regression Equation Specification Error Test (RESET), McLeod and Li test, Keenan test, Neural Network test, White’s information matrix, and the one proposed by [4]), when applied to independent random walks, tend to identify spurious (non-existing) nonlinear relationships (it is noteworthy that [5] studied the spurious regression phenomenon under stochastic nonstationarity when the logarithms of independent integrated of order one, I(1), variables are used; logarithmic transformations are commonly used in applied studies to deal with nonlinearity). The author of [6] extend these results by studying the behavior of two additional tests: the Brock-Dechert-Scheinkman (BDS) test and another one proposed by [7]; he finds that the former also yields results that do not make sense, whilst the latter proves to have good power properties even in small samples. The author of [8] studies the properties of the nonparametric Phillip’s unit root test applied to polynomials of integrated processes and concludes, broadly speaking, that the tests does not possess an asymptotic nuisance-parameter-free distribution, except under very specific conditions. To the best of our knowledge, the “nonlinear relationship-spurious inference” literature (briefly sketched earlier) focuses on statistical tests rather than polynomial regressions. The latter are used to linearly relate the dependent variable to a kth order polynomial on an independent variable, x. Such regressions therefore fit (through ordinary least squares, OLS) a nonlinear relationship between a polynomial on the independent variable and the conditional mean of y. These specifications can be traced back to the nineteenth century, to impute series (see [9]). Despite its old age, polynomial regressions remain widely used in a large number of scientific fields, which include epidemiology/disease progression [10], geophysics [11], physics [12], political analysis [13], psychology [14], and, of course, statistics. Splines regression models (cubic splines, for example) can be used to smooth/impute series. In empirical economics, polynomial specifications can be found in many subfields, such as, financial economics [15,16], labor economics [17,18], agricultural economics [19], macroeconomics (exchange rates, [20]) and environmental economics [21,22]. An evocative example can be found in the empirical research dealing with the Kuznets curve and the environmental Kuznets curve; the inverse U-shaped relationship between the variables is typically specified as the dependent variable regressed on the independent and its square (see [23,24]; it is noteworthy that Kuznets’ specifications usually employ even-order polynomials). Even though polynomial regressions remain an important empirical tool, we could not find in the literature any attempt to study their properties when the variables behave as independent nonstationary processes. This might be so because the effect of nonstationarity is rather intuitive, and econometricians, at least those familiar with the spurious regression, could speculate that t-ratios diverge and the R2does not collapse. However, many researchers in diverse fields seem to be unaware of this possibility. In this paper, we confirm that an inference drawn from a polynomial regression, when the variables are generated as independent integrated processes, is misleading (when the variables cointegrate, inference
Econometrics 2013,1238 drawn from such a specification is no longer misleading). We provide evidence that generalizes Phillip’s results in two new directions: (i) we allow the exponent of the variables, both explanatory and dependent in a bivariate regression, to take any natural number; (ii) we allow for an arbitrary (natural number) order for the polynomial in xin a k-variate regression. The main objective of this work is to warn practitioners about the considerable risks of spurious inference when the powers of a nonstationary variable are used as regressors. This paper is organized in a very simple manner. The next section presents the data-generating processes (DGPs) and the main results, divided in two theorems. A small Monte Carlo shows that the asymptotics are a sufficiently accurate representation of the finite sample behavior of the regressions. 2. Asymptotics of Polynomial Regressions The variables, both dependent and independent, are generated as independent driftless unit roots: zt=zt1+uz,t (1) for z=x, y. The innovations, uy,t and ux,t, are independent of each other and obey the conditions stated by Phillips ([2], p. 313, Assumption 1). We use these variables to estimate the following specification: ym t=↵+xk t+ut(2) where m, k 2N. A word on notation; the symbol, D !, denotes weak convergence, and, for simplicity, Wz⌘Wz(r), for z=x, y, denotes a Wiener standard process. The stochastic integral, R1 0, is written as R. Theorem 1. Let {yt}1 t=1 and {xt}1 t=1 be independently generated by Equation (1). Estimate by OLS specification Equation (2). Then, as T!1: 1. Tm 2ˆ↵D !m yRwm yRw2k xRwk xwm yRwk x Rw2k x(Rwk x)2 2. T1 2(mk)ˆ D !m y k xRwk xwm yRwk xRwm y Rw2k x(Rwk x)2⌘m y k xe 3. T1 2tˆ D !Rwk xwm yRwk xRwm y h⇣Rw2k x(Rwk x)2⌘⇣Rw2m y(Rwm y)2(Rwk xwm yRwk xRwm y)2⌘i1 2 4. R2D !e 2Rw2k x(Rwk x)2 Rw2m y(Rwm y)2 Proof: See Appendix A. Note that all these results are an extension of [2]. It is noteworthy to mention that, for k=m=1, our results are exactly those of [2]. This implies that, no matter what power does the practitioner applies to the variables, the spurious regression phenomenon remains identical. That said, a more interesting specification should allow for a more complete polynomial of the independent variable, as in: yt=0+1xt+2x2 t+···+kxk t+ut(3) where k2N. In this case, OLS estimates still generate a spurious regression:
Econometrics 2013,1239 Theorem 2. Let {yt}1 t=1 and {xt}1 t=1 be independently generated by Equation (1). Estimate by OLS specification Equation (3). Then, as T!1: 1. 2 6 6 6 6 6 6 6 6 6 6 6 6 6 6 6 6 4 T 1 2ˆ 0 ˆ 1 T 1 2ˆ 2 . . . T 1 2(k1) ˆ k 3 7 7 7 7 7 7 7 7 7 7 7 7 7 7 7 7 5 D !y 0 B B B B B B B B B B B B B B B B @ 10 0··· 0 0x0··· 0 002 x··· 0 . . .. . .. . .··· . . . 00 0··· k x 1 C C C C C C C C C C C C C C C C A 1 ⇥ 0 B B B B B B B B B B B B B B B B B B @ 1RwxR!2 x··· Rwk x RwxR!2 xR!3 x··· R!k+1 x R!2 xR!3 x··· ··· R!k+2 x . . ........... . . Rwk x··· ··· ··· R!2k x 1 C C C C C C C C C C C C C C C C C C A 1 0 B B B B B B B B B B B B B B B B @ Rwy R!x!y R!2 x!y . . . R!k x!y 1 C C C C C C C C C C C C C C C C A 2. S2=Op(T), where S2=T1PT t=1 ⇣ytPk i=0 ˆ ixi t⌘2 3. tˆ i=Op ⇣T1 2⌘for i=0,1,2,...,k Proof: See Appendix B. Table 1. Rejection rates of t-ratios. T Specification (2) Specification (3) km With k = 4 123 1234 100 1 0.77 0.71 0.71 2 0.71 0.66 0.65 0.46 0.35 0.33 0.31 3 0.72 0.66 0.66 250 1 0.85 0.82 0.82 2 0.81 0.78 0.78 0.64 0.56 0.52 0.50 3 0.82 0.78 0.78 500 1 0.89 0.87 0.87 2 0.86 0.84 0.84 0.73 0.67 0.64 0.63 3 0.88 0.84 0.84 Rejection rates of the t-ratio associated with: (i) for specification Equation (2), ˆ ; (ii) for specification Equation (3), all ’s. Data-generating process (DGP) parameters: uz,t ⇠iidN(0,1), for z=x, y. The code of this Monte Carlo experiment is available as supplementary material.
Econometrics 2013,1240 Note the linear pattern in the order of convergence of the parameters; whilst the constant term, ˆ 0, diverges at rate T1 2,ˆ 1neither diverges, nor collapses, ˆ 2collapses at rate T1 2, and so on. Nonetheless, all the t-ratios of the estimated parameters diverge at the usual rate T1 2. In both theorems, the convergence rate of the t-ratios associated with the estimates diverge. This implies that, for a sufficiently large sample, the null hypothesis that the parameters are equal to zero will eventually be rejected. Finite sample evidence suggests that this actually occurs in even rather small samples of 100–500 observations (Table 1). 3. Concluding Remarks In this paper, we extended the results of what is known as spurious inference by studying the asymptotic and finite-sample behavior of the t-ratios in an OLS-estimated regression, where the dependent variable and/or the explanatory variable are nonlinearly transformed by means of a polynomial. When the variables are independent and stochastically nonstationary, the inference based on OLS estimates is misleading. Our results concern pure integrated of order one processes, but provide a natural guide to future research; near-integration, integrated of order two, and broken linear trend processes should be further studied. This result should be understood as a call to be cautious whenever practitioners estimate polynomial regressions. Acknowledgments The first draft of the article was written while Carlos Vladimir Rodr´ ıguez-Caballero was visiting the Center for Research and Teaching in Economics (CIDE). He gratefully acknowledges Alejandro L´ opez-Feldman for his support. Conflicts of Interest The authors declare no conflict of interest. Appendix A. Proof of Theorem 1 Proof. In order to get all of the results, we use the asymptotic results provided in [25]: 1. T1 2(k+2) P⇠k z,t1 D !k zR1 0[!z(r)]kdr 2. T1 2(k+m+2) P⇠k x,t1⇠m y,t1 D !k xm yRwk xwm y
Econometrics 2013,1241 We now define the rates of convergences of OLS estimates (Pis short for PT t=1). 2 6 4 ˆ↵ ˆ 3 7 5=2 6 4 TPxk t Pxk tPx2k t 3 7 5 12 6 4Pym t Pxk tym t 3 7 5 =2 6 6 4 O(T)Op ⇣T1 2(k+2)⌘ Op ⇣T1 2(k+2)⌘Op Tk+13 7 7 5 12 6 6 4 Op ⇣T1 2(m+2)⌘ Op ⇣T1 2(m+k+2)⌘3 7 7 5 By simple algebra, we get: 2 6 4 ˆ↵ ˆ 3 7 5=2 6 6 4 Op Tm 2 Op ⇣T1 2(mk)⌘3 7 7 5 Therefore: 2 6 4 Tm 2ˆ↵ T1 2(mk)ˆ 3 7 5D !2 6 4 1k xRwk x k xRwk x2k xRw2k x 3 7 5 12 6 4 m yRwm y m yk xRwk xwm y 3 7 5 =1 2k x⇣Rw2k x(Rwk x)2⌘2 6 4 2k xRw2k xk xRwk x k xRwk x1 3 7 5⇥ 2 6 4 m yRwm y m yk xRwk xwm y 3 7 5 =1 2k x(Rw2k x(Rwk x)2)2 6 4 m y2k xRwm yRw2k xRwk xwm yRwk x m yk x(Rwk xwm yRwk xRwm y) 3 7 5 Finally: 2 6 4 Tm 2ˆ↵ T1 2(mk)ˆ 3 7 5D !2 6 6 6 6 4 m yRwm yRw2k xRwk xwm yRwk x Rw2k x(Rwk x)2 m y k xRwk xwm yRwk xRwm y Rw2k x(Rwk x)2 3 7 7 7 7 5 as T!1 (A1) which proves results 1and 2in Theorem 1.
Econometrics 2013,1242 Let ˜ ⌘Rwk xwm yRwk xRwm y Rw2k x(Rwk x)2; following [2], to get tˆ , we define S2=T1P⇣ym tˆ↵ˆ xk t⌘2 . Then: TmS2=T(m+1) Xh(ym t¯y)ˆ xk t¯xi2 =T(m+1) X(ym t¯y)2ˆ 2T(m+1) Xxk t¯x2 TmS2D !2m y"Zw2m y✓Zwm y◆2 ˜ 2 Zw2k x✓Zwk x◆2!# (A2) as T!1 Then, we use Equations (A1) and (A2) to get: T1 2tˆ =ˆ T1 2Sˆ =ˆ T1 2ShP(xk t¯x)2i1 2 =⇣T1 2(mk)⌘ˆ T⇣T(k+1) P(xk t¯x)2⌘1 2 T(Tm 2S) T1 2tˆ D ! m y k x ˜ k xhRw2k x(Rwk x)2i1 2 m yhRw2m y(Rwm y)2˜ 2⇣Rw2k x(Rwk x)2⌘i1 2as T!1 Then, after simple algebra, we get: T1 2tˆ D !Rwk xwm yRwm yRwk x h⇣Rw2k x(Rwk x)2⌘⇣Rw2m y(Rwm y)2(Rwk xwm yRwk xRwm y)2⌘i1 2 as T!1 proving result 3of Theorem 1. Finally, the asymptotic nonstandard distribution of R2is given by: R2=P(ˆym t¯y)2 P(ym t¯y)2 =ˆ 2T(mk)T(k+1) P(xk t¯x)2 T(m+1) P(ym t¯y)2 R2D !˜ 2hRw2k x(Rwk x)2i Rw2m y(Rwm y)2,as T!1 This proves the last result of Theorem 1.
Econometrics 2013,1243 B. Proof of Theorem 2. Proof. Polynomial specification Equation (3) has the following OLS estimators: 2 6 6 6 6 6 6 6 6 6 6 6 6 6 6 6 6 4 ˆ 0 ˆ 1 ˆ 2 . . . ˆ k 3 7 7 7 7 7 7 7 7 7 7 7 7 7 7 7 7 5 = 0 B B B B B B B B B B B B B B B B B B @ TPxtPx2 t··· Pxk t PxtPx2 tPx3 t··· Pxk+1 t Px2 tPx3 t··· ··· Pxk+2 t . . ........... . . Pxk t··· ··· ··· Px2k t 1 C C C C C C C C C C C C C C C C C C A 10 B B B B B B B B B B B B B B B B @ Pyt Pxtyt Px2 tyt . . . Pxk tyt 1 C C C C C C C C C C C C C C C C A or ˆ B=⌃ 1 xx ⌃xy, for short. To obtain the rates of convergences of the OLS estimates, note that ⌃1 xx is a Hankel matrix. The orders of convergence of each element in such a matrix are given by: 0 B B B B B B B B B B B B B B B B B B @ O(T)Op ⇣T3 2⌘Op T2··· Op ⇣T1 2(k+2)⌘ Op ⇣T3 2⌘Op T2Op ⇣T5 2⌘··· Op ⇣T1 2(k+3)⌘ Op T2Op ⇣T5 2⌘··· ··· Op ⇣T1 2(k+4)⌘ . . ........... . . Op ⇣T1 2(k+2)⌘··· ··· ··· Op Tk+1 1 C C C C C C C C C C C C C C C C C C A The Hankel matrix can be inverted using some results from linear algebra theory (spectral decomposition). Furthermore, while it would be possible to analyze some interesting properties of Hankel matrices given by [26] or [27]inter alia, there are some numerical algorithms, like [28], or [29] for polynomial regressions, that work with a Hankel matrix. That said, we are not interested in computing the exact inverse, but rather, using cases with k=1,2,3,.... For specification Equation (3), it is straightforward to see that: ⌃1 xx = 0 B B B B B B B B B B B B B B B B B B @ O(T1)Op ⇣T 3 2⌘Op T2··· Op ⇣T 1 2(k+2)⌘ Op ⇣T 3 2⌘Op T2Op ⇣T 5 2⌘··· Op ⇣T 1 2(k+3)⌘ Op T2Op ⇣T 5 2⌘··· ··· Op ⇣T 1 2(k+4)⌘ . . ........... . . Op ⇣T 1 2(k+2)⌘··· ··· ··· Op T(k+1) 1 C C C C C C C C C C C C C C C C C C A