Simple and honest confidence intervals in nonparametric regression
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Armstrong, Timothy B.; Kolesár, Michal Article Simple and honest confidence intervals in nonparametric regression Quantitative Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Armstrong, Timothy B.; Kolesár, Michal (2020) : Simple and honest confidence intervals in nonparametric regression, Quantitative Economics, ISSN 1759-7331, The Econometric Society, New Haven, CT, Vol. 11, Iss. 1, pp. 1-39, https://doi.org/10.3982/QE1199 This Version is available at: https://hdl.handle.net/10419/217182 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/
Quantitative Economics 11 (2020), 1–39 1759-7331/20200001 Simple and honest confidence intervals in nonparametric regression Timothy B. Armstrong Department of Economics, Yale University Michal Kolesár Department of Economics, Princeton University We consider the problem of constructing honest confidence intervals (CIs) for a scalar parameter of interest, such as the regression discontinuity parameter, in nonparametric regression based on kernel or local polynomial estimators. To ensure that our CIs are honest, we use critical values that take into account the possible bias of the estimator upon which the CIs are based. We show that this approach leads to CIs that are more efficient than conventional CIs that achieve coverage by undersmoothing or subtracting an estimate of the bias. We give sharp efficiency bounds of using different kernels, and derive the optimal bandwidth for constructing honest CIs. We show that using the bandwidth that minimizes the maximum mean-squared error results in CIs that are nearly efficient and that in this case, the critical value depends only on the rate of convergence. For the commoncaseinwhichtherateofconvergenceisn−2/5, the appropriate critical value for 95% CIs is 218, rather than the usual 196 critical value. We illustrate our results in a Monte Carlo analysis and an empirical application. Keywords. Confidence intervals, regression discontinuity, nonparametric regression. JEL classification. C14, C21. 1. Introduction This paper considers the problem of constructing confidence intervals (CIs) for a scalar parameter T(f)of a function f, which can be a conditional mean or a density. The scalar parameter may correspond, for example, to a conditional mean, or its derivatives at a point, the regression discontinuity or the regression kink parameter, or the value of a density or its derivatives at a point. A popular approach to estimation of T(f) is to use kernel or local polynomial estimators. These estimators are both simple to implement, Timothy B. Armstrong: [email protected] Michal Kolesár: [email protected] We thank Don Andrews, Sebiastian Calonico, Matias Cattaneo, Max Farrell, Christoph Rothe, and numerous seminar and conference participants for helpful comments and suggestions. We thank Kwok Hao Lee for research assistance. All remaining errors are our own. The research of the first author was supported by National Science Foundation Grant SES-1628939. The research of the second author was supported by National Science Foundation Grant SES-1628878. ©2020 The Authors. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at http://qeconomics.org.https://doi.org/10.3982/QE1199
2Armstrong and Kolesár Quantitative Economics 11 (2020) and highly efficient in terms of their mean squared error (MSE) properties (Fan (1993), Cheng, Fan, and Marron (1997)). CIs are typically formed by undersmoothing (choosing the bandwidth to shrink more quickly than the MSE optimal bandwidth) or biascorrection (subtracting an estimate of the estimator’s bias). In this paper, we propose a simple alternative approach to forming CIs based on these estimators that is more efficient than both undersmoothing and bias-correction in the sense that it leads to shorter CIs while maintaining coverage over the same parameter space Ffor f(which typically places bounds on derivatives of f). In particular, one simply adds and subtracts the estimator’s standard error times a critical value that is larger than the usual normal quantile z1−α/2, and takes into account the possible bias of the estimator.1Asymptotically, these CIs correspond to fixed-length CIs as defined in Donoho (1994), and so we refer to them as fixed-length CIs. We show that the critical value depends only on (1) the order of the derivative that one bounds to define the parameter space F; and (2) the criterion used to choose the bandwidth. In particular, if the MSE optimal bandwidth is used with a local linear estimator, computing our CI at the 95% coverage level amounts to replacing the usual critical value z0975 =196 with 218. When the criterion for bandwidth choice is the length of the resulting CI, we show that the resulting bandwidth is in fact larger than the MSE optimal bandwidth. This contrasts with the work of Hall (1992)andCalonico, Cattaneo, and Farrell (2018)onoptimality of undersmoothing. Importantly, these papers restrict attention to CIs that use the usual critical value z1−α/2. It then becomes necessary to choose a small enough bandwidth so that the bias is asymptotically negligible relative to the standard error, since this is the only way to achieve correct coverage. Our results imply that rather than choosing a smaller bandwidth, it is better to use a larger critical value that takes into account the potential bias; this also ensures correct coverage regardless of the bandwidth sequence. While the fixed-length CIs shrink at the optimal rate, undersmoothed CIs shrink more slowly. We also show that under smoothness assumptions needed to implement biascorrection, our CIs shrink at a faster rate than bias-corrected CIs, once the standard error is adjusted to take into account the variability of the bias estimate (Calonico, Cattaneo, and Titiunik (2014) show that doing so is important for maintaining coverage). The oversmoothing relative to the MSE optimal bandwidth is relatively modest: under a range of conditions most commonly used in practice, a fixed-length CI centered at the MSE optimal bandwidth is 99% efficient relative to using the CI optimal bandwidth. Therefore, a practically attractive implementation of our CIs is to simply center them around an estimator with MSE optimal bandwidth, rather than reoptimizing the bandwidth for length and coverage of the CI. A key requirement that underlies our results is the notion of honesty: as in Li (1989), we require that the CIs cover the true parameter asymptotically at the nominal level uniformly over the parameter space F. Furthermore, we allow this parameter space to grow with the sample size. The notion of honesty is closely related to the use of the minimax 1An R package implementing our CIs in regression discontinuity designs is available at https://github. com/kolesarm/RDHonest.
Quantitative Economics 11 (2020) Simple and honest confidence intervals 3 criterion used to derive the MSE efficiency results: in both cases, one requires good performance uniformly over the parameter space F. The requirement that the CIs be honest is necessary for good finite-sample performance. In contrast, approaches to inference based on pointwise-in-fasymptotics, such as using bandwidths that optimize the pointwise-in-fasymptotic MSE can lead to arbitrarily poor finite-sample behavior, as we discuss further in Section 4.1. To illustrate the practical importance of this point, we conduct a Monte Carlo study in which we show that commonly used CIs based on plugin bandwidths that attempt to estimate this pointwise-in-foptimal bandwidth exhibit severe undercoverage, even when combined with undersmoothing or bias-correction. When the parameter space places a bound Mon a derivative of f, our CIs require this bound to be specified explicitly. While this may appear to be a disadvantage of our particular approach, due to impossibility results of Low (1997), Cai and Low (2004), and Armstrong and Kolesár (2018a), this cannot be avoided, regardless of how one forms the CI, without making further restrictions on the function f. In particular, these papers show that, without additional assumptions on the parameter space, one cannot use a data-driven method to estimate Mand maintain coverage over the whole parameter space—any other method that appears to avoid making this choice must do so implicitly. For example, an apparent advantage of undersmoothing is that it leads to correct coverage for any fixed smoothness constant M. However, as we discuss in detail in Section 4.2, a more accurate description of undersmoothing is that for each sample size n, it implicitly chooses a constant Mnunder which coverage is controlled. Given a sequence of undersmoothed bandwidths, we show how Mncan be calculated explicitly. One can then obtain a shorter CI with the same coverage properties by computing a fixed-length CI for the corresponding Mn. Regardless of how one chooses M,thefixedlength CIs we propose are more efficient than undersmoothed or bias-corrected CIs that use the same (implicit or explicit) choice of M. In fact, it follows from the calculations in Donoho (1994)andArmstrong and Kolesár (2018a) that our CIs, when constructed using a length-optimal or MSE-optimal bandwidth, are highly efficient among all honest CIs: no other approach to inference can substantively improve on their length, while still maintaining coverage. As an alternative to choosing Ma priori, one can place additional conditions on the function fthat allow for an upper bound on Mto be estimated. To maintain efficiency of the resulting CI, however, care must be taken in doing so: if Mis a bound on the pth derivative, and one imposes a bound ˜ Mon the (p +1)th derivative in order to estimate M, then the optimal CI will be based on a different estimator and will depend on the new bound ˜ M. To avoid such issues, we propose a regularity class that relates a global polynomial approximation to smoothness of the function fnear the point of interest, and we show formally that, for this class, one can obtain a valid and highly efficient CI using a global polynomial rule of thumb suggested by Fan and Gijbels (1996). However, given the additional assumptions required by this (or any) data driven choice of M,werecommend that this approach be used as a starting point for sensitivity analysis allowing for other choices of M. Another approach to data-driven choices of Mis to use “self-similarity” conditions, as suggested by Giné and Nickl (2010), which relate the maximum and minimum bias
4Armstrong and Kolesár Quantitative Economics 11 (2020) at different bandwidths. Bull (2012)andChernozhukov, Chetverikov, and Kato (2014) have obtained rate optimal confidence bands under such conditions, which, like the CIs considered here, use a critical value based on an upper bound on the bias. While these results for confidence bands could be extended to cover the problem of constructing CIs for a scalar parameter, obtaining sharp critical values appears to be very difficult. Indeed, the results of Armstrong (2019) show that the sharp form of such CIs must depend to first order on auxiliary constants used to define self-similarity. Nonetheless, our approach of bounding local smoothness using a global polynomial approximation is inspired by the self-similarity approach taken by this literature, and we see it as being in the same spirit. Schennach (2015) also used an upper bound on the bias based on an estimated smoothness constant. While the coverage of the resulting CIs is pointwise-inf, it is plausible that the CIs are honest under additional auxiliary conditions, similar in spirit to self-similarity. In addition to calculating the relative efficiency of CIs constructed using different bandwidths, our results allow us to calculate the relative efficiency of CIs constructed using different kernels. In particular, we show that the relative efficiency of kernels for the CIs we propose is the same as the relative efficiency of the estimates in terms of MSE. Thus, relative efficiency calculations for MSE, such as the ones in Fan (1993), Cheng, Fan, and Marron (1997), and Fan, Gasser, Gijbels, Brockmann, and Engel (1997) for estimation of a nonparametric mean at a point (estimation of f(x0)for some x0) that motivate much of empirical practice in the applied regression discontinuity literature, translate directly to CI construction. Despite their importance in motivating empirical practice, however, such results are subject to a technical critique about how the parameter space is specified: rather than placing a bound on a derivative of f(a Hölder condition), currently available relative efficiency results place assumptions directly on the error of a Taylor approximation at a particular point, so that some “nonsmooth” functions are in fact not ruled out.2To address this, we derive the minimax performance of local polynomial estimators under Hölder restrictions on f. These results confirm that the local polynomial estimators used in empirical practice are also highly efficient under Hölder restrictions on f. Furthermore, while we focus on asymptotic CIs and relative efficiency, these results include a derivation of the finite-sample worst-case bias of local polynomial estimators under Hölder restrictions, which was used by Kolesár and Rothe (2018) to form finite-sample valid CIs in a fixed-design regression setting. These findings may be of independent interest. The requirement of honesty is also important to ensure that our concept of optimality is well-defined and consistent. As discussed above, it allows us to consider bandwidth or kernel efficiency for constructing CIs. In addition, it also allows us to formally show that using local polynomial regression of an order that is too high given the amount of smoothness imposed is suboptimal. In contrast, under pointwise-in-fasymptotics, high-order local polynomial estimates are superefficient at every point in the parameter space (see Chapter 1.2.4 in Tsybakov (2009), and Brown, Low, and Zhao (1997)). 2See Imbens and Wager (2019), as well as our discussion in Section 3.2.1 for an elaboration of this critique.
Quantitative Economics 11 (2020) Simple and honest confidence intervals 5 To illustrate the implementation of the honest CIs, we reanalyze the data from Ludwig and Miller (2007), who, using a regression discontinuity design, find a large and significant effect of receiving technical assistance to apply for Head Start funding on child mortality at a county level. However, this result is based on CIs that ignore the possible bias of the local linear estimator around which they are built, and an ad hoc bandwidth choice. We find that, if one bounds the second derivative globally by a constant Musing a Hölder class, the uncertainty associated with the effect size is much larger than originally reported, unless one is very optimistic about the constant M, allowing fto only be linear or nearly-linear. Our results build on the literature on estimation of linear functionals in normal models with convex parameter spaces, as developed by Donoho (1994), Ibragimov and Khas’minskii (1985) and many others. As with the results in that literature, our setup gives asymptotic results for problems that are asymptotically equivalent to the Gaussian white noise model, including nonparametric regression (Brown and Low (1996)) and density estimation (Nussbaum (1996)). Our main results build on the “renormalization heuristics” of Donoho and Low (1992), who show that many nonparametric estimation problems have renormalization properties that allow easy computation of minimax MSE optimal kernels and rates of convergence. Our results hold under essentially the same conditions, which apply in many classical nonparametric settings. The CIs we consider in this paper are applications of the fixed-length CIs proposed in the context of inference on linear functionals T(f)in Gaussian nonparametric regression by Donoho (1994), which have also been studied recently in Armstrong and Kolesár (2018a), and in contemporaneous and subsequent work by Kolesár and Rothe (2018) and Imbens and Wager (2019). In contrast to the finite-sample approach taken in these papers, we focus on asymptotic results, and we also allow T(f)to be nonlinear. Instead of imposing the nonparametric regression model, we require a renormalization condition (see equation (4) below) that allows us to apply the “renormalization heuristics” of Donoho and Low (1992); we are thus able to cover settings such as density estimation or estimation of a bidder valuation in first-price auctions (see Appendix C in the Online Supplemental Material). Our asymptotic approach allows for simplifications that deliver our main relative efficiency results. These efficiency results are different from and complementary to the asymptotic form of the efficiency bounds given in Donoho (1994)and Armstrong and Kolesár (2018a): whereas we consider relative efficiency of estimators and fixed-length CIs based on different kernels and bandwidths, Donoho (1994)and Armstrong and Kolesár (2018a) bound the scope for efficiency gains from CIs that do not fall into this class. Donoho (1994)andArmstrong and Kolesár (2018a) found that the scope for further improvement is small, which motivates our focus on this class of estimators and CIs. See Remark 2.3 for further discussion. The rest of this paper is organized as follows. Section 2gives the main results. Section 3applies our results to inference at a point, sharp and fuzzy RD, and it discusses practical implementation issues, including a rule of thumb for choosing M.Section4 gives a theoretical comparison of our fixed-length CIs to other approaches, and Section 5compares them in a Monte Carlo study. Finally, Section 6presents an empirical application based on Ludwig and Miller (2007). Appendix Agivesproofsoftheresultsin
6Armstrong and Kolesár Quantitative Economics 11 (2020) Section 2. Additional results are collected in the Appendices in the Online Supplemental Material (Armstrong and Kolesár (2020)). 2. General results We are interested in a scalar parameter T(f)of a function f, which is typically a conditional mean or a density. The function fis assumed to lie in a function class F=F(M), which places “smoothness” conditions on f,whereMindexes the level of smoothness. We focus on classical nonparametric function classes, in which Mcorresponds to a bound on a derivative of fof a given order. We allow M=Mnto grow with the sample size n. We have available a class of estimators ˆ T(h;k), indexed by a bandwidth h=hn>0 and a kernel k.Let se(h;k) denote the standard error of ˆ T(h;k), an estimate of its standard deviation sdf(ˆ T(h;k)). We assume that a central limit theorem applies to ˆ T(h;k), so that in large samples, the t-statistic [ˆ T(h;k) −T(f)]/ se(h;k) will be approximately normal with variance 1and mean given by the ratio of bias to standard deviation, tf=(Ef[ˆ T(h;k) −T(f)])/sdf(ˆ T(h;k)). Since tfdepends on the unknown function f, this ratio is unknown. Note, however, that we can bound |tf|by the worst-case ratio of bias to standard deviation (bias-sd ratio), tF=supf∈F|Ef[ˆ T(h;k)−T(f)]|/sdf(ˆ T(h;k)). Therefore, if this bias-sd ratio can be computed up to asymptotically negligible terms, we can construct an honest CI as ˆ T(h;k) ±cv1−α(t) · se(h;k) (1) where the approximate bias-sd ratio tsatisfies t=tF(1+o(1)),andcv1−α(t) is the 1−α quantile of the folded normal distribution |N(t1)|, or, equivalently, the square root of the 1−αquantile of a χ2distribution with 1degree of freedom, and noncentrality parameter t2, which is readily available in statistical software. For easy reference, we list these critical values in Table 1for selected values of t. Because the quantiles of a χ2distribution are increasing in its noncentrality parameter, replacing tfwith an upper bound Table 1. Critical values cv1−α(·). α rt001 005 01 002576 1960 1645 6/70408 2764 2113 1777 4/5052842 2181 1839 2/30707 3037 2362 2008 1/2103327 2646 2284 153826 3145 2782 204326 3645 3282 Note:Criticalvaluescv1−α(t) and cv1−α(1/r −1),fortheFLCIsin(1)and(8), corresponding to the 1−αquantiles of the |N(t1)|and |N(1/r −11)|distributions, where tis the bias-sd ratio, and ris the rate exponent. For t≥2,cv1−α(t) ≈ t+z1−α/2up to 3decimal places for these values of α.
Quantitative Economics 11 (2020) Simple and honest confidence intervals 7 that is valid for all f∈Fyields a CI that is honest over F.TheCIin(1) is an approximate version of a fixed-length confidence interval (FLCI) studied in Donoho (1994), who replaces se(h;k) with sdf(ˆ T(h;k)) in the definition of this CI, and assumes sdf(ˆ T(h;k)) is constant over f, in which case its length will be fixed. We thus refer to CIs of this form as “fixed-length,” even though se(h;k) is random. To motivate our main regularity condition (4) below that will facilitate studying the performance of these FLCIs and allow for an easy computation of the bias-sd ratio t, suppose that the standard deviation and the worst-case bias of the estimator ˆ T(h;k), biasˆ T(h;k)=sup f∈FEfˆ T(h;k) −T(f) scale as powers of h. In particular, suppose that, for some γb>0,γs<0,B(k) > 0and S(k) > 0, biasˆ T(h;k)=hγbMB(k)1+o(1)sdfˆ T(h;k)=hγsn−1/2S(k)1+o(1)(2) where the o(1)term in the second equality is uniform over f∈F.WeshowinAppendix B in the Online Supplemental Material that this condition will hold whenever the renormalization heuristics of Donoho and Low (1992) can be formalized. This includes most classical nonparametric problems, such as estimation of a density or a conditional mean, or its derivative, evaluated at a point (which may be a boundary point). In Section 3.2.1, we show that (2)holdswithγb=p,andγs=−1/2under mild regularity conditions when ˆ T(h;k) is a local polynomial estimator of a conditional mean at a point, and F(M) consists of functions with pth derivative bounded by M. Remark 2.1. The second condition in (2) implies that the standard deviation does not depend on the underlying function fasymptotically. In certain settings, such as density estimation (see Appendix C.1 in the Online Supplemental Material), this may require choosing a localized sequence of parameter spaces Fn, similar to local asymptotic minimax results in parametric settings (e.g., Section 8.7 in van der Vaart (1998)). While we allow for such dependence, we keep any dependence of Fon nimplicit in our notation in the main text. Similarly, the quantities B(k) and S(k) generally depend on F(which if the parameter space is localized includes the localization point), as well as on other nuisance parameters, such as the variance of the regression errors. To prevent notational clutter, we keep this dependence implicit. Under (2), we can use the ratio t=hγb−γsMB(k)/(n−1/2S(k)) of the leading worstcase bias and standard deviation terms to compute the critical value cv1−α(t) in (1). Analogously to the two-sided case, honest one-sided 1−αCIs based on ˆ T(h;k) can be constructed by subtracting the standard error times a 1−αquantile of the distribution N(t1). This is asymptotically equivalent to the CI ˆ T(h;k) −hγbMB(k)−z1−αhγsn−1/2S(k)∞(3) which subtracts the maximum bias, in addition to subtracting z1−αtimes the standard deviation, from ˆ T(h;k).
8Armstrong and Kolesár Quantitative Economics 11 (2020) Remark 2.2. One could also form honest two-sided CIs by simply adding and subtracting the worst case bias, in addition to adding and subtracting the standard error times z1−α/2=cv1−α(0),the1−α/2quantile of a standard normal distribution, forming the CI as ˆ T(h;k) ±(hγbMB(k)+z1−α/2· se(h;k)). However, since the estimator ˆ T(h;k) cannot simultaneously have a large positive and a large negative bias, such CI will be conservative, and longer than the CI given in equation (1). To discuss the optimal choice of bandwidth hand compare efficiency of different kernels kin forming oneand two-sided CIs, and compare the results to the bandwidth and kernel efficiency results for estimation, it will be useful to introduce notation for a generic performance criterion. Let R( ˆ T)denote the worst-case (over F) performance of ˆ Taccording to a given criterion, and let ˜ R(bs) denote the value of this criterion when ˆ T−T(f)∼N(bs2). For FLCIs, we can take their half-length as the criterion, which leads to RFLCIαˆ T(h;k)=infχ:Pfˆ T(h;k) −T(f)≤χ≥1−αfor all f∈F ˜ RFLCIα(bs) =infχ:PZ∼N(01)|sZ +b|≤χ≥1−α=s·cv1−α(b/s) To evaluate one-sided CIs, one needs a criterion other than length, which is infinite. A natural criterion is expected excess length, or quantiles of excess length. We focus here on the quantiles of excess length. For CI of the form (3), its worst-case βquantile of excess length is given by ROCIαβ(ˆ T(h;k)) =supf∈Fqfβ(T(f ) −ˆ T(h;k) +hγbMB(k) + z1−αhγsn−1/2S(k)),whereqfβ(Z) is the βquantile of a random variable Z.Theworstcase βquantile of excess length based on an estimator ˆ Twhen ˆ T−T(f)is normal with variance s2and bias ranging between −band bis ˜ ROCIαβ(bs) =2b+(z1−α+zβ)s.Finally, to evaluate ˆ T(h;k) as an estimator we use the maximum root mean squared error (RMSE) under Fas the performance criterion: RRMSE(ˆ T)=sup f∈FEfˆ T−T(f)2˜ RRMSE(b s) =b2+s2 The key regularity condition that we impose on the class of estimators ˆ T(h;k) is that their performance can be approximated in large samples by the performance of a normally distributed estimator with bias and standard deviation that scale as powers of h, Rˆ T(h;k)=˜ RhγbMB(k)hγsn−1/2S(k)1+o(1)(4) For the performance criteria above, if the estimator ˆ T(h;k) satisfies an appropriate central limit theorem, and equation (2) holds, condition (4) will hold so long as the estimator is centered, so that, up to asymptotically negligible terms, its maximum and minimum bias over Fsum to zero, supf∈FEf(ˆ T(h;k) −T(f))=−inff∈FEf(ˆ T(h;k) − T(f))(1+o(1)).3Heuristically, this follows because if (ˆ T(h;k) −Efˆ T(h;k))/sdf(ˆ T) is 3This centering condition holds automatically by a symmetry argument for kernel or local polynomial estimators if fis a conditional mean or a density, T(f) is its value or its derivative at a point, or a regres-
Quantitative Economics 11 (2020) Simple and honest confidence intervals 15 with the weight wn +given by wn +(x;hk) =e 1Q−1 n+m1(x)k+(x/h) k+(u) =k(u)I{u≥0} and Qn+=n i=1k+(xi/h)m1(xi)m1(xi). The weights wn −,GrammatrixQn−and kernel k−are defined similarly. Let σ2 +(x) =σ2(x)I{x≥0},andσ2 −(x) =σ2(x)I{x<0}. Fuzzy RD In a fuzzy RD design, the treatment diis not entirely determined by whether the running variable xiexceeds a cutoff. Instead, the cutoff induces a jump in the treatment probability. This fits into our framework if we let f=(f1f2)comprise two regression functions, corresponding to the reduced-form regression of the outcome on the running variable, and the first-stage regression of the treatment on the running variable: yi=f1(xi)+ui1 di=f2(xi)+ui2 i=1n Eui=0var(ui)=Ω(xi) (13) with ui=(ui1ui2). The parameter of interest is given by the ratio T(f)=L1(f)/L2(f ) of sharp RD parameters Lj(f) =limx↓0fj(x) −limx↑0fj(x) in the reduced-form (j=1)and first-stage regression (j=2). If the regression functions of the potential outcomes and potential treatments are continuous at zero, and a monotonicity condition holds, then T(f)measures the average treatment effect for individuals with xi=0who are compliers (see Hahn, Todd, and van der Klaauw (2001)). We consider estimating T(f)by its sample analog, replacing L1and L2with sharp RD local linear estimates, which are for simplicity assumed to be based on the same bandwidth, ˆ T(h;k) =ˆ L1(h;k)/ ˆ L2(h;k),where ˆ L(h;k) =ˆ L1(h;k) ˆ L1(h;k)= iwn +(x;hk) −wn −(x;hk)yi di with the weights wn +and wn −defined as in (12). 3.2 Theoretical results We now discuss the conditions under which the key regularity condition (4)holdsin each application. We also discuss kernel efficiency results, and gains from imposing global, rather than just local, smoothness on f. 3.2.1 Inference at a point To state the results, it will be convenient to define the equivalent kernel k∗ q(u) =e 1X mq(t)mq(t)k(t)dt−1 mq(u)k(u) (14) where the integral is over X=Rif 0is an interior point, and over X=[0∞)if 0is a (left) boundary point. We assume the following conditions on the design points and regression errors ui. Assumption 3.1. For some d>0,the sequence {xi}n i=1satisfies 1 nhnn i=1g(xi/hn)→d· Xg(u)dufor any bounded function gwith finite support and any sequence hnwith 0< liminfnhn(nM2)1/(2p+1)<lim supnhn(nM2)1/(2p+1)<∞.
16 Armstrong and Kolesár Quantitative Economics 11 (2020) Assumption 3.2. The random variables {ui}n i=1are independent with Eui=0,Eu2+η i≤ 1/η for some η>0,and var(ui)=σ2(xi)for some variance function σ2(x) that is continuous at x=0with σ2(0)>0. Assumption 3.1 requires that the empirical distribution of the design points is smooth around 0. When the support points are treated as random, the constant dtypically corresponds to their density at 0. Because the estimator is linear in yi, its variance does not depend on f, sdˆ Tq(h;k)2= n i=1 wn q(xi)2σ2(xi)=S(k)2 nh 1+o(1) S(k) = σ2(0)X k∗ q(u)2du d (15) where the second equality holds under Assumptions 3.1 and 3.2,asweshowinAppendix B.3 in the Online Supplemental Material. The condition on the standard deviationinequation(2)thusholdswithγs=−1/2,andS(k) given in the preceding display. Appendix D.3 in the Online Supplemental Material gives the constant Xk∗ q(u)2dufor selected kernels. On the other hand, the worst-case bias will be driven primarily by the function class F. We consider inference under two popular function classes. First, the Taylor class of order p, FTp(M) =f:f(x)− p−1 j=0 f(j)(0)xj/j !≤M|x|p/p!x∈X This class consists of all functions for which the approximation error from a (p −1)th order Taylor approximation around 0can be bounded by 1 p!M|x|p. It formalizes the idea that the pth derivative of fat zero should be bounded by some constant M.Using this class of functions to derive optimal estimators goes back at least to Legostaeva and Shiryaev (1971), and it underlies much of existing minimax theory concerning local polynomial estimators (see Fan and Gijbels (1996, Chapters 3.4–3.5)). While analytically convenient, the Taylor class may not be attractive in some empirical settings because it allows fto be nonsmooth and discontinuous away from 0. We therefore also consider inference under Hölder classes (for simplicity, we focus on Hölder classes of integer order) FHölp(M) =f:f(p−1)(x) −f(p−1)x≤Mx−xxx∈X This class is the closure of the family of ptimes differentiable functions with the pth derivative bounded by M, uniformly over X, not just at 0. It formalizes the intuitive notion that fshould be p-times differentiable with a bound on the pth derivative. The case p=1corresponds to the Lipschitz class of functions.
Quantitative Economics 11 (2020) Simple and honest confidence intervals 17 Theorem 3.1. Suppose that Assumption 3.1 holds and that k(·)is bounded with bounded support and q≥p−1.Then,for any bandwidth sequence hnwith nhn→∞ and 0<liminfnhn(nM2)1/(2p+1)<lim supnhn(nM2)1/(2p+1)<∞, biasFTp(M)ˆ Tq(hn;k)=Mhp n p!BT pq(k)1+o(1)BT pq(k) =Xupk∗ q(u)du and biasFHölp(M)ˆ Tq(hn;k)=Mhp n p!BHöl pq (k)1+o(1) BHöl pq (k) =p∞ t=0u∈X|u|≥t k∗ q(u)|u|−tp−1dudt Thus,the first part of equation (2)holds with γb=pand B(k) =Bpq(k)/p!,where Bpq(k) =BHöl pq (k) for FHölp(M),and Bpq(k) =BT pq(k) for FTp(M). If,in addition,Assumption 3.2 holds,then equation (4)holds for the RMSE,FLCI, and OCI performance criteria,with γband B(k) given above and γsand S(k) given in equation (15). The theorem verifies the regularity conditions needed for the results in Section 2, and implies that r=2p/(2p+1)for FTp(M) and FHölp(M).Ifp=2,thenweobtain r=4/5.ByTheorem2.1(i), the optimal rate of convergence of a criterion Ris R( ˆ T(h∗ R;k)) =O((n/M1/p)−p/(2p+1)). As we will see from the relative efficiency calculation below, the optimal order of the local polynomial regression is q=p−1for the kernels considered here. The theorem allows q≥p−1, so that we can examine the efficiency of local polynomial regressions that are of order that is too high relative to the smoothness class. Allowing for q<p−1is not meaningful, as in this case, the maximum bias is infinite.6 Under the Taylor class FTp(M), the least favorable (bias-maximizing) function is given by f(x)=M/p!·sign(wn q(x))|x|p. In particular, if the weights are not all positive, it will be discontinuous away from the boundary. The first part of Theorem 3.1 then follows by taking the limit of the bias under this function. Assumption 3.1 ensures that this limit is well-defined. Under the Hölder class FHölp(M), the least favorable function takes the form of a pth order spline. See Appendix B.3 in the Online Supplemental Material for details. These results imply that given a kernel kand order of a local polynomial q,the RMSE-optimal bandwidth for FTp(M) and FHölp(M) is given by h∗ rmse =1 2pn S(k)2 M2B(k)21 2p+1=⎛ ⎜ ⎜ ⎝ σ2(0)p!2 2pndM2X k∗ q(u)2du Bpq(k)2⎞ ⎟ ⎟ ⎠ 1 2p+1 (16) 6The smoothness classes FTp(M) and FHölp(M) do not restrict derivatives of order p−1and lower, so that, in order to achieve a finite worst-case bias, the estimator needs to be unbiased for polynomials of order p−1,whichrequiresq≥p−1.
18 Armstrong and Kolesár Quantitative Economics 11 (2020) where Bpq(k) =BHöl pq (k) for FHölp(M),andBpq(k) =BT pq(k) for FTp(M).Forkernels given by polynomial functions over their support, k∗ qalso has the form of a polynomial, and BT pq and BHöl pq can be computed analytically. Appendix D.3 in the Online Supplemental Material gives these constants for selected kernels. Kernel efficiency It follows from Theorem 2.1(ii) that the optimal equivalent kernel minimizes S(k)rB(k)1−r, independently of the performance criterion. Under the Taylor class FTp(M), this is equivalent to minimizing X k∗(u)2dup ·Xupk∗(u)du (17) The solution to this problem follows from Sacks and Ylvisaker (1978, Theorem 1) (see also Cheng, Fan, and Marron (1997)). We give details of the solution in Appendix D.2 in the Online Supplemental Material. Table 2compares the asymptotic relative efficiency of local polynomial estimators based on the uniform, triangular, and Epanechnikov kernels to the optimal Sacks–Ylvisaker kernels. Fan et al. (1997)andCheng, Fan, and Marron (1997) conjectured that minimizing (17) yields a sharp bound on kernel efficiency. It follows from Theorem 2.1(ii) that this conjecture is correct, and Table 2matches the kernel efficiency bounds in these papers. Table 2shows that the choice of the kernel does not matter very much, so long as the local polynomial is of the right order. However, if the order is too high, q>p−1, the efficiency can be quite low, even if the bandwidth used was optimal for the function class or the right order, FTp(M), especially on the boundary. If the bandwidth picked is optimal for FTq−1(M), it will shrink at a lower rate than optimal under FTp(M), and the resulting rate of convergence will be lower than r.Consequently, the relative asymptotic efficiency will be zero. A similar point in the context of pointwise asymptotics was made in Sun (2005, Remark 5, p. 8). The solution to minimizing S(k)rB(k)1−runder FHölp(M) is only known in special cases. When p=1, the optimal estimator is a local constant estimator based on the triangular kernel. When p=2, the solution is given in Fuller (1961)andZhao (1997)for Table 2. Relative efficiency of local polynomial estimators for the function class FTp(M). Boundary Point Interior Point Kernel Order p=1p=2p=3p=1p=2p=3 Uniform I{|u|≤1}009615 09615 105724 09163 09615 09712 204121 06387 08671 07400 07277 09267 Triangular (1−|u|)+01 1 106274 09728 1 09943 204652 06981 09254 08126 07814 09741 Epanechnikov 3 4(1−u2)+009959 09959 106087 09593 09959 1 204467 06813 09124 07902 07686 09672 Note: Efficiency is relative to the optimal equivalent kernel k∗ SY . The functional T(f) corresponds to the value of fat a point.
Quantitative Economics 11 (2020) Simple and honest confidence intervals 19 Table 3. Relative efficiency of local polynomial estimators for the function class FHölp(M). Boundary Point Interior Point Kernel Order p=1p=2p=3p=1p=2p=3 Uniform I{|u|≤1}009615 09615 107211 09711 09615 09662 205944 08372 09775 08800 09162 09790 Triangular (1−|u|)+01 1 107600 09999 1 09892 206336 08691 1 09263 09487 1 Epanechnikov 3 4(1−u2)+009959 09959 107471 09966 09959 09949 206186 08602 09974 09116 09425 1 Note:Forp=12, efficiency is relative to the optimal kernel, for p=3, efficiency is relative to the local quadratic estimator with triangular kernel. The functional T(f)corresponds to the value of fat a point. the interior point problem, and in Gao (2018) for the boundary point problem. See Appendix D.2 in the Online Supplemental Material for details. When p≥3, the solution is unknown. Therefore, for p=3, we compute efficiencies relative to a local quadratic estimator with a triangular kernel. Table 3calculates the resulting efficiencies for local polynomial estimators based on the uniform, triangular, and Epanechnikov kernels. Relative to the class FTp(M), the bias constants are smaller: imposing smoothness away from the point of interest helps to reduce the worst-case bias. Furthermore, the loss of efficiency from using a local polynomial estimator of order that is too high is smaller. Finally, local linear regression with a triangular kernel achieves high asymptotic efficiency under both FT2(M) and FHöl2(M), both at the interior and at a boundary, with efficiency at least 97%, giving a theoretical justification to this popular choice in empirical work. Gains from imposing smoothness globally The Taylor class FTp(M), only restricts the pth derivative locally to the point of interest, while the Hölder class FHölp(M) restricts the pth derivative globally. How much can one tighten a confidence interval or reduce the RMSE due to this additional smoothness? It follows from Theorem 3.1 and from arguments underlying Theorem 2.1 that the performance of using a local polynomial estimator of order p−1with kernel kHand optimal bandwidth under FHölp(M) relative to using a local polynomial estimator of order p−1with kernel kTand optimal bandwidth under FTp(M) is given by inf h>0RFHölp(M)ˆ T(h;kH) inf h>0RFTp(M)ˆ T(h;kT)=⎛ ⎜ ⎜ ⎝X k∗ Hp−1(u)2du X k∗ Tp−1(u)2du ⎞ ⎟ ⎟ ⎠ p 2p+1BHöl pp−1(kH) BT pp−1(kT)1 2p+11+o(1) (18) where RF(ˆ T) denotes the worst-case performance of ˆ Tover F. If the same kernel is used, the first term equals 1, and the efficiency ratio is determined by the ratio of the
20 Armstrong and Kolesár Quantitative Economics 11 (2020) Table 4. Gains from imposing global smoothness. Boundary Point Interior Point Kernel p=1p=2p=3p=1p=2p=3 Uniform 10855 0764 1 1 0848 Triangular 10882 0797 1 1 0873 Epanechnikov 10872 0788 1 1 0866 Optimal 10906 1 0995 Note: The table gives the relative asymptotic risk of local polynomial estimators of order p−1and a given kernel under the class FHölp(M) relative to the risk under FTp(M) given in equation (18). “Optimal” refers to using the optimal kernel under a given smoothness class. bias constants Bpp−1(k).Table4computes the resulting efficiency gain for common kernels. In general, the gains are greater for larger p, and greater at the boundary. For estimation at a boundary point with p=2, for example, imposing global smoothness of freduces CI length by about 13–15%, depending on the kernel, and about 10% if the optimal kernel is used. 3.2.2 Sharp regression discontinuity We focus on the most empirically relevant case in which the regression function fis assumed to lie in the class FHöl2(M) on either side of the cutoff: f∈FSRD(M) =f+(x)I{x≥0}−f−(x) I{x<0}: f+f−∈FHöl2(M) Inference on T(f)is then equivalent to inference on the difference between two regression functions evaluated at boundary points, and the results follow by a slight extension of the results for estimation at a boundary point in Section 3.2.1. It follows from the results in Section 3.2.1 that if Assumptions 3.1 and 3.2 hold (with the requirement that σ2(x) is continuous 0replaced by rightand left-continuity of σ2 +(x) and σ2 −(x)), then the variance of the estimator does not depend on fand satisfies sdˆ T(h;k)2= n i=1 ˜ wn(xi)2σ2(xi)=S(k)2 nh 1+o(1) S(k)2=∞ 0 k∗ 1(u)2duσ2 +(0)+σ2 −(0) d with ddefined in Assumption 3.1,and ˜ wn(xi)=wn +(xi)+wn −(xi).Theorem3.1 and arguments in Appendix B.3 in the Online Supplemental Material imply that the bias of ˆ T(h;k) is maximized at f(x)=−Mx2/2·(I{x≥0}−I{x<0}),solongasthekernelk(·) takes on nonnegative values. The worst-case bias therefore satisfies biasˆ T(h;k)=−M 2 n i=1 ˜ wn(xi)x2 i=Mh2B(k)1+o(1)B(k)=−∞ 0 u2k∗ 1(u)du
Quantitative Economics 11 (2020) Simple and honest confidence intervals 21 It follows that for the RMSE, FLCI, and OCI criteria, equation (4)holdswithγb=2, γs=−1/2,andB(k) and S(k) given in the displays above. Thus, the RMSE-optimal bandwidth is given by h∗ rmse =⎛ ⎜ ⎜ ⎜ ⎝∞ 0 k∗ 1(u)2du ∞ 0 u2k∗ 1(u)du2·σ2 +(0)+σ2 −(0) 4dnM2⎞ ⎟ ⎟ ⎟ ⎠ 1/5 (19) The kernel efficiency results are analogous to those in Section 3.2.1. In principle, one could allow the bandwidths on either side of the cutoff to be different. We show in Appendix D.1 in the Online Supplemental Material, however, that the loss in efficiency resulting from constraining the bandwidths to be the same is quite small unless the ratio of variances on either side of the cutoff, σ2 +(0)/σ2 −(0),isquitelarge. 3.2.3 Fuzzy regression discontinuity We assume that f=(f1f2)lies in the class FFRD(M1M2)=FSRD(M1)×FSRD(M2), so that both the reduced-form and the firststage regression functions are assumed to have a bounded second derivative on either side of the cutoff.7 Since the estimator is nonlinear, to ensure that (4) holds, it will be necessary to consider a sequence of parameter spaces FFRDn(M1M2)localized around a particular value L∗of L(f) =(L1(f ) L2(f))with a nonzero jump in the first-stage regression L∗ 2= 0. This allows us to apply a version of the delta method to ˆ L(h;k). We defer details to Appendix B.4 in the Online Supplemental Material, where we show that under Assumption 3.1 and a version of Assumption 3.2, the distribution of ˆ T(h;k) −T(f)can in large samples be approximated by a normal distribution with variance avarˆ T(h;k)=S(k)2 nh = n i=1 ς2xi;T(f) L2(f)2˜ wn(xi;hk)21+o(1) and mean bounded by abiasˆ T(h;k)=M1h2B(k) =−M1+T(f)M2 2L2(f) n i=1 ˜ wn(xi;hk)x2 i1+o(1) where ˜ wn(xi;hk) =wn +(xi)+wn −(xi),ς2(xi;T)=(1−T)Ω(xi)(1−T), B(k) =−∞ 0 u2k∗ 1(u)du1+T(f)M2/M1 L2(f) S(k)2=∞ 0 k∗ 1(u)2du d ς2 +0;T(f)+ς2 −0;T(f) L2(f)2 7WhileweallowtheboundsM1and M2to change with sample size, we assume that their ratio M1/M2is fixed for simplicity.
22 Armstrong and Kolesár Quantitative Economics 11 (2020) ς2 +(0;T)=limx↓0ς2(x;T),andς2 −(0;T)=limx↑0ς2(x;T). It then follows that for the FLCI, OCI, and a truncated version of the RMSE criterion, equation (4)holdswithM=M1,γb=2,γs=−1/2,andB(k) and S(k) given in the preceding display. The RMSE-optimal bandwidth is therefore given by h∗ rmse =⎛ ⎜ ⎜ ⎜ ⎝∞ 0 k∗ 1(u)2du ∞ 0 u2k∗ 1(u)du2·ς2T(f) 4dnM1+T(f)M2⎞ ⎟ ⎟ ⎟ ⎠ 1/5 (20) Since S(k) and B(k) depend on the kernel kthrough the same quantities as for inference at a boundary point, the kernel efficiency results are analogous to those in Section 3.2.1. Because the optimal bandwidth depends on T(f), implementing a feasible version of it requires replacing it with an initial estimate. An alternative approach to the construction of two-sided CIs for T(f) that does not require localization or the use of initial estimates is an Anderson and Rubin (1949) style construction studied by Noack and Rothe (2019). In particular, Noack and Rothe (2019) proposed constructing, for each T0, an auxiliary CI for the jump in the mean of yi−diT0at the cutoff, using an approach similar to that we use for inference in sharp RD. The CI for T(f)is then constructed by collecting all T0’s for which the auxiliary CI contains zero. This approach also has the additional advantage that it can allow for weak identification while it yields asymptotically equivalent CIs under strong identification.8See Noack and Rothe (2019) for a more detailed discussion. 3.3 Practical implementation We now discuss some practical issues that arise when implementing our CIs for inference at a point, and in sharp and fuzzy RD studied in the previous subsections. To focus the discussion, we consider smoothness classes FHöl2(M),FSRD(M),andFFRD(M1M2) that constrain the second derivative globally, so that, in the discussion below, p=2.In other words, for inference at a point, we assume that the conditional mean fequation (10) is (almost everywhere) twice differentiable with the second derivative bounded by M; for sharp RD, we assume that that fis twice differentiable on either side of the cutoff, with the second derivative bounded by M; and for fuzzy RD, we assume that f1and f2 in in equation (13) are twice differentiable on either side of the cutoff, with the second derivative bounded by M1and M2, respectively. These assumptions imply optimality of the estimators defined in Section 3.1 based on local linear regression (q=1), which is the most popular method in practice; they also imply that both the Epanechnikov and the triangular kernel are nearly optimal. 8Because we require that the sequence of parameter spaces FFRDn(M1M2)be localized around a value of L∗with L∗ 2= 0, we rule out sequences in which the jump in the first-stage regression is arbitrarily close to zero (the term “weak identification” refers to such sequences). As a result, the CI we propose, unlike the CI proposed by Noack and Rothe (2019), is not honest over the original parameter space FFRD(M1M2).
Quantitative Economics 11 (2020) Simple and honest confidence intervals 23 3.3.1 Choice of MAppropriate choice of the smoothness constant is key to implementing our method. Since the smoothness classes we consider are convex, the results of Low (1997), Cai and Low (2004)andArmstrong and Kolesár (2018a) imply that, to maintain honesty over the whole function class, a researcher must choose Ma priori, rather than attempting to use a data-driven method.9We therefore recommend that, whenever possible, problem-specific knowledge be used to decide what choice of Mis reasonable a priori, and that one consider a range of plausible values by way of sensitivity analysis.10 If one imposes additional restrictions on fthat make the parameter space for fnonconvex, a data-driven method for choosing Mmay be feasible.11 In Apprendix E in the Online Supplemental Material, we consider a restriction which relates Mto a global polynomial approximation to the regression function. In particular, the restriction formalizes the notion that the second derivative in a neighborhood of zero is bounded by the maximum second derivative of a ˜ pth order global polynomial approximation. Heuristically, such restriction will hold if the local smoothness of fis no smaller than its smoothness at large scales. This restriction allows us to calibrate Mbased on the following rule of thumb. For inference at a point, let ˘ f(x)be an estimate of fbased on a global polynomial regression of order ˜ p,andlet[xminxmax]denote the support of xi.Put ˆ Mrot =supx∈[xminxmax]|˘ f(p)(x)|. This rule of thumb is similar to the suggestion of Fan and Gijbels (1996, Chapter 4.2), with the important distinction that their rule of thumb was designed to estimate the pointwise-in-foptimal bandwidth. We discuss the difference between this bandwidth and h∗ rmse in Section 4. In sharp RD, the rule of thumb is analogous, except we define ˘ f(p)(x) to be the global polynomial estimate of order ˜ pin which the intercept and all coefficients are allowed to be different on either side of the discontinuity (i.e., as regressors, we use 1xix˜ p i, and their interactions with the indicator I{xi≥0}). For fuzzy RD, we use an analogous approach to separately calibrate the reduced-form and firststage smoothness parameters M1and M2based on the reduced-form and first-stage regressions. As a default choice, we set ˜ p=p+2=4. In Appendix E in the Online Supplemental Material, we give a formal analysis of this rule, showing that the resulting CIs are honest and nearly optimal (over a regularity class that imposes the additional restriction f discussed above). In contrast, we expect that calibrating Mbased on local smoothness estimates may be difficult to justify, since estimating a local derivative of fis a harder problem than the initial problem of estimating its value at a point. We investigate the 9These negative results contrast with more positive results for estimation. See, for example, Lepski (1990) who, in the context of estimating the value of the regression function at a point, proposes a data-driven method that automates the choice of both pand M. 10As is well known, if the final bandwidth choice is influenced by such sensitivity analysis, the resulting CI may undercover, even if the estimator is unbiased. In this case, one can combine our method with the bandwidth snooping adjustment of Armstrong and Kolesár (2018b). 11An alternative to restricting the parameter space is to change the notion of coverage. For example, in the context of constructing confidence bands for a regression function f(x),Hall and Horowitz (2013) proposed bands that have an average coverage property in that the bands achieve coverage of f(x) for a random subset of values of x. This subset may vary with the unknown regression function and the realized sample.
24 Armstrong and Kolesár Quantitative Economics 11 (2020) finite-sample performance of FLCIs based on ˆ Mrot in a Monte Carlo exercise in Section 5. 3.3.2 Computation of RMSE-optimal bandwidth Given a choice of M, one can compute a feasible version ˆ h∗ rmse of the RMSE-optimal bandwidth by plugging this choice into the expressions (16), (19), and (20), along with consistent estimates of d,andofthe variance at 0(for fuzzy RD, one also needs a preliminary estimate of T(f)). In the simulation exercise and empirical application below, we use an alternative approach based on directly minimizing the finite-sample RMSE over the bandwidth h.Todescribeit,let ˜ wn(xi;hk) denote the weights wn 1(xi;hk) givenin(11) if the parameter of interest is the conditional mean at a point, and let ˜ wn(xi;hk) =wn +(xi)+wn −(xi)if the parameter of interest is the sharp or fuzzy RD parameter. For inference at a point, or for sharp RD, the finite-sample RMSE takes the form RMSE(h;M)2=M2 4n i=1 ˜ wn(xi;hk)x2 i2 + n i=1 ˜ wn(xi;hk)σ2(xi) (21) Since σ2(xi)is typically unknown, one needs to replace it by an estimate. For inference at a point, the simplest choice is to use some estimate ˆσ2(xi)=ˆσ2that assumes homoskedasticity of the variance function. For sharp RD, one can use the estimate ˆσ2(xi)=ˆσ2 +(0)I{x≥0}+ ˆσ2 −(0)I{x<0},where ˆσ2 +(0)and ˆσ2 −(0)aresomepreliminary variance estimates based on observations above and below the cutoff. We use the bandwidth ˆ h∗ rmse˜ Mthat minimizes equation (21)forM=˜ M, the chosen smoothness constant. This method was considered previously in Armstrong and Kolesár (2018a). Since the estimate in fuzzy RD is nonlinear, its moments, and hence the finitesample RMSE do not exist. However, one can still employ an analogous approach minimizing the finite-sample analog of the asymptotic RMSE. As the asymptotic bias and the asymptotic standard deviation both scale with the jump in the first-stage regression at the cutoff, L2(f ), this scaling does not affect the optimum, we can equivalently minimize the asymptotic RMSE times L2(f ), ARMSE(h;M1M2)2=M1+T(f)M22 4n i=1 ˜ wn(xi;hk)x2 i2 + n i=1 wn q(xi;h;k)2ς2xi;T(f) with ς2(x;T)=(1−T)Ω(x)(1−T). Since Ω(xi)is unknown, one can again replace it with ˆ Ω2(xi)=ˆ Ω2 +(0)I{x≥0}+ ˆ Ω2 −(0)I{x<0},where ˆ Ω2 +(0)and ˆ Ω2 −(0)are some preliminary variance estimates for observations above and below the cutoff. As a preliminary estimate of T(f), one can take the estimate ˆ T(ˆ h0;k),where ˆ h0minimizes the above expression at T(f)=0.Onecanalsouse ˆ h0directly as a simple bandwidth selector, which, while not RMSE optimal, has the advantage that it does not depend on the choice of M2.
Quantitative Economics 11 (2020) Simple and honest confidence intervals 31 Design 3 Design 1 Design 2 −10 −05 00 05 10 −10−05000510 x f(x) Figure 4. Monte Carlo simulation designs 1–3 and M=2. For each design, we implement the optimal FLCI centered at a local linear estimate with a triangular kernel and MSE optimal bandwidth, as described in Section 3.3,for each choice of M∈{26},andwithMcalibrated using the rule-of-thumb (ROT) described in Section 3.3. The implementations with M∈{26}allow us to gauge the effect of using an appropriately calibrated M, compared to a choice of Mthat is either too conservative or too liberal by a factor of 3. The ROT calibration chooses Mautomatically, but requires additional conditions in order to have correct coverage (see Section 3.3). In addition to these FLCIs, we consider seven other CIs (Appendix F in the Online Supplemental Material considers one more method). The first five are different implementations of the robust bias-corrected (RBC) CIs proposed by CCT (discussed in Section 4). Implementing these CIs requires two bandwidth choices: a bandwidth for the local linear estimator, and a pilot bandwidth that is used to construct an estimate of its bias. The first two CIs use bandwidth choices justified by pointwise-in-fasymptotics. The first CI uses a plug-in estimate of h∗ pt defined in (24), as implemented by Calonico, Cattaneo, and Farrell (2018), and an analogous estimate for the pilot bandwidth. The second CI, also implemented by Calonico, Cattaneo, and Farrell (2018), uses bandwidth estimates for both bandwidths that optimize the pointwise asymptotic coverage error (CE) among CIs that use usual z1−α/2critical value. This CI can be considered a particular form of undersmoothing. The third CI sets both the pilot bandwidth and the main three times continuously differentiable, depending on the neighborhood definition, this assumption is arguably violated. The results in the Appendix are nearly identical to those reported here, implying that the performance of the RBC method is not driven by this lack of smoothness.
32 Armstrong and Kolesár Quantitative Economics 11 (2020) bandwidth to the plug-in estimate of h∗ pt. For the next three CIs, we consider bandwidths justified by uniform-in-fasymptotics. For the fourth and fifth CIs, we set both the main and the pilot bandwidth to h∗ rmse with M=2,andM=6, respectively. For the sixth CI, we set both bandwidths to ˆ h∗ rmseˆ Mrot . Finally, we consider a conventional CI centered at a plug-in bandwidth estimate of h∗ pt, using the rule-of-thumb estimator of Fan and Gijbels (1996, Chapter 4.2). All CIs are computed at the nominal 95% coverage level. Table 6reports the results. The FLCIs perform well when the correct Mis used. As expected, they suffer from undercoverage if Mis chosen too small, or suboptimal length when Mis chosen too large. The ROT choice of Mappears to do a reasonable job of having good coverage and length in these designs without requiring knowledge of the true smoothness constant. However, as discussed in Section 3.3, this ROT choice imposes additional restrictions on the parameter space, so one must take care in extrapolating these results to other designs. As predicted by the theory in Section 4, the RBC CIs also have good coverage when implemented using the h∗ rmse bandwidth, and they are less sensitive to the choice of M than the corresponding FLCIs, at the expense of being on average about 25% longer. RBC CIs with bandwidth given by ˆ h∗ rmseˆ Mrot also achieve good coverage, but they are again about 25% longer than the corresponding FLCIs. The CIs based on bandwidths justified by pointwise-in-fasymptotics (rows 1,2,3, and 7for each design in the table) all have very poor coverage for at least one of the designs. Our analysis in Section 4suggests that this is due to the tuning parameter choices required by these bandwidths. Indeed, looking at the average of the bandwidth over the Monte Carlo draws (also reported in Table 6), it can be seen that the bandwidths tend to be much larger than those that estimate h∗ rmse. This is even the case for the CE bandwidth, which is intended to minimize coverage errors. Overall, the Monte Carlo analysis suggests that default approaches to nonparametric CI construction (bias-correction or undersmoothing relative to plug-in bandwidths) can lead to severe undercoverage when implemented using bandwidths justified by pointwise-in-fasymptotics. Bias-corrected CIs such as the one proposed by CCT can have good coverage if one starts from the minimax RMSE bandwidth, although they will be wider than FLCIs proposed in this paper. 6. Empirical illustration To illustrate the implementation of feasible versions of the CIs (22), we use a subset of the dataset from Ludwig and Miller (2007). In 1965, when the Head Start federal program launched, the Office of Economic Opportunity provided technical assistance to the 300 poorest counties in the United States to develop Head Start funding proposals. Ludwig and Miller (2007) used this cutoff in technical assistance to look at intent-to-treat effects of the Head Start program on a variety of outcomes using as a running variable the county’s poverty rate relative to the poverty rate of the 300th poorest county (which had poverty rate equal to approximately 592%). We focus here on their main finding, the effect on child mortality due to causes addressed as part of Head Start’s health services. See Ludwig and Miller (2007)fora
Quantitative Economics 11 (2020) Simple and honest confidence intervals 33 Table 6. Monte Carlo simulation: inference at a point. M=2M=6 Method Bandwidth Bias SE E[h]Cov RL Bias SE Em[h]Cov RL Design 1 RBC h=ˆ h∗ pt,b=ˆ b∗ pt 0063 0035 075 556073 0157 0036 062 01061 RBC h=ˆ hce,b=ˆ bce 0030 0041 045 858085 0059 0045 034 724076 RBC h=b=ˆ h∗ pt 0025 0042 075 931088 0042 0047 062 891078 RBC h=b=ˆ h∗ rmse20001 0061 036 945127 0002 0061 036 945101 RBC h=b=ˆ h∗ rmse60000 0076 023 942158 0000 0075 023 942126 RBC h=b=ˆ h∗ rmseˆ Mrot 0000 0078 022 939164 0000 0097 014 934163 Conventional ˆ h∗ pt,rot 0032 0036 056 766076 0049 0046 031 774077 FLCI, M=2ˆ h∗ rmse20021 0043 036 949100 0065 0043 036 752080 FLCI, M=6ˆ h∗ rmse60009 0054 023 966125 0028 0053 023 947100 FLCI, M=ˆ Mrot ˆ h∗ rmseˆ Mrot 0008 0056 022 956129 0010 0069 014 963130 Design 2 RBC h=ˆ h∗ pt,b=ˆ b∗ pt 0043 0035 077 759072 0129 0035 077 46058 RBC h=ˆ hce,b=ˆ bce 0028 0040 049 874083 0074 0041 044 541069 RBC h=b=ˆ h∗ pt 0026 0041 077 909087 0077 0042 077 530070 RBC h=b=ˆ h∗ rmse20002 0061 036 945127 0006 0061 036 944101 RBC h=b=ˆ h∗ rmse60000 0076 023 942158 0000 0075 023 942126 RBC h=b=ˆ h∗ rmseˆ Mrot 0001 0068 030 940143 0000 0083 020 938138 Conventional ˆ h∗ pt,rot 0032 0032 078 744067 0073 0040 044 530066 FLCI, M=2ˆ h∗ rmse20020 0043 036 951100 0061 0043 036 781080 FLCI, M=6ˆ h∗ rmse60009 0054 023 966125 0028 0053 023 947100 FLCI, M=ˆ Mrot ˆ h∗ rmseˆ Mrot 0013 0048 030 943113 0020 0059 020 943110 Design 3 RBC h=ˆ h∗ pt,b=ˆ b∗ pt −0043 0035 077 757072 −0123 0035 074 99059 RBC h=ˆ hce,b=ˆ bce −0026 0040 049 881083 −0063 0043 043 642071 RBC h=b=ˆ h∗ pt −0024 0042 077 908087 −0066 0043 074 603071 RBC h=b=ˆ h∗ rmse2−0002 0061 036 945127 −0007 0061 036 944101 RBC h=b=ˆ h∗ rmse60000 0076 023 942158 0000 0075 023 942126 RBC h=b=ˆ h∗ rmseˆ Mrot 0000 0074 025 942154 0000 0092 016 936154 Conventional ˆ h∗ pt,rot −0032 0033 072 747069 −0065 0042 039 620070 FLCI, M=2ˆ h∗ rmse2−0020 0043 036 950100 −0060 0043 036 781080 FLCI, M=6ˆ h∗ rmse6−0009 0054 023 965125 −0027 0053 023 947100 FLCI, M=ˆ Mrot ˆ h∗ rmseˆ Mrot −0010 0052 025 956122 −0013 0065 016 961122 Note:Legend: SE—average standard error; E[h]—average (over Monte Carlo draws) bandwidth; Cov—coverage of CIs (in %); RL—relative (to optimal FLCI) length. Bandwidth descriptions:ˆ h∗ pt—plugin estimate of pointwise MSE optimal bandwidth (bw); ˆ b∗ pt—analog for estimate of the bias; ˆ hce—plugin estimate of coverage error optimal bw; ˆ bce—analog for estimate of the bias; The implementation of Calonico, Cattaneo, and Farrell (2018) is used for all four bws. ˆ h∗ rmse2,ˆ h∗ rmse6—RMSE optimal bw, assuming M=2,andM=6, respectively. ˆ h∗ pt,rot —Fan and Gijbels (1996)ruleofthumb; ˆ h∗ rmseˆ Mrot —RMSE optimal bw, using rule-of-thumb for M.50,000 Monte Carlo draws.
34 Armstrong and Kolesár Quantitative Economics 11 (2020) 2 4 6 −40 −20 0 20 Poverty rate minus 59.1984 Mortality rate Figure 5. Average county mortality rate per 100,000 for children aged 5–9over 1973–83 due to causes addressed as part of Head Start’s health services (labeled “Mortality rate”) plotted against poverty rate in 1960 relative to the 300th poorest county. Each point corresponds to an average for 25 counties. Data are from Ludwig and Miller (2007). detailed description of this variable. Relative to the dataset used in Ludwig and Miller (2007), we remove one duplicate entry and one outlier, which after discarding counties with partially missing data leaves us with 3103 observations, with 294 of them above the poverty cutoff. Figure 5plots the data (to reduce the noise in the outcome variable, we plot bin averages of size 25). To estimate the discontinuity in mortality rates, Ludwig and Miller (2007) used a uniform kernel16 and consider bandwidths equal to 9,18,and36.This yields point estimates equal to −190,−120,and−111, respectively, which are large effects given that the average mortality rate for counties not receiving technical assistance was 215 per 100,000.Thep-values reported in the paper, based on bootstrapping the t-statistic (which ignores any potential bias in the estimates), are 0036,0081,and0027. The standard errors for these estimates, obtained using the nearest neighbor method (with J=3)are104,070,and052. These bandwidth choices are optimal in the sense that they minimize the RMSE expression (21)ifM=0040,00074,and00014, respectively. Thus, for these bandwidths to be optimal, one has to be very optimistic about the smoothness of the regression function. In comparison, the rule of thumb method for estimating Mdiscussed in Section 3.3 yields ˆ Mrot =0299, implying h∗ rmse estimate 40, and the point estimate −317.Forthese smoothness parameters, the critical values based on the finite-sample bias-sd ratio are 16Ludwig and Miller (2007) stated that the estimates were obtained using a triangular kernel. However, due to a bug in the code, the results reported in the paper were actually obtained using a uniform kernel.
Quantitative Economics 11 (2020) Simple and honest confidence intervals 35 given by 2165,2187,2107,and2202, respectively, which is very close to the asymptotic value cv095(1/2)=2181. The resulting 95% confidence intervals are given by (−41430353) (−27200323) (−2215−0013) and (−63520010) respectively. The p-values based on these estimates are given by 0100,0125,0047,and 0051.Thesep-values are larger than those reported in the paper, as they take into account the potential bias of the estimates. Using a triangular kernel helps to tighten the confidence intervals by a few percentage points in length, as predicted by the relative asymptotic efficiency results from Table 3, yielding (−41380187) (−29270052) (−2268−0095) and (−5980−0322) The underlying optimal bandwidths are given by 116,231,458,and49, respectively. The p-values associated with these estimates are 0074,0059,0033,and0028,tightening the p-values based on the uniform kernel. These results indicate that unless one is very optimistic about the smoothness of the regression function, the uncertainty associated with the magnitude of the effect of Head Start assistance on child mortality is much higher than originally reported. This is due mainly to the relatively large bandwidths used by Ludwig and Miller (2007), which imply an optimistic bound on the smoothness of the regression function if we assume that such bandwidths are close to optimal for MSE. Interestingly, while the more conservative smoothness bound in our benchmark specification leads to much wider CIs, the point estimate is larger in magnitude, so that one still finds a statistically significant effect at a 5or 10% level, depending on the kernel. Appendix A: Proofs of theorems in Section 2 A.1 Proof of Theorem 2.1 Parts (ii) and (iii) follow from part (i) and simple calculations. To prove part (i), note that, if it did not hold, there would be a bandwidth sequence hnsuch that liminf n→∞ Mr−1nr/2Rˆ T(hn;k)<S(k) rB(k)1−rinf ttr−1˜ R(t1) By equation (7), the bandwidth sequence hnmust satisfy liminfn→∞ hn(nM2)1/[2(γb−γs)]> 0and limsupn→∞ hn(nM2)1/[2(γb−γs)]<∞.Thus,byequation(6), Mr−1nr/2Rˆ T(hn;k)=S(k)rB(k)1−rtr−1 n˜ R(tn1)+o(1) where tn=hγb−γs nB(k)/(n−1/2S(k)). This contradicts the display above. A.2 Proof of Theorem 2.2 The second statement (relative efficiency) is immediate from (6). For the first statement (coverage), fix ε>0and let sdn=n−1/2(h∗ rmse)γsS(k) so that sdn/ se(h∗ rmse;k) p →1uni-
36 Armstrong and Kolesár Quantitative Economics 11 (2020) formly over f∈F. Note that, by Theorem 2.1 and the fact that t∗ RMSE =1/r −1, ˜ RFLCIα+εˆ Th∗ rmse;k=sdn·cv1−α−ε(1/r −1)1+o(1) and similarly for ˜ RFLCIα−ε(ˆ T(h∗ rmse;k)). Since cv1−α(1/r −1)is strictly decreasing in α, it follows that there exists η>0such that, with probability approaching 1 uniformly over f∈F, RFLCIα+εˆ Th∗ rmse;k< seˆ Th∗ rmse;k·cv1−α(1/r −1) <(1−η)RFLCIα−εˆ Th∗ rmse;k Thus, liminf ninf f∈FPT(f)∈ˆ Th∗ rmse;k± seˆ Th∗ rmse;k·cv1−α(1/r −1) ≥liminf ninf f∈FPT(f)∈ˆ Th∗ rmse;k±RFLCIα+εˆ Th∗ rmse;k≥1−α−ε and limsup n inf f∈FPT(f)∈ˆ Th∗ rmse;k± seˆ Th∗ rmse;k·cv1−α(1/r −1) ≤limsup n inf f∈FPT(f)∈ˆ Th∗ rmse;k±RFLCIα−εˆ Th∗ rmse;k(1−η)≤1−α+ε where the last inequality follows by definition of RFLCIα−ε(ˆ T(h∗ rmse;k)). Taking ε→0 gives the result. References Abadie, A. and G. W. Imbens (2006), “Large sample properties of matching estimators for average treatment effects.” Econometrica, 74 (1), 235–267. [25] Abadie, A., G. W. Imbens, and F. Zheng (2014), “Inference for misspecified models with fixed regressors.” Journal of the American Statistical Association, 109 (508), 1601–1614. [25] Anderson, T. W. and H. Rubin (1949), “Estimation of the parameters of a single equation in a complete system of stochastic equations.” The Annals of Mathematical Statistics,20 (1), 46–63. [22] Armstrong, T. B., and M. Kolesár (2020), “Supplement to ‘Simple and honest confidence intervals in nonparametric regression’.” Quantitative Economics Supplemental Material, 11, https://doi.org/10.3982/QE1199.[6] Armstrong, T. B. (2019), “Adaptation bounds for confidence bands under self-similarity.” Available at arXiv:1810.09762.[4] Armstrong, T. B. and M. Kolesár (2018a), “Optimal inference in a class of regression models.” Econometrica, 86 (2), 655–683. [3,5,11,23,24,25]
Quantitative Economics 11 (2020) Simple and honest confidence intervals 37 Armstrong, T. B. and M. Kolesár (2018b), “A simple adjustment for bandwidth snooping.” Review of Economic Studies, 85 (2), 732–765. [23] Brown, L. D. and M. G. Low (1996), “Asymptotic equivalence of nonparametric regression and white noise.” The Annals of Statistics, 24 (6), 2384–2398. [5] Brown, L. D., M. G. Low, and L. H. Zhao (1997), “Superefficiency in nonparametric function estimation.” The Annals of Statistics, 25 (6), 2607–2625. [4] Bull, A. D. (2012), “Honest adaptive confidence bands and self-similar functions.” Electronic Journal of Statistics, 6, 1490–1516. [4] Cai, T. T. and M. G. Low (2004), “An adaptation theory for nonparametric confidence intervals.” The Annals of Statistics, 32 (5), 1805–1840. [3,23] Calonico, S., M. D. Cattaneo, and M. H. Farrell (2018), “On the effect of bias estimation on coverage accuracy in nonparametric inference.” Journal of the American Statistical Association, 113 (522), 767–779. [2,31,33] Calonico, S., M. D. Cattaneo, and M. H. Farrell (2019), “Coverage error optimal confidence intervals for local polynomial regression.” Available at arXiv:1808.01398.[30] Calonico, S., M. D. Cattaneo, and R. Titiunik (2014), “Robust nonparametric confidence intervals for regression-discontinuity designs.” Econometrica, 82 (6), 2295–2326. [2,28] Cheng, M.-Y., J. Fan, and J. S. Marron (1997), “On automatic boundary corrections.” The Annals of Statistics, 25 (4), 1691–1708. [2,4,11,18] Chernozhukov, V., D. Chetverikov, and K. Kato (2014), “Anti-concentration and honest, adaptive confidence bands.” The Annals of Statistics, 42 (5), 1787–1818. [4] Donoho, D. L. (1994), “Statistical estimation and optimal recovery.” The Annals of Statistics, 22 (1), 238–270. [2,3,5,7,11] Donoho, D. L. and M. G. Low (1992), “Renormalization exponents and optimal pointwise rates of convergence.” The Annals of Statistics,20 (2), 944–970. [5,7,9] Fan, J. (1993), “Local linear regression smoothers and their minimax efficiencies.” The Annals of Statistics, 21 (1), 196–216. [2,4,11] Fan, J., T. Gasser, I. Gijbels, M. Brockmann, and J. Engel (1997), “Local polynomial regression: Optimal kernels and asymptotic minimax efficiency.” Annals of the Institute of Statistical Mathematics, 49 (1), 79–99. [4,11,18] Fan, J. and I. Gijbels (1996), Local Polynomial Modelling and Its Applications. Monographs on Statistics and Applied Probability. Chapman & Hall/CRC, New York, NY. [3,16,23,26,32,33] Fuller, A. T. (1961), “Relay control systems optimized for various performance criteria.” In Automatic and Remote Control: Proceedings of the First International Congress of the International Federation of Automatic Control, Vol. 1 (J. F. Coales, J. R. Ragazzini, and A. T. Fuller, eds.), 510–519, Butterworths, London. [18]
38 Armstrong and Kolesár Quantitative Economics 11 (2020) Gao, W. Y. (2018), “Minimax linear estimation at a boundary point.” Journal of Multivariate Analysis, 165, 262–269. [19] Giné, E. and R. Nickl (2010), “Confidence bands in density estimation.” The Annals of Statistics, 38 (2), 1122–1170. [3] Hahn, J., P. E. Todd, and W. van der Klaauw (2001), “Identification and estimation of treatment effects with a regression-discontinuity design.” Econometrica, 69 (1), 201–209. [14,15] Hall, P. (1992), “Effect of bias estimation on coverage accuracy of bootstrap confidence intervals for a probability density.” The Annals of Statistics, 20 (2), 675–694. [2,28] Hall, P. and J. Horowitz (2013), “A simple bootstrap method for constructing nonparametric confidence bands for functions.” The Annals of Statistics, 41 (4), 1892–1921. [23] Ibragimov, I. A. and R. Z. Khas’minskii (1985), “On nonparametric estimation of the value of a linear functional in Gaussian white noise.” Theory of Probability & Its Applications, 29 (1), 18–32. [5] Imbens, G. and S. Wager (2019), “Optimized regression discontinuity designs.” Review of Economics and Statistics, 101 (2), 264–278. [4,5,25] Kolesár, M. and C. Rothe (2018), “Inference in regression discontinuity designs with a discrete running variable.” American Economic Review, 108 (8), 2277–2304. [4,5,25] Legostaeva, I. L. and A. N. Shiryaev (1971), “Minimax weights in a trend detection problemofarandomprocess.”Theory of Probability & Its Applications, 16 (2), 344–349. [16] Lepski, O. V. (1990), “On a problem of adaptive estimation in Gaussian white noise.” Theory of Probability & Its Applications, 35 (3), 454–466. [23] Li, K.-C. (1989), “Honest confidence regions for nonparametric regression.” The Annals of Statistics, 17 (3), 1001–1008. [2] Low, M. G. (1997), “On nonparametric confidence intervals.” The Annals of Statistics,25 (6), 2547–2554. [3,23] Ludwig, J. and D. L. Miller (2007), “Does head start improve children’s life chances? Evidence from a regression discontinuity design.” Quarterly Journal of Economics,122 (1), 159–208. [5,32,34,35] Noack, C. and C. Rothe (2019), “Bias-aware inference in fuzzy regression discontinuity designs.” Unpublished manuscript, University of Mannheim. [22] Nussbaum, M. (1996), “Asymptotic equivalence of density estimation and Gaussian white noise.” The Annals of Statistics, 24 (6), 2399–2430. [5] Sacks, J. and D. Ylvisaker (1978), “Linear estimation for approximately linear models.” The Annals of Statistics, 6 (5), 1122–1137. [18] Schennach, S. M. (2015), “A bias bound approach to nonparametric inference.” Working Paper CWP71/15, Cemmap. [4]
Quantitative Economics 11 (2020) Simple and honest confidence intervals 39 Sun, Y. (2005), “Adaptive estimation of the regression discontinuity model.” Unpublished manuscript, Univesity of California, San Diego. [18] Tsybakov, A. B. (2009), Introduction to Nonparametric Estimation. Springer, New York, NY. [4] van der Vaart, A. W. (1998), Asymptotic Statistics. Cambridge University Press, New York, NY. [7] Zhao, L. H. (1997), “Minimax linear estimation in a white noise problem.” The Annals of Statistics, 25 (2), 745–755. [18] Co-editor Andres Santos handled this manuscript. Manuscript received 29 August, 2018; final version accepted 28 August, 2019; available online 17 September, 2019.