scieee AI-readable full text Open interactive document viewer

Tests for the equality of conditional variance functions in nonparametric regression

Pardo Fernández, Juan Carlos; Jiménez Gamero, María Dolores; El Ghouch, Anouar

Abstract

In this paper we are interested in checking whether the conditional variances are equal in k ≥ 2 location-scale regression models. Our procedure is fully nonparametric and is based on the comparison of the error distributions under the null hypothesis of equality of variances and without making use of this null hypothesis. We propose four test statistics based on empirical distribution functions (Kolmogorov-Smirnov and Cramér-von Mises type test statistics) and two test statistics based on empirical characteristic functions. The limiting distributions of these six test statistics are established under the null hypothesis and under local alternatives. We show how to approximate the critical values using either an estimated version of the asymptotic null distribution or a bootstrap procedure. Simulation studies are conducted to assess the finite sample performance of the proposed tests. We also apply our tests to data on household expenditures.

Full text

Electronic Journal of Statistics Vol. 9 (2015) 1826–1851 ISSN: 1935-7524 DOI: 10.1214/15-EJS1058 Tests for the equality of conditional variance functions in nonparametric regression Juan Carlos Pardo-Fern´andez∗ Department of Statistics and Operational Research, Universidade de Vigo Campus Universitario, 36310 Vigo, Spain e-mail: [email protected] Mar´ıa Dolores Jim´enez-Gamero∗ Department of Statistics and Operational Research, Universidad de Sevilla Facultad de Matem´aticas, Calle Tarfia s.n., 41012 Sevilla, Spain e-mail: [email protected] and Anouar El Ghouch† Institute of Statistics, Biostatistics and Actuarial Sciences, Universit´e catholique de Louvain, Voie du Roman Pays 20, B-1348 Louvain-la-Neuve, Belgium e-mail: [email protected] Abstract: In this paper we are interested in checking whether the conditional variances are equal in k≥2 location-scale regression models. Our procedure is fully nonparametric and is based on the comparison of the error distributions under the null hypothesis of equality of variances and without making use of this null hypothesis. We propose four test statistics based on empirical distribution functions (Kolmogorov-Smirnov and Cram´er-von Mises type test statistics) and two test statistics based on empirical characteristic functions. The limiting distributions of these six test statistics are established under the null hypothesis and under local alternatives. We show how to approximate the critical values using either an estimated version of the asymptotic null distribution or a bootstrap procedure. Simulation studies are conducted to assess the finite sample performance of the proposed tests. We also apply our tests to data on household expenditures. MSC 2010 subject classifications: Primary 62G10, 62G08; secondary 62G20, 62G09. Keywords and phrases: Asymptotics, bootstrap, comparison of curves, empirical characteristic function, empirical distribution function, kernel smoothing, local alternatives, regression residuals. Received July 2014. ∗These authors were financially supported by grant MTM2014-55966-P (ERDF support included) of the Spanish Ministry of Economy and Competitiveness. †This author was financially supported by IAP research network P6/03 of the Belgian Government (Belgian Science Policy), and from the contract ‘Projet d’Actions de Recherche Concert´ees’ (ARC) 11/16-039 of the ‘Communaut´efran¸caise de Belgique’, granted by the ‘Acad´emie universitaire Louvain’. 1826 Equality of conditional variance functions 1827 1. Introduction When comparing k(k≥2) populations it is interesting not only comparing the means, but also other characteristics, like the variances. For example, in quality control, it is important to check the uniformity and the stability of the production process under different experimental and practical conditions. In biomedical research, detecting variation in gene expression levels is important for many reasons, for example, to identify experimental and environmental factors that affect a biological process; for a concrete example, see e.g. [19]. Equality of variances, when satisfied, can also be used to develop more powerful and simple ANOVA-type test statistics. Without controlling for the effect of covariates, there are a substantial number of tests available in the literature for the equality of (unconditional) variances of two or more populations. The standard procedures include the classical F-test and Levene’s test (see [17]) which is known to be robust to the violation of normality; see [11] for a recent review and some interesting examples and applications. In this paper, we are interested in the comparison of conditional variances. We assume that in each population, along with the variable of interest or response variable, Y, it is also observed another variable, X, the covariate, so that the mean and the variance of the response variable depend on the values of X.More specifically, let (Xj,Y j), 1 ≤j≤k,bekindependent random vectors satisfying general nonparametric regression models Yj=mj(Xj)+σj(Xj)εj,(1) where mj(x)=E(Yj|Xj=x) is the regression function, σ2 j(x)=Var(Yj| Xj=x) is the conditional variance function and εjis the regression error, which is assumed to be independent of Xj. Note that, by construction, E(εj)=0and Var(εj)=1. The covariate Xjis continuous with density function fj. Since the objective is to compare the variance functions, it is reasonable to assume that the covariates have common support, say R. The regression functions, the variance functions, the distribution of the errors and the distribution of the covariates are completely unknown and no parametric models are assumed for them. Thus, our approach is completely nonparametric. In this conditional setting, the hypothesis of equality of variances is stated in terms of the conditional variance functions, H0:σ2 1(x)=σ2 2(x)=···=σ2 k(x) for all x∈R, or, equivalently, H0:σj(x)/σ0(x)=1,for 1 ≤j≤k, where σ2 0(x) is the common variance that can be expressed as σ2 0(x)= k  j=1 πj(x)σ2 j(x), for some positive functions π1,...,π ksatisfying k j=1 πj(x) = 1. The alternative hypothesis is H1:σj(x)/σ0(x)=1,for some j∈{1...,k}. 1828 J. C. Pardo-Fern´andez et al. We will develop several test statistics and study their distribution under H0and under local alternatives converging to the null hypothesis at the rate n−1/2,n being the total sample size. Specifically, we consider the following local alternative hypothesis H1,n :σj(x)/σ0(x)=1+n−1/2δj(x),for 1 ≤j≤k, for some functions δj. To be more precise, in the previous expression we should have written σn,j(x)/σn,0(x) instead of σj(x)/σ0(x), as this function depends on n. However, to short the notation we suppress this explicit dependence on n. Observe that as nincreases H1,n becomes closer and closer to H0. Also, when δj(x)=0,for 1 ≤j≤k,H1,n reduces to H0. Statistical literature concerning the problem of testing for common features in several regression models has mainly focused on testing for common regression curves or testing for common error distributions. The problem of testing for the equality of regression curves in nonparametric settings has been extensively treated; see for example [3,16,22,25,24,27], and [12] for a recent review. On the other hand, testing for the equality of error distributions has been addressed in [23]. To the best of our knowledge, the comparison of conditional variance functions has not been studied before. Most papers dealing with testing on the conditional variance function focus on homoscedasticity assumption (see for example [18,4] and the references therein) or, more in general, on the parametric form of the conditional variance function (see for example [5,15] and the references therein). In order to construct a test for testing H0, several approaches are possible. Here we follow the ideas in [25,24] for testing the equality of the regression functions, m1,...,m k, which consist of comparing the distributions of the errors of the regression models. Specifically, let εj=Yj−mj(Xj) σj(Xj)(2) be the regression error in population j,1≤j≤k. Define ε0j=Yj−mj(Xj) σ0(Xj)=εj σj(Xj) σ0(Xj)(3) to be the error under the null hypothesis, 1 ≤j≤k.LetFεj(t)=P(εj≤t)and Fε0j(t)=P(ε0j≤t) be the cumulative distribution function (CDF) of εjand ε0j, respectively. The following theorem shows that H0is true if and only if the distributions of εjand ε0jcoincide. The proof can be found in the Appendix. Theorem 1. Assume that σjis continuous on Rand 0<E(ε4 j)<∞,for 1≤j≤k. (i) H0is true if and only if the random variables εjand ε0jhave the same distribution for all 1≤j≤k. (ii) Let p1,...,p kbe such that pj>0,1≤j≤k,andk j=1 pk=1.Let Fε(t)=k j=1 pjFεj(t)and Fε0(t)=k j=1 pjFε0j(t). Assume also that Equality of conditional variance functions 1829 E(ε4 1)=···=E(ε4 k).ThenH0is true if and only if Fε(t)=Fε0(t),for all t∈R. The assertions in the previous result can be interpreted in terms of the CDF or in terms of any other function characterizing a probability law, such as the characteristic function (CF). In this paper we will consider both cases, that is, to test H0we will compare consistent estimators of the CDFs and CFs of the random variables εjand ε0j,1≤j≤k. With this aim, the paper is organized as follows. In Section 2we introduce the test statistics and explain the testing procedures. Sections 3and 4contain the main asymptotic results concerning the empirical CDF-based test statistics and the empirical CF-based test statistics, respectively, and discuss some practical considerations. In Section 5we explain how the critical values of the proposed test statistics can be approximated. Investigating the finite sample performance of our tests is the topic of Section 6. A data example follows in Section 7and conclusions are given in Section 8. All proofs of the theoretical results are deferred to the Appendix. The following notation will be used along the paper: P0denotes probability assuming that H0is true; E0denotes expectation assuming that H0is true; P∗ denotes the conditional probability law, given the data; all limits in this paper are taken when n→∞;L −→ denotes convergence in distribution; P −→ denotes convergence in probability; if x∈Rk,withx=(x1,...,x k), then diag(x)is the k×kdiagonal matrix whose (i, i)entryisxi,1≤i≤k; for any complex number z=a+ib,Re(z)=ais its real part, Im(z)=bis its imaginary part, ¯z=a−ibis its conjugate and |z|is its modulus; Nk(μ, Σ) denotes the k-variate normal distribution with mean vector μand variance-covariance matrix Σ; an unspecified integral denotes integration over the whole real line R;sup tstands for supt∈R;I(S) denotes the indicator function of a set S. 2. The test statistics As in the Introduction, let (Xj,Y j), 1 ≤j≤k,bekindependent random vectors satisfying general nonparametric regression models (1). For 1 ≤j≤k,letεj and ε0jbe as defined in (2)and(3), respectively. As justified in Theorem 1, to test for H0we will compare consistent estimators of the CDFs and CFs of the random variables εjand ε0j,1≤j≤k, and also consistent estimators of the CDFs Fεand Fε0and of their associated CFs. Since neither εjnor ε0jare observable, the inference must be based on residuals. Next we construct them. Let (Xjl,Y jl), 1 ≤l≤nj, be independent and identically distributed (iid) observations from (Xj,Y j), 1 ≤j≤k,andletn=k j=1 nj. Along the paper it will be assumed that nj/n →pj>0, 1 ≤j≤k. In order to estimate the errors, we first need to estimate the regression functions, mj(x)=E(Yj|Xj=x), the variance functions, σ2 j(x)=E[{Yj−mj(x)}2|Xj=x], and the common variance function under H0,σ2 0(x). With this aim we use nonparametric estimators based on kernel smoothing techniques. Let Kdenote a nonnegative kernel function 1830 J. C. Pardo-Fern´andez et al. defined on R,let0<h n≡h→0 be the bandwidth or smoothing parameter and Kh(x)=h−1K(x/h). We use the following estimators for the functions mj, σ2 jand σ2 0: ˆmj(x)= nj  l=1 wjl(x)Yjl,ˆσ2 j(x)= nj  l=1 wjl(x)Y2 jl −ˆm2 j(x),ˆσ2 0(x)= k  j=1 πj(x)ˆσ2 j(x). The quantities wjl are either the local-linear weights given by wjl(x)=Kh(Xjl −x)S2,nj(x)−(Xjl −x)S1,nj S0,nj(x)S2,nj(x)−S2 1,nj(x), with Sk,nj(x)=nj l=1(Xjl−x)kKh(Xjl−x), k=0,1,2, or the Nadaraya-Watson weights wjl(x)= Kh(Xjl −x) nj v=1 Kh(Xjv −x). Both are particular cases of local-polynomial weighting (see [8]). Under the model assumptions that will be stated in the next section, the results in this article are valid for local-linear and for Nadaraya-Watson (local-constant) estimators. Note that we have implicitly assumed that the functions π1,...,π kdo not depend on unknowns. The theory also apply to the case where they depend on unknowns, replacing πjby ˆπjin the expression ˆσ2 0(x), whenever ˆπjconverges to πifast enough. Later we will discuss this issue in more detail. Based on these estimators, for each population j,1≤j≤k, we construct two samples of residuals, ˆεjl =Yjl −ˆmj(Xjl) ˆσj(Xjl)and ˆε0jl =Yjl −ˆmj(Xjl) ˆσ0(Xjl),(4) 1≤l≤nj. Then we can construct the corresponding empirical CDFs (ECDFs), ˆ Fεj(t)= 1 nj nj  l=1 I(ˆεjl ≤t)and ˆ Fε0j(t)= 1 nj nj  l=1 I(ˆε0jl ≤t), and empirical CFs (ECFs), ˆϕεj(t)= 1 nj nj  l=1 exp(itˆεjl)and ˆϕε0j(t)= 1 nj nj  l=1 exp(itˆε0jl), respectively. These ECDFs are consistent kernel-based nonparametric estimators of the population CDFs Fεj(t)andFε0j(t), respectively (see Theorem 2 below). Analogously, the above ECFs are consistent kernel-based nonparametric estimators of the population CFs ϕεj(t)=E{exp(itεj)}and ϕε0j(t)= E{exp(itε0j)}, respectively (see Theorem 6below). We can also consider the following ECDFs ˆ Fε(t)= 1 n k  j=1 nj  l=1 I(ˆεjl ≤t)and ˆ Fε0(t)= 1 n k  j=1 nj  l=1 I(ˆε0jl ≤t), Equality of conditional variance functions 1831 and ECFs ˆϕε(t)= 1 n k  j=1 nj  l=1 exp(itˆεjl)and ˆϕε0(t)= 1 n k  j=1 nj  l=1 exp(itˆε0jl), which estimate the functions Fε(t)=k j=1 pjFεj(t), Fε0(t)=k j=1 pjFε0j(t), ϕε(t)=k j=1 pjϕεj(t)andϕε0(t)=k j=1 pjϕε0j(t), respectively. To test for H0, we will construct Kolmogorov-Smirnov type statistics and Cram´er-von Mises type statistics to compare the ECDFs, and weighted L2distances to compare the ECFs. More precisely, the considered statistics are T1 KS = k  j=1 √njsup t|ˆ Fεj(t)−ˆ Fε0j(t)|, T1 CM = k  j=1 nj{ˆ Fεj(t)−ˆ Fε0j(t)}2dˆ Fε0j(t), T2 KS =√nsup t|ˆ Fε(t)−ˆ Fε0(t)|, T2 CM =n{ˆ Fε(t)−ˆ Fε0(t)}2dˆ Fε0(t), T1= k  j=1 njˆϕεj(t)−ˆϕε0j(t) 2w(t)dt, T2=n|ˆϕε(t)−ˆϕε0(t)|2w(t)dt, where wis a positive weight function that is needed to guarantee consistency (see Section 4). Note that in the case of T1and T2,|·|represents the modulus of a complex number. In Section 3we will study the asymptotic properties of the statistics T1 KS,T1 CM,T2 KS and T2 CM and in Section 4we will deal with T1 and T2. 3. Asymptotics for ECDF-based test statistics This section studies some asymptotic properties of the ECDF-based test statistics T1 KS,T1 CM,T2 KS and T2 CM. To derive them we will need some commonly assumed regularity assumptions. First let us define Fj(t|x)=P(Yj≤t|Xj=x) and Fj(x)=P(Xj≤x), for 1 ≤j≤k. Assumption (A1): For 1 ≤j≤k, (i) Xjis absolutely continuous with compact support Rand density fj. (ii) fj,mj,σjand πjare twice continuously differentiable on R. (iii) infx∈Rfj(x)≥c>0andinf x∈Rσj(x)≥d>0, for some c, d ∈R. (iv) E(ε4 j)<∞. (v) nh4 n→0andnh3+2δ n(log h−1 n)−1→∞, for some δ>0. 1832 J. C. Pardo-Fern´andez et al. (vi) The kernel Kis a symmetric density function with compact support and twice continuously differentiable. Assumption (A2): For 1 ≤j≤k,Fj(t|x) is continuous in (x, t) and differentiable with respect to t,∂ ∂tFj(t|x)=F j(t|x) is continuous in (x, t)and supx,t |t2F j(t|x)|<∞. The same holds for all other partial derivatives of Fj(t|x) with respect to xand tup to order two. From now on we will name Assumption A to be the set of Assumptions (A1)– (A2). Assumption A (skipping (A1)(iv)) was also considered in [25]toderive asymptotic properties of some ECDF-based tests designed to detect differences between the conditional mean functions. This assumption is mainly needed to guarantee the uniform consistency of the estimators ˆ fj,ˆσj,ˆmjand ˆσ0. Also note that Assumption (A2), which will only be needed for the asymptotics related to ECDF-based tests, implies that εjhas a density, denoted by fεj. We first give the following result that justifies the use of the test statistics T1 KS,T1 CM,T2 KS and T2 CM for testing H0. Theorem 2. Suppose that Assumption A holds. Then, ˆ Fε0j(t)=Fε0j(t)+op(1) and ˆ Fεj(t)=Fεj(t)+op(1), uniformly in t,1≤j≤k. Corollary 3. Suppose that Assumption A holds. Then, 1 √nT1 KS P −→ k  j=1 √pjsup t|Fεj(t)−Fε0j(t)|, 1 nT1 CM P −→ k  j=1 pj{Fεj(t)−Fε0j(t)}2dFε0j(t), 1 √nT2 KS P −→ sup t|Fε(t)−Fε0(t)|, 1 nT2 CM P −→ {Fε(t)−Fε0(t)}2dFε0(t). Observe that all considered test statistics converge in probability to nonnegative quantities. Under the assumptions in Theorem 1, such quantities are 0 if and only if H0is true. Therefore it seems reasonable to reject the null hypothesis for large values of these test statistics. Now, to determine what a large value means in each case, we must calculate the null distribution of the test statistic, or at least an approximation to it. Since the null distributions are unknown, we study their asymptotic null distributions. Theorem 4. Suppose that Assumption A holds. Then, under H1,n, √njˆ Fεj(t)−ˆ Fε0j(t)=1 2tfεj(t)(p1/2 jΔj+Zn,j)+op(1), uniformly in t,whereΔj=2E[δj(Xj)],and Zn,j =√nj k  v=1 1 nv nv  l=1 I(v=j)−πv(Xvl)fj(Xvl) fv(Xvl)ε2 vl −1.(5) Equality of conditional variance functions 1833 The next Corollary, derived mainly by applying the multivariate Central Limit Theorem to Zn=(Zn,1,...,Z n,k), gives the asymptotic distribution of our ECDF-based test statistics under H0and H1,n. Corollary 5. Suppose that Assumption A holds. Then, under H1,n, T1 KS L −→ 1 2 k  j=1 |Zj+p1/2 jΔj|sup t|tfεj(t)|, T1 CM L −→ 1 4 k  j=1 (Zj+p1/2 jΔj)2t2f2 εj(t)dFεj(t), T2 KS L −→ 1 2sup y|Z(t)+Δ(t)|, T2 CM L −→ 1 4{Z(t)+Δ(t)}2dFε(t), where Z(t)=k j=1 p1/2 jtfεj(t)Zjand Δ(t)=k j=1 pjtfεj(t)Δj,with(Z1,..., Zk)∼Nk(0,Σ),Σ=(σjv)being the k×k-matrix whose elements are σjv =(pjpv)1/2 k  l=1 E{(ε2 l−1)2} pl×(6) ×Eπl(Xl)fj(Xl) fl(Xl)−I(l=j)πl(Xl)fv(Xl) fl(Xl)−I(l=v) 1≤j, v ≤k. Let Tdenote any of the test statistics T1 KS,T1 CM,T2 KS and T2 CM.SinceH0 can be seen as a special case of H1,n with δj=0,1≤j≤k, the asymptotic distribution of Tunder the null hypothesis trivially follows by setting Δj=0. For example, under H0, T1 KS L −→ 1 2 k  j=1 |Zj|sup t|tfεj(t)|,and T2 KS L −→ 1 2sup t|Z(t)|. Let α∈(0,1) be arbitrary but fixed. As an immediate consequence of Theorem 1 and Corollaries 3and 5, the test that rejects H0when T≥tα,wheretαis the 1−αpercentile of the null distribution of Tor any consistent estimator of it, is consistent against all fixed alternatives. It is also able to detect local alternatives converging to the null at the rate n−1/2, whenever Δj=0forsome1≤j≤k. So far we have assumed that the weight functions π1,...,π kare known. Nevertheless we did not make any restriction on them except the fact that they are positive and sum up to one. In our simulation study, see Section 6, we take πj=nj/n. This simple choice is shown to work reasonably well for all the investigated examples. Another possibility is to choose the πj’s from the data. For example, as for the problem of testing the equality of regression curves, see 1834 J. C. Pardo-Fern´andez et al. [25,24], one may take πj(x)=pjfj(x)/fmix(x),with fmix(x)=k j=1 pjfj(x). For this choice, since f1,...f kare unknown, the functions π1,...,π kmust be estimated. A careful reading of the proofs reveals that all the results in this paper continue to be true whenever π1,...,π kare replaced by estimators ˆπ1,...,ˆπk satisfying supx∈R|πj(x)−ˆπj(x)|=op(n−1/4), 1 ≤j≤k. 4. Asymptotics for ECF-based test statistics In order to study the limit behaviour of the test statistics T1and T2we also need some regularity conditions. Recall that to derive the asymptotic properties for the ECDF-based test statistics we assumed that the regression errors have a twice differentiable CDF. Analogously, to derive the asymptotic properties for the ECF-based test statistics we need that the regression errors has a twice differentiable CF, which is tantamount to assume that the regression errors has finite second order moment. But this assumption is implicit in the the definition of the regression models (1). As a consequence, the assumptions required to derive the asymptotics for ECF-based test statistics will be weaker than those assumed in Section 3, in the sense that no restriction on the distribution of the errors will be imposed, such as the existence of a density. Specifically, we mainly need to assume that Assumption (A1) holds. The motivation behind the test statistics T1and T2is in the following result. Theorem 6. Suppose that Assumption (A1) holds and that w≥0is such that t2w(t)dt < ∞.Then,n−1Ti=τi+op(1),i=1,2,where τ1= k  j=1 pjϕεj(t)−ϕε0j(t) 2w(t)dt, τ2=|ϕε(t)−ϕε0(t)|2w(t)dt. Thus, T1and T2converge in probability to non-negative quantities. Since two distinct CFs can be equal in a finite interval (see, for example, [10], p. 479), a general way to ensure that τ1>0andτ2>0 whenever σr=σs, for some 1≤r, s ≤k,r=s, is to take w(t)>0, for all t∈R. For instance, one can take was the pdf of a normal law. Now, the reasoning made just after Corollary 3can be repeated for the test statistics T1and T2. So our next goal is to determine the asymptotic distribution of T1and T2. With this aim we first give a result that provides an asymptotic approximation for √nj{ˆϕεj(t)−ˆϕε0j(t)},1≤j≤k.Let ϕ εj(t)= ∂ ∂tϕεj(t)= ∂ ∂tReϕεj(t)+i∂ ∂tImϕεj(t)=iE[εjexp(itεj)], which exists because E(|εj|)<∞,1≤j≤k. Theorem 7. Suppose that Assumption (A1) holds. Then, under H1,n, √njˆϕεj(t)−ˆϕε0j(t)=1 2tϕ εj(t)(p1/2 jΔj−Zn,j)+tR1j(t)+t2R2j(t), where supt|Rsj(t)|=op(1),s=1,2,andZn,j and Δj,1≤j≤k, are defined as in Theorem 4. Equality of conditional variance functions 1841 Table 3 Observed rejection frequencies in 1000 simulated data sets for the tests based on the test statistics T1and T2. The critical values obtained by bootstrap model (n1,n 2)h: cv 0.10 0.20 0.30 cv 0.10 0.20 0.30 T1T2 (L1) (50,50) 0.046 0.046 0.052 0.045 0.040 0.051 0.045 0.047 (100,50) 0.035 0.032 0.035 0.031 0.041 0.044 0.041 0.041 (100,100) 0.040 0.038 0.040 0.041 0.048 0.053 0.046 0.052 (L2) (50,50) 0.070 0.067 0.076 0.069 0.055 0.061 0.056 0.058 (100,50) 0.043 0.045 0.043 0.041 0.040 0.041 0.034 0.037 (100,100) 0.058 0.056 0.052 0.060 0.063 0.063 0.057 0.059 (P1) (50,50) 0.562 0.566 0.553 0.556 0.361 0.358 0.371 0.367 (100,50) 0.693 0.686 0.679 0.679 0.466 0.480 0.488 0.490 (100,100) 0.895 0.896 0.887 0.894 0.543 0.541 0.546 0.547 (P2) (50,50) 0.928 0.926 0.926 0.917 0.726 0.716 0.738 0.738 (100,50) 0.968 0.971 0.961 0.956 0.850 0.857 0.869 0.859 (100,100) 0.997 0.997 0.995 0.993 0.937 0.936 0.935 0.938 (P3) (50,50) 0.427 0.431 0.424 0.440 0.298 0.294 0.300 0.301 (100,50) 0.581 0.580 0.565 0.566 0.367 0.356 0.376 0.381 (100,100) 0.743 0.737 0.744 0.742 0.432 0.421 0.412 0.400 (P4) (50,50) 0.328 0.326 0.331 0.319 0.194 0.192 0.206 0.203 (100,50) 0.381 0.380 0.383 0.383 0.238 0.240 0.249 0.248 (100,100) 0.573 0.574 0.581 0.572 0.306 0.308 0.309 0.313 Table 4 Observed rejection frequencies in 1000 simulated data sets for the tests based on the critical values obtained from the asymptotic null distribution of T1for several error distributions and cross-validation bandwidths model (n1,n 2) errors: N(0,1) t5/5/3t7/7/5Exp(1) −1 (L2) (100,100) 0.065 0.065 0.063 0.082 (200,100) 0.052 0.086 0.083 0.088 (200,200) 0.051 0.067 0.070 0.061 (400,200) 0.050 0.083 0.070 0.045 (400,400) 0.041 0.064 0.046 0.051 (L3) (100,100) 0.100 0.077 0.066 0.088 (200,100) 0.062 0.082 0.086 0.090 (200,200) 0.066 0.073 0.073 0.070 (400,200) 0.055 0.077 0.074 0.041 (400,400) 0.051 0.065 0.051 0.063 (P4) (100,100) 0.494 0.374 0.378 0.264 (200,100) 0.542 0.348 0.390 0.289 (200,200) 0.796 0.532 0.614 0.391 (400,200) 0.859 0.582 0.664 0.432 (400,400) 0.968 0.776 0.876 0.596 (P5) (100,100) 0.917 0.669 0.729 0.518 (200,100) 0.936 0.740 0.831 0.587 (200,200) 0.992 0.872 0.946 0.755 (400,200) 0.997 0.927 0.975 0.847 (400,400) 0.999 0.966 0.996 0.939 1842 J. C. Pardo-Fern´andez et al. Table 5 p-values for testing for the equality of the conditional variance functions for the data set concerning expenditures of Dutch households asymptotic bootstrap hT 1 CM T1T1 CM T2 CM T1 KS T2 KS T1T2 0.20 0.011 0.395 0.207 0.101 0.123 0.135 0.428 0.735 0.25 0.028 0.380 0.321 0.296 0.378 0.350 0.387 0.688 0.30 0.028 0.368 0.349 0.269 0.386 0.302 0.387 0.615 0.35 0.045 0.384 0.417 0.280 0.344 0.263 0.377 0.509 0.40 0.049 0.435 0.443 0.274 0.450 0.247 0.435 0.470 0.45 0.078 0.509 0.588 0.247 0.439 0.241 0.536 0.435 0.50 0.083 0.597 0.602 0.239 0.477 0.237 0.647 0.419 7. Application to data To illustrate our testing procedure we will use a data set concerning monthly expenditures of Dutch households. The variable ‘log of the total monthly expenditure’ is considered as a covariate and ‘log of the expenditure on food’ is considered as the response. See [7,25] for more details on these data. In the latter paper, the equality of the regression curves of households of 2, 3 and 4 members was tested and the equality between the regression curves of 3-member households (43 observations) and 4-member households (73 observations) was accepted. Here we move one step further in the comparison of the regression models and test for the equality of the conditional variance functions. Table 5 shows the p-values obtained from the asymptotic null distribution for T1and T1 CM and from the bootstrap for the six test statistics with fixed bandwidths ranging from 0.20 and 0.50 (the support of the covariates is approximately between 9.5 and 11.5). The results are quite homogeneous, as all test statistics, except the asymptotic version of T1 CM, lead to the acceptance of the equality of the conditional variance functions. As we have seen in the simulations presented in Section 6, the approximation of the asymptotic null distribution of the T1 CM is not satisfactory, specially for small sample sizes. Since here we are working with 43 and 73 observations, the results for this test statistic are not reliable, and we should only consider its bootstrap version. 8. Conclusions In this paper, we have constructed and studied six tests for the equality of k conditional variances. To do so, we compare the ECDF and ECF of the error terms estimated nonparametrically under H0and H1. Under some regularity conditions, the proposed tests are consistent against any fixed alternative and are able to detect contiguous alternatives converging to the null at a rate n−1/2. The assumptions needed to derive these properties are weaker for the ECFbased test statistics. Specifically, no requirement is imposed on the distributions of the errors. An approximation of the asymptotic null distribution has been proposed and the performance of each test has been evaluated by means of some simulations. The proposed approximation works, in the sense of providing Equality of conditional variance functions 1843 type I errors close to the nominal values, specially when the sample sizes are large (at least 200). For smaller sample sizes it is recommended to approximate the null distribution through a bootstrap mechanism. The estimation of the conditional variance functions has been also studied in the econometric literature when the data present dependence structure (see for example [9,28,20]). The proposed tests could be extended to this setting by assuming mixing conditions on the data and using appropriate results for the nonparametric estimators for the variance and regression functions in the same line of [6], who tested for a constant variation coefficient in regression models with stationary data and used the results about kernel estimators with dependent data given in [13]. Although the results in this paper are presented for local-constant and locallinear weights variance estimators, in practice, it is well-known that the locallinear variance estimator may take negative values. In our simulations and application to real data we used the local-constant estimator to estimate the conditional variance functions, as it guarantees the positiveness of the estimate. There are other possibilities to obtain positive estimators, such us, for instance, the local-exponential estimator studied by [28]. Under suitably adapted conditions, the procedures and the results in this paper can be extended for the local-exponential and other conditional variance estimators. The above extensions, as well as others motivated by recent applications (for instance, in casual inference, [14]) constitute fields of future research. Appendix We now sketch the proofs of the results stated in Sections 1–4. With this aim we first give some preliminary results, some of them are of independent interest. A.1. Preliminary results Under Assumption (A1), and consequently under Assumption A, we have that, for 1 ≤j≤k,sup x∈R|ˆmj(x)−mj(x)|=op(n−1/4 j),supx∈R|ˆσj(x)−σj(x)|= op(n−1/4 j),and supx∈R|ˆ fj(x)−fj(x)|=op(n−1/4 j).This, together with some routine calculations, show that sup x∈R ˆσ2 j(x)−σ2 j(x)−1 njfj(x) nj  s=1 Kh(Xjs −x){Yjs −mj(x)}2−σ2 j(x) =op(n−1/2).(9) Also, from the equality ˆσj(x)−σj(x)= ˆσ2 j(x)−σ2 j(x) 2σj(x)−{ˆσj(x)−σj(x)}2 2σj(x), 1844 J. C. Pardo-Fern´andez et al. it follows that, sup x∈R ˆσj(x)−σj(x)−ˆσ2 j(x)−σ2 j(x) 2σj(x) =op(n−1/2).(10) So, by the definition of ˆσ0(x)andσ0(x), we also have that supx∈R|ˆσ0(x)− σ0(x)|=op(n−1/4)and sup x∈R ˆσ0(x)−σ0(x)− k  j=1 πj(x)ˆσ2 j(x)−σ2 j(x) 2σ0(x) =op(n−1/2).(11) Lemma 9. Suppose that Assumption (A1) holds. Then, (i) For 1≤j≤k, ˆσj(x)−σj(x) σj(x)fj(x)dx =1 2nj nj  s=1 ε2 js −1+op(n−1/2). (ii) ˆσ0(x)−σ0(x) σ0(x)fj(x)dx = k  v=1 1 2nv nv  s=1 πv(Xvs)fj(Xvs) fv(Xvs) σ2 v(Xvs) σ2 0(Xvs)ε2 vs −1+op(n−1/2). Proof. From (9)and(10), we get ˆσj(x)−σj(x) σ0(x)fj(x)dx =1 2nj nj  s=1 Kh(Xjs −x){Yjs −mj(x)}2/σ2 j(x)−1dx +op(n−1/2). Part (i) follows from the above equality by making the change of variable Ujs =(Xjs −x)/h and applying a Taylor’s development. Part (ii) can be proved similarly by using (9)and(11). Lemma 10. Let ˜ϕεj(t)= 1 njnj l=1 exp(itεjl),ˆϕεj(t)= 1 njnj l=1 exp(itˆεjl),and similarly define ˜ϕε0j(t)and ˆϕε0j(t).Suppose that Assumption (A1) holds. Then, for 1≤j≤k, (i) ˆϕεj(t)= ˜ϕεj(t)+i t nj nj  l=1 exp(itεjl)mj(Xjl)−ˆmj(Xjl) σj(Xjl) Equality of conditional variance functions 1845 +i t nj nj  l=1 exp(itεjl)σj(Xjl)−ˆσj(Xjl) σj(Xjl)εjl +tRj,1(t)+t2Rj,2(t), with supt|Rj,s(t)|=op(n−1/2),s=1,2. (ii) ˆϕε0j(t)= ˜ϕε0j(t)+i t nj nj  l=1 exp(itε0jl)mj(Xjl)−ˆmj(Xjl) σ0(Xjl) +i t nj nj  l=1 exp(itε0jl)σ0(Xjl)−ˆσ0(Xjl) σ0(Xjl)ε0jl +tR0j,1(t)+t2R0j,2(t), with supt|R0j,s(t)|=op(n−1/2),s=1,2. Proof. Using a Taylor’s development, we get ˆϕεj(t)−˜ϕεj(t)=i t nj nj  l=1 (ˆεjl −εjl)exp(itεjl)+t2Rj(t)1 nj nj  l=1 (ˆεjl −εjl)2, with supt|Rj(t)|=Op(1).Part (i) follows from the equality ˆεj−εj=mj(Xj)−ˆmj(Xj) ˆσj(Xj)+σj(Xj)−ˆσj(Xj) ˆσj(Xj)εj =mj(Xj)−ˆmj(Xj) σj(Xj)+{mj(Xj)−ˆmj(Xj)}{σj(Xj)−ˆσj(Xj)} σj(Xj)ˆσj(Xj) +σj(Xj)−ˆσj(Xj) σj(Xj)εj+{σj(Xj)−ˆσj(Xj)}2 σ(Xj)ˆσj(Xj)εj. Similarly, one can prove (ii). Lemma 11. Let gbe a bounded function. Suppose Assumption (A1) holds. Then, it √nj nj  l=1 εjl exp(itεjl)g(Xjl)ˆσv(Xjl)−σv(Xjl) σv(Xjl) =t 2ϕ εj(t)√nj nv nv  s=1 (ε2 vs −1)g(Xvs)fj(Xvs) fv(Xvs)+tRj,v(t), with supt|Rj,v(t)|=op(1),1≤j, v ≤k. Proof. From (9)and(10), it √nj nj  l=1 exp(itεjl)εjlg(Xjl)ˆσv(Xjl)−σv(Xjl) σv(Xjl)=it 2√njHjv +tR1(t), 1846 J. C. Pardo-Fern´andez et al. where supt|R1(t)|=op(1) and Hjv(t)= 1 njnv nj  l=1 nv  s=1 Uv(Xjl,ε jl;Xvs,ε vs;t), with Uv(X1,ε 1;X2,ε 2;t) =ε1exp(itε1)g(X1) fv(X1)Kh(X1−X2){mv(X2)+ε2σv(X2)−mv(X1)}2 σ2 v(X1)−1. •If j=v, then, for every t,Hjv(t) is a two sample U-statistic of degree (1,1) with kernel Uv(Xjl,ε jl;Xvs,ε vs;t). Its H´ajek projection, H jv(t), is given by H jv(t)=−iϕ εj(t)1 nv nv  s=1 (ε2 vs −1)g(Xvs)fj(Xvs) fv(Xvs)+R jv(t) where supt|R jv(t)|=Op(h2). Moreover, Var{Hjv(t)−H jv(t)}≤ 1 njnv E{U2 h(Xj,ε j;Xv,ε v;t)}=O(n−1 jn−1 vh−1). Therefore, √njHjv(t)=−iϕ εj(t)√nj nv nv  s=1 (ε2 vs −1)g(Xvs)fj(Xvs) fv(Xvs)+Rjv(t),(12) with supt|Rjv(t)|=op(1). •If j=v,then Hjj(t)=K(0) n2 jh nj  l=1 εjl exp(itεjl)(ε2 jl −1) g(Xjl) fj(Xjl)+nj−1 2nj Hj(t), where, for every t,Hj(t) is a one sample U-statistic of degree 2 with kernel Uj(Xjl,ε jl;Xjs,ε js;t)+Uj(Xjs,ε js;Xjl,ε jl;t). Arguments very similar to those employed for the case j=vcan be used to show that √njHj(t)=−2iϕ εj(t)1 √nj nj  s=1 (ε2 js −1)g(Xjs)+Rj(t), with supt|Rj(t)|=op(1). Since √nj K(0) n2 jh nj  l=1 εjl exp(itεjl)(ε2 jl −1) g(Xjl) fj(Xjl)≤M √nh2 1 nj nj  l=1 |εjl|3, for some positive constant M, we conclude that Hjj(t) also satisfies (12)with j=v. This proves the result. Equality of conditional variance functions 1847 A.2. Proofs of main results ProofofTheorem1.(i)The direct implication is trivial. To prove the converse implication, assume that ε0jand εjhave the same distribution. They will also share the same moments. Now, because of the independence of εjand Xj,we get E(ε2 0j)=E(ε2 j)⇒E{σ2 j(Xj)/σ2 0(Xj)}=1,and E(ε4 0j)=E(ε4 j)⇒E{σ4 j(Xj)/σ4 0(Xj)}=1. Hence, E[(σ2 j(Xj) σ2 0(Xj)−1)2]=0,for1≤j≤k, and so we deduce that H0holds. (ii) Let ε(respectively, ε0) be a random variable with CDF Fε(respectively, Fε0). Note that ε(respectively, ε0) is the mixture of the random variables {εj,1≤j≤k}(respectively, {εj0,1≤j≤k}) with probabilities {pj,1≤ j≤k}.Asforpart(i), using the fact that E(ε2 j)=1andE(ε4 j)=E(ε4 1)>0, for 1 ≤j≤k, we obtain E(ε2 0)=E(ε2)⇒ k  j=1 pjE{σ2 j(Xj)/σ2 0(Xj)}=1,and E(ε4 0)=E(ε4)⇒ k  j=1 pjE{σ4 j(Xj)/σ4 0(Xj)}=1. Hence, k j=1 pjE[(σ2 j(Xj) σ2 0(Xj)−1)2]=0.Since pj>0, we conclude that H0is true. ProofofTheorem2.From the proof of Theorem 1 in [1], ˆ Fε0j(t)= 1 nj nj  l=1 I(ε0jl ≤t)+tfε0j(t)ˆσ0(x)−σ0(x) σ0(x)fj(x)dx +fε0j(t)ˆmj(x)−mj(x) σ0(x)fj(x)dx +op(n−1/2), (13) and ˆ Fεj(t)= 1 nj nj  l=1 I(εjl ≤t)+tfεj(t)ˆσj(x)−σj(x) σj(x)fj(x)dx +fεj(t)ˆmj(x)−mj(x) σj(x)fj(x)dx +op(n−1/2), (14) uniformly in t,wherefε0jdenotes the density corresponding to Fε0j. The desired results follow directly from (13)and(14). ProofofTheorem4.From the proof of Lemma 1 in [1], we have that 1 nj nj  l=1 I(εjl ≤t)= 1 nj nj  l=1 I(ε0jl ≤t)+Fεj(t)−Fε0j(t)+op(n−1/2),(15) 1848 J. C. Pardo-Fern´andez et al. uniformly in t. Observe that σ0(x) σj(x)=1−n−1/2σ0(x) σj(x)δj(x)=1−n−1/2δj(x)+n−1σ0(x) σj(x)δ2 j(x). Using this and a Taylor’s development leads to Fε0j(t)=EFεjtσ0(Xj) σj(Xj)=Fεj(t)−n−1/2tfεj(t)E[δj(Xj)] + o(n−1/2), (16) uniformly in t, sup t|fεj(t)−fε0j(t)|=O(n−1/2),and sup t|tfεj(t)−tfε0j(t)|=o(1).(17) From (13)–(17), after some routine calculations, using the fact that σj(Xj)/ σ0(Xj)=1+n−1/2δj(Xj), we obtain that, √njˆ Fεj(t)−ˆ Fε0j(t)=p1/2 jtfεj(t)E[δj(Xj)] + t 2fεj(t)Zn,j +op(1), uniformly in t,whereZn,j is defined in (5). Proof of Theorem 6.First observe that ˆϕεj(t)−ˆϕε0j(t) =ˆϕεj(t)−˜ϕεj(t)−ˆϕε0j(t)−˜ϕε0j(t)+˜ϕεj(t)−˜ϕε0j(t).(18) We have that |ˆϕεj(t)−˜ϕεj(t)|2w(t)≤2 nj l (ˆεjl −εjl)2t2w(t)dt =op(1), and, similarly, |ˆϕε0j(t)−˜ϕε0j(t)|2w(t)=op(1).On the other hand, ˜ϕεj(t)−˜ϕε0j(t) 2w(t)dt =1 n2 j nj  r,s=1 {Iw(εjr −εjs)+Iw(ε0jr −ε0js)−2Iw(εjr −ε0js)}, with Iwas defined in (7), is a V-statistic of degree 2 with a bounded kernel and thus (see [26]) it converges to its expected value ϕεj(t)−ϕε0j(t) 2w(t)dt.We conclude that 1 nT1 p −→ τ1.The limit of 1 nT2can be similarly derived. Proof of Theorem 7.Using a Taylor’s development, we get ˜ϕε0j(t)−˜ϕεj(t)=i t nj nj  l=1 (ε0jl −εjl)exp(itεjl)+t2R0j(t)1 nj nj  l=1 (ε0jl −εjl)2, Equality of conditional variance functions 1849 with supt|R0j(t)|=Op(1).This, together with the fact that ε0j−εj=σj(Xj) σ0(Xj)−1εj=n−1/2δj(Xj)εj,(19) leads to √nj˜ϕε0j(t)−˜ϕεj(t)=p1/2 jtϕ εj(t)E[δj(Xj)] + tR(1) 0j(t)+t2R(2) 0j(t),(20) with supt|R(s) 0j(t)|=op(1),s=1,2. Now using (11), (10), (19) and Lemmas 11 and 10, we obtain that ˆϕεj(t)−˜ϕεj(t)−ˆϕε0j(t)−˜ϕε0j(t) =i t nj nj  l=1 exp(itεjl)ˆmj(Xjl)−mj(Xjl) σj(Xjl)σj(Xjl) σ0(Xjl)−1 +t 2ϕ εj(t) k  v=1 1 nv nv  s=1 (ε2 vs −1)πv(Xvs)fj(Xvs) fv(Xvs)σ2 j(Xvs) σ2 0(Xvs)−1 −t 2ϕ εj(t)1 √nj Zn,j +tR(2) 0j,1(t)+t2R(2) 0j,2(t) =−t 2ϕ εj(t)1 √nj Zn,j +tR(2) 0j,1(t)+t2R(2) 0j,2(t),(21) where Zn,j is given by (5)andsup t|R(2) 0j,s(t)|=op(n−1/2), s=1,2. Combining (18), (20)and(21), we conclude that √njˆϕεj(t)−ˆϕε0j(t)=p1/2 jtϕ εj(t)E[δj(Xj)]−t 2ϕ εj(t)Zn,j +tR1j(t)+t2R2j(t), with supt|Rsj(t)|=op(1), s=1,2. Acknowledgements The authors thank the Editor and an anonymous referee for their comments and suggestions. References [1] Akritas, M. G. and Van Keilegom, I. (2001). Non-parametric estimation of the residual distribution. Scandinavian Journal of Statistics,28 549–567. MR1858417 [2] Alba-Fern´ andez, V., Jim´ enez-Gamero, M. D. and Mu˜ nozGarc´ ıa, J. (2008). A test for the two-sample problem based on empirical characteristic functions. Computational Statistics & Data Analysis, 52 3730–3748. MR2427377 [3] Delgado, M. A. (1993). Testing the equality of nonparametric regression curves. Statistics and Probability Letters,17 199–204. MR1229937 1850 J. C. Pardo-Fern´andez et al. [4] Dette, H. and Marchlewski, M. (2010). A robust test for homoscedasticity in nonparametric regression. Journal of Nonparametric Statistics,22 723–736. MR2682218 [5] Dette, H., Neumeyer, N. and Van Keilegom, I. (2007). A new test for the parametric form of the variance function in non-parametric regression. Journal of the Royal Statistical Society, Series B,69 903–917. MR2368576 [6] Dette, H., Pardo-Fern´ andez, J. C. and Van Keilegom, I. (2009). Goodness-of-fit tests for multiplicative models with dependent data. Scandinavian Journal of Statistics,36 782–799. MR2573308 [7] Einmahl, J. and Van Keilegom, I. (2008). Specification tests in nonparametric regression. Journal of Econometrics,143 88–102. MR2384434 [8] Fan, J. and Gijbels. I. (1996). Local Polynomial Modelling and Its Applications. Chapman & Hall, London. MR1383587 [9] Fan, J. and Yao, Q. (1998). Efficient estimation of conditional variance functions in stochastic regression. Biometrika,85 645–660. MR1665822 [10] Feller, W. (1971). An Introduction to Probability Theory and Its Applications, Vol. 2. Wiley, New York. MR0270403 [11] Gastwirth, J. L., Gel, Y. and Miao, W. (2009). The impact of Levene’s test of equality of variances on statistical theory and practice. Statistical Science,24 343–360. MR2757435 [12] Gonz´ alez-Manteiga, W. and Crujeiras, R. M. (2013). An updated review of Goodness-of-Fit tests for regression models. Test,22 361–411. MR3093195 [13] Hansen, B. E. (2008). Uniform convergence rates for kernel estimation with dependent data. Econometric Theory,24 726–748. MR2409261 [14] Imbens, G. W. and Lemieux, T. (2008). Regression discontinuity designs: A guide to practice. Journal of Econometrics,142 615–635. MR2416821 [15] Koul, H. L. and Song, W. (2010). Conditional variance model checking. Journal of Statistical Planning and Inference,140 1056–1072. MR2574668 [16] Kulasekera, K. B. (1995). Comparison of regression curves using quasiresiduals. Journal of the American Statistical Association,90 1085–1093. MR1354025 [17] Levene, H. (1960). Robust tests for equality of variances. In Contributions to Probability and Statistics (I.Olkin,S.G.Ghurye,W.Hoeffding, W. G. Madow and H. B. Mann, editors), 278–292. Stanford University Press, Stanford. MR0120709 [18] Liero, H. (2003). Testing homoscedasticity in nonparametric regression. Journal of Nonparametric Statistics,15 31–51. MR1958958 [19] Mathur, S. K. and Dolo, S. (2008). A new efficient statistical test for detecting variability in the gene expression data. Statistical Methods in Medical Research,17 405–419. MR2526661 [20] Mishra, S., Su, L. and Ullah, A. (2010). Semiparametric estimator of time series conditional variance. Journal Business and Economic Statistics, 28 256–274. MR2681200