On the eigenvalues associated with the limit null distribution of the Epps-Pulley test of normality
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Ebner, Bruno; Henze, Norbert Article — Published Version On the eigenvalues associated with the limit null distribution of the Epps-Pulley test of normality Statistical Papers Provided in Cooperation with: Springer Nature Suggested Citation: Ebner, Bruno; Henze, Norbert (2022) : On the eigenvalues associated with the limit null distribution of the Epps-Pulley test of normality, Statistical Papers, ISSN 1613-9798, Springer, Berlin, Heidelberg, Vol. 64, Iss. 3, pp. 739-752, https://doi.org/10.1007/s00362-022-01336-6 This Version is available at: https://hdl.handle.net/10419/308594 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/
Statistical Papers (2023) 64:739–752 https://doi.org/10.1007/s00362-022-01336-6 REGULAR ARTICLE On the eigenvalues associated with the limit null distribution of the Epps-Pulley test of normality Bruno Ebner1 ·Norbert Henze1 Received: 20 September 2021 / Revised: 19 April 2022 / Accepted: 16 June 2022 / Published online: 5 July 2022 © The Author(s) 2022 Abstract The Shapiro–Wilk test (SW) and the Anderson–Darling test (AD) turned out to be strong procedures for testing for normality. They are joined by a class of tests for normality proposed by Epps and Pulley that, in contrast to SW and AD, have been extended by Baringhaus and Henze to yield easy-to-use affine invariant and universally consistent tests for normality in any dimension. The limit null distribution of the Epps– Pulley test involves a sequences of eigenvalues of a certain integral operator induced by the covariance kernel of a Gaussian process. We solve the associated integral equation and present the corresponding eigenvalues. Keywords Test for normality ·Integral operator ·Fredholm determinant Mathematics Subject Classification Primary 62F03 ; Secondary 65C60 ·65R20 1 Introduction Let X,X1,X2...be a sequence of independent and identically distributed (i.i.d) random variables with unkown distribution. To test the hypothesis H0that the distribution of Xis some unspecified normal distribution, there is a myriad of testing procedures, among which the tests of Shapiro–Wilk (SW) and Anderson–Darling (AD) deserve special mention, see, e.g., the monographs of D’Agostino and Stephens (1996) and Thode (2002). There is, however, a further test which was proposed by Epps and Pulley (1983). This test, which is based on the empirical characteristic function, comes as a serious competitor to SW and AD, as shown in simulation studies (see, e.g., BarBBruno Ebner [email protected] Norbert Henze [email protected] 1Institute of Stochastics, Karlsruhe Institute of Technology (KIT), Englerstr. 2, 76133 Karlsruhe, Germany 123
740 B. Bruno , N. Henze inghaus et al. 1989; Betsch and Ebner 2020). Baringhaus and Henze (1988) extended the approach of Epps and Pulley to test for normality in any dimension. By now, the BHEP-test (an acronym coined by Csörgö (1989) after earlier developers of the idea) is known to be an affine-invariant and universally consistent test of normality in any dimension, and limit distributions of the test statistic have been obtained under H0 as well as under fixed and contiguous alternatives to normality (see the review article Ebner and Henze 2020). In this paper, we revisit the limit null distribution of the Epps–Pulley test statistic in the univariate case. The test statistic involves a positive tuning parameter β, and, based on X1,...,Xn, is denoted by Tn,β .Itisgivenby Tn,β =n∞ −∞ ψn(t)−e−t2/2 2 ϕβ(t)dt, where ψn(t)=n−1n j=1exp itYn,jis the empirical characteristic function of the scaled residuals Yn,1,...,Yn,n. Here, Yn,j=S−1 n(Xj−Xn),j=1,...,n, and Xn=n−1n j=1Xj,S2 n=n−1n j=1(Xj−Xn)2denote the sample mean and the sample variance of X1,...,Xn, respectively. Moreover, ϕβ(t)=1 β√2πexp −t2 2β2,t∈R, is the density of the centred normal distribution with variance β2. A closed-form expression of Tn,β that is amenable to computational purposes is Tn,β =1 n n j,k=1 exp−β2 2Yn,j−Yn,k2 −2 1+β2 n j=1 exp− β2Y2 n,j 2(1+β2)+n 1+2β2. The limit null distribution of Tn,β ,asn→∞, is that of T∞:=∞ −∞ Z2(t)ϕ β(t)dt. Here, Z(·)is a centred Gaussian element of the Hilbert space L2=L2(R,B,ϕ β(t)dt) (of equivalence classes) of Borel-measurable real-valued functions that are squareintegrable with respect to ϕβ(t)dt, and the covariance function of Z(·)is given by K(s,t)=exp −(s−t)2 2−1+st +(st)2 2exp −s2 2−t2 2,s,t∈R (1.1) 123
On the eigenvalues associated... 741 (see Henze and Wagner 1997). The kernel Kis the starting point of this paper. Writing ∼for equality in distribution, it is well-known that T∞∼∞ j=0 λjN2 j, where λ0,λ 1,...is the sequence of nonzero eigenvalues associated with the integral operator A:L2→L2defined by (Af)(s):=∞ −∞ K(s,t)f(t)ϕβ(t)dt, and N0,N1,...is a sequence of i.i.d. standard normal random variables. In the next section, we obtain the eigenvalues of Aby numerical methods. In Sect. 3the sum of powers of the largest eigenvalues is compared to normalized cumulants. The difference should be close to 0 if the eigenvalues have been computed correctly. Section 4 demonstrates that the results can be applied to fit a Pearson system of distributions, and that the fit is reasonable to approximate critical values of the Epps-Pulley test. The article ends by some concluding remarks. Finally, Appendix A extends the results to the cases in which no parameters, only the mean and only the variance, are estimated. 2 Solution of a Fredholm integral equation To obtain the values λ0,λ 1,λ 2,..., that determine the distribution of T∞, one has to solve the integral equation ∞ −∞ K(s,t)f(t)ϕβ(t)dt=λf(s), s∈R. In general, this task is considered a hard problem, and solutions for kernels associated with testing problems involving composite hypotheses are very sparse, see Stephens (1976,1977) for the classical tests of normality and exponentiality that are based on the empirical distribution function. In what follows, we use a result of Zhu et al. (1997) to obtain the eigenvalues of Aby a stable numerical method. To this end, let K0(s,t)=exp −(s−t)2 2,s,t∈R, and φ1(s)=s2 √2exp −s2 2, φ2(s)=sexp −s2 2, 123
742 B. Bruno , N. Henze φ3(s)=exp −s2 2,s∈R. Notice that K(s,t)=K0(s,t)− 3 j=1 φj(s)φj(t), s,t∈R.(2.1) The first step is to solve the eigenvalue problem for the covariance kernel K0.The associated eigenvalue problem, which leads to the kernel K0, was solved in the context of machine learning in Chapter 4 of Zhu et al. (1997). Here, we use the formulation given in Rasmussen and Williams (2008), Sect. 4.3.1. The eigenvalues of K0are given by λ(0) k=2 1+4β2+2β2+1⎛ ⎝ 2β2 1+4β2+2β2+1⎞ ⎠ k ,k=0,1,2,..., with corresponding normalized eigenfunctions ψk(x)=hkexp −β−2+4 4β−(4β)−1x2Hkβ−4+4β−21/4 x/√2, k=0,1,2,... (see also the errata to Rasmussen and Williams (2008) on the books homepage). Here, h−2 k=(4β2+1)−1/42kk!, and Hk(x)=(−1)kexp(x2)dk dxkexp(−x2)is the kth order Hermite polynomial. Remark 2.1 Note that for the special case β=1 the eigenvalues λ(0) kcoincides with the formula given in (6) of Baringhaus (1996). In that article, the limit distribution of a modified statistic T(0) n,β , which originates from Tn,β by replacing ψn(t)with ψ(0) n(t)= n−1n j=1exp itXj, i.e., the problem is to test for standard normality and thus no estimation of parameters is involved, is analyzed, cf. Appendix A. The corresponding covariance kernel is K1(s,t)=K0(s,t)−φ3(s)φ3(t),s,t∈R, and explicit formulae for the eigenvalues and eigenfunctions are given, see Baringhaus (1996), p. 3878. To solve the eigenvalue problem of Afiguring in (1.1), we adapt the methodology in Stephens (1976). Define aj,k=∞ −∞ ψj(x)φk(x)ϕβ(x)dx, Sk(λ) =1+∞ j=0 a2 j,k 1/λ −λ(0) j , 123
On the eigenvalues associated... 743 Sk,(λ) =∞ j=0 aj,kaj, 1/λ −λ(0) j ,λ>0. With this notation, we can formulate our main result. Theorem 2.2 The eigenvalues of Aare the reciprocals of the solutions λ>0of the equation D(λ) =d(λ)S2(λ) S1(λ)S3(λ) −S2 1,3(λ)=0,(2.2) where d(λ) =∞ k=01/λ −λ(0) kis the Fredholm determinant connected to the eigenvalue problem of K0. Moreover, none of the reciprocals of the eigenvalues λ(0) k of K0solve equation (2.2). Proof Since aj,1aj,2=aj,2aj,3=0 holds for all j=0,1,2,..., we use Theorem 2.2 of Sukhatme (1972) to see that the Fredholm determinant for the eigenvalue problem takes the form D(λ) =d(λ) det ⎛ ⎝ S1(λ) 0S1,3(λ) 0S2(λ) 0 S1,3(λ) 0S3(λ) ⎞ ⎠=d(λ)S2(λ) S1(λ)S3(λ) −S2 1,3(λ). Hence, the reciprocals of the roots of D(λ) are the eigenvalues of A. By direct calculation, it follows that aj,2=0if jis even, and we have aj,1=aj,3=0if jis odd. Consequently, none of the reciprocals of the eigenvalues λ(0) kis a root of D(λ) and thus a solution of the eigenvalue problem associated with the kernel K. According to Theorem 2.2, the eigenvalues of Aare the roots of S2(λ) and of S1(λ)S3(λ) −S2 1,3(λ). The reciprocals of the roots have been obtained numerically, and the largest twenty eigenvalues are displayed in Table 1for different values of β. Note that, since these values tend to be very small, the reciprocal approach used here leads to numerically stable procedures to find the roots of the Fredholm determinant. 3 Accuracy of the numerical solutions The accuracy of the values presented in Table 1may be judged by a comparison with results of Henze (1990). That paper gives the first four cumulants of the distribution of T∞in the special case β=1. The m-th cumulant of T∞is κm(β) =2m−1(m−1)!∞ j=1 λm j=2m−1(m−1)!∞ −∞ Km(x,x)ϕβ(x)dx,m≥1, 123
744 B. Bruno , N. Henze Table 1 Eigenvalues of Afor different tuning parameters β, here E–j stands for 10−j λj\β0.25 0.5 1 2 3 λ04.07235E−04 1.01443E−02 7.42748E−02 1.54164E−01 1.59960E−01 λ13.96229E−05 2.98027E−03 4.48104E−02 1.29257E−01 1.45877E−01 λ28.87536E−07 2.13968E−04 8.41907E−03 4.99665E−02 7.56703E−02 λ37.41169E−08 5.45396E−05 4.58684E−03 3.98239E−02 6.69664E−02 λ42.36367E−09 5.42325E−06 1.07998E−03 1.70946E−02 3.68745E−02 λ51.81032E−10 1.27337E−06 5.51939E−04 1.31547E−02 3.19372E−02 λ66.72990E−12 1.46554E−07 1.45739E−04 6.00412E−03 1.82413E−02 λ74.87430E−13 3.26023E−08 7.12110E−05 4.49725E−03 1.55292E−02 λ81.97617E−14 4.08130E−09 2.01821E−05 2.14175E−03 9.10854E−03 λ91.37638E−15 8.73898E−10 9.53839E−06 1.56980E−03 7.64324E−03 λ10 5.90134E−17 1.15555E−10 2.83684E−06 7.71666E−04 4.57714E−03 λ11 3.99302E−18 2.40495E−11 1.30684E−06 5.55566E−04 3.79317E−03 λ12 1.78040E−19 3.30498E−12 4.02503E−07 2.79917E−04 2.31011E−03 λ13 1.17813E−20 6.72882E−13 1.81702E−07 1.98509E−04 1.89295E−03 λ14 5.40732E−22 9.51525E−14 5.74665E−08 1.01995E−04 1.16876E−03 λ15 3.51545E−23 1.90367E−14 2.55205E−08 7.13642E−05 9.46891E−04 λ16 1.64986E−24 2.75201E−15 8.24056E−09 3.71939E−05 5.90055E−04 λ17 1.05731E−25 5.42770E−16 3.61045E−09 2.56128E−05 4.70590E−04 λ18 6.60890E−28 1.07411E−17 1.18543E−09 2.86884E−06 8.19769E−05 λ19 7.53533E−29 3.77859E−18 5.13526E−10 3.67008E−06 1.22699E−05 where K1(x,y):=K(x,y)and Km(x,y)=∞ −∞ Km−1(x,z)K(z,y)ϕβ(z)dz for m≥2 (see e.g., Chapter 5 of Shorack and Wellner (1986)). We have κ1(1)=E(T∞)=∞ j=1 λj=1−√3 2=0.133974596 ..., (3.1) κ2(1)=V(T∞)=2∞ j=1 λ2 j=2√5 5+5 6−155√2 128 =0.015236301 ... and thus ∞ j=1 λ2 j=1 √5+5 12 −155 128√2=0.0076181509 ... 123
On the eigenvalues associated... 745 Furthermore, κ3(1)=E(T∞−κ1(1))3=8∞ j=1 λ3 j=0.00400343 ..., (3.2) κ4(1)=E(T∞−κ1(1))4=48 ∞ j=1 λ4 j=0.001654655 ... and thus ∞ j=1 λ3 j=0.0005004285291 ... and ∞ j=1 λ4 j=0.00003447197917 ... From Table 2, we see that the corresponding sums of the first 20 numerical values of the eigenvalues (as well of their squares and cubes) agree approximately with the values figuring in (3.1) and (3.2), respectively, in most cases up to five significant digits. The results of Henze (1990) have been partially generalized in Henze and Wagner (1997), Theorem 2.3, for the first three cumulants and a fixed tuning parameter β, and they thus lead to general formulae in the univariate case. For the sake of completeness, we restate the formulae of the first two cumulants here. For the first cumulant, we have κ1(β) =1−(2β2+1)−1/21+β2 2β2+1+3β4 (2β2+1)2, and the second cumulant is κ2(β) =2 1+4β2+2 1+2β21+2β4 (1+2β2)2+9β8 4(1+2β2)4 −4 1+4β2+3β41+3β4 2(1+4β2+3β4)+3β8 2(1+4β2+3β4)2. The formula for the third cumulant is found in Henze and Wagner (1997), Theorem 2.3, for the case d=1. Table 2exhibits the normalized cumulants, together with the corresponding sums of the first 20 eigenvalues taken from Table 1. We stress that by now no formula for the fourth cumulant is known in the literature for general tuning parameter β. 4 Pearson system fit for approximation of critical values The first four cumulants can directly be used in packages that implement the Pearson system of distributions (see Sect. 4.1 of Johnson et al. (1994)). In the statistical computing language R(see R Core Team 2021), we use the package PearsonDS 123
746 B. Bruno , N. Henze Table 2 Sums over different powers of the first 20 eigenvalues and corresponding theoretical cumulants for different values of β. The entry denoted by ∗could not be computed due to numerical instabilities β0.25 0.5 1 2 3 19 j=0λj4.47822E−04 1.34000E−02 1.33975E−01 4.19722E−01 5.83761E−01 19 j=0λ2 j1.67411E−07 1.11838E−04 7.61814E−03 4.50863E−02 6.02196E−02 19 j=0λ3 j6.75980E−11 1.04392E−06 5.00428E−04 6.01902E−03 8.02464E−03 19 j=0λ4 j2.75053E−14 1.06687E−08 3.44719E−05 8.52858E−04 1.16350E−03 κ1(β) 4.47822E−04 1.34000E−02 1.33975E−01 4.19753E−01 5.84700E−01 κ2(β)/2 1.66500E−07 1.11838E−04 7.61814E−03 4.50863E−02 6.02202E−02 κ3(β)/8∗1.07115E−06 5.00429E−04 6.01903E−03 8.02468E−03 (see Becker and Klößner 2017) to approximate critical values of the Epps–Pulley test statistic. The Epps–Pulley test is implemented in the R-Package mnt (see Butsch and Ebner 2020) by using the function BHEP. Table 3shows simulated empirical critical values of the Epps–Pulley statistic for sample sizes n∈{10,25,50,100,200}and levels of significance α∈{0.1,0.05,0.01}. For each combination of nand β,the entries corresponding to different values of αare based on 106replications under the null hypothesis. Each entry in a row named ’∞’ is the calculated (1 −α)-quantile of the fitted Pearson system using the cumulants given in Table 2. We conclude that, for larger sample sizes, the simulated critical values are close to the approximated counterparts of the Pearson system. Moreover, we have corroborated the results of Henze (1990) for the special case β=1, and we have extended these results for general β>0. 5 Conclusions We have solved the eigenvalue problem of the integral operator associated with the covariance kernel Kof the limiting Gaussian process that occurs in the limit null distribution of the Epps–Pulley test statistic. In view of a comparison with the first three known cumulants from the literature, Table 2shows that the eigenvalues obtained by numerical methods are very close to the corresponding theoretical values. In Sect. 5 of Ebner and Henze (2021), the authors present a Monte Carlo based approximation method to find stochastic approximations of the eigenvalues. A comparison of Table 1and Table 1of Ebner and Henze (2021) reveals that there are some significant differences for some values of β, which can be explained by the approximation of the eigenvalues by a Monte Carlo method in Ebner and Henze (2021). This observation is of particular interest, since the largest eigenvalue is used in the derivations of approximate Bahadur efficiencies. Recent results concerning this topic for the Epps–Pulley test are presented in Ebner and Henze (2021) and, for other normality tests based on the empirical distribution function, in Miloševi´cetal.(2021). We point out the difficulties encountered if one tries to generalize our findings to the multivariate case, i.e. to obtain the eigenvalues associated with the limit null 123