scieee AI-readable full text Open interactive document viewer

A note on interval estimation for the mean of inverse Gaussian distribution

Arefi, M.; Mohtashami Borzadaran, G. R.; Vaghei, Y.

Abstract

In this paper, we study the interval estimation for the mean from inverse Gaussian distribution. This distribution is a member of the natural exponential families with cubic variance function. Also, we simulate the coverage probabilities for the confidence intervals considered. The results show that the likelihood ratio interval is the best interval and Wald interval has the poorest performance.

Full text

Statistics & Operations Research Transactions SORT 32 (1) January-June 2008, 49-56 Statistics & Operations Research Transactions A note on interval estimation for the mean of inverse Gaussian distribution c Institut d’Estad´ ıstica de Catalunya [email protected] ISSN: 1696-2281 www.idescat.net/sort M. Arefi1, G.R. Mohtashami Borzadaran2,∗and Y. Vaghei3 Abstract In this paper, we study the interval estimation for the mean from inverse Gaussian distribution. This distribution is a member of the natural exponential families with cubic variance function. Also, we simulate the coverage probabilities for the confidence intervals considered. The results show that the likelihood ratio interval is the best interval and Wald interval has the poorest performance. MSC: 62F10, 62F12, 62E20 Keywords: Wald Interval, Score Interval, Likelihood ratio, coverage probability 1 Introduction Interval estimation for natural exponential family (NEF) is discussed in many references for more than sixty years. Santner (1998), Agresti and Coull (1998), Brown, Cai and DasGupta (2001, 2002) considered Wald score, Agresti-Coull, likelihood ratio and Jeffreys intervals. Also, a good contribution for comparison of these intervals for NEF with quadratic variance function is given by Brown, Cai and DasGupta (2003). Several alternative intervals, coverage probability and length expansion of the Wald, score, *Corresponding author: E-mail addresses: [email protected], [email protected]. The authors acknowledge the support received from the “Ordered and Spatial Data Center of Excellence of Ferdowsi University of Mashhad”. 1Department of Mathematical Sciences, Isfahan University of Technology, Isfahan 84156, Iran. E-mail: [email protected] 2Department of Statistics, School of Mathematical Sciences, Ferdowsi University of Mashhad, Mashhad-Iran. 3Department of Statistics, University of Birjand, Birjand-Iran. Received: November 2006 Accepted: December 2007 50 A note on interval estimation for the mean of inverse Gaussian distribution likelihood ratio, Agresti-Coull and Jeffreys intervals are studied, simulated and approximated in Brown, Cai and DasGupta (2003) for normal, gamma, generalized hyperbolic secant (GHS), binomial, negative binomial and Poisson as the member of NEF with quadratic variance function. In this paper, we extend several intervals for inverse Gaussian distribution as a member of NEF with cubic variance function (CVF). Also, we show the coverage probabilities of these intervals via a simulation study, while noting that the theoretical approach for it is difficult, because finding the roots of a polynomial of degree 3 equation is not always easy. Applying the results for other NEFCVF, especially for Abel and Takacs distributions, is a topic for future study. This paper is organized as follows: In Section 2, we introduce the natural exponential family with quadratic and cubic variance functions. In Section 3, we study the confidence intervals for the mean of the inverse Gaussian as the most famous member of the natural exponential family with cubic variance function. In Section 4, the coverage probabilities for the confidence intervals are simulated in two case: with the fixed sample size (n), fixed λand the variable parameter μ, and with the fixed parameters and the variable sample size. In Section 5, the proposed confidence intervals have been applied for the mean of ball bearings. A brief conclusion is provided in Section 6. 2 Natural exponential family The distributions in a natural exponential family (NEF) have the following form f(x;ξ)=exp(ξx−ψ(ξ))h(x), where ξis called the natural parameter. The mean, variance and cumulant generating function for this family are ⎧ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎩ μ=E(X)=ψ(ξ), σ2=Var(X)=ψ(ξ), Φ(t)=ψ(t+ξ)−ψ(ξ). In a subclass, on the real line, the variance function is a polynomial function with degree lower than or equal to 3. We get twelve different types, the first six appear in Morris (1982,1983), Shanbhag (1972, 1979) and Brown (1986). The others introduced in Letac and Mora (1990), where the inverse Gaussian distribution is one of the most well-known members of NEF-CVF. The variance ψ(ξ) of NEF-CVF depends on ξonly through the mean μ,as σ2=Var(X)=a0+a1μ+a2μ2+a3μ3, M. Arefi, G.R. Mohtashami Borzadaran and Y. Vaghei 51 where a0,a1,a2,anda3are suitable constants. The NEF-QVF is the special case with a3=0. NEF-QVF consists of six important distributions: normal, gamma, generalized hyperbolic secant (GHS), binomial, negative binomial, and Poisson (see Morris (1982,1983) and Brown (1986)). Also, NEF-CVF consists of these six distributions: Ressel, inverse Gaussian, Abel, Takacs, strict arcsine, and large arcsine. We focus our attentions on the inverse Gaussian distribution, IG(μ, λ), with density f(x)=λ 2πx3exp −λ(x−μ)2 2μ2x,x>0,μ>0,λ>0. For this distribution, we have: ξ=−λ 2μ2,ψ(ξ)=−(−2λξ),h(x)=λ 2πx3exp(−λ 2x). Hence, Var(X)=μ3 λwith a0=a1=a2=0, and a3=1 λ. 3 Confidence intervals for the mean of inverse Gaussian distribution Based on a random sample of size nfrom IG(μ, λ)(λknown), we want to calculate some of the confidence intervals, at the confidence level 1 −α, for mean μ. We also know that for this distribution the MLE of μis the sample mean, that is ˆμ=¯ X. We can construct the following confidence intervals for μ. Definition 1 The Wald interval (CIW)is based on Slutsky’s statistic Wn=√n(ˆμ−μ) ˆσ as CIW:ˆμ±k(nλ)−1/2ˆμ3/2, where k =z1−α/2. Definition 2 The score interval (CIS)is based on the Central Limit theorem with Z=√n(ˆμ−μ) σ. We construct this interval by inverting Rao’s equal-tailed score test of H0:μ=μ0. Hence, we accept H0based on Rao’s score test if and only if μ0is in the interval. We solve a cubic equation in terms of μ, under the following relation: −k≤√nλ(ˆμ−μ) μ3≤k. 52 A note on interval estimation for the mean of inverse Gaussian distribution By substituting μ=1/x2in the above equations, we have ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ x3−x ˆμ−k √nλˆμ≤0,(1) x3−x ˆμ+k √nλˆμ≥0.(2) Suppose that, Q =1 3ˆμ,R 1=−k 2√nλˆμ, and R2=k 2√nλˆμ, then we have the following cases: •Case 1: If 4nλ−27ˆμk2<0,thenQ 3−R2 1<0. Hence, the relations (1) and (2) imply ⎧ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎩ x≤[(R2+R2 2−Q3)1/3+(R2−R2 2−Q3)1/3]=M, x≥−[(R2+R2 2−Q3)1/3+(R2−R2 2−Q3)1/3]=−M. By substituting μ=1/x2in the above equations, we obtain a one-sided interval for μas 1 M2≤μ. •Case 2: If 4nλ−27ˆμk2≥0,thenQ 3−R2 1≥0. The roots of the equations (1) and (2) are as ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ x1=−2√Qcos θ1 3, x2=−2√Qcos θ1+2π 3, x3=−2√Qcos θ1+4π 3, ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ x 1=−2√Qcos θ2 3, x 2=−2√Qcos θ2+2π 3, x 3=−2√Qcos θ2+4π 3, where θ1=arccos R1 √Q3and θ2=arccos R2 √Q3. Based on the above roots and by substituting μ=1 x2, the Score interval is calculated as ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ 1 √x 3≤μ, 1 √x2≤μ≤1 x 2 . M. Arefi, G.R. Mohtashami Borzadaran and Y. Vaghei 53 Definition 3 The likelihood ratio interval (CILR)is constructed by inverting the likelihood ratio test under H0:μ=μ0(see Rao (1973) and Serfling (1980)). The likelihood ratio Λn=L(μ0) supμL(μ)for inverse Gaussian distribution is given by Λn=exp(−nλˆμ 2μ2+nλ μ−nλ 2ˆμ). By solving the equation −2logΛn≤χ2 α,1=k2in terms of μ, the likelihood ratio interval for μis calculated as nλˆμ nλ+knλˆμ≤μ≤nλˆμ nλ−knλˆμ, where 0<k<nλ ˆμ. Note: If λin inverse Gaussian distribution is unknown, then we substitute a point estimation instead of λin the above equations (for example MLE ˆ λ=n⎡⎢⎢⎢⎢⎢⎣ n  i=1 (xi−x)2 xix2⎤⎥⎥⎥⎥⎥⎦ −1. 4 The coverage probability via a simulation study In this section, we estimate by simulation the coverage probabilities of the considered confidence intervals for the mean. Suppose that, the parameters μand λof inverse Gaussian distribution are fixed values. Based on a random sample X1,X2,...,Xnof IG(μ, λ), we calculate the confidence intervals with ˆ λ=n⎡⎢⎢⎢⎢⎢⎣ n  i=1 (xi−x)2 xix2⎤⎥⎥⎥⎥⎥⎦ −1 ,at the confidence level 1 −α=0.95, using 1000 repetitions for each simulation set. The estimated coverage probability (ECP) is the proportion of confidence intervals containing the fixed value of the mean. The estimation of the coverage probability is calculated by simulation in two cases: with fixed parameters and variable sample size (λ=231.711, μ=72.219 and 5≤n≤100), and with fixed sample size (n=23), fixed λ=10 and variable 0 <μ≤6 (see Fig. 1 and Fig. 2). From Fig. 1 and Fig. 2, we can refer the following results for the confidence intervals in inverse Gaussian distribution: •The averages of the estimated coverage probabilities (ECP) for the considered confidence intervals, are organized as ECPW≤ECPS≤ECPLR. 54 A note on interval estimation for the mean of inverse Gaussian distribution •The averages of the estimated coverage probabilities for Wald, score, and likelihood ratio intervals are lower than the confidence level. Figure 1: Coverage probability with fixed parameters (λ, μ)and variable n from 5 to 100. Figure 2: Coverage probability with fixed n, fixed λand variable parameter μ. M. Arefi, G.R. Mohtashami Borzadaran and Y. Vaghei 55 •When parameter μgoes to zero, the estimated coverage probability decreases. •When the sample size increases, the average of coverage probabilities tends to the confidence level. 5 Application of the confidence intervals from an inverse Gaussian distribution Twenty three ball bearings were used in the life test study and yielded the following in millions of revolutions to failure, (see Pavur et al., 1992). Data Data Data Data Data Data 17.88 51.96 93.12 28.92 54.12 98.64 33.00 55.56 105.12 41.52 67.80 105.84 42.12 68.64 127.92 45.60 68.64 127.92 48.48 68.88 173.40 51.84 84.12 They showed that the data was from an inverse Gaussian population. We use the data as well to demonstrate our confidence intervals. We estimate ˆμ=72.219, and ˆ λ=231.711. The confidence intervals, at the confidence level 1 −α=0.95, for the mean are calculated as: ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ CIW=[55.74,88.70], CIS=[59.80,98.44], CILR =[58.80,93.57]. The lengths of the three considered confidence intervals are: LW=32.95,LS=38.64,LLR =34.77. Hence, LW<LLR <LS.Based on a random sample of size n=23 with λ=231.711 and μ=72.219, the estimated the coverage probabilities (ECP) for Wald, score, and likelihood ratio intervals are calculated by simulation as (see Fig. 1): ECPW=0.922,ECPS=0.933,ECPLR =0.935. Because of Wald interval is uniformly poor, we prefer to use of the likelihood ratio interval. The likelihood ratio interval is better than the score interval of the viewpoint its coverage probability and length (ECPS<ECPLR and LLR <LS). Also, we can use of the above confidence intervals for testing hypothesis for the mean of ball bearings. 56 A note on interval estimation for the mean of inverse Gaussian distribution 6 Conclusion In this paper, we calculate Wald score, and likelihood ratio intervals for inverse Gaussian distribution. Also, we estimate the coverage probabilities by simulation. The simulation study shows that the Wald confidence interval has the lowest estimated coverage probability. The likelihood ratio interval is better, having an estimated coverage probability very close to the nominal value. Finally we apply the results for an example (ball bearings). Acknowledgements The authors are grateful to the referees for their comments and suggestions which improved the quality of the paper. References Agresti, A. and Coull, B. A. (1998). Approximate is better than exact for interval estimation of binomial proportions. The American Statistician, 52, 119-126. Brown, L. D. (1986). Fundamental of statistical exponential families with applications in statistical decision theory. Lecture Notes-Monograph Series, Institute of Mathematical Statistics,Hayward, California. Brown, L.D., Cai, T. and DasGupta, A. (2001). Interval estimation for a binomial proportion (with discussion). Statistical Science, 16, 101-133. Brown, L. D., Cai, T. and DasGupta, A. (2002). Confidence intervals for a binomial proportion and Edgeworth expansions. The Annals of Statistics, 30, 160-201. Brown, L. D., Cai, T. and DasGupta, A. (2003). Interval estimation in exponential families. Statistics Sinica, 13, 19-49. Letac, G. and Mora, M. (1990). Natural real exponential families with cubic variance functions. The Annals of Statistics, 18, 1-37. Morris, C. N. (1982). Natural exponential families with quadratic variance functions. The Annals of Statistics, 10, 65-80. Morris, C. N. (1983). Natural exponential families with quadratic variance functions: statistical theory. The Annals of Statistics, 11, 515-529. Pavur, R. J., Edgeman, R. L. and Scott, R. C. (1992). Quadratic statistics for the goodness-of-fit test of the inverse Gaussian distribution. IEEE Transactions in Reliability, 41, 118-123. Rao, C. R. (1973). Linear statistical inference and its applications. Wiley, New York. Santner, T. J. (1998). A note on teaching binomial confidence intervals. Teaching Statistics, 20, 20-23. Serfling, R. J. (1980). Approximation theorems of mathematical statistics. Wiley, New York. Shanbhag, D. N. (1972). Some characterizations based on the Bhattacharyya matrix. Journal of Applied Probability, 9, 580-587. Shanbhag, D. N. (1979). Diagonality of the Bhattacharyya matrix as a characterization. Theory of Probability and Its Applications, 24, 430-433.