A characterization of distributions based on linear regression of order statistics and record values
Abstract
We obtain the family of distributions for which the regression of one order statistic on another, not necessarily adjacent, is linear. As a consequence, we present a characterization of uniform distributions on an interval. We also characterize the distributions that appear when we impose the condition of linearity of regression for record values.
Full text
Sankhy¯a : The Indian Journal of Statistics 1997, Volume 59, Series A, Pt. 3, pp. 311-323 A CHARACTERIZATION OF DISTRIBUTIONS BASED ON LINEAR REGRESSION OF ORDER STATISTICS AND RECORD VALUES By FERNANDO L´ OPEZ BL´ AQUEZ and J. LUIS MORENO REBOLLO Universidad de Sevilla, Spain SUMMARY. We obtain the family of distributions for which the regression of one order statistic on another, not necessarily adjacent, is linear. As a consequence, we present a characterization of uniform distributions on an interval. We also characterize the distributions that appear when we impose the condition of linearity of regression for record values. 1. Introduction Ferguson (1967) characterized the distributions for which the regression of an order statistic on an adjacent one is linear. He pointed out that it is not known which distributions would be characterized if non-adjacent order statistics are considered. In the concluding remarks of his paper, Nagaraja (1988) affirms that the problem remains still unsolved. Another reference about this problem can be found in Arnold, Balakrishnan, and Nagaraja (1992) (p. 155). Nagaraja (1977, 1988) obtains a characterization based on the linear regression of two adjacent record values. As in the case of order statistics, the problem is open if the condition of linearity of regression for nonadjacent record values is imposed. After reviewing previous results in section 2, we give the characterization of the distributions which are characterized when the regression of two order statistics, not necessarily adjacent, is linear. In Section 4, we deal with a similar problem for record values. Paper received. April 1995; revised November 1996. AMS (1980) subject clasification. Primary 62E10; secondary: 62G30. Key words and phrases. Characterization of distributions, order statistics, record values, linear regression.
312 fernando l´ opez bl´ aquez and j. luis moreno rebollo 2. Previous Results Let us denote by X(i:n)the ith order statistic of a simple random sample (s.r.s.) of size nfrom a r.v. X, with c.d.f. Fand density f. For any three fixed integers, i,kand n, such that 1 ≤i < i +k≤n, and a real number xsatisfying 0< F(x)<1, the conditional density of X(i+k:n)given X(i:n)=xis f(i,k,n,x)(y) = Cikn (1−F(x))n−i(F(y)−F(x))k−1(1 −F(y))n−if(y),if y > x = 0,otherwise with Cikn =(n−i)! (k−1)!(n−i−k)!. Denote by DFthe set of real numbers such that 0 < F(x)<1 and the conditional expectation Rikn(x, F) = EX(i+k:n)/X(i:n)=xexists. In that case, for any xin DF, Rikn(x, F) = Cikn (1 −F(x))n−iZ∞ x y(F(y)−F(x))k−1(1 −F(y))n−i−kf(y)dy. Note also that, X(i+k:n)given X(i:n)=xis distributed as the k-th order statistic of a simple random sample of size n−ifrom a left-truncated at x random variable, Y(x), with c.d.f. Fx(y) = F(y)−F(x) 1−F(x),if y ≥x. Then Rikn(x, F) = E[Y(x) (k:n−i)]. It will be useful to consider the quantile function of F, defined as Q(u) = inf{x:F(x)≥u}, u ∈(0,1). In the following lemma we quote without proof some properties of the quantile function (e.g. see Port (1994), p. 98). Lemma 2.1. Let Fbe a distribution function and Qits quantile function, then (a) Qis non-decreasing. (b) Qis continuous at any point of (0,1) save, perhaps, at a countable set of points at which Qis left-continuous. (c) If Fis continuous, then Qis strictly increasing in (0,1).
regression of order statistics and record values 313 (d) If Qis continuous in (0,1), then Fis strictly increasing in {x: 0 < F(x)<1}. (e) If Uis a U(0,1) distribution, then the c.d.f. of Q(U)is F. (f) Moreover, (X(1:n), . . . , X(n:n))d ≡(Q(U(1:n)), . . . , Q(U(n:n))). (g) If Fis absolutely continuous, the conditional expectation of X(i+k:n)given X(i:n)=x, whenever it exists, is: EX(i+k:n)|X(i:n)=x=Ci,k,n (1 −u)n−iZ1 u Q(v)(v−u)k−1(1 −v)n−i−kdv. . . . (2.1) with u=F(x). Lemma 2.2. Let Fbe a continuous c.d.f., then (a) Rikn(., F)is continuous and non-decreasing. (b) If Fis strictly increasing, then Rikn(., F)also is. Proof. (a) The continuity of Rikn(., F) follows from the continuity of F. To show that Rikn(., F) is non-decreasing, let us choose x1, x2∈DF, such that x1≤x2. Consider the r.v.’s Y(xm) (k:n−i)=Q(F(xm) + (1 −F(xm))U(k:n−i)), m = 1,2, . . . (2.2) where U(k:n−i)is distributed as the kth order statistic of a s.r.s. of size n−i from a U(0,1) distribution. Then, the r.v.’s Y(xm) (k:n−i),m= 1,2, are distributed as the kth order statistic of a s.r.s of size n−ifrom the left-truncated distribution Fxm,m= 1,2. As Qis non-decreasing, from (2.2), we have Y(x1) (k:n−i)≤Y(x2) (k:n−i). . . (2.3) and taking expected values in (2.3), it follows that Rikn(x1)≤Rikn(x2). . . . (2.4) (b) Note that, if Fis strictly increasing the inequalities (2.3) and (2.4) are both strict. Lemma 2.3. If Fis an absolutely continuous c.d.f. such that Rikn(x, F) = bx +a, for any x∈DF, . . . (2.5)
314 fernando l´ opez bl´ aquez and j. luis moreno rebollo then (a) Fis strictly increasing in {x: 0 < F(x)<1}. (b) b > 0. (c) If b6= 1, the c.d.f. Fµ(x) = F(x−µ), with µ=a/(1 −b), satisfies Rikn(x, Fµ) = bx (d) If b= 1, then a > 0. Proof. (a) According to Lemma 2.1(d), it suffices to show that Qis continuous in (0,1). For that, define the function Sikn(u, Q) = Ci,k,n (1 −u)n−iZ1 u Q(v)(v−u)k−1(1 −v)n−i−kdv. As Qis a quantile function, from Lemma 2.1(b), Qis continuous a.e. Thus Sikn is continuous at any u∈(0,1), and from (2.5) we have, Sikn(u, Q) = bQ(u) + a, for any point of continuity of Q,. . . (2.6) But as the LHS of (2.6) is continuous, the continuity of Qin (0,1) follows. (b) As Fis strictly increasing, from Lemma 2.2(b), we conclude b > 0. (c) This is a consequence of the following relation Rikn(x, Fµ) = Rikn(x+µ, F )−µ. (d) If Xis a r.v. for which (2.5) holds with b= 1, we have a=ERikn X(i:n), F−EX(i:n) =EEX(i+k:n)−X(i:n)|X(i:n)=EX(i+k:n)−X(i:n)>0. Lemma 2.4. Consider the polynomial equation Pk(z) = 1 bPk(n−i), b > 0. . . (2.7) with Pk(z) = z(z−1) · · · (z−k+ 1). Then, (a) The real roots of (2.7) are at most double. (b) If z0is a complex root of (2.7), with Im(z0)6= 0, then z0is simple. (c) There is a unique simple real root of (2.7) in (k−1,+∞).Moreover, (c.1) If 0< b < 1, there is a unique real root in (n−i, +∞). (c.2) If b > 1, there is a unique root in (k−1, n −i). (c.3) If b= 1,z=n−iis the unique root in (k−1,+∞).
regression of order statistics and record values 315 (d) Pk(D)(texp(αt)) = (P0 k(α)+ Pk(α)t) exp(αt),with Dthe derivative operator (with respect to the variable t), and αa real number. Proof. (a), (b) Note that if k > 2, Pkis a polynomial of degree k, which has a local extremum (maximum or minimum) in each one of the open intervals (j, j + 1), j= 0, . . . , k −2, in other words, P0 kis a polynomial of degree k−1 with k−1 simple real roots. From this fact (a) and (b) are immediate. (c) As Pk(z) is continuous, strictly increasing in (k−1,+∞), Pk(k−1) = 0, and limz→∞ Pk(z) = +∞, then, for any α > 0,there is a unique (simple) real root in (k−1,+∞) of the equation Pk(z) = α. In particular, consider α=b−1Pk(n−i)>0. Note that, if 0 < b < 1 then, Pk(n−i)<1 bPk(n−i) , therefore the root is in the interval (n−i, +∞). If b > 1, we have, 0 = Pk(k−1) <1 bPk(n−i)< Pk(n−i), and therefore the root is in (k−1, n −i). If b= 1, as k−1< n −i, then 0 = Pk(k−1) < Pk(n−i), thus the root is in (k−1,+∞). (d) The proof follows easily by using induction. Lemma 2.5. Consider the functions, Q1(v) = (1 −v)r−(n−i), Q2(v) = (1 −v)r−(n−i)log(1 −v), Q3(v) = (1 −v)r−(n−i)log(1 −v) cos(slog(1 −v)), s 6= 0, Q4(v) = (1 −v)r−(n−i)log(1 −v) sin(slog(1 −v)), s 6= 0. If r≤k−1, then Z1 u Qj(v)(v−u)k−1(1 −v)n−i−kdv, j = 1,2,3,4. . . (2.8) is not convergent for any u∈(0,1). Proof. The proof is straightforward by using the classical criteria for the convergence of integrals. 3. Linear Regression of Order Statistics In this section, we characterize the distributions for which the regression of two order statistics is linear. Our main result is presented in the following theorem.
316 fernando l´ opez bl´ aquez and j. luis moreno rebollo Theorem 3.1. Let X be a r.v. with distribution function Fwhich is ktimes differentiable in DF, such that EX(i+k:n)|X(i:n)=bX(i:n)+a. . . . (3.1) Then, except for location and scale parameters, F(x) = 1− | x|δ,for x∈[−1,0] ,if 0< b < 1. . . (3.2) F(x) = 1 −exp(−x),for x∈[0,∞),if b= 1 . . . (3.3) F(x) = 1 −xδ,for x∈[1,∞),if b > 1. . . (3.4) where, δ= (r−(n−i))−1and r is the unique real root greater than k−1of the polynomial equation Pk(x) = 1 bPk(n−i). . . . (3.5) Proof. Let Fbe a c.d.f. for which (3.1) holds and Qits quantile function. Expression (3.1) can be rewritten as Ci,k,n Z1 u Q(v)(v−u)k−1(1 −v)n−i−kdv = (bQ(u) + a)(1 −u)n−i. . . . (3.6) From Lemma 2.3(a), Fis strictly increasing in DF, then f(x) = F0(x)>0, for x∈DF, and as Fis k-times differentiable, it follows that Q=F−1is also k times differentiable in (0,1). Differentiating ktimes both sides of (3.6), we obtain the ordinary differential equation, (1 −u)kH(k)(u) = (−1)k1 bPk(n−i)H(u)−a(1 −u)n−i, . . . (3.7) with, H(u) = Q(u)(1 −u)n−i. The change of variables t= log(1 −u) transforms (3.7) into the linear differential equation Pk(D)(G(t)) = 1 bPk(n−i){G(t)−aexp(t(n−i))}, . . . (3.8) where, Dis the derivative operator and G(t) = H(1 −et). Let us distinguish two cases: b6= 1 and b= 1. Case A: b6= 1. According to Lemma 2.3(c), we can assume, w.l.o.g., that a= 0. For obtaining the general solution of (3.8), we must solve the associated polynomial equation (3.5).
regression of order statistics and record values 317 Let r1,· · · , rsand rs+1,· · · , rdbe the simple and double real roots of (3.5) respectively, and zh=rh+ish,h > d, (sh6= 0) the (simple) complex roots. Then, the complete set of solutions of the homogeneous linear differential equation (3.8) is the linear space generated by {exp (rht), h = 1, . . . , s;texp (rht), h =s+ 1, . . . , d; texp (rht) cos (sht), t exp (rht) sin (sht), h > d}. Hence, if Qis the quantile function of a r.v. Xfor which (3.1) holds, it must satisfy the following conditions: (i) Qbelongs to the linear space of functions generated by (1 −u)rh−(n−i), h = 1, . . . , s;. . . (3.9) (1 −u)rh−(n−i)log(1 −u), h =s+ 1, . . . , d;. . . (3.10) (1 −u)rh−(n−i)log(1 −u) cos(shlog(1 −u)), h > d;. . . (3.11) (1 −u)rh−(n−i)log(1 −u) sin(shlog(1 −u)), h > d. . . . (3.12) (ii) The conditional expectation E[X(i+k:n)|X(i:n)] must exist, or equivalently, the integral Z1 u Q(v)(v−u)k−1(1 −v)n−i−kdv . . . (3.13) must exist, for any u∈(0,1). (iii) Qis monotone strictly increasing. Let us analyze the functions in the basis of the linear space of solutions described in (i). The functions (3.11) and (3.12) change their signs infinitely many times in a neighborhood of u= 1, so that they are not monotone, in other words, they do not satisfy (iii). From Lemma 2.5, the integral (3.13) does not converge for the functions of the form (3.9) or (3.10) which have rh≤k−1. But, according to Lemma 2.4(c), there exist only one simple root of (3.5) greater than k−1, namely r, then the quantile functions which are solutions of our problem are of the form Q(u) = A(1 −u)1 δ, . . . (3.14) where, δ= (r−(n−i))−1and Aa real number chosen in such a way that (iii) holds, that is to say, A/δ < 0. We have two subcases. Firstly, if b < 1, Lemma 2.4(c.1) implies that δ > 0, and w.l.o.g., it can be assumed A=−1. In this case, it follows easily that Q is the quantile function of (3.2). Secondly, if b > 1, Lemma 2.4(c.2) implies that δ < 0, and choosing A= 1 we obtain the quantile of the c.d.f. (3.4).
318 fernando l´ opez bl´ aquez and j. luis moreno rebollo Case B: b= 1. From Lemma 2.3(d), ais positive. We must solve the nonhomogeneous linear differential equation (3.8). It can be shown, using Lemma 2.4(d), that a particular solution of (3.8) is G0(t) = −aPk(n−i) P0 k(n−i)texp((n−i)t) The general solution of the homogeneous linear differential equation associated to (3.8) is obtained by solving the polynomial equation Pk(z) = Pk(n−i). A similar argument to the one used in CASE A shows that the quantile functions which are solutions of our problem are of the form Q(u) = −aPk(n−i) P0 k(n−i)log(1 −u) + A, and, w.l.o.g., choosing A= 0 and a=P0 k(n−i) Pk(n−i), we obtain the quantile function of the c.d.f. given in (3.3). We obtain the following dual result easily. Theorem 3.2 Let Xbe a r.v. with distribution function Fwhich is k-times differentiable in DF, such that EX(i:n)|X(i+k:n)=cX(i+k:n)+d, . . . (3.15) then, except for location and scale parameters, F(x) = xθ,for x∈[0,1] ,if 0< c < 1. . . (3.16) F(x) = exp(x),for x∈(−∞,0],if c= 1 . . . (3.17) F(x) =|x|θ,for x∈(−∞,1],if c > 1. . . (3.18) where, θ= (r−(i+k−1))−1and ris the unique real root greater than k−1of the polynomial equation Pk(x) = 1 cPk(i+k−1). Proof. Let Xbe a r.v. for which (3.15) holds, then the r.v. Y=−X satisfies EY(n−i+1:n)|Y(n−i−k+1) =y=−(c(−y) + d) = cy −d, thus the c.d.f. of Ybelongs to one of the three families described in Theorem 3.1.
regression of order statistics and record values 319 Note that, depending on the values of the slopes, we characterize in Theorem 3.2 three families of distributions which are just the same as those obtained by Ferguson (1967) in the adjacent case. Combining the results stated in Theorems 3.1 and 3.2 gives the following characterizations of uniform distributions. Corollary 3.1. If Fis ktimes differentiable, and for certain integers i,n, such that 1≤i < i +k≤n, the conditional expectations EX(i+k:n)|X(i:n)and EX(i:n)|X(i+k:n)are both linear, then Fis the c.d.f. of a uniform distribution on a finite interval. Corollary 3.2. Let Fbe a c.d.f. which is (n−1)-times differentiable. F is uniform on a finite interval iff EX(i:n)|X(j:n)is linear for any i, j such that 1≤i, j ≤n. 4. Linear Regression of Record Values Let {Xn}n≥1be a sequence of i.i.d. r.v.’s with common continuous distribution function F. The sequence of record times, {L(n)}n≥0, is defined recursively as L(0) = 1, L(n) = min{j:j > L(n−1) and Xj> XL(n−1)}, n > 1, and the sequence of record values is defined as {XL(n)}n≥0. If Q, the quantile function of F, is strictly increasing and Wfollows a standard negative exponential distribution then (XL(0), . . . , XL(j))d ≡(Q(1 −exp(−WL(0))), . . . , Q(1 −exp(−WL(j)))), and the joint density of WL(i)and WL(i+k), with 0 ≤i < i +kis fi,k(w, s) = Di,kwi(s−w)k−1exp(−s) if 0 < w < s < ∞ = 0,otherwise . . . (4.1) with Di,k =1 i!(k−1)! Theorem 4.1. Let Fbe a strictly increasing c.d.f. which is k-times differentiable, such that EXL(i+k)|XL(i)=bXL(i)+a, . . . (4.2)