scieee AI-readable full text Open interactive document viewer

Construction of multivariate distributions : a review of some recent results

Sarabia, José María; Gómez-Déniz, Emilio

Abstract

The construction of multivariate distributions is an active field of research in theoretical and applied statistics. In this paper some recent developments in this field are reviewed. Specifically, we study and review the following set of methods: (a) Construction of multivariate distributions based on order statistics, (b) Methods based on mixtures, (c) Conditionally specified distributions, (d) Multivariate skew distributions, (e) Distributions based on the method of the variables in common and (f) Other methods, which include multivariate weighted distributions, vines and multivariate Zipf distributions.

Full text

Statistics & Operations Research Transactions SORT 32 (1) January-June 2008, 3-36 Statistics & Operations Research Transactions Construction of multivariate distributions: a review of some recent results c Institut d’Estad´ ıstica de Catalunya [email protected] ISSN: 1696-2281 www.idescat.net/sort Jos´ e Mar´ ıa Sarabia1and Emilio G´ omez-D´ eniz2 Abstract The construction of multivariate distributions is an active field of research in theoretical and applied statistics. In this paper some recent developments in this field are reviewed. Specifically, we study and review the following set of methods: (a) Construction of multivariate distributions based on order statistics, (b) Methods based on mixtures, (c) Conditionally specified distributions, (d) Multivariate skew distributions, (e) Distributions based on the method of the variables in common and (f) Other methods, which include multivariate weighted distributions, vines and multivariate Zipf distributions. MSC: 60E05, 62E15, 62H10 Keywords: Order statistics, Rosenblatt transformation, mixture distributions, conditionally specified distributions, skew distributions, variables in common, multivariate weighted distributions, vines, multivariate Zipf distributions, associated random variables 1 Introduction The construction, study and applications of multivariate distributions is one of the classical fields of research in statistics, and it continues to be an active field of research. In recent years several books containing theory about multivariate nonnormal distributions have been published: Hutchinson and Lai (1990), Joe (1997), Arnold, Castillo and Sarabia (1999), Kotz, Balakrishnan and Johnson (2000), Kotz and Nadarajah (2004), Nelsen (2006). In the discrete case specifically, we cannot ignore the books of Kocherlakota and Kocherlakota (1992) and Johnson, Kotz and Balakrishnan (1997) and the review papers by Balakrishnan (2004, 2005). 1Department of Economics, University of Cantabria, Santander, Spain. 2Department of Quantitative Methods, University of Las Palmas de Gran Canaria, Gran Canaria, Spain. Received: November 2007 4 Construction of multivariate distributions: a review of some recent results In this paper some recent methods for constructing multivariate distributions are reviewed. Reviews on constructions of discrete and continuous bivariate distributions are given by Lai (2004 and 2006). One of the problems of this work is the impossibility of producing a standard set of criteria that can always be applied to produce a unique distribution which could unequivocally be called the multivariate version (Kemp and Papageorgiou, 1982). In this sense, there is no satisfactory unified scheme of classifying these methods. In the bivariate continuous case Lai (2004 and 2006) has considered the following clusters of methods, •Marginal transformation method •Methods of construction of copulas •Mixing and compounding •Variables in common and trivariate reduction techniques •Conditionally specified distributions •Marginal replacement •Geometric approach •Constructions of extreme-value models •Limits of discrete distributions •Some classical methods •Distributions with a given variance-covariance matrix •Transformations Some of these methods have merited considerable attention in the recent literature and they will not be revised here. For instance, a detailed study on the construction of copulas is provided by Nelsen (2006) and also in the review paper by Mikosch (2006). The choice of the methods revised in this paper responds to a general interest and our own research experience. Therefore, and as is obvious, this revision cannot be considered as exhaustive regarding multivariate distributions. The contents of this paper are as follows. In Section 2 we study multivariate distributions based on order statistics. Section 3 reviews methods based on mixtures. Conditionally specified distributions are studied in Section 4. Section 5 reviews multivariate skew distributions. Some recent distributions based on the method of the variables in common are studied in Section 6. Finally, other methods of construction (multivariate weighted distributions, vines and multivariate Zipf distributions.) are briefly commented in Section 7. 2 Multivariate Distributions based on Order Statistics Order statistics and related topics (especially extreme value theory) have received a lot of attention recently, see Arnold et al. (1992), Castillo et al. (2005), David and Nagaraja Jos´ e Mar´ ıa Sarabia and Emilio G´ omez-D´ eniz 5 (2003) and Ahsanullah and Nevzorov (2005). In this section we review multivariate distributions beginning with the idea of order statistics. 2.1 An extension of the multivariate distribution of subsets of order statistics Let X1,...,Xnbe a sample of size ndrawn from a common probability density function (pdf) f(x) and cumulative distribution function (cdf) F(x), and let X1:n≤ ··· ≤ Xn:n denote the corresponding order statistics. Now, let Xn1:n,...,Xnp:nbe a subset of porder statistics, where 1 ≤n1<···<np≤n,p=1,2,...,n. The joint pdf of Xn1:n,...,Xnp:nis n! Qp+1 j=1(nj−nj−1−1)!        p Y j=1 f(xj)       p+1 Y j=1F(xj)−F(xj−1)nj−nj−1−1,(1) for x1≤ ··· ≤ xp, where x0=−∞,xp+1= +∞,n0=0 and np+1=n+1. Beginning with the idea of fractional order statistics (Stigler (1977) and Papadatos (1995)) Jones and Larsen (2004) proposed generalizing (1) by considering real numbers a1,...,ap+1>0 instead of integers n1,...,np+1, to obtain the joint pdf, gF(x1,...,xp)=Γ(a1+···+ap+1) Qp+1 j=1Γ(aj)       p Y j=1 f(xj)       p+1 Y j=1F(xj)−F(xj−1)aj−1,(2) on −∞ =x0≤x1≤ ··· ≤ xp≤xp+1=∞. Two particular cases merit our attention. If F∼ U[0,1] is a uniform distribution on [0,1], (2) becomes gU(u1,...,up)=Γ(a1+···+ap+1) Qp+1 j=1Γ(aj) p+1 Y j=1 (uj−uj−1)aj−1,(3) defined on 0 =u0≤u1≤ ··· ≤ up≤up+1=1, which is the generalization of uniform order statistics. If {Uj},j=1,2,...,pis distributed as (3), then {Xj=F−1(Uj)}, j=1,2,...,pis distributed as (1). Another important relation is obtained from the Dirichlet distribution. Let (V1,...,Vk) a Dirichlet distribution with joint pdf Γ(a1+···+ap+1) Qp+1 j=1Γ(aj) p Y j=1 vak−1 p 1− p X j=1 vj ap+1−1 ,(4) defined on the simplex vj≥0, j=1,2,...,p,v1+···+vp≤1. In this case Ui=V1+···+Vi,i=1,2,...,pand Xi=F−1(V1+···+Vi), i=1,2,...,p. 6 Construction of multivariate distributions: a review of some recent results In the univariate case, family (2) becomes gF(x)=Γ(a1+a2) Γ(a1)Γ(a2)f(x)Fa1−1(x)[1 −F(x)]a2−1,(5) which was also proposed by Jones (2004) and it is a generalization of the r-order statistics. The idea of this author is to begin with a symmetric distribution f(a1=a2=1 in (5)) and enlarge this family with parameters a1and a2, controlling skewness and tail weight. If B∼ Be(a1,a2) is a beta distribution with parameters a1and a2, family (5) can be obtained by the simple transformation X=F−1(B). As a last comment in this section, we mention the concept of generalized order statistics introduced by Kamps (1995), as a unified model for ordered random variables, which includes among others the usual order statistics, record values and k-record values as special cases. 2.1.1 An example with the normal distribution In this section we include an example with the normal distribution. If p=2, F= Φ, f=φ, where Φand φare the cdf and the pdf of the standard normal distribution, respectively, general expression (2) becomes gΦ(x,y;a)=Γ(a1+a2+a3) Γ(a1)Γ(a2)Γ(a3)φ(x)φ(y)[Φ(x)]a1−1[Φ(y)−Φ(x)]a2−1[1 −Φ(y)]a3−1,(6) on x<y, and a1,a2,a3>0. Both marginals distributions Xand Yare like (5) with parameters (a1,a2+a3) and (a1+a2,a3), respectively. The local dependence function is given by γ(x,y)=∂2loggΦ(x,y,a) ∂x∂y=(a2−1)φ(x)φ(y) [Φ(y)−Φ(x)]2, if x<y. Figure 1 shows two examples of the bivariate distribution (6). 2.2 Multivariate distribution involving the Rosenblatt construction As a multivariate version of Jones’ (2004) univariate construction defined in equation (5), Arnold, Castillo and Sarabia (2006) have proposed multivariate distributions based on an enriching process using a representation of a p-dimensional random vector with a given distribution due to Rosenblatt (1952). Consider an initial family of p-dimensional joint distribution functions F(x1,...,xp). We assume that these Jos´ e Mar´ ıa Sarabia and Emilio G´ omez-D´ eniz 7 Figure 1: Joint pdf and contour plots of bivariate ordered normal distribution (6) with a1=2, a2=3and a3=4(above) and a1=0.2, a2=3and a3=2(below). distributions are absolutely continuous with respect to Lebesgue measure on Rpand we will be denoted by f(x1,...,xp), the corresponding joint density function. We next introduce the notation F1(x1)=Pr(X1≤x1), F2(x2|x1)=Pr(X2≤x2|X1=x1), . . . Fp(xp|x1,...,xp−1)=Pr(Xp≤xp|X1=x1,...,Xp−1=xp−1), 8 Construction of multivariate distributions: a review of some recent results and associated conditional densities f1(x1),f2(x2|x1),..., fp(xp|x1,...,xp−1).(7) In the spirit of the Rosenblatt transformation, these authors proposed the multivariate distribution of X=(X1,...,Xp) defined by X1=F−1 1(V1), X2=F−1 2(V1|X1), . . . Xp=F−1 p(Vp|X1,...,Xp−1), where V1,...,Vprepresent independent beta distributions Vi∼ Be(ai,bi). The resulting joint density for Xis that given by g(x1,...,xp;a,b)=f(x1,...,xp) p Y i=1 fBe(ai,bi)(Fi(xi|x1,...,xi−1)),(8) where fBe(ai,bi),i=1,...,pdenotes the density of a beta random variable with parameters (ai,bi). It is clear from (8) that the initial joint density f(x1,...,xp) is included in (8) as a special case setting ai=bi=1, i=1,...,p. All the conditional densities (7) are of the form (5). The proposed method is quite general, and several new p-dimensional parametric families have been proposed, including: Frank-beta distribution, the FarlieGumbel-Morgenstern-beta family, the normal-beta family, the Dirichlet-beta family and thePareto-beta family. The familiesof distributionsobtained inthisway areveryflexible and easy to estimate. Details can be found in Arnold, Castillo and Sarabia (2006). 3 Methods Based on Mixtures The use of mixtures to obtain flexible families of densities has a long history, especially in the univariate case. The advantages of the mixtures mechanism are diverse. The new classes of distributions obtained by mixing are more flexible than the original, overdispersed with tails larger than the original distribution and often providing betterfits. The extension of a mixture to the multivariate case is usually simple, and the marginal distributions belong to the same family. On the other hand, simulation and Bayesian estimation of mixtures are quite direct. Since the introduction of simulation-based methods for inference (particularly the Gibbs sampler in a Bayesian framework), complicated densities such as those having mixture representation have been satisfactorily handled. Jos´ e Mar´ ıa Sarabia and Emilio G´ omez-D´ eniz 9 3.1 Common mixture models In this model we assume conditional independence among components and a common parameter shared by all components. The joint cdf is given by FX(x)=Z∞ −∞        p Y k=1 Fk(xk|θ)       dFΘ(θ),x=(x1,...,xp),(9) or in terms of joint densities, fX(x)=Z∞ −∞        p Y k=1 fk(xk|θ)       dFΘ(θ),x=(x1,...,xp).(10) In these models θacts as a frailty parameter. In the joint cdf (9), if each component Fk(xk|θ) is stochastically increasing in θ, then X1,...,Xpare associated and cov(u(X1,...,Xp),v(X1,...,Xp)) ≥0,(11) for all increasing functions u,vfor which the covariance exists. 3.2 A more general model In this situation we have the general models, FX(x)=ZRp       p Y k=1 Fk(xk|θk)       dFΘ(θ),x=(x1,...,xp),(12) or in terms of joint densities, fX(x)=ZRp       p Y k=1 fk(xk|θk)       dFΘ(θ),x=(x1,...,xp),(13) where θ=(θ1,...,θp). In the following sections we include some recent multivariate distributions proposed in the literature obtained by using previous formulations. 3.3 Multivariate discrete distributions The study of the variability of multivariate counts arises in many practical situations. In ecology the counts may be the different species of animals in different geographical 10 Construction of multivariate distributions: a review of some recent results areas whilst in insurance, the number of claims of different policyholders in the portfolio. 3.3.1 Multivariate Poisson-lognormal distribution The multivariate Poisson-lognormal distribution (Aitchison and Ho, 1989) is one of the most relevant models. This distribution is the mixture of conditional independent Poisson distributions, where the mean vector θ=(θ1,...,θp) follows a p-dimensional lognormal distribution with probability density function, g(θ;µ, Σ)=(2π)−p/2(θ1···θp)−1|Σ|−1/2exp"−1 2(logθ−µ)⊤Σ−1(logθ−µ)#. The probability mass function is given by (formula (13)): Pr(X1=x1,...,Xp=xp)=ZRp + p Y i=1 f(xi|θi)g(θ;µ, Σ)dθ, x1,...,xp=0,1,..., (14) where Rp +denotes the positive orthant of p-dimensional real space Rp. Although it is not possible to obtain a closed expression for the probability mass function in (14), its moments can be easily obtained by conditioning E(Xi)=exp µi+1 2σii!=αi,(15) Var(Xi)=αi+α2 iexp(σii)−1,(16) cov(Xi,Xj)=αiαjexp(σii)−1,i,j=1,...,p,i,j.(17) From (15) and (16) it is obvious than the marginal distributions are overdispersed and from (17) the model admits both negative and positive correlations. Other versions of this model can be viewed in Tonda (2005). 3.3.2 Multivariate Poisson-generalized inverse Gaussian distribution Departing from the Sichel distribution (Poisson-generalized inverse Gaussian distribution) in Sichel (1971) investigated bivariate extensions of that distribution. In Stein et al. (1987) one of them is studied in order to obtain the estimation of the parameters via the likelihood method. Conditionally given λ, the Xiare independently distributed as Poisson random variables with parameter λξi, where ξiis a scale factor (i=1,...,p) and the parameter Jos´ e Mar´ ıa Sarabia and Emilio G´ omez-D´ eniz 11 λfollows a generalized inverse Gaussian distribution with parameters 1,w>0 and γ∈R. The multivariate mixture is: Pr(X1=x1,...,Xp=xp)=KPxi+yw1/2(w+Pξi)1/2 Kγ(w) × w w+2Pξi!(Pxi+y)/2p Y i=1 ξxi i xi!, x1,...,xp=0,1, . . . , α, ξ1,...ξp>0,−∞ < γ < ∞, where Kν(z) denotes the modified Bessel function of the second kind of order νand argument z. The basic moments and the correlation matrix are E(Xi)=ξiRγ(w), Var(Xi)=ξ2 iKγ+2(w)/Kγ(w)+E(Xi)[1 −E(Xi)], corr(Xi,Xj)=(1+1/ξig(w, γ))(1 +1/ξjg(w, γ)−1/2, i,j=1,...,p,i,j, where Rν(z)=Kν+1(z)/Kν(z) and g(w, γ)=Rγ+1(w)−Rγ(w). The correlations between marginals are positive. 3.3.3 Multivariate negative binomial-inverse Gaussian distribution G´ omez-D´ eniz et al. (2008) have considered a new distribution by mixing a negative binomial distribution with an inverse Gaussian distribution, where the parameterization ˆp=exp(−λ) was considered. This new formulation provides a tractable model with attractive properties, which makes it suitable for application in disciplines where overdispersion is observed. The multivariate negative binomial-inverse Gaussian distribution can be considered as the mixture of independent NB(ri,ˆp=e−λ), i=1,2,...,pcombined with an inverse Gaussian distribution for λ. The joint probability mass function is given by (formula (10)), Pr(X1=x1,X2=x2,...,Xp=xp) = p Y i=1 ri+xi−1 xi!˜x X j=0 (−1)j ˜x j!exp         ψ µ 1−s1+2(˜r+j)µ2 ψ         , where x1,x2,...,xp=0,1,2,...;µ, ψ, r1,...,rp>0 and r=r1+···+rp,˜x=x1+···+xp and the moments, 18 Construction of multivariate distributions: a review of some recent results coefficient range ρ(X,Y)∈(−1,0). The marginal distributions of (28) are Pr(X=x)=kλx 1 x!exp(λ2λx 3),x=0,1,2,... Pr(Y=y)=kλy 2 y!exp(λ1λy 3),y=0,1,2,..., which are not Poisson except in the independence case. Wesolowski (1996) has characterized this distribution using a conditional distribution and the other conditional expectation. The second example corresponds to the normal case. Again, using Theorem 2, the most general bivariate distribution with normal conditionals is given by fX,Y(x,y;m)=exp           (1,x,x2) m00 m01 m02 m10 m11 m12 m20 m21 m22  1 y y2           .(29) Distributions with densities of form (29) are called normal conditional distributions. Note that (29) is an eight parameter family of densities, and m00 is the normalizing constant. The conditional expectations and variances are: E(Y|X=x)=−m01 +m11x+m21x2 2(m02 +m12x+m22x2),(30) Var(Y|X=x)=−1 2(m02 +m12x+m22x2),(31) E(X|Y=y)=−m10 +m11y+m12y2 2(m20 +m21y+m22y2),(32) Var(X|Y=y)=−1 2(m20 +m21y+m22y2).(33) The normal conditional distributions give rise to models where the mij constants satisfy one of the two sets of conditions (a) m22 =m12 =m21 =0; m20 <0; m02 <0; m2 11 <4m02m20. (b) m22 <0; 4m22m02 >m2 12; 4m20m22 >m2 21. Models satisfying conditions (a) are the classical bivariate normal models with normal marginals and conditionals, linear regressions and constant conditional variances. More interesting are the models satisfying conditions (b). These models have normal conditional distributions, non-normal marginals, and the regression functions are either constant or non-linear given by (30) and (32). Each regression function is Jos´ e Mar´ ıa Sarabia and Emilio G´ omez-D´ eniz 19 Figure 2: Joint pdf and contour plots of two bivariate distributions with normal conditionals. bounded (in contrast with the bivariate normal model) and the conditional variance functions are also bounded and non constant. They are given by (31) and (33). Another unexpected property of (29) is the multimodality, where one, two and three modes are possible (see Arnold et al., 2000). Figure 2 presents two models of kind (b), one unimodal and the other one bimodal. 4.4 Multivariate extensions Previous results can be extended to higher dimensions. Technical details appear in Arnold, Castillo and Sarabia (1999 and 2001). As an important model, we consider 20 Construction of multivariate distributions: a review of some recent results the case in which the families of conditional densities Xi|X(i)=x(i),i=1,...,p are exponential families, where X(i)denotes the p-dimensional vector Xwith the ith coordinate deleted. In this situation, the most general joint density with exponential conditionals must be of the form fX(x)= p Y i=1 ri(xi)exp         ℓ1 X i1=0 ℓ2 X i2=0··· ℓp X ip=0 mi1,i2,...,ip p Y i=1 qiij(xj)         . For example, the p-dimensional distribution with normal conditionals is of the form fX(x)=exp        X i∈Tp mi p Y i=1 xij i         ,(34) where Tpis the set of all vectors of 0’s, 1’s and 2’s of dimension p. Densities of the form (34) have normal conditional densities for Xigiven X(i)=x(i)for every x(i),i=1,...,p. The classical p-variate normal density is a special case of (34). 4.5 Applications of the conditionally specified models Applications of these conditional models are contained in the book by Arnold, Castillo and Sarabia (1999). These applications include modelling of bivariate extremes, conditional survival models, multivariate binary response models with covariates (Joe and Liu, 1996) and Bayesian analysis using conditionally specified models. The use of this kind of distribution in risk analysis and economics in general is quite recent. Some applications have been provided by Sarabia, G´ omez and V´ azquez (2004) and Sarabia et al. (2005). The class of bivariate income distribution with lognormal conditionals has been studied by Sarabia et al. (2007). In the risk theory context, Sarabia and Guill´ en (2008) have proposed flexible bivariate joint distributions for modelling the couple (S,N), where Nis a count variable and S=X1+···+XNis the total claim amount. 5 Multivariate Skew Distributions The skew-normal (SN) distribution, its different variants and their corresponding multivariate versions, have received considerable attention over the last few years. Two recent reviews of these classes appear in the book edited by Genton (2004) and the paper by Azzalini (2005). To introduce the multivariate version it is necessary to know the univariate case and its properties, which the multivariate version is based on. A random variable Xis said to have a skew-normal distribution with parameter λ, if the probability Jos´ e Mar´ ıa Sarabia and Emilio G´ omez-D´ eniz 21 density function is given by f(x;λ)=2φ(x)Φ(λx),−∞ <x<∞.(35) A random variable with pdf (35) will be denoted as X∼ SN(λ). Parameter λcontrols the skewness of the distribution and varies in (−∞,∞). The linearly skewed version of this distribution is given by, f(x;λ0, λ1)∝φ(x)Φ(λ0+λ1x),−∞ <x<∞,(36) and λ0, λ1∈R, which we will denote by X∼ SN(λ0, λ1). The next two properties hold for distribution (35) and allow us to understand multivariate extensions: •Hidden truncation mechanism. Let (X0,X1) be a bivariate normal distribution with standardized marginals and correlation coefficient δ. Then, the variable, {X1|X0>0}(37) is distributed as a SN(λ(δ)) distribution, where λ(δ)=δ √1−δ2. •Convolution representation. If X0and X1are independent N(0,1) random variables, and −1< δ < 1, then Z=δ|X0|+(1 −δ2)1/2X1(38) is a SN(λ(δ)). A general treatment of the hidden truncation mechanism (37) is found in Arnold and Beaver (2002). A multivariate version of the basic model (35) has been considered by Azzalini and Dalla Valle (1996) and Azzalini and Capitanio (1999). This multivariate version of the SN distribution is defined as f(x)=2φp(x−µ;Σ)Φ(α⊤w−1(x−µ)),(39) where φp(x−µ;Σ) is the joint pdf of a multivariate normal distribution Np(µ, Σ), µ∈Rp is a location parameter, Σis a positive definite covariance matrix, α∈Rpis a parameter whichcontrolsskewness andwisa diagonal matrixcomposed by thestandard deviations of Σ. If we set α=0 in (39), we obtain a classical Np(µ, Σ) distribution. Similar to the univariate case, we can obtain (39) using a hidden truncation mechanism (37) and convolution representation (38). 22 Construction of multivariate distributions: a review of some recent results Let X0and X1be random variables of dimensions 1 and psuch that X0 X1!∼ N1+p(0,Σ∗),Σ∗= 1δ⊤ δ˜ Σ!, where ˜ Σis a correlation matrix and δ=(1 +α⊤˜ Σα)−1/2˜ Σα. Then, the p-dimensional random variable Z={X1|X0>0}, has the joint pdf f(z)=2φp(z;˜ Σ)Φ(α⊤z),(40) which is an affine transformation of (39). For the convolution representation let X0∼ N(0,1) and X1∼ Np(0,R) be independent random variables, where Ris a correlation matrix. Let ∆ = diag{δ1,...,δp}, −1< δj<1, j=1,2,...,pand Ipthe identity matrix of order pand 1pthe pdimensional vector of all 1s. Then, Z= ∆1p|X0|+(Ip−∆2)1/2X1, is distributed in the form (40). The relationship between (R,∆) and (˜ Σ, α) can be found in Azzalini and Capitanio (1999). 5.1 An alternative multivariate skew normal distribution An alternative class of multivariate-normal distributions was considered by Gupta, Gonz´ alez-Far´ ıas and Dom´ ınguez-Molina (2004). Previous multivariate versions (40) were obtained by conditioning that one random sample be positive; these authors condition that the same number of random variables be positive and then, in the univariate case both families are the same. A random vector of dimension pis said to have a multivariate skew normal distribution (according to Gupta et al., 2004) if its pdf is given by fp(x;µ, Σ,D)=φp(x;µ, Σ)Φp(D(x−µ);0,I) Φp(0;0,I+DΣD⊤),(41) Jos´ e Mar´ ıa Sarabia and Emilio G´ omez-D´ eniz 23 where µ∈Rp,Σ>0, D(p×p), and φp(·;ξ, Ω) and Φp(·;ξ, Ω) denote the pdf and the cdf, respectively of a Np(ξ, Ω) distribution. As an extension to (41), Gonz´ alez-Far´ ıas et al. (2003, 2004) introduced the closed skew-normal family of distributions. This family is closed under conditioning, linear transformations and convolutions. It is defined as fp(x;µ, Σ,D, ν, ∆)=φp(x;µ, Σ)Φq(D(x−µ);ν, ∆) Φq(0q;ν;∆ + DΣD⊤),(42) where x, µ, ν ∈Rp,Σ∈Rp×Rp,D∈Rq×Rp,∆∈Rq×Rqand Σand ∆are positive definite matrices and 0q=(0,...,0) ∈Rq. The closed skew-normal distributions can be generated by conditioning the first components of a normal random vector in the event that the remaining components are greater than certain given values. 5.2 Conditional specification Arnold, Castillo and Sarabia (2002) have discussed the problem of identifying pdimensional densities with skew-normal conditionals. They address the question of identifying joint densities for a p-dimensional random vector Xthat has the property that for each x(i)∈Rp−1we have Xi|X(i)=x(i)∼ SN(λ(i) 0(x(i)), λ(i) 1(x(i))),i=1,2,...,p.(43) An important parametric family of densities takes the form f(x1,...,xp;λ)∝ p Y i=1 φ(xi)ΦX s∈Sp λs p Y i=1 xsi i ,(44) whereSpdenotes the set of all vectors of 0’s and 1’s of dimension p. Inthe bivariate case, we obtain the following bivariate distribution with linearly skewed-normal conditionals, f(x,y;λ)∝φ(x)φ(y)Φ(λ00 +λ10x+λ01y+λ11xy).(45) Note (45) does not belong to class (39) except when λ00 =λ11 =0. The normalizing constant is complicated in general, except when λ00 =λ10 =λ10 =0, in which is equals 2 and the density is explicitly given by f(x,y;λ)=2φ(x)φ(y)Φ(λxy).(46) This model has normal marginals and skew-normal conditionals of type (35), and bimodality is possible. 24 Construction of multivariate distributions: a review of some recent results Model (44) can be viewed as a generalized hidden truncation model, defining ˜ X=(X0,X1,...,Xp), with Xi’s i.i.d. N(0,1), in which we retain only those ˜ Xfor which X0≤X s∈Sk λs p Y i=1 Xsi i, and the resulting conditional density of (X1,...,Xp) will then be given by (44). More about skew conditionals models can be found in Sarabia (2002) and Arnold, Castillo and Sarabia (2007a, 2007b). 5.3 Balakrishnan skew-normal distribution Balakrishnan (2002) as a discussant of Arnold and Beaver (2002) generalized the SN distribution as fn(x;λ)=φ(x)[Φ(λx)]n cn(λ),x∈R,(47) where nis an integer and cn(λ)=R∞ −∞ φ(x)[Φ(λx)]ndx. This distribution is known as Balakrishnan skew-normal distribution. If we set n=0 and n=1 in (47) the above density reduces to the N(0,1) distribution and the SN distribution, respectively. Gupta and Gupta (2004) have studied some properties of (47). Several multivariate versions are possible. If we think of an extension by conditionals, in the simpler bivariate case, we obtain the joint pdf, fn(x,y;λ)=˜cn(λ)φ(x)φ(y)[Φ(λxy)]n,(x,y)∈R2.(48) For this distribution, both conditionals are like (47), but the marginal distributions are not.Yadegari et al. (2008) have considered the extension of (47) given by fn,m(x;λ)=1 cn,m(λ)[Φ(λx)]n[1 −Φ(λx)]mφ(x),x∈R,(49) where cn,m(λ)=Pm i=0m i(−1)icn+i(λ). A natural extension of (49) to the multivariate case is fn,m(x;λ)=1 cn,m(λ)[Φ(λ⊤x)]n[1 −Φ(λ⊤x)]mφp(x),x∈Rp.(50) For m=0 and n=1 this distribution reduces to the multivariate SN distribution. Jos´ e Mar´ ıa Sarabia and Emilio G´ omez-D´ eniz 25 5.4 Extensions and applications The initial formulation (35) gives rise to an important number of extensions and variants. One of these variants appears replacing the normality assumption with alternative symmetric distribution. An interesting class of skewed densities is provided by the following elementary, but useful, result (Azzalini, 2005). Lemma 1 If f0is a p-dimensional pdf such that f0(x)=f0(−x)for x ∈Rp, G is a onedimensional differentiable cdf such that G′is a density symmetric about zero, and w is real-valued function such that w(−x)=−w(x)for all x ∈Rp, then f(x)=2f0(x)G{w(x)},x∈Rp,(51) is a genuine pdf on Rp. Different choices for f0,Gand win (51) give rise to a huge number of variants of skewed densities. In a more general setting, Wang et al. (2004) have shown that any p-dimensional multivariate pdf g(x) admits for any fixed location parameter λ∈Rpa unique skew-symmetric representation g(x)=2f(x−λ)π(x−λ),x∈Rp,(52) where f:Rp→R+is a symmetric pdf (in the sense of previous lemma) and π:Rp→ [0,1] is a skewing function such that π(−x)=π(x). Conversely, any function gof the kind (52) is a valid pdf. Multivariate distribution such as skew-Cauchy (Arnold and Beaver, 2000), skew-t (Branco and Dey, 2001; Azzalini and Capitanio, 2003) and other skew-elliptical distributions can be represented using previous formulations (51)-(52). Finally, we mention some applications of the distribution described in this section: compositional data, financial market and insurance (Vernic, 2005), selective sampling, stochastic frontier models and modelling of environmental data. 6 The Variables in Common Method This method, also known as “trivariate reduction”, is a popular and old technique used for building dependent variables, both in continuous and discrete cases. Our attention focuses on the bivariate case. The method consists of building a pair of dependent random variables starting from three (or more) random variables. These initial random variables are usually independent. The functions that connect initial variables are generally elementary functions, or are given by the structure of the variables that we want to generate. A broad definition can be 26 Construction of multivariate distributions: a review of some recent results (X=υ1(eX,cXY), Y=υ2(eY,˜cXY), where eX,eYrepresent two sets containing the specific variables of Xand Yrespectively, and cXY,˜cXY sets containing the common or latent variables. According to Marshall and Olkin (2007), many of the couples (X,Y) here presented are associated (formula (11)), and then only positive correlations are possible. Over the last few years, several new dependent distributions using this method have been proposed. All the models presented in this section can be extended to higher dimensions. We present some relevant models. 6.1 Bivariate generalized Poisson distribution Let Xi,i=1,2,3 be mutually independent random variables. An usual trivariate reduction scheme is defined as X=X1+X3, Y=X2+X3. A disadvantage of this model is that only positive correlations are possible. If the Xi’s are discrete, the joint pgf is gX,Y(u,v)=gX1(u)gX2(v)gX3(uv). If the Xiare Poisson random variables, we obtain the classical bivariate Poisson distribution, which is often used for obtaining compound bivariate Poisson distributions. Ifwe consider forthe Xi’s random variables a generalizedPoisson distribution,weobtain the model considered by Vernic (1997, 2000). 6.2 Bivariate beta distribution In a Bayesian context, when we work with independent or correlated binomial distributions, a density defined over {0≤xi≤1;i=1,...,p}on the unit cube is needed. Olkin and Liu (2003) proposed the following method for constructing this distribution. Let Xi∼ G(ai,1), i=1,2,3 be independent gamma variables with unit scale parameters, and define X=X1 X1+X3 , Y=X2 X2+X3 . Jos´ e Mar´ ıa Sarabia and Emilio G´ omez-D´ eniz 27 Now, we have correlated beta distributions Be(a1,a3) and Be(a2,a3) over 0 ≤x,y≤1 with joint pdf, f(x,y;a1,a2,a3)=xa1−1ya2−1(1 −x)a2+a3−1(1 −y)a1+a3−1 B(a1,a2,a3)(1 −xy)a1+a2+a3,(53) where B(a1,a2,a3)=Q3 i=1Γ(ai)/Γ(P3 i=1ai). The bivariate density (53) is positively likelihood ratio dependent and hence positive quadrant dependent. Sarabia and Castillo (2006) have considered a generalization of (53) under a conditional specification. With this specification, they obtain a broad class of distributions, where an important submodel is f(x,y;a1,a2,b1,m)=xa1−1(1 −x)b1−1ya2−1(1 −y)a1+b1−a2−1 n(a1,a2,b1,m)(1 −mxy)a1+b1,(54) where a1,b1,a1+b1−a2>0, m≤1 and where 1/nis the normalizing constant. This model contains the Olkin and Liu (2003) proposal for m=1 and Xis stochastically increasing or decreasing with Y, so, consequently signρ(X,Y)=sign(m). Then, if 0 <m≤1 we have positive correlation and if m<0, negative correlations. The marginal distributions are of the Gauss hypergeometric type. 6.3 Bivariate t distribution The usual bivariate spherically symmetric distribution on n1degrees of freedom is defined as (Fang et al., 1990) X=X1/pX3/n1, Y=X2/pX3/n1, where X1,X2,X3are mutually independent random variables with distributions X1,X2∼ N(0,1) standard normal and X3∼χ2 n1chi-squared distribution on n1degrees of freedom. Note that the marginal distributions are both (dependent) Student tdistributions on n1 degrees of freedom. If we need a bivariate distribution with Student tmarginals with different degrees of freedom ν1and ν2, one possibility is defined, X=X1/pX3/ν1, Y=X2/p(X3+X4)/ν2, 34 Construction of multivariate distributions: a review of some recent results Joe, H. and Liu, Y. (1996). A model for a multivariate binary response with covariates based on conditionally specified logistic regressions. Statistics and Probability Letters, 31, 113-120. Johnson, N. L., Kotz, S. and Balakrishnan, N. (1997). Discrete Multivariate Distributions. John Wiley and Sons, New York. Jones, M. C. (2002). A dependent bivariate tdistribution with marginals on different degrees of freedom. Statistics and Probability Letters, 56, 163-170. Jones, M. C. (2004). Families of distributions arising from distributions of order statistics (with discussion), Test, 13, 1-43. Jones, M. C. and Larsen, P. V. (2004). Multivariate distributions with support above the diagonal. Biometrika, 91, 975-986. Kamps, U. (1995). A concept of generalized order statistics. Journal of Statistical Planning and Inference, 48, 1-23. Katti, S. K. (1966). Interrelations among generalized distributions and their components. Biometrics, 22, 44-52. Kemp, C. D. and Papageorgiou, H. (1982). Bivariate Hermite distributions. Sankhya, Ser. A, 44, 269-280. Kocherlakota, S. and Kocherlakota, K. (1992). Bivariate Discrete Distributions. Dekker, New York. Kotz, S., Balakrishnan, N. and Johnson, N. L. (2000). Continuous Multivariate Distributions, vol. 1: Models and Applications. John Wiley and Sons, New York. Kotz, S. and Nadarajah, S. (2004). Multivariate T-Distributions and Their Applications. Cambridge University Press, Cambridge. Kurowicka, D. and Cooke, R. (2006). Uncertainty Analysis with High Dimensional Dependence Modelling. John Wiley and Sons, New York. Lai, C. D. (2004). Constructions of continuous bivariate distributions. Journal of the Indian Society for Probability and Statistics, 8, 21-43. Lai, C. D. (2006). Constructions of discrete bivariate distributions. In: Advances on Distribution Theory, Order Statistics and Inference, Birkh¨ auser, Boston., N. Balakrishnan, E. Castillo, J.M. Sarabia Editors, 29-58. Lee, M.-L. T. (1996). Properties and applications of the Sarmanov family of bivariate distributions. Communications in Statistics, Theory and Methods, 25, 1207-1222. Marshall, A. W. and Olkin, I. (1967). A multivariate exponential distribution. Journal of the American Statistical Association, 62, 30-44. Marshall, A. W. and Olkin, I. (2007). Life Distrubutions. Structure of Nonparametric, Semiparametric, and Parametric Families. Springer, New York. Mikosch, T. (2006). Copulas: tales and facts. Extremes, 9, 3-20. Navarro, J., Ruiz, J. M. and Del Aguila, Y. (2006). Multivariate weighted distributions: a review and some extensions. Statistics, 40, 51-64. Nelsen, R. B. (2006). An Introduction to Copulas. Second Edition. Springer, New York. Øigard, T. A. and Hanssen, A. (2002). The multivariate normal inverse Gaussian heavy-tailed distribution: simulation and estimation. In: IEEE International Conference on Acoustics, Speech and Signal Processing, 2, 1489-1492. Olkin, I. and Liu, R. (2003). A bivariate beta distribution. Statistics and Probability Letters, 62, 407-412. Papadatos, N. (1995). Intermediate order statistics with applications to nonparametric estimation. Statistics and Probability Letters, 22, 231-238. Protassov, R. S. (2004). EM-based maximum likelihood parameter estimation for multivariate generalized hyperbolic distributions with fixed λ.Statistics and Computing, 14, 67-77. Rao, C. R. (1965). On discrete distributions arising out of methods of ascertainment. Sankhya Series A, 27, 311-324. Rosenblatt, M. (1952). Remarks on a multivariate transformation. The Annals of Mathematical Statistics, 23, 470-472. Jos´ e Mar´ ıa Sarabia and Emilio G´ omez-D´ eniz 35 Sarabia, J. M. (2002). Discussion of “Skewed Multivariate Models Related to Hidden Truncation and/or Selective Reporting”, by B. C. Arnold and R. J. Beaver. Test, 11, 49-52. Sarabia, J. M. and Castillo, E. (2006). Bivariate distributions based on the generalized three-parameter beta distribution. In: Advances on Distribution Theory, Order Statistics and Inference, Birkh¨ auser, Boston., N. Balakrishnan, E. Castillo, J. M. Sarabia Editors, 85-110. Sarabia, J. M., Castillo, E., G´ omez, E. and V´ azquez, F. (2005). A class of conjugate priors for log-normal claims based on conditional specification. Journal of Risk and Insurance, 72, 479-495. Sarabia, J. M., Castillo, E., Pascual, M. and Sarabia, M. (2007). Bivariate income distributions with lognormal conditionals. Journal of Economic Inequality, 5, 371-383. Sarabia, J. M. and G´ omez-D´ eniz, E. (2008). Univariate and multivariate Poisson-beta distributions with applications. Submitted. Sarabia, J. M., G´ omez-D´ eniz, E. and V´ azquez-Polo, F. (2004).On the use of conditional specification models in claim count distributions: an application to bonus-malus systems. ASTIN Bulletin, 34, 85-98. Sarabia, J. M. and Guill´ en, M. (2008). Joint modelling of the total amount and the number of claims by conditionals. Submitted. Sarhan, A. M. and Balakrishnan, N. (2007). A new class of bivariate distributions and its mixture. Journal of Multivariate Analysis, 98, 1508-1527. Sarmanov, O. V. (1966). Generalized normal correlation and two-dimensional Frechet classes. Doklady, Soviet Mathematics, 168, 596-599. Shaw, W. T. and Lee, K. T. A. (2008). Bivariate Student tdistributions with variable marginal degrees of freedom and independence. Journal of Multivariate Analysis, In press. Sichel, H. S. (1971). On a family of discrete distributions particularly suited to represent long-tailed frequency data. In: Proceedings of the third Symposium on Mathematical Statistics, Eds. N.F. Laubscher, Pretoria, South Africa: Council for Scientific and Industrial Research, 51-97. Stein, G., Zucchini, W. and Juritz, J. (1987). Parameter estimation for the Sichel distribution and its multivariate extension. Journal of the American Statistical Association, 82, 938-944. Stigler, S. M. (1977). Fractional order statistics, with applications. Journal of the American Statistical Association, 72, 544-550. Tonda, T. (2005). A class of multivariate discrete distributions based on an approximate density in GLMM. Hiroshima Math. J., 35, 327-349. Vernic, R. (1997). On the bivariate generalized Poisson distribution. ASTIN Bulletin, 27, 23-31. Vernic, R. (2000). A multivariate generalization of the generalized Poisson distribution. ASTIN Bulletin, 30, 57-67. Vernic, R. (2005). Multivariate skew-normal distributions with applications in insurance. Insurance: Mathematics and Economics, 38, 413-426. Walker, S. G. and Stephens, D. A. (1999). A multivariate family on (0,∞)p.Biometrika, 86, 703-709. Wang, J., Boyer, J. and Genton, M. G. (2004). A skew-symmetric representation of multivariate distributions. Statistica Sinica, 14, 1259-1270. Wesolowski, J. (1996). A New conditional specification of the bivariate Poisson conditionals distribution. Statistica Neerlandica, 50, 390-393. Yadegari, I., Gerami, A. and Khaledi, M. J. (2008). A generalization of the Balakrishnan skew-normal distribution. Statistics and Probability Letters, In Press. Yeh, H. C. (2002). Six multivariate Zipf distributions and their related properties. Statistics and Probability Letters, 56, 131-141. Yeh, H. C. (2004). Some properties and characterizations for generalized multivariate Pareto distributions. Journal of Multivariate Analysis, 88, 47-60. Yeh, H. C. (2007). Three general multivariate semi-Pareto distributions and their characterizations. Journal of Multivariate Analysis, 98, 1305-1319. Discussion of “Construction of multivariate distributions: a review of some recent results” by Jos´ e Mar´ ıa Sarabia and Emilio G´ omez-D´ eniz M. C. Pardo Deparment of Statistics and O.R (I), Complutense University of Madrid, 28040-Madrid, SPAIN, Phone: 34 91 3944473, fax: 34 91 3944606. E-mail: [email protected], http://www.mat.ucm.es/˜mcapardo I congratulate the authors on this excellent review. In this review paper they present a nice overview on construction of multivariate distributions. Due to it being such an active field of research, new models are constantly being discovered. However, they have been able to present some of the most recent methods in a very clear manner, and many of those omitted can be found in the references mentioned. It is my pleasure to comment on this article. I agree with professors Sarabia and G´ omez-D´ eniz that is not possible to mention all the methods for constructing distributions that exist. However, owed perhaps to my own field of research, I miss the Maximum Entropy Principle used to construct probability distributions. Therefore, my discussion focuses on presenting the practical usefulness of this method. Let (X, βX,P)be the statistical space associated with the random variable X, where βXis the σ-field of Borel subsets A⊂ X and {P}is a family of probability distributions defined on the measurable space (X,βX).We assume that the probability distributions P are absolutely continuous with respect to σ-finite measure µon (X,βX).The Shannon entropy is defined by H=−ZX f(x)log f(x)dµ(x)(1) where f(x)=dP dµ(x). The Maximum Entropy Principle states that, maximizing entropy subject to a set of constraints can be regarded as deriving a distribution that is consistent with the information specified in the constraints while making minimal assumptions about the form of the distribution other than those embodied in the constraints. Numerous distributions have been obtained in this manner (Kapur, 1994; Ebrahimi, 2000, Asadi et al., 2004). For example, the normal distribution may be obtained as the distribution on the real line that has maximum entropy subject to having specified mean and variance; see Rao (1965, p. 132). An earlier result by Goldman (1955) characterized N0, σ2 as the MED with specified value of EZ2being Za continuous random variable with support (−∞,∞); the MED with specified value of Eh(Z−a)2iwas shown to be Na, σ2by Lisman and van Zuylen (1972). More generally, if a set of moments 40 E[Xr],r=1,...,R,is specified, the distribution on the real line that has the maximum entropy subject to these constraints has probability density f(x)∝exp R X i=1 αrxr for suitable constants αr,r=1,..., R. Hosking (2007) derived the distribution that has maximum entropy conditional on having specified values of its first r Lmoments. Note that L-moments are now widely used in the environmental sciences to summarize data and fit frequency distributions. This maximum entropy distribution has a polynomial density-quantile distribution (PDQ distribution). Some special cases of the PDQ distribution are: On a finite interval, the MED is the uniform distribution; on a semi-infinite interval, the MED with the first L-moments specified is the exponential distribution and on an infinite interval, the MED with the first two L-moments specified is the logistic distribution. Maximum entropy distributions conditional on specified Lmoments of orders {1,2,3}and {1,2,4}generate families of distributions that generalize the logistic distribution and may be useful for modelling data. There is a lot of work devoted to the maximum entropy characterization of the most well-known univariate probability distributions. Although available literature is significantly less for the multivariate distributions, the book of Kapur (1989) considers several usual multivariate distributions and Zografos (1999) considered the cases of Pearson’s types II and VII multivariate distributions (t-distribution and generalized Cauchy distribution are obtained from an application of Pearson’s types VII distribution). Aulogiaris and Zografos (2004) considered symmetric Kotz type and Burr multivariate distributions. Later Bhattacharya (2006) derived appropriate constraints which establish the maximum entropy characterization of the Liouville distributions among all multivariate distributions. Amongst discrete distributions, the geometric distribution with support 1,2, ... is the MED given a specified arithmetic mean. The Riemann zeta distribution, also called the discrete Pareto distribution, is the MED for a specific geometric mean. In linguistics, it is called the Zipf distribution.It has also been used to model numbers of insurance policies, the distribution of surnames and scientific productivity. The Good type-I distribution is the MED when the arithmetic and geometric means are both specified. If x=1,2,...,n and there is no restriction on the probabilities, then the MED is a discrete rectangular distribution. Given finite support and specified arithmetic or geometric means, or both arithmetic and geometric mean, the MEDs are the right-truncated geometric, righttruncated Rienmann zeta, and right-truncated Good type-I distribution, respectively. Kemp (1997) obtained a discrete analogue of the normal distribution as the distribution that is characterized by maximum entropy, specified mean and variance, and integer support on (−∞,∞).Binomial and Poisson distributions are also MEDs of suitable defined sets (Harremo¨ es, 2001). 41 The Maximum Entropy Principle has applications in many domains, but was originally motivated by statistical physics (Jaynes, 1957), which attempts to relate macroscopic measurable properties of physical systems to a description at an atomic or molecular level. Applications in econometrics can be seen in several works of Theil (see for example, Theil and Fiebig, 1984). A popular method for estimation of spectral densities was given by Burg (1975) based on Maximum Entropy method. Many works and books following this idea have appeared, see for instance Golan et al. (1996). In Finance, this principle is applied to infer a probability density from option prices. Buchen and Kelly (1996) showed that, with a set of well-spread simulated exact-option prices, the MED approximates a risk-neutral distribution to a high degree of accuracy. Guo (2001), motivated by the characteristic that a call price is a convex function of the option’s strike price, suggests a simple convex-spline procedure to reduce the impact of noise on observed option prices before inferring the MED. Apart from density estimation, many statistical problems have been studied on the basis of the Maximum Entropy Principle. Using sample quantiles, Men´ endez et al. (1997) proposed a point estimation procedure as well as a goodness-of-fit test statistic based on the he Maximum Entropy Principle. But there are many important different entropy measures (see Chapter 2 of Pardo, 2006), and in a similar manner the Maximum Entropy Principle associated with these others entropies can be defined. Men´ endez et al. (1997) generalized the previous work using a general family of entropies that contains (1). References Asadi, M., Ebrahimi, N., Hamedani, G. G. and Soofi, E. S. (2004). Maximum dynamic Entropy Models. IEEE Transactions on Information Theory, 50, 177-183. Aulogiaris, G. and Zografos, K. (2004). A Maximum Entropy characterization of symmetric Kotz type and multivariate Burr distributions. TEST, 13, 65-83. Bhattacharya, B. (2006). Maximum entropy characterizations of the multivariate Liouville distributions. Journal of Multivariate Analysis, 97, 1272-1283. Ebrahimi, N. (2000). The maximum entropy method for life time distributions. Sankhy¯a, A, 62, 236-243. Burg, J. P. (1975). Maximum Entropy Spectral Analysis. PhD. dissertation, Department of Geophysics, Standford University. Buchen, P. W. and Kelly, M. (1996). The Maximum Entropy Distribution of an Asset Inferred from OPtion Prices. Journal of Financial and Quantitative Analysis, 31, 143-159. Golan, A., Judge, G. G. and Miller, D. (1996). Maximum Entropy Econometrics: Robust Estimation with Limited Data. Wiley, New York. Goldman, S. (1955). Information Theory, Prentice-Hall, New York. Guo, W. (2001). Maximum Entropy in Option Pricing: A convex-Spline Smoothing Method. The Journal of Future Markets, 21, 819-832. Harremo¨ es, P. (2001). Binomial and Poisson Distributions as Maximum Entropy Distributions. IEEE Transactions on Information Theory, 47, 2039-2041. 42 Hosking, J. R. M. (2007). Distributions with maximum entropy subject to constraints on their L-moments or expected order statistics. Journal of Statistical Planning and Inference, 137, 2840-2891. Jaynes, E. T. (1957). Information Theory and Statistical Mechanics. Physical Review, 106, 620-630. Kapur, J. N. (1989). Maximum Entropy Models in Engineering. Wiley, New York. Kapur, J. N. (1994). Measures of Information and their applications. Wiley, New York. Kemp, A. W. (1997). Characterizations of a discrete normal distributions. Journal of Statistical Planning and Inference, 63, 223-229. Lisman, J. H. C. and van Zuylen, M. C. A. (1972). Note on the generation of most probable frequency distriburions. Statistica Neerlandica, 26, 19-23. M´ enendez, M. L., Morales, D. and Pardo, L. (1997a). Maximum entropy principle and statistical inference on condensed data. Statistics and Probability Letters, 34, 85-93. M´ enendez, M. L., Pardo, J. A. and Pardo, M. C. (1997b). Estimators based on sample quantiles using (h, φ)−entropy measures. Applied Mathematics Letters, 11, 99-104. Pardo, L. (2006). Statistical Inference Based on Divergence Measures. Chapman & Hall. Rao, C. R. (1965). Linear Statistical Inference and Its Application. Wiley, New York. Theil, H. and Fiebig, D. G. (1984). Maximun Entropy Estimation of Continuous Distributions. Ballinger, Cambridge. Zografos, K. (1999). On maximum entropy characterization of Pearson’s type II and VII multivariate distributions. Journal of Multivariate Analysis, 71, 67-75. Jorge Navarro Facultad de Matem´ aticas, Universidad de Murcia, 30100 Murcia, Spain The construction of Multivariate Distributions which can be fitted to multivariate data sets is a very relevant topic of research in probability and statistics. First of all I would like to warmly congratulate Professors Sarabia and G´ omez-D´ eniz for an excellent and stimulating review of some recent results on this topic. The different methods presented can be classified in two groups: (i) multivariate distributions arising out from univariate distributions, and (ii) multivariate distributions obtained from other multivariate distributions. In the first group we can include the techniques based on (a) order statistics, (b) mixtures, (c) conditional specification and (e) the method of the variables in common, while, in the second group, we can include the methods of (d) skew distributions and (f) weighted distributions. The distribution of order statistics (OS) or other generalizations such as the Generalized Order Statistics (GOS) only depends on the univariate parent distribution from which the sample of IID (independent and identically distributed)random variables is obtained. Two possible extensions can be considered here. If we consider (or we have) a sample X1,X2,...,Xnof INID (independent non-necessarily identically distributed) random variables, then the joint distribution only depend on the univariate distributions Fi(x)=Pr(Xi≤x)i=1,2,...,n. In this case, the joint distribution of the OS and the joint distribution of a subset of OS can be represented in terms of permanents (see Balakrishnan (2007)). The second option is to consider the OS obtained from a random vector (X1,X2,...,Xn), where the possible dependence between the random variables is modelled through their joint distribution. This case has special interest when (X1,X2,...,Xn) represent the lifetimes of some components in a system. This case will be included in the second group (ii) since we obtain multivariate distributions (that of subsets of OS) from a parent multivariate distribution. In the three cases, it is of special interest to study the distribution of the kfirst OS (X1:n,X2:n,...,Xk:n) (for k<n) since in many practical situations, when we put-on-test some devices (with lifetimes Xi,i=1,2,...,n), at the end of the test period we only have information about the ‘early failures’ (see, e.g., Balakrishnan, Ng and Panchapakesan (2006)). In other cases we only have information about the series system (X1:n) or the parallel system (Xn:n). It these cases, it is interesting to note how multivariate models can also be used to obtain new relevant univariate models (see Navarro, Ruiz and Sandoval (2006)). With respect to the methods based on mixtures, first we must note that they are not the usual mixtures used to represent heterogeneous populations obtained by mixing some groups with different characteristics (e.g.a mixture of two multivariate normal