scieee AI-readable full text Open interactive document viewer

POWER GENERALIZATION OF CHEBYSHEV’S INEQUALITY – MULTIVARIATE CASE

Budny, Katarzyna

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Budny, Katarzyna Article POWER GENERALIZATION OF CHEBYSHEV’S INEQUALITY – MULTIVARIATE CASE Statistics in Transition New Series Provided in Cooperation with: Polish Statistical Association Suggested Citation: Budny, Katarzyna (2019) : POWER GENERALIZATION OF CHEBYSHEV’S INEQUALITY – MULTIVARIATE CASE, Statistics in Transition New Series, ISSN 2450-0291, Exeley, New York, NY, Vol. 20, Iss. 3, pp. 155-170, https://doi.org/10.21307/stattrans-2019-029 This Version is available at: https://hdl.handle.net/10419/207949 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc-nd/4.0/ STATISTICS IN TRANSITION new series, September 2019 155 STATISTICS IN TRANSITION new series, September 2019 Vol. 20, No. 3, pp. 155–170, DOI 10.21307/stattrans-2019-029 Submitted – 04.03.2019; Paper ready for publication – 14.05.2019 POWER GENERALIZATION OF CHEBYSHEV’S INEQUALITY – MULTIVARIATE CASE Katarzyna Budny 1 ABSTRACT In the paper some multivariate power generalizations of Chebyshev’s inequality and their improvements will be presented with extension to a random vector with singular covariance matrix. Moreover, for these generalizations, the cases of the multivariate normal and the multivariate t distributions will be considered. Additionally, some financial application will be presented. Key words: multivariate Chebyshev’s inequality, Mahalanobis distance, multivariate normal distribution, multivariate t distribution. 1. Introduction Chebyshev’s inequality yields a bound on the probability of a univariate random variable taking values close to the mean expressed by its variance. Pearson (1919) proposed its univariate power generalization presenting bounds by the central moments of a random variable of even orders. Theorem 1.1. (Pearson, 1919). If we take a random variable R:  with finite central moments of 2s order   s2  , then for all 0      ss s EP 22 2     . (1.1) There also exist multivariate generalizations of Chebyshev’s inequality (see, e.g. Olkin and Pratt, 1958, Marshall and Olkin, 1960, Osiewalski and Tatar, 1999). In the paper we present one of those providing upper bounds on the probability that the Mahalanobis distance of a random vector from its mean is greater or equal than the fixed value. These bounds will be given by the power transformations and will constitute the multivariate extension of (1.1). There are many applications of the Mahalanobis distance in statistical analysis. In particular, this is used in classification methods and in cluster analysis. The multivariate power generalization of Chebyshev’s inequality presented below can be exploited to detect outliers. 1 Department of Mathematics, Cracow University of Economics, Poland. E-mail: [email protected]. ORCID ID: https://orcid.org/0000-0002-3683-0327. 156 K. Budny: Power generalization of Chebyshev’s… 2. Multivariate power generalization of Chebyshev’s inequality We begin by recalling the inequality which is given by the measure of multivariate kurtosis. Theorem 2.1. (Mardia, 1970) Let n R:X be a random vector with nonsingular covariance matrix  and finite fourth–order moments. Then, for any 0  the following inequality holds           2 ,2 1    X XXXX n TEEP    , (2.1) where                 2 1 ,2 XXXXX EEE T n  is Mardia’s kurtosis of a random vector (Mardia, 1970). Chen (2007, 2011) proposed a tight upper bound (see Navarro, 2014) in the case of a random vector for which only mean and covariance matrix are known. Theorem 2.2. (Chen, 2007, Chen, 2011) Assume that n R:X is a random vector with positive covariance matrix  . Then, for all 0  we get           n EEP T XXXX 1  . (2.2) Budny (2014) obtained the multivariate power generalization of Chebyshev’s inequality. Theorem 2.3. (Budny, 2014) Suppose that n R:X is a random vector with nonsingular covariance matrix  . Let us consider any 0s such that                 s T ns EEEI XXXXX 1 , exists. Then, for all 0            s ns TI EEP   X XXXX , 1   . (2.3) Remark 2.1. (Budny, 2014) Observe that theorems 2.1 and 2.2 can be considered as the special cases of theorem 2.3. Taking 1s we get (2.2) and for 2s we obtain (2.1). Budny (2016), following Navarro (2016), extended (2.3) to the case of a random vector with singular covariance matrix by using the spectral decomposition. STATISTICS IN TRANSITION new series, September 2019 157 Assume that n R:X is a random vector with covariance matrix  , mrank ,   nm ,...,1 . Let  T PP be a spectral decomposition of a covariance matrix, i.e. P is an orthogonal matrix such that n TT IPPPP  and   0,...,0,,...,diag 1m   is the diagonal matrix with the ordered eigenvalues 0...... 11  nmm  . Hence, the Moore-Penrose generalized inverse matrix of  is of the form T PCP   , where   0,...,0,,...,diag 11 1 m C  . Let us consider any 0s such that                 s T ms EEEI XXXXX  , exists. Theorem 2.4. (Budny, 2016) Under the above assumptions, for any 0  , we have           s ms TI EEP   X XXXX ,    . (2.4) We will denote by S the set of all 0s such that   X ms I, exists. Let us define, for fixed 0  , the function  RSBd : :     s ms I sBd  X ,  , Ss . It is easily seen that for Sss  21, if 21 ss  , then the following conditions are equivalent:     21 sBdsBd       12 1 2 1 , ,ss ms ms I I         X X  and     21 sBdsBd       12 1 2 1 , , 0ss ms ms I I          X X  . Summarizing, we get following remark. 158 K. Budny: Power generalization of Chebyshev’s… Remark 2.2. For Sss  21, if 21 ss  , then the upper bound   2 sBd of           XXXX EEP T is better than   1 sBd for all     12 1 2 1 , ,ss ms ms I I         X X  . On the contrary, the upper bound   1 sBd is better than   2 sBd for all                         12 1 2 1 , , ,0 ss ms ms I I X X  . In particular, if we consider 1 1s , 2 2s and nrank , then the upper bound   2 sBd is better than   1 sBd for all   n nX ,2    . Conversely, the upper bound   1 sBd is better than   2 sBd for all           n nX ,2 ,0   . 3. The case of the multivariate normal distribution Budny (2016) proposed the form of the multivariate power generalization of Chebyshev’s inequality for a normally distributed random vector for all   0\Ns . In the next theorem we extend this result to the case of any real 0s . Theorem 3.1. Let n R:X be a normally distributed random vector with mean  and covariance matrix  ,   ,  n N~X . Suppose that mrank ,   nm ,...,1 . Then, for all 0  and 0s we obtain          XX  T P                     2 2 2 m s m s  . (3.1) Proof: The proof is similar to that presented for theorem 3.1 in Budny (2016). A slight change is that we consider sth uncorrected moment (sth moment about zero) of a chi-square distribution with m degrees of freedom for any real 0s (not only for   0\Ns ). Hence, for 0s :                 2 2 2 ,m s m I s ms X (Johnson, Kotz and Balakrishnan, 1994, p. 420) and it completes the proof. STATISTICS IN TRANSITION new series, September 2019 159 Remark 3.1. For   0\Ns the inequality (3.1) takes the following form            s Tsmmm P   12...2  XX  (Budny, 2016). Remark 3.2. On account of remark 2.2, if 21 ss  , then the upper bound   2 sBd is better than   1 sBd for all 12 1 12 22 2ss s m s m                     . Conversely, the upper bound   1 sBd is better than   2 sBd for all                              12 1 12 22 2,0 ss s m s m  . Particularly if we take 1 1s and 2 2s , then the upper bound   2 sBd is better than   1 sBd for all 2 m  . Conversely, the upper bound   1 sBd is better than   2 sBd for all   2,0  m  . Example 3.1. Let us consider normally distributed random vector n R:X with mean  and covariance matrix  ,   ,  n N~X . Assume that 3rank  m . A random variable       XX  T has a chi-square distribution with m degrees of freedom (Kotz, Balakrishnan and Johnson, 2000, p. 110, Budny, 2016), hence we know the exact value of P . From remark 3.2 for 1 1s and 2 2s we get that the upper bound   2 sBd is better than   1 sBd for all 52  m  and the upper bound   1 sBd is better than   2 sBd for all   5,0  (see Figure 3.1). 160 K. Budny: Power generalization of Chebyshev’s… Figure 3.1. The upper bounds ( 1s , 2s ) and exact value of P for   ,~  n NX , rank 3 . In turn Figure 3.2 shows the upper bounds of         XX 1  T P for various values of s . STATISTICS IN TRANSITION new series, September 2019 161 Figure 3.2. The upper bounds ( 5.0s , 1s , 2s , 4s ) and exact value of P for   ,~  n NX , rank 3 . 4. The case of the multivariate t distribution A n variate random vector n R:X is said to have multivariate t distribution with degrees of freedom  , mean  and nonsingular covariance matrix R 2   , 2  , denoted by   nt ,,R   , if its joint probability density function (pdf) is given by           2/ 1 2/1 2/ 1 1 2 πν 2n T nxRx R n xf                              n Rx . If   nt ,,~ RX   , then the random variable     n T     XRX 1 has a central F-distribution with n ,  degrees of freedom,    ,~ nF (Lin, 1972). 162 K. Budny: Power generalization of Chebyshev’s… It follows that for any 2  s we get                                     22 22     n ss n n E s s (4.1) (Johnson, Kotz and Balakrishnan, 1995, p. 349). The power generalization of Chebyshev’s inequality for multivariate t distribution is established by our next theorem. Theorem 4.1. Assume that   nt ,,~ RX   , rank nR . Then, for any 0  the inequality   3.2 takes the following form                                        22 22 2 1        n ss n P s TXX (4.2) for any 0s such that 2  s . Proof: We first observe that 11 2  R    . From this it is obvious that     ss s ns EnI           2 ,X . (4.3) Substituting (4.1) into (4.3) yields                                22 22 2 ,    n ss n Is ns X . (4.4) This establishes the inequality (4.2). Remark 4.1. For   0\Ns , 2  s from   4.4 we get                svvv snnnv Is ns 2...42 12...22 ,  X . (4.5) Hence, the inequality (4.2) is of the form                  svvv snnn P s T 2...42 12...22 XX 1               . STATISTICS IN TRANSITION new series, September 2019 169 6. Applications in finance Let us take a random vector t r of n assets returns on a specific day t with mean  (sample mean vector of historical returns) and covariance matrix  (sample covariance matrix of historical returns). Kritzman and Li (2010) propose to use the Mahalanobis distance as a measure of financial turbulence, which is understood as occurrence of unusual multivariate financial data. They defined (the so-called “the turbulence index”) turbulence for a particular time t as:       t T tt rrd 1  . In the examples presented in section 3 and 4, for any 0  , we know the exact value of         t T tt rrdP 1  . In the general case, this probability may not be easy to compute and if we are able to calculate the upper bounds   2.5 , then we can estimate the exact value of P. Other financial applications of the Mahalanobis distance were presented by Stöckl and Hanke (2014). Acknowledgement The publication was financed from the funds granted to the Faculty of Finance and Law at Cracow University of Economics, within the framework of the subsidy for the maintenance of research potential. REFERENCES BUDNY, K., (2014). A generalization of Chebyshev's inequality for Hilbert-spacevalued random elements. Statistics and Probability Letters, 88, pp. 62–65. BUDNY, K., (2016). An extension of the multivariate Chebyshev’s inequality to a random vector with a singular covariance matrix, Communications in Statistics – Theory and Methods, 45 (17), pp. 5220–5223. CHEN, X., (2007). A new generalization of Chebyshev inequality for random vectors. Available at: <https://arxiv.org/abs/0707.0805>[Accessed 5 July 2007]. CHEN, X., (2011). A new generalization of Chebyshev inequality for random vectors. Available at: <https://arxiv.org/abs/0707.0805v2>[Accessed 24 June 2011]. JOHNSON, N.L., KOTZ, S., BALAKRISHNAN, N., (1994). Continuous univariate distribution. Vol. 1, 2nd ed. John Wiley & Sons Inc. JOHNSON, N.L., KOTZ, S., BALAKRISHNAN, N., (1995). Continuous univariate distribution, Vol. 2, 2nd ed. John Wiley & Sons Inc. 170 K. Budny: Power generalization of Chebyshev’s… KRITZMAN, M., Li, Y., (2010). Skulls, financial turbulence, and risk management. Financial Analysts Journal, 66 (5), pp. 30–41. KOTZ, S., BALAKRISHNAN, N., JOHNSON, N.L., (2000). Continuous multivariate distribution, Vol. 1: Models and applications, 2nd ed. John Wiley & Sons Inc. LIN, P., (1972). Some characterizations of the multivariate t distribution, Journal of Multivariate Analysis, 2, pp. 339–344. LOPERFIDO, N., (2014). A probability inequality related to Mardia’s kurtosis. In: C. Perna, M. Sibillo (eds.). Mathematical and statistical methods of actuarial science and finance, Springer: Springer International Publishing Switzerland 201. pp. 129–132. MARDIA, K.V., (1970). Measures of multivariate skewness and kurtosis with applications, Biometrika, 57 (3), pp. 519–530. MARSHALL, A., OLKIN, I., (1960). Multivariate Chebyshev inequalities, The Annals of Mathematical Statistics, 31, pp. 1001–1014. NAVARRO, J., (2014). Can the bounds in the multivariate Chebyshev inequality be attained? Statistics and Probability Letters, 91, pp. 1–5. NAVARRO, J., (2016). Avery simple proof of the multivariate Chebyshev’s inequality. Communications in Statistics – Theory and Methods, 45 (12), pp. 3458–3463. OLKIN, I., PRATT, J.W., (1958). A multivariate Tchebycheff inequality. The Annals of Mathematical Statistics, 29, pp. 226–234. OSIEWALSKI, J., TATAR, J., (1999). Multivariate Chebyshev inequality based on a new definition of moments of a random vector, Przegląd Statystyczny (Stat. Rev.), 2, pp. 257–260. PEARSON, K., (1919). On generalised Tchebycheff theorems in the mathematical theory of statistics, Biometrika,12(3–4), pp. 284–296. STÖCKL, S., HANKE, M., (2014). Financial applications of the Mahalanobis distance, Applied Economics and Finance, 1 (2), pp. 78–84.