ON ASYMMETRY OF PREDICTION ERRORS IN SMALL AREA ESTIMATION
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Żądło, Tomasz Article ON ASYMMETRY OF PREDICTION ERRORS IN SMALL AREA ESTIMATION Statistics in Transition New Series Provided in Cooperation with: Polish Statistical Association Suggested Citation: Żądło, Tomasz (2017) : ON ASYMMETRY OF PREDICTION ERRORS IN SMALL AREA ESTIMATION, Statistics in Transition New Series, ISSN 2450-0291, Exeley, New York, NY, Vol. 18, Iss. 3, pp. 413-432, https://doi.org/10.21307/stattrans-2016-078 This Version is available at: https://hdl.handle.net/10419/207866 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/
STATISTICS IN TRANSITION new series, September 2017 413 STATISTICS IN TRANSITION new series, September 2017 Vol. 18, No. 3, pp. 413–432, DOI: 10.21307/stattrans-2016-078 ON ASYMMETRY OF PREDICTION ERRORS IN SMALL AREA ESTIMATION Tomasz Żądło1 ABSTRACT The mean squared error reflects only the average prediction accuracy while the distribution of squared prediction error is positively skewed. Hence, assessing or comparing accuracy based on the MSE (which is the mean of squared errors) is insufficient and even inadequate because we should be interested not only in the average but in the whole distribution of prediction errors. This is the reason why we propose to use different than MSE measures of prediction accuracy in small area estimation. In the prediction accuracy comparisons we take into account our proposal for the empirical best predictor, which is a generalization of the predictor presented by Molina and Rao (2010). The generalization results from the assumption of a longitudinal model and possible changes of the population and subpopulations in time. Key words: empirical best predictor, prediction errors, small area estimation. 1. Introduction Nowadays, estimates of the population and large subpopulations characteristics are not sufficient for decision-makers. They require accurate estimates for subpopulations with small or even zero sample sizes. However, because of cost constraints, it is not possible to increase sample sizes continuously to make the estimation of smaller and smaller subpopulations possible using classical methods. The problem is solved using small area estimation methods "borrowing strength" from other subpopulations or time periods. There are three aims of our paper, with the first the main one and the second and the third the supplementary aims. Firstly, our observation that the distribution of squared prediction errors has strong positive asymmetry (values of the third standardized moments obtained in the simulation study based on real data are presented in Table 1) has become a focus of our attention. It implies that their mean (known as MSE – Mean 1 University of Economics in Katowice, Faculty of Management, Katowice, Poland. E-mail: tomasz.zadl[email protected].
414 T. Żądło On asymmetry of prediction… Squared Error) does not have to be a good measure of prediction accuracy in terms of the average and, what is in our opinion even more important, their whole distribution should be studied. Hence, the main purpose of the paper is our proposal for assessing the prediction accuracy using new univariate and multivariate prediction measures based on quantiles of the distribution of absolute prediction errors. It will be shown that even if the accuracies of two predictors in terms of the average are similar, the accuracy comparison based on right tails of distributions of absolute prediction errors can give different results (which will be presented in e.g. Figure 4). Secondly, in the accuracy comparisons resulting from the main aim we will include a generalization of the Empirical Best Predictor (EBP) proposed by Molina and Rao (2010). They proposed the predictor under a model assumed for data from surveys conducted in one period. We will propose a predictor assuming a longitudinal model. It means that in the case of longitudinal surveys we will be able to use information from previous periods to increase the prediction accuracy in the period of interest. Thirdly, in the proposed longitudinal model we will take into account that the population and subpopulations may change in time. It will cover many longitudinal models known from small area estimation (which are special cases of the general linear mixed models) including models studied by: Saei and Chambers (2003), who assume mutually independent two random effects (domain-specific and time-specific) and random components (and the generalization, where AR(1) process is assumed for time-specific random effects), Saei and Chambers (2003), where domain-and-time-specific random effects with independent distributions in domains and AR(1) model in time are taken into account, Stukel and Rao (1999) and Nissinen (2009) p. 22 with mutually independent two random effects (domain-specific and element-specific) and random components, Nissinen (2009) p. 60, who assumes independent domain-specific random effects and autocorrelated (assuming AR(1)) random components in time, Molina, Morales, Pratesi, Tzavidis (2010) pp. 143-180 with independent domain-specific random effects, independent for domains and autoccorelated (assuming AR(1)) in time domain-and-time-specific random effects and heteroscedastic random components. In the simulation study the properties of the proposed predictor (in terms of MSE and the proposed accuracy measures) will be studied under the proposed model taking into account the model misspecification as well.
STATISTICS IN TRANSITION new series, September 2017 415 2. Alternative prediction accuracy measures In the case of positive asymmetry usually the mean is not the only measure used to describe the distribution. The MSE is the mean of squared errors (which have positive or even strong positive asymmetry) and it is usually used as the only accuracy measure. Moreover, a better predictor is usually defined as the one with smaller MSE. Żądło (2013) proposed a new measure of prediction accuracy Quantile of Absolute Prediction Error defined for the problem of prediction in the dth domain as follows: () inf : d QAPE p x P U x p, (1) where ˆ ddd U is the prediction error of ˆd , which is the predictor of d in the dth domain. It means that (1) is the quantile of order p of d U. It means that at least p100% of realizations of absolute prediction errors in the dth domain are smaller or equal to ()QAPE p . In Żądło (2013) it was used to measure prediction accuracy of the empirical best linear unbiased predictor. Żądło (2015) proposes multivariate versions of (1), which allow us to measure and compare accuracy in the case of simultaneous prediction in all of domains. It can be treated as the alternative to the average mean squared error studied, e.g. by Fabrizi and Trivisano (2010). Let prediction errors in D domains be denoted by ˆ ddd U , where 1,2,...,dD. Let us define the multivariate version of QAPE as follows: 1 () inf : D d d M QAPE p x P U x Dp . (2) It means that it is the quantile of order p of a distribution of a mixture of random variables 1,..., ,..., dD UU U with equal weights. It means that at least p100% of realizations of absolute prediction errors in all domains are smaller or equal to () M QAPE p . Let relative prediction errors be denoted by ˆ ddd d dd U W , where 1,2,...,dD. Let us define relative M QAPE as follows: 1 () inf : D d d rMQAPE p x P W x Dp . (3) It means that it is the quantile of order p of a distribution of a mixture of random variables 1,..., ,..., dD WW W with equal weights. It means that at least p100% of realizations of moduli of relative prediction errors in all domains are smaller or equal to ()rMQAPE p . The estimation of (1), (2) and (3) is possible using a well-known parametric bootstrap method studied, e.g. by González-Manteiga et al. (2007, 2008) and
416 T. Żądło On asymmetry of prediction… Molina and Rao (2010). Using the method, the estimator of the MSE is given by the mean of squared bootstrap realizations of prediction errors. Similarly, by computing quantiles of bootstrap realizations of: moduli of prediction errors in one of domains we can estimate (1), moduli of prediction errors in all of domains we can estimate (2) and moduli of relative prediction errors in all of domains we can estimate (3). 3. Model and predictor We consider longitudinal data in periods 1,2,...,tM , where the population of size t N in the period t is denoted by t . The population is divided into D disjoint subpopulations (domains) dt each of size dt N, where 1,2,...,dD. A sample in the period t of size t n is denoted by t s . Let dt t dt ss and dt dt s n. The d*th domain of interest in the period of interest t* will be denoted by **dt . Let rdt dt dt s , rdt dt dt NNn , 1 M t t , N , 1 , M dt d t dd N , 1 M rdt rd t , rd rd N , 1 M t t s s , s n , 1 M dt d t s s , dd s n. Let id M be the number of periods when the ith population element belongs to the dth domain and id m – the number of periods when the ith population element (which belongs to the dth domain) is observed. Let rid id id M Mm . It is assumed that the population may change in time and that one population element may change its domain affiliation in time. Hence, sets of population elements d (where 1,2,...,dD) may overlap. The assumption that one population element may change its domain affiliation in time is very important in practice of longitudinal surveys. For example, let us consider the population of households and the division of the population into domains made according to the household size. In this case we should assume that some households can change their sizes in time, which causes the change of the domain affiliation. If a human population is under the study one may be interested in its characteristics for subpopulations defined according to some social or economic criteria. In the case of business surveys the population of firms may be divided into subpopulations according to some economic or financial criteria, what can imply even stronger changes of domains affiliations. Values of the variable of interest (or the variable of interest after a transformation) are realizations of idj Y’s for the ith population element, which belongs to the dth domain in the period ij t, where 1,2,...,iN ; 1,2,..., id j M;
STATISTICS IN TRANSITION new series, September 2017 417 1,2,...,dD. The vector 1 id idj M Y id Y will be called the profile and the vector 1 id idj m Y sid Y will be called the sample profile. Let the vector 1 rid idj M Y rid Y be the profile for non-observed realizations of random variables. Let us introduce assumptions of the following longitudinal model, which is a special case of the general linear mixed model (e.g. Datta and Lahiri, 2000). The difference is introduced in the sizes of matrices, which allows us to take into account longitudinal data and possible changes in population and subpopulations in time. We assume that 2 2 () () () () (,) D D Cov YXβZve vGδ eRδ ve 0 , (4) where is the superpopulation model, 11() d dD iN col col id YY , 11() d dD iN col col id ee , where id e is the 1 id M vector of random components, 11() d dD iN col col id XX , where id X is the known matrix of size id M p, β is the 1 p vector of unknown parameters, Z is the known matrix of size 11 ND id id M h , v is the vector of random effects of size 1h , δ is the vector of q unknown in practice parameters called variance components. Let us consider the following decomposition of the vector Y: T TT sr YYY, (5) where s Y is the vector of size 11 1 ND id id m of random variables, whose realizations are known, and r Y is the vector of size 11 1 ND rid id M of random variables, which are not observed in the longitudinal survey. Then, 22 () () () () () () ss sr s rs rr r DD Vδ Vδ Y YVδ Vδ Vδ Y, (6) where under (4): () () () T Vδ ZGδZ Rδ. (7) Let us consider the problem of predicting any given function of the random vector Y denoted by () Y or shortly by . Among predictors ˆ of , the Best Predictor (BP) is defined as the one, which minimizes (e.g. Molina and Rao 2010):
418 T. Żądło On asymmetry of prediction… 2 ˆˆ () ( )MSE E . (8) Hence, it is given by: ˆ(| ) B Ps E Y, (9) which means that it may be obtained as a conditional expected value of assuming that the conditional distribution of | rs YY is known. We assume that the conditional distribution of | rs YY can be derived (the example is presented in Remark 1 in this section). In practice, it depends on the vector of unknown parameters, which will be denoted by τ. If we replace the parameters by their estimators, we obtain the Empirical Best Predictor (EBP) denoted by ˆ E BP . Hence, the value of the EBP of () Y can be obtained through the Monte Carlo approximation algorithm presented below (for prediction in surveys conducted in one period see Molina and Rao 2010). (a) We estimate τ based on the realization of s Y and we obtain the value of the estimator denoted by ˆ τ. (b) Assuming that the distribution of | rs YY can be derived, we generate L vectors r Y (denoted by ()l r Y, where 1,2,...,lL ) from the distribution of |, rs YY where the unknown vector τ is replaced by ˆ τ. (c) We make L vectors denoted by ()l Y, where () () T lTlT sr YYY and 1,2,...,lL, what means that L vectors ()l Y include the same realization of s Y and different realizations of r Y. (d) The value of the EBP of () Y is obtained as follows: 1() 1 ˆ() L l EBP l L Y. Due to the estimation of an unknown in practice vector of model parameters denoted by τ, the resulting predictor generally is not unbiased and it does not minimize the MSE (as the BP) but its value should be very close to the BP. Its MSE estimator, which takes into account the uncertainty resulting from the estimation of τ, can be obtained using parametric bootstrap method as in Molina and Rao (2010), where their model is replaced by (4). Remark 1. If we additionally assume that the vector Y (which may be the vector of the variable of interest after a transformation) is normally distributed, which can be written as follows ~( ,())NYXβVδ (where under (4) ()Vδ is given by (7)), then T TT τβδ and 11 | ~ () () , () () () () rs r rsss ss rr rsss sr N Y Y Xβ V δV δ Y Xβ V δ V δV δV δ . (10)
STATISTICS IN TRANSITION new series, September 2017 419 Hence, in the step (b) of the procedure presented above, vectors ()l r Y, where 1,2,...,lL are generated based on (10), where parameters are replaced by their estimates. The idea of using EBPs was presented earlier by Molina and Rao (2010) but for studies conducted in one period. They study the general case assuming the general linear mixed model for studies conducted in one period. In the special case of their considerations Y is the vector of the variable of interest after the following transformation: ()TYY , where Y is the variable of interest, and () ln( )Tc YY , (11) where c is a constant. Then, they study the following model T id id d id Yve xβ , (12) where 1,2,...,dD; 1,2,...,iN, 2 ~(0, ) iid dv vN , 2 ~(0, ) iid id e eN , id e and d v are independent, 22 T ev δ. 4. Simulation study – real data We consider data on 378N Polish poviats (NUTS-4 level) from years 2010-2012 (M=3). We have excluded one observation because of the lack of the data and one outlying observation (Warsaw). The problem of prediction of totals of the sold production of industry in 16D domains (voivodships – NUTS-2 level) for companies with at least 10 employees is considered. The number of companies with at least 10 employees is the auxiliary variable. In the first period a sample of 38 poviats is drawn at random with probabilities proportional to the values of the auxiliary variable. Sample sizes in the domains are random and equal from 0 to 5 (with mean 2.375). The balanced panel survey is considered – elements sampled in the first period are observed until the end of the longitudinal survey (which gives 114 observations in 3 periods). Because empirical best predictors are studied, the distribution of the variable of interest must be assumed. We consider the transformation of the variable of interest given by (11) and logarithmic transformation of the auxiliary variable. To test the distribution of the variable of interest, we use the transformation of residuals based on the Cholesky decomposition of the inverse of variancecovariance matrix (see, e.g. Jacqmin-Gadda et al. 2007). For the model chosen based on the AIC and BIC criteria (more details will be presented in the next paragraph) p-values for Shapiro-Wilk, Jarque-Bera and adjusted Jarque-Bera tests obtained for the sample equal 0.2297; 0.6046 and 0.446 respectively. For the considered model but without the transformations of both variables, p-values for the tests of normality were smaller than 12 10 . But if we test normality based on the whole population data (based on 3 378 1134MN observations) in both cases (with and without transformations of variables) we should reject the
420 T. Żądło On asymmetry of prediction… null hypothesis on normality. That is why the problem of model misspecification will be taken into account in the simulation study as well. To choose the appropriate model we consider different models: classic and mixed linear models with and without the auxiliary variable, with and without constant, nested-error models and models with random slopes. For mixed models with random slope we consider time-specific, domain-specific, time-and-domainspecific and finally profile-specific random effects. In models with nested errors we consider one random effect (time-specific, domain-specific, time-and-domainspecific and profile-specific) or two random effects (firstly: domain-specific and profile-specific; secondly: domain-specific and domain-and-time-specific). The model with the smallest both AIC and BIC criteria was the following model studied earlier by Stukel and Rao (1999) and Nissinen (2009) p. 22: 10idt idt d id idt Yx uve , (13) where idt Y is the variable of interest after transformation (11), idt x is the auxiliary variable after logarithmic transformation, 1,2,...,iN , 1,2,...,dD, 1,2,...,tM, d u, id v and idt e are mutually independent with zero expected values and variances given by 2 u , 2 v and 2 e respectively. Permutation test (with the test statistic given by the loglikelihood) was used to test the significance of the model parameters – at the significance level 0.05 tested parameters were significantly different from zero. Good properties of these tests are presented by Krzciuk and Żądło (2014a, 2014b). The model-based simulation study was prepared using R software (R Core Team 2016). To mimic the real data, values of the variable of interest after transformation (11) are generated based on the model (13) with one auxiliary variable and the constant, where the parameters of the model are replaced by REML estimates obtained based on all of the observations (sampled and unsampled) of the real data. Hence, both random effects and random components are generated with zero expected values and variances 2 u , 2 v and 2 e equal REML estimates based on (13) and the whole population data. Random effects and random components d u, id v and idt e are generated independently from: normal distributions, shifted exponential distributions (the third standardized moment is equal to 2), shifted gamma distributions (with the value of third standardized moment equal to 4) and shifted Pareto distributions (with the value of third standardized moment equal to 5). It means that in the case of the normal and the shifted exponential distributions, the assumed values of the mean and the variance give explicitly values of the parameters of the distributions used in the simulation study. The case of the shifted gamma and the shifted Poisson distributions is more interesting because it
STATISTICS IN TRANSITION new series, September 2017 427 section), but where the real auxiliary variable is replaced by the artificial one. Values of the auxiliary variable were generated independently from shifted gamma distribution assuming real values of the mean, variance and the third standardized moment for each year. In this case the maximum gain in accuracy of our EBP measured by MSE is 25.7% and measured by QAPE is higher than 10%. It means that in the case of longitudinal surveys we should use auxiliary variable, which is weakly autocorrelated but even in this case the gain in accuracy will not be very large. In the second scenario (results in Table 2 in the column “without x”) we do not use the uxiliary variable both in the model and at the estimation stage. Hence, we compare prediction accuracy of empirical best predictors only under random parts of models (13) and (12). The accuracy measured by MSE of our EBP is higher by 12.3% compared with EBP-MR (by less than 5% in terms of QAPE ). The reason is that model (13) chosen based on AIC and BIC for real longitudinal data is quite similar to the model (12) assumed by Molina and Rao (2010). In both models we have domain-specific random effects, although in the case of (13) it additionally implies non-zero covariances between observations within domains in different periods. The main difference between the models is the profile (element)-specific random effect in model (13), but results in the last column of Table 2 show that it does not imply a large gain in prediction accuracy. It means that the larger gain in accuracy can be obtained when the longitudinal model explains the variability of the variable of interest considerably better than the model assumed for one period. To sum up, in this section based on the Monte Carlo analysis we have identified two reasons of the relatively small gain in accuracy, which was presented in the previous section, comparing our predictor with the predictor proposed by Molina and Rao (2010). Firstly, it has been autocorrelation in time of the auxiliary variable. Secondly, we have presented similarity of the proposed longitudinal model and the model studied by Molina and Rao (2010). Moreover, we have shown that in the studied cases the maximum gain in accuracy comparing these two predictors can be even higher than 25% in terms of MSE. 6. Real data application In this section we consider values of the same predictors and estimators, the same data and the same sample as discussed in section 4. However, in this case their values are computed once based on the real data (they are not generated as in the simulation studies presented in section 4). Because the whole population data are available, we are able to compare estimates with real values of D=16 domains totals (see Figure 5). The largest differences between estimates and real values for the considered sample are observed for SYNT-REG and EBLUP, whereas the values of EBP-MR and the proposed EBP are very similar.
428 T. Żądło On asymmetry of prediction… Figure 5. Values of estimates and real domain totals 7. Conclusions In the paper the problem of assessing and comparing the prediction accuracy is studied. Because of strong positive asymmetry of absolute prediction error, it is shown that prediction accuracy measures alternative to the MSE should be used. These measures allow us to assess the prediction accuracy not limited to the average values and to obtain more stable results of accuracy comparisons, especially in the case of the model misspecification. In the accuracy comparisons based on the Monte Carlo simulation studies our proposal for the empirical best predictor is taken into account. Although its prediction accuracy was only slightly better for the considered data compared with the empirical best predictor proposed by Molina and Rao (2012), we present how to obtain a substantial gain in accuracy. The considerations are also supported by real data application. 0 500 1000 1500 2000 0 500 1000 1500 2000 estimates of domain totals real domain totals SYNT-REG EBLUP EBP-MR EBP
STATISTICS IN TRANSITION new series, September 2017 429 REFERENCES BRACHA, CZ., (1996). Teoretyczne podstawy metody reprezentacyjnej, PWN, Warszawa. DATTA, G. S., LAHIRI, P., (2000). A Unified Measure of Uncertainty of Estimated Best Linear Unbiased Predictors in Small Area Estimation Problems, Statistica Sinica, Vol. 10, pp. 613–627. FABRIZI, E., TRIVISANO, C., (2010). Robust linear mixed models for Small Area Estimation, Journal of Statistical Planning and Inference, Vol. 140, pp. 433–443. GONZÁLEZ-MANTEIGA, W., LOMBARDÍA, M.J., MOLINA, I., MORALES, D., SANTAMARÍA, L., (2007). Estimation of the mean squared error of predictors of small area linear parameters under a logistic mixed model, Computational Statistics and Data Analysis, Vol. 51, pp. 2720–2733. GONZÁLEZ-MANTEIGA, W., LOMBARDÍA, M. J., MOLINA, I., MORALES, D., SANTAMARÍA, L., (2008). Bootstrap mean squared error of small-area EBLUP, Journal of Statistical Computation and Simulation, Vol. 78, 443–462. JACQMIN-GADDA, H., SIBILLOT, S., PROUST, C., MOLINA J.-M., THIÉBAUT, R., (2007). Robustness of the Linear Mixed Model to Misspecified Error Distribution, Computational Statistics & Data Analysis, Vol. 51, pp. 5142–5154. JIANG, J., (1996). REML Estimation: Asymptotic Behavior and Related Topics, The Annals of Statistics, Vol. 24, pp. 255–286. KRZCIUK M., ŻĄDŁO T., (2014a). On some tests of variance components for linear mixed models, Studia Ekonomiczne, Vol. 189, pp. 77–85. KRZCIUK M., ŻĄDŁO T., (2014b). On some tests of fixed effects for linear mixed models, Studia Ekonomiczne, Vol. 189, pp. 49–57. MOLINA, I., RAO, J. N. K., (2010). Small Area Estimation of Poverty Indicators, The Canadian Journal of Statistics, Vol. 38, pp. 369–385. NISSINEN, K., (2009). Small Area Estimation With Linear Mixed Models For Unit-Level Panel and Rotating Panel Data, University of Jyväskylä Printing House, Jyväskylä. R CORE TEAM, (2016). R: A Language and Environment For Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria.
430 T. Żądło On asymmetry of prediction… ROYALL, R. M., (1976). The Linear Least Squares Prediction Approach to TwoStage Sampling, Journal of the American Statistical Association, Vol. 71, pp. 657–473. STUKEL, D. M, RAO, J. N. K., (1999). On Small-Area Estimation Under TwoFold Nested Error Regression Models, Journal of Statistical Planning and Inference, Vol. 78, pp. 131–147. ŻĄDŁO, T., (2013). On Parametric Bootstrap and Alternatives of MSE, Proceedings of 31st International Conference Mathematical Methods in Economics 2013, College of Polytechnics Jihlava, pp. 1081–1086. ŻĄDŁO, T., (2015). Statystyka małych obszarów w badaniach ekonomicznych. Podejście modelowe i mieszane, Wydawnictwo Uniwersytetu Ekonomicznego w Katowicach.
STATISTICS IN TRANSITION new series, September 2017 431 APPENDIX Figure 6. Values of QAPE0.75(.)/QAPE0.75(EBP) for different predictors and different distributions of random effects and random components (each boxplot presents values for D=16 domains) Figure 7. Values of QAPE0.90(.)/QAPE0.90(EBP) for different predictors and different distributions of random effects and random components (each boxplot presents values for D=16 domains) SYNT-REG EBLUP EBP-MR after transformation: normal 1.0 1.5 2.0 2.5 3.0 after transformation: ex p onential SYNT-REG EBLUP EBP-MR 1.0 1.5 2.0 2.5 3.0 after transformation: g amma after transformation: Pareto SYNT-REG EBLUP EBP-MR after transformation: normal 12345 after transformation: ex p onential SYNT-REG EBLUP EBP-MR 12345 after transformation: g amma after transformation: Pareto
432 T. Żądło On asymmetry of prediction… Figure 8. Values of QAPE0.95(.)/QAPE0.95(EBP) for different predictors and different distributions of random effects and random components (each boxplot presents values for D=16 domains) Figure 9. Values of rMQAPE(p) for p=(0.5, 0.75, 0.9, 0.95), different predictors and different distributions of random effects and random components – selected results SYNT-REG EBLUP EBP-MR after transformation: normal 123456 after transformation: ex p onential SYNT-REG EBLUP EBP-MR 123456 after transformation: g amma after transformation: Pareto SYNT-REG EBLUP EBP-MR EBP after transformation: normal 20 40 60 80 100 after transformation: ex p onential SYNT-REG EBLUP EBP-MR EBP 20 40 60 80 100 after transformation: g amma after transformation: Pareto