Robustness of randomisation tests as alternative analysis methods for repeated measures design
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Oladugba, Abimibola Victoria; Obasi, Ajali John; Asogwa, Oluchukwu Chukwuemeka Article Robustness of randomisation tests as alternative analysis methods for repeated measures design Statistics in Transition new series (SiTns) Provided in Cooperation with: Polish Statistical Association Suggested Citation: Oladugba, Abimibola Victoria; Obasi, Ajali John; Asogwa, Oluchukwu Chukwuemeka (2022) : Robustness of randomisation tests as alternative analysis methods for repeated measures design, Statistics in Transition new series (SiTns), ISSN 2450-0291, Sciendo, Warsaw, Vol. 23, Iss. 4, pp. 77-90, https://doi.org/10.2478/stattrans-2022-0043 This Version is available at: https://hdl.handle.net/10419/301890 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-sa/4.0/
STATISTICS IN TRANSITION new series, December 2022 Vol. 23, No. 4, pp. 77–90, DOI 10.2478/stattrans-2022-0043 Received – 09.09.2021; accepted – 01.08.2022 Robustness of randomisation tests as alternative analysis methods for repeated measures design Abimibola Victoria Oladugba 1 , Ajali John Obasi 2 , Oluchukwu Chukwuemeka Asogwa 3 ABSTRACT Randomisation tests (R-tests) are regularly proposed as an alternative method of hypothesis testing when assumptions of classical statistical methods are violated in data analysis. In this paper, the robustness in terms of the type-I-error and the power of the R-test were evaluated and compared with that of the F-test in the analysis of a single factor repeated measures design. The study took into account normal and non-normal data (skewed: exponential, lognormal, Chi-squared, and Weibull distributions), the presence and lack of outliers, and a situation in which the sphericity assumption was met or not under varied sample sizes and number of treatments. The Monte Carlo approach was used in the simulation study. The results showed that when the data were normal, the R-test was approximately as sensitive and robust as the F-test, while being more sensitive than the F-test when data had skewed distributions. The R-test was more sensitive and robust than the F-test in the presence of an outlier. When the sphericity assumption was met, both the R-test and the F-test were approximately equally sensitive, whereas the R-test was more sensitive and robust than the F-test when the sphericity assumption was not met. Key words: randomisation test, repeated measures design, sensitivity, robustness, Monte Carlo. 1. Introduction Research in many areas of application as affirmed by Ma et al. (2012) normally involves study plans in which measurements or responses are repeatedly obtained from an experimental unit (EU). According to Davis (2002), repeated measurements refer broadly to data in which the response of each experimental unit or subject is observed on multiple treatment conditions or time points. Repeated measures design (RMD) 1 Corresponding Author. Department of Statistics, University of Nigeria, Nsukka, Nigeria. E-mail: [email protected]. ORCID: https://orcid.org/0000-0002-6402-8833. 2 Department of Statistics, University of Nigeria, Nigeria. ORCID: https://orcid.org/0000-0002-4761-9682. 3 Department of Mathematics/Computer Science/Statistics and Informatics, Alex Ekwueme Federal University Ndufu Alike Ikwo, Nigeria. ORCID: https://orcid.org/0000-0001-7297-9201. © A. V. Oladugba, A. John Obasi, O. C. Asogwa. Article available under the CC BY-SA 4.0 licence
78 Abimibola Victoria Oladugba et al.: Robustness of randomisation tests… is an experimental design that involves multiple measures of the same variable(s) taken on the same EU either under different treatment conditions or over two or more time periods (Kreuger and Tian, 2004).The major advantage of RMD is that it uses exactly the same individuals or subjects in all treatment conditions thereby eliminating the influence of individual differences from the analysis and also being economical in the use of resources and enabling the subjects to be their own control as measurements are taken under both control and other experimental conditions (Reed III, 2003; Howitt and Cramer, 2011). An approach to RMD data analysis is the repeated measures analysis of variance (RM ANOVA) that is based on the F-test statistic which has assumptions that must be met to ensure valid results are obtained from the analysis and therefore is limited in its application (Dragset, 2009). The assumptions include random sampling of EU from the population, normality of responses, and equality of all pairwise differences in variance between experimental conditions called sphericity (Girden, 1992, Lindman, 1992). The F-test is a statistical test in which the sampling distribution of the test statistic has an F-distribution when the null hypothesis is true (Oladugba et al., 2014). In statistical analysis, if the assumptions for any parametric test cannot be satisfied, there is risk of passing invalid inference if such test is deployed. So, researchers either transform the response data so that the resulting variable meets the conditions of the intended test to be used or resort to a different test such as the non-parametric test, which is not affected by the assumptions of the parametric test (Zimmerman and Zumbo, 1990) but transformation of data according to Sawilowsky et al. (1989) can have poor power properties. Also, the use of ranks in nonparametric tests leads to loss of information, thus the researchers cannot rely with high confidence level on ranking or transformation of data as an alternative to the F-test when its assumptions are not met (Gleason, 2013). Randomization test (R-test) or permutation test can provide excellent solutions in the presence of unsuitable conditions for the use of the F-test or when the researchers want to maintain the use of the original data. The R-test is a way of hypothesis testing that can be deployed for analysis of experimental data when assumptions of parametric tests are not tenable (Edgington, 1995; Kherad-Pajouh and Renaudi, 2014). It provides an efficient approach to hypothesis testing. In other words, the R-test is perceived as an alternative method to data analysis in conditions when assumptions of parametric procedures are not met (Craig and Fisher, 2019; Berry et al., 2018). R-test performs well in conditions not favourably for the F-test and is as sensitive and robust as the F-test when parametric test assumptions are met (Mundry, 1999; Mewhort, 2005; Mewhort et al., 2010). Since the validity of any statistical inference depends largely on satisfaction of the assumptions of the underlying model, researchers should not anticipate any statistical
STATISTICS IN TRANSITION new series, December 2022 79 test to be the most appropriate in any situation but rather subject proposed statistical test to scrutiny to ensure it is better than other alternatives in terms of sensitivity and robustness (Peres-Neto and Olden, 2001). The sensitivity of a test is the ability of a test to make right decision vis-à-vis rejection or acceptance of a hypothesis also known as power of a test; it is greatly influenced by sample size and presence of outliers (Cohen, 1988) and assumption of sphericity (constant variance) for RMD (Dragset, 2009) while robustness, on the other hand, refers to the ability of a test to yield correct conclusion or perform optimally in terms of controlling the type-I-error (p) that is not to falsely detect an effect when some of the distributional assumptions are not met or under unfavourable conditions (Vorapongsathorn et al., 2004). Hence, this paper used the R-test to analyse the RMD and compared the results to that of the F-test in order to find out which was more sensitive and robust under the conditions that data are normal and non-normal (exponential, lognormal, Chi-square, and Weibull distributions), in the absence and presence of outliers, when sphericity assumption was met or not in variant number of treatments and sample sizes. 2. Materials and methods 2.1. Material The data presented in Table 1 were obtained from Gravetter and Wallnau (2007). The responses generated from the study were based on the time (in seconds) lapsed until participants reported they felt nothing called latency when a stimulus (of 500-milligram weight) was gently placed on a region of the body. The study compared the adaptation for four regions of the body for a sample of 7 participants. Table 1. Data on sensory adaption experiment Area of stimulation (Treatment) Subjects Back of hand Lower back Middle of Palm Chin below lower Lip 1 6.5 4.6 10.2 12.1 2 5.8 3.5 9.7 11.8 3 6.0 4.2 9.9 11.5 4 6.7 4.7 8.1 10.7 5 5.2 3.6 7.9 9.9 6 4.3 3.5 9.0 11.3 7 7.4 4.8 10.8 12.6 2.2. The F-test method for analysis of single factor RMD The F-test procedure for hypothesis testing in analysis of RMD involves computing the F-statistic associated with the problem. In this section, the model, ANOVA table
80 Abimibola Victoria Oladugba et al.: Robustness of randomisation tests… presented in Table 2 and the F-test procedures for analysing single factor RMD are defined as follows. The model for this design is defined as: ij i j ij y i = 1, 2, …, n; j = 1, 2, …, t ∑τ 0; βi ~ N(0, 2 ); εij ~ N(0, 2 ) where, yij is the response from the ith subject at treatment j; μ is the grand mean; τj is the fixed effect of the jth treatment (the treatments are assumed to have fixed effects thus the zero-sum constraint); βi is the random effect for ith subject and εij is a random error component specific to ith subject at jth treatment. Table 2. ANOVA table for single factor RMD Source of Variation SS df Mean Square F0 Subject SSB n -1 MSS Treatments SST t – 1 MST 𝑀𝑆 𝑀𝑆 Error SSE (t - 1)(n - 1) MSE Total SST tn – 1 where MSS = ; MST = ; MSE = . The sums of squares are then defined as follows: SST = ∑∑𝑦. 𝑦.. 2; SSS = ∑∑𝑦. 𝑦.. 2; SSE = ∑∑𝑦 𝑦.𝑦. 𝑦..2 2.3. Randomization test procedure The hypothesis to be tested is: Ho: the different treatments had the same effect vs H1: there is a differential effect of at least one treatment α = 0.05 Test statistic Here, F-statistic was used as the test statistic. It summarizes the differences between means and eliminates the effects of between-subject variability. Procedure With repeated measures, we permute the data within subject. If there is no effect of treatments, then the set of scores from any subject can be exchanged across treatments. The steps are as follows: Compute the F-statistic for the original data, and denote that as Fcal. Permute the data within each subject, and do it for every subject.
STATISTICS IN TRANSITION new series, December 2022 81 Calculate an F-statistic for each of the permuted data. If this F-statistic is greater than Fcal, increment the counter. Repeat the preceding three steps B times, where B ≥ 10,000. Divide the value in the counter by B to obtain the probability of obtaining an F-statistic as large as Fcal if the null hypothesis were true. Denote this value as empirical type-I-error (p-value). Reject the null hypothesis of no difference due to treatment if p-value is less than our chosen level of significance. 2.4. Randomization test procedure for RMD The R-test for analysing single factor RMD involves the following procedures. Compute a test statistic that sufficiently explains the experimental data (the F-statistic in this case) for the data in Table 1. Afterwards, the data are rearranged within the subject repeatedly and the test statistic is recomputed for all resultant data permutations. Randomization test uses the obtained results from all data permutations and the original result of the experiment to form a reference set which is used to decide the significance of the test. The fraction of the data permutation in the reference set having test statistic values greater than or equal to the value obtained from the original results before data were permuted is the type-I-error (significance or probability value). In permuting data in RMD, Edgington (1995) proposed two schemes, namely systematic and random permutation schemes. In this paper, the random permutation scheme was adopted and carried out in the following way. Firstly, the data are arranged in a table with k columns and n rows, where k is the number of treatments and n is the number of subjects. An index number 1 to n was assigned to the subjects and 1 to k to the treatments, so that each measurement has associated with it a compound index number, the first part which indicates the subject and the second indicates the treatments. Accordingly, index (2, 3) for instance referred to the measurement for the second person under the third treatment. Then a random number generation algorithm was used to randomly determine for each subject independently of the other subjects which of the k measurements is to be assigned to the first treatment, which of the remaining k-1 measurements to the second treatment, and so on. The random determination of order of measurements within each subject performed over all subjects constitutes a single permutation or arrangement of the data. The arrangement is repeated for a large number of times like 10,000 permutations, and for each permutation, the test statistic is computed. The p-value is computed as the number of the test statistic value, including the obtained test statistics values that are as large as the obtained test statistics value.
82 Abimibola Victoria Oladugba et al.: Robustness of randomisation tests… 2.5. Outlier detection and sphericity assumption Outliers were randomly injected into the dataset in Table 1, and Tukey’s method of outlier detection as explained by Songwon (2006) was used in detecting them. One of the ways to test for sphericity in RMD is the use of Mauchly’s test. Mauchly’s test tests the hypothesis that the variances of the differences between any two conditions are equal. Thus, if the significance level of Mauchly’s test is less than or equal to the alpha level, sphericity is violated. Mauchly's test of sphericity in SPSS version 22 was used to verify this condition. 2.6. Monte Carlo Simulation In order to analyse RMD with the R-test and the F-test so as to check their robustness, a Monte Carlo simulation was conducted using RMD in Table 1 with n = 7 subjects and t = 4 treatments. Three variables were manipulated: (i) sample sizes (n); (ii) number of treatments (t); and (iii) distribution structure of the data (normal, exponential, lognormal, Chi-square and Weibull distributions). The performance of the two tests was investigated with three sample conditions n = 5, 7, and 9, and three treatment conditions t = 3, 4, and 5, under 5 distributional structures of the data in the presence and absence of outliers and when sphericity assumption is met or not, respectively. The R statistical package was used to implement the Monte Carlo technique sampling of 10,000 permutations from the possible (t!)n permutations for the R-test. In the simulation, the experiment was repeated 1000 times for each distribution. In each repetition, the resulting tables of data set were analysed appropriately using the F-test and the R-test methods to obtain the rate of type-I-error and power. The percentage of significant tests out of 10,000 iterations was considered as the rejection rate. The comparison procedures were considered in two scenarios. Firstly, in the scenario that the null hypothesis (H0: μi = 0) is true, the rejection rate of the null hypothesis was regarded as the type-I-error rate for each test. The test that had the closest type-I-error to the nominal α = 0.05 was considered as the more robust of the two. Secondly, in the scenario that the alternative hypothesis (Ha: μi ≠ 0) was true, the rejection rate of the null hypothesis was considered as the power for each test. The test that had larger power was taken to be more sensitive than the other. 2.7. Distribution structure of data Data were simulated from five theoretic distributions. The normal distribution was used to test condition under which normality assumption holds. The skewed distributions used include Chi-square, exponential, lognormal, and Weibull distributions; this represents condition under which the distribution assumption (normality) does not hold. The probability density function of the five distributions is defined as follows.
STATISTICS IN TRANSITION new series, December 2022 83 (a) Normal distribution The normal distribution has probability density function (pdf) as f(x) = √𝑒 , , 0 , 0x The parameters (μ and σ2) of the normal distribution were estimated using the maximum likelihood estimators (MLE). For the normal distribution, data were simulated using mean, 𝑥 = 7.7250 and variance, σ2= 9.1180, of the experimental data. (b) Exponential distribution The exponential distribution has pdf with parameter θ is given by 𝑓𝑥 𝑒 , 0, 0x Data were simulated to follow the exponential distribution using the MLE of the exponential distribution parameters obtained as 𝜃 = 7.7250 as fitted using fitdistrplus package in R statistical computing. (c) Chi-square distribution The pdf of Chi-square distribution with parameter n, is given as 2 2 1/2 2 () , 0 2() n n x n xe fx x Using fitdistrplus package in R statistical computing, the parameter of the Chi-square distribution, n = 4.559 ~ 5, was used for simulation of data where n is the mean of Chi-square distribution. (d) Lognormal distribution The pdf for the two-parameter (μ and σ2) lognormal distribution is f (X|μ,𝜎) = 𝑒 , X > 0, -∞ < μ < ∞, σ > 0 The MLE of μ and σ2 were obtained as: 𝜇 = ∑𝒍𝒏 𝑿𝒊 𝒏 𝒊𝟏 = 7.7880 and 𝜎2 = ∑𝒍𝒏 𝑿𝒊∑𝒍𝒏 𝑿𝒊 𝒏 𝒊𝟏𝒏 𝒏 𝒊𝟏 𝟐 = 12.0120 (e) Weibull distribution The two-parameter Weibull distribution has pdf given as 1() (/,) k kx kx fxk e , 0x, 0, 0k
84 Abimibola Victoria Oladugba et al.: Robustness of randomisation tests… Data were simulated using the MLE of the Weibull distribution parameters obtained as k = 7.020 and λ = 9.10 as fitted using fitdistrplus package in R statistical computing. 3. Results In Tables 3, 4, 5 and 6, the simulation results (type-I-error and power) for the F-test and the R-test based on the three manipulated variables (sample size, number of treatments and distribution structure of the data) are presented. Following from the methods mentioned in Section 2 as implemented in R Statistical package, for each sample size, the optimal values of the type-I-error and power were recorded. The sample size was denoted as n, the values in bracket indicate the number of treatment (t) that produced optimal type-I-error and highest power as the number of treatments were varied. The values in bold are either the optimal type-I-error or the highest power for each of the test. Table 3 shows the type-I-error of the F-test and the R-test for the data in the absence of outliers. The results indicated that as n increased, type-I-error decreased for data with normal distribution, for Chi-square, lognormal, exponential and Weibull, it initially increased but afterwards decreased for the F-test while the R-test produced type-I-error that increased as n increased under the normal distribution but reduced as n increased for Chi-square, exponential and Weibull, while for data with lognormal distribution, the type-I-error decreased as n decreased. On the other hand, the power values for the normal data decreased initially but later increased as the sample size increased, it increased initially and subsequently decreased for Chi-square and Weibull distributions for the F-test and increased for exponential and lognormal data distributions. The R-test on the other hand had increasing power as n increased for exponential, Weibull, and Chi-square but had an increasing trend for lognormal although with a slight initial decrease at n = 7. When outliers were introduced, the typeI-error and power values are presented in Table 4. The results indicated that type-Ierror for the F-test under all the data distributions had a decreasing trend as n increased but an increasing trend for Weibull distribution. Furthermore, the power for the F-test exhibited a slight decreasing trend for Chi-square and lognormal, while it increased for normal, exponential and Weibull as n increased. On the other hand, the power values of the R-test for all data distribution were increasing as sample size increased. The results for when sphericity assumption was met are displayed in Table 5. The type-I-error for the F-test in this table revealed that as n increased, normal and exponential data distributions initially had a slight increasing trend but substantially increased afterwards for Weibull data distribution while a decreasing trend was observed for Chi-square and lognormal data distribution. The results of sphericity