scieee AI-readable full text Open interactive document viewer

Recent Trends in the Use of Statistical Tests for Comparing Swarm and Evolutionary Computing Algorithms: Practical Guidelines and a Critical Review

Carrasco, Jacinto,García López, Salvador,Rueda García, María Del Mar,Das, Swagatam,Herrera Triguero, Francisco

Abstract

Spanish Ministry of Economy, Industry and Competitiveness

Full text

Recent Trends in the Use of Statistical Tests for Comparing Swarm and Evolutionary Computing Algorithms: Practical Guidelines and a Critical Review J. Carrasco1, S. García1, M.M. Rueda2, S. Das3, and F. Herrera1 1Department of Computer Science and AI, Andalusian Research Institute in Data Science and Computational Intelligence, University of Granada, Granada, Spain 2Department of Statistic and Operational Research, University of Granada, Andalusian Research Institute of Mathematics, Granada, Spain 3Electronics and Communication Sciences Unit, Indian Statistical Institute, 203 B.T.Road, Kolkata 700108, West Bengal, India February 24, 2020 Abstract A key aspect of the design of evolutionary and swarm intelligence algorithms is studying their performance. Statistical comparisons are also a crucial part which allows for reliable conclusions to be drawn. In the present paper we gather and examine the approaches taken from different perspectives to summarise the assumptions made by these statistical tests, the conclusions reached and the steps followed to perform them correctly. In this paper, we conduct a survey on the current trends of the proposals of statistical analyses for the comparison of algorithms of computational intelligence and include a description of the statistical background of these tests. We illustrate the use of the most common tests in the context of the Competition on single-objective real parameter optimisation of the IEEE Congress on Evolutionary Computation (CEC) 2017 and describe the main advantages and drawbacks of the use of each kind of test and put forward some recommendations concerning their use. Keywords statistical tests, optimisation, parametric, non-parametric, Bayesian 1 Introduction Over the few last years the comparison of evolutionary optimisation algorithms and statistical analysis have undergone some changes. The classic paradigm consisted of the application of classic frequentist tests on the final results over a 1 arXiv:2002.09227v1 [cs.NE] 21 Feb 2020 set of benchmark functions, although different trends have been proposed since then [1], such as: •The first amendment after the popularisation of the use of statistical tests was the proposal of non-parametric tests that consider the underlying distribution of the analysed results, improving the robustness of the drawn conclusions [2]. •With this new perspective, some non-parametric tests [3], whose assumptions were less restrictive than the previous ones, were suggested for the comparison of computational intelligence algorithms [4]. •However, there are other approaches and considerations, like the convergence of the solution [5], robustness with respect to the seed and results over different runs and the computation of the confidence interval and confidence curves [6]. •As has already occurred in several research fields, a Bayesian trend [7] has emerged with some criticisms to the well known Null Hypothesis Statistical Tests (NHST) and there are some interesting proposals of Bayesian tests analogous to the classic frequentist tests [8]. Inferential statistics make predictions and obtain conclusions from data, and these predictions are the basis for the performance comparisons made between algorithms [9]. The procedure followed to reach relevant information is detailed below: 1. The process begins with the results of the runs from an algorithm in a single benchmark function. These results assume the role of a sample from an unknown distribution whose parameters we can just estimate to compare it with another algorithm’s distribution. 2. Depending on the nature and purpose of the test, the results of the algorithms involved in the comparison will be aggregated in order to compute a statistic. 3. The statistic is used as an estimator of a characteristic, called parameter, of the distribution of interest, either the distribution of the results of our algorithm or the distribution of the algorithms’ performance difference when we are comparing the results from a set of algorithms. This procedure allows us to get information from experimental results, although it must be followed in order to obtain impartial conclusions which can also be reached by other researchers. Statistical tests should be considered as a toolbox to collect relevant information, not as a set of methods to confirm the previously stated conclusion. There are two main approaches: 2 Frequentist These are the most common tests. Both Parametric and NonParametric tests follow this approach. Here, a non-effect hypothesis (in the sense of non-practical differences between algorithms) H0and an alternative hypothesis H1are set up and a test, known as Null Hypothesis Statistical Test (NHST) is performed. With the results of this test, we determine if we should reject the null hypothesis in favour of the alternative one or if we do not have enough evidence to reject the null hypothesis. This decision is made according to two relevant concepts in the Frequentist paradigm [10, 3, 11]: •αor confidence coefficient: While estimating the difference between the populations, there is a certain confidence level about how likely the true statistic is to lie in the range of estimated differences between the samples. This confidence level is denoted as (1 −α), where α∈ [0,1], and usually α= 0.05. This is not the same situation as if the probability of the parameter of interest lay in the range of estimated differences, but the percentage of times that the true parameter would lie in this interval if repeated samples from the population had been extracted. •p-value: Given a sample Dand considering a null hypothesis H0, the associated p-value is the probability of obtaining a new sample as far from the null hypothesis as the collected data. This means that if the gathered data is not consistent with the assumed hypothesis, the obtained probability will be lower. Then, in the context of hypothesis testing, αrepresents the established threshold which the p-value is compared with. If the p-value is lower than this value, the null hypothesis is rejected. •The confidence interval, the counterpart of the p-values, represents the certainty about the difference between the samples at a fixed confidence level. Berrar proposes the confidence curves, as a graphic representation that generalise the confidence intervals for every confidence level [6]. Bayesian Here, we do not compute a single probability but a distribution of the parameter of interest itself. With this approach, we avoid some main drawbacks of NHST, although they require a deeper understanding of the underlying statistics and conclusions are not as direct as in frequentist tests. The main concepts of these families of statistical tests can be explained using Figure 1. In this figure we have plotted the density of distribution of the results of the different runs of two algorithms (DYYPO and TLBO-FL, which are presented in subsection 7.2) for a single benchmark. The density represents the relative probability of the random variable of the results of these algorithms for each value in the x-axis. In this scenario, a Parametric NHST would set up a null hypothesis about the means of both populations (plotted with a dotted 3 Figure 1: Comparison between real and fitted Gaussian distribution of results. line), assuming that their distribution follows a Gaussian distribution. We have also plotted the Gaussian distributions with the mean and standard deviation of each population. In this context, there is a difference between the two means and the test could reject the null hypothesis. However, if we consider the estimated distributions, they are overlapped. This overlapping lead to the concept of effect size, which indicates not only if the population means can be considered as different, but quantifies this difference. In subsubsection 3.4.1 and Section 4 we will go in-depth in the study of this issue. Nonetheless, there is a difference between the estimated Gaussian distributions and the real densities. This is the reason why Non-Parametric Tests arose as an effective alternative to their parametric counterparts, as they do not suppose the normality of the input data. In the Bayesian procedure, we would make an estimation of the distribution of the differences between the algorithms’ results as we have made with the results themselves in Figure 1. Then, with the estimated distribution of the parameter of interest, i. e. the difference between the performances of the algorithms, we could extract the desired information. Bayesian paradigm makes statements about the distribution of the difference between the two algorithms, which can reduce the burden of the researchers the NHST do not find significant differences between the algorithms. Other significant concepts in the description of statistical tests come with the comparison of the tests themselves. Type I error occurs when H0is rejected but it should not be. In our scenario, this means that the equivalence of the algorithms has been discarded although there is not a significant difference between them. Otherwise, type II error arises when there is a significant difference but this has not been detected. We denote alpha as the probability of making a type I error (often called significance level or size of test) and beta as the prob4 ability of make a type II error. The power of a test is 1−β, i. e. the probability of rejecting H0when it is false, so we are interested in comparing the power of the tests, because with the same probability of making a type I error, the more powerful a test is, the more differences that will be regarded as significant. There is some notation that is used along this work and it is summarised in Table 1. Notation Description Location H0Null hypothesis Section 3 H1Alternative hypothesis Section 3 HjWhen comparing multiple algorithms, each one of the multiple comparisons. Subsection 3.3 Hi,j When comparing multiple algorithms, the comparison between the algorithm iand j. Subsection 3.3 αSignificance level Section 3 µMean of the performance of an algorithm Section 3.1 kNumber of compared algorithms Sections 3, 5 Nm(µ, Σ) Multivariate Gaussian distribution, with parameters µand Σ Section 3 nNumber of benchmark functions Sections 3,5 θProbability of stated hypothesis Section 5 Dir Dirichlet distribution Section 5 DP Dirichlet Process Section 5 αPrior distribution of DP Section 5 EThe expectation with respect to the Dirichlet Process. Subsection 5.3 δzDirac’s delta centered in zSection 5 I[precondition]Indicator function. Takes 1 when precondition is fulfilled, 0 otherwise. Section 5 Table 1: General notations used in this paper In this paper, we gather the different points of view of the statistical analysis of results in the computational intelligence research field with the associated theoretical background and the properties of the tests. This allows comparisons to be made regarding the appropriateness of the use of each testing paradigm, depending on the specific situation. Moreover, we also adapt the calculation of these confidence curves with a non-parametrical procedure [12], as it is more convenient in the context of the comparison of evolutionary optimisation algorithms. We have included in this paper the performance and the analyses of the tests in the context of the 2017 IEEE Congress on Evolutionary Computation (CEC) Special Session and Competition on Single Objective Real Parameter Numerical Optimisation [13]. This case study allows for a clear illustration of the use of these tests, their behaviour with different distributions of data and the conclusions that can be made with their outputs. The test-suite proposed for this 5 competition can be considered to be a relevant benchmark for the comparison of any newly proposed algorithms. Moreover, the results achieved and final rankings of this competition are supposed to indicate what is the state-of-the-art in evolutionary computation and against which algorithms our proposals should be compared. The code for the tests and the analysis, tables and plots are included as a vignette in the developed Rpackage rNPBST 1[14]. We have also developed a shiny application to facilitate the use of the aforementioned tests. This application processes the results of two or more algorithms and performs the selected test. The results are exported in TEXformat and an HTML table. Cases in which there is a plot associated with the test are highlighted. This paper is organised as follows. In Section 2 we include a survey on the main statistical analyses proposed in the literature for evaluating the classification and optimisation algorithms. Section 3 contains a depiction of the well-known parametric tests and non-parametric tests respectively. In Section 4 we describe the main criticisms and proposals concerning the traditional tests. Section 5 introduces Bayesian test concepts and notations. Due to the relevance of the Multi-Objective problem in the optimisation scientific community, we have gathered the tests that address this issue in Section 6 regardless of their statistical nature. In Section 7, we describe the setting of the CEC’2017 Special Session, which is performed in Section 8 and summarised in Section 9. The lessons learnt and test considerations are presented in Section 10. Section 11 summarises the conclusions obtained in the analyses. 2 Survey on Statistical Analyses Proposed In this section, we provide a chart with an extensive survey on the different statistical proposals made for the comparison of machine learning and optimisation algorithms. The survey is included in Table 2. This table includes a categorisation of the methodologies according to their statistic nature, a brief description of the underlying idea and the considered scenario of the comparison of the data. The proposals are sorted by the year of publication. This order highlights past and present trends in the statistical comparison. The first proposals were made from a frequentist and fundamentally parametric point of view. Later, with the work of Demšar [15] non-parametric tests arose as the alternative to parametric ones in certain circumstances. In recent years, some Bayesian proposals were made, although they still depend on the results’ distributions and are focused on the comparison of classifiers. Moreover, most of the proposals are oriented to compare classifiers. The tests and guidelines made for the comparison of optimisation algorithms are consistent and are constantly used in the literature. However, there is not a specific test that takes account of the multiple runs in the same benchmark 1https://github.com/JacintoCC/rNPBST 6 function and estimates the correlation between these runs, as the tests suggested for cross-validation setups do. 7 Citation Kind of test Description [16] Non-parametric - Classification Cochran Test for the distribution of mistakes on the classification of nsamples. [17] Parametric - Classification McNemar Test and t-test variants for different scenarios of the comparison of two classifiers accuracy. [18] Parametric - Classification Proposal of 5×2 cv F test. [19] Parametric - Optimisation ANOVA test for parameter analysis in genetic algorithms. [20] Frequentist - Classification Introduces corrections for multiple testing (Bonferroni, Tukey, Dunnet, Hsu). [21] Parametric - Classification Proposal using variance estimators of cross-validation results. [22] Non-parametric - Classification Proposal of multiple comparisons using Cochran test. [23] Parametric - Optimisation Guide on the use of ANOVA test in exploratory analyses of genetics algorithms. [24] Parametric - Classification Computation of the sample size for a desired confidence interval width of the True Positive Rate and True Negative Rate. [25] Parametric - Classification Proposal of MultiTest algorithm that order the competitors using post-hoc test. [15] Frequentist - Classification Review of the use of previous parametric tests, Wilcoxon Signed-Rank and Friedman test. [26] Frequentist - Information Retrieval Comparison of pairwise non-parametric tests with Student’s paired t-test. [2] Frequentist - Classification Criticisms and tips on the use of statistical tests. [27] Non-parametric - Classification Extensive proposal of non-parametric tests and associated post-hoc methods. [28] Parametric - Classification k-fold cross validated paired t-test for AUC values. [4] Non-parametric - Optimisation Study of the preconditions for a safe use of the parametric tests and proposal of non-parametric methods for an optimisation scenario. [29] Non-parametric - Classification Study of test application prerequisite in a classification context and with different measures. [30] Non-parametric - Classification Study on the use of statistical tests in neural networks’ results. [31] Non-parametric - Classification Newly proposed post-hoc methods and non-parametric comparison (Li, Holm, Holland, Finner, Hochberg, Hommel and Rom post-hoc procedures). [32] Parametric - Classification Repeated McNemar’s test. [33] Frequentist - Classification Decomposition of the variance of the k-fold CV for prediction error estimation. [34] Permutational - Classification Presentation of two permutational tests that study if the algorithm has learnt the data structure and if it uses the attributes distribution and dependencies. [35] Parametric - Optimisation A multicriteria comparison algorithm (MCStatComp) using the aggregation of the criteria through a non-dominance analysis. [36] Non-parametric - Optimisation Tutorial on the use of non-parametric tests and post-hoc procedures. [37] Non-parametric Review on the use of non-parametric tests and post-hoc tests, and the impact of normality. [38] Non-parametric - Classification Multi2Test algorithm, which orders the algorithms using non-parametric pairwise tests. [39] Non-parametric - Classification Survey on the use of statistical tests in Bioinformatics field. [40] Bayesian - Classification Bayesian Poisson binomial test for pairwise comparison of classifiers. [41] Bayesian - Classification Hierarchical study through Bayesian inference. [42] Parametric - Classification Inclusion of a certain level of allowed error in paired t-test. [43] Parametric - Classification Comparison using McNemar test with different measures. [44] Permutation - Classification Permutation (bootstrap) tests for a cross-validation setup. [45] Parametric - Classification Blocked 3×2cross validation estimator of variance. [5] Non-parametric - Optimisation Analysis of convergence using Page test. [46] Non-parametric - Classification Proposal of Page test for parameter trend study. [47] Bayesian Bayesian version of Wilcoxon Signed Rank test. [48] Bayesian Bayesian test without prior information using Imprecise Dirichlet Process. [49] Bayesian - Classification Bayesian test for two classifiers on multiple data sets accounting the correlation of cross-validation. [50] Bayesian Proposal of Multiple Measures tests. [51] Bayesian Presentation of the Bayesian Friedman test. [52] Parametric - Information retrieval Confidence interval for F1measure using blocked 3×2cross validation. [53] Non-parametric - Classification Generalisation of Wilcoxon rank-sum test for interval data. [54] Non-parametric - Classification Proposal off Mann-Whitney U test for two classifiers and the Kruskal-Wallis H test for multiple classifiers with the associated post-hoc corrections. [55] Frequentist - Classification Proposal of Wald and Score tests for precision comparison. [56] Bayesian - Classification Bayesian hierarchical model for the joint comparison of multiple classifiers on multiple data sets with the cross-validation results. [6] Parametric - Classification Presentation of the confidence curves as the confidence interval generalisation. [57] Non-parametric - Classification Proposal of exact computation of Friedman test. [8] Bayesian Extensive tutorial on the use of Bayesian tests. [58] Non-parametric - Classification New proposals of non-parametric tests that introduce weights. [59] Frequentist - Optimisation Application of Deep Statistical Comparison of Multi-Objective Optimisation algorithms for an ensemble of quality indicators. [60] Bayesian Bayesian analysis based on a model over the algorithms’ rankings. [61] Parametric - Classification Methodology for the definition of the sample sizes. [62] Frequentist - Optimisation Extension of Deep Statistical Comparison, a two-step comparison that select the appropiate parametric or non-parametric test according to the normality of the data. Table 2: Survey on different statistical proposals for results analysis 8 3 Frequentist tests In this section, we describe the classic frequentist tests: the properties of the parametric tests and their assumptions, and the different non-parametric tests and their application in the context of the comparison of single-objective and bound-constrained evolutionary optimisation algorithms. Although there are other tests, algorithms and proposals, as reflected in Table 2, the tests presented in this section represent the core of the statistical comparison methodology. 3.1 Parametric Statistical Tests Parametric tests make the assumption that our sample comes from a distribution that belongs to a known family, usually the Gaussian family, and it is described with a little number of parameters [3]. t-test This classic test is used to compare two samples. Null hypothesis consists in the equivalence of the means of both populations. The main assumptions made by t-test is that the samples have been extracted randomly and the distribution of the populations of the samples are normal. The required input of this test is the group of observations of the different runs of the pair algorithms that will be compared for a single problem. Analysis of Variance When we are interested in the comparison of kdistinct algorithms, we need another test, because repeating the t-test for every pair of algorithms would increment the type I error. The Analysis of Variance (ANOVA) test of the null hypothesis consists in the equivalence of all the means: H0:µ1=· · · =µk, against H1:∃i6=j, µi6=µj. ANOVA test deals with the variance within a group, between groups or a combination of the two types. Here the input consists of a matrix where each column represents an algorithm and each row is a single benchmark function, while the cells contain the mean of the performance for all runs of each algorithm in each benchmark [19]. These tests are very relevant in the statistical comparison, although they have troublesome prerequisites in the field of comparison of optimisation algorithms. Then, non-parametric statistical tests were proposed to address this issue with a known methodology. 3.2 Non-Parametrical Statistical Tests According to Pesarin [63], Pis a non-parametric family of distributions if it is not possible to find a finite-dimensional space Θin which there is a one-to-one relationship between Θand P. This means that we do not have to assume that the underlying distribution belongs to a known family of distributions. Consequently, the prerequisites for non-parametric tests such as symmetry or 9 4 Known Criticisms to Null Hypothesis Statistical Tests The need of the use of statistical tests in the analyses and experimentation is unquestioned. However, certain controversy exists regarding the types of these tests and the implications of their results, mainly from a Bayesian perspective [75]. The most repeated criticisms about the NHST are: •They do not provide the answer that we expected [76, 8]. Commonly, when we are using statistical tests to compare optimisation algorithms, we want to prove that our algorithm outperforms existing algorithms. Then, once an NHST is carried out, a p-value is obtained and this is often misunderstood as the probability of the null hypothesis being true. As we previously mentioned, the p-value is the probability of getting a sample as extreme as the observed one, i. e. the probability of getting that data if the null hypothesis is true, P(D|H0), instead of P(H0|D). •We can almost always reject H0if we get enough data, making a little difference significant through a high number of experiments. This implies that a test could determine a statistically significant difference but without practical implications. On the other hand, a significant difference may not be detected when there is insufficient data. Usually when an NHST is made, all hypotheses with p-value p≤0.05 are rejected, and those with p > 0.05 are considered not significant. This issue could lead to the inclusion in the experimentation of a new proposal of many benchmark functions to cause the rejection of the null hypothesis although the differences between the algorithms were random and very small. •There is an important misconception related to the reproducibility of the experiments, as a lower p-value does not mean a higher probability of obtaining the same results if the experiments are replicated as is described by Berrar and Lozano [73]. This is because, if the null hypothesis is rejected, the probability of obtaining a significant p-value is determined by the power of the test, which depends on the αlevel, the true effect size in the target distribution and the size of the test set, but not on the p-value. •A crucial issue is that we have no information when an NHST does not reject the null hypothesis and not finding a significant difference does not mean that the algorithms performance is equivalent. 5 Bayesian Paradigm and Distribution Estimation In recent years, the Bayesian approach has been proposed as an alternative of frequentist statistics for comparing algorithms performance in optimisation and 16 classification problems [8]. In the following we also describe some differences with respect to the NHST and their problems addressed in the previous section. 5.1 Bayesian Parameter Estimation and Notation In this Bayesian approach, a set of candidate values of a parameter, which includes the possibility of no difference between samples, is set up [7]. Then we compute the relative credibility of the values in the candidate set according to the observations and using Bayesian inference. This approach is preferable in optimisation comparison to Bayesian model comparison approach, because this kind of comparison does not have a null value that brings the NHST objections back to Bayesian analysis. Bayesian analysis is executed in three steps: First step Establish a mathematical model of the data. In a parametric model it will be the likelihood function P(Data|θ). Second step Establish the prior distribution P(θ). The common procedure is to select a prior distribution whose posterior distribution is known. Third step Use Bayes’ rule to obtain the posterior distribution P(θ|Data) from the combination of likelihood and prior distribution. This means that we can see the distribution of the difference of performance between the algorithms, which may reveal that their results differ but that there is no algorithm that outperforms the other. For instance, if one algorithm gets better results in one problem, but worse in another one, using Frequentist statistics we would not reject the hypothesis of equivalence, but we could not know if their results are similar in all the observations. The tests described in this section are oriented to obtain the posterior distribution of the difference of performance between two algorithms, noted as z. We set a prior Dirichlet Process (DP) with base measure α. This process is a probability measure on the set of probability measures on a determined space (as we are interested in the distribution of z, this space will be R). This means that if we make a sample from a DP, we do not obtain a number, like we would do if we make a sample from a uni-dimensional Gaussian distribution, but a probability measure on R. Moreover, as a property of the DP, if P∼DP(α), for any measurable partition of RB={B1, . . . , Bm}, we obtain P(B1), . . . , P(Bm)∼Dir(α(B1), . . . , α(Bm)). This procedure could be thought as if we started with a distribution of the probability of an algorithm outperforming another algorithm. Then, using the observations we change this distribution: so that while before it was a wider distribution with no certainty regarding the results, now it is closer to the observations made. 17 5.2 Bayesian Sign and Signed-Rank Tests This test is presented by Benavoli et al. [47] as a Bayesian version of sign test. We consider a DP to be the prior distribution of the scalar variable z. The DP is determined by the prior strength parameter s > 0and the prior mean α∗=α s, which is a probability measure for DPs. We usually use the measure δz0, i. e. a Dirac’s delta centered on the pseudo-observation z0. Then the posterior probability density function of the difference between algorithms Zfollows the expression [8]: P(z) = w0δz0(z) + n X j=1 wjδzj(z), where (w0, w1, . . . , wn)∼Dir(s, 1,...,1),that is a combination of Dirac delta functions centered on the observations with Dirichlet distributed weights. In the sign test, we compute θl, θeand θr, the probabilities of the mean difference between algorithms being in the intervals (−∞,−r),[−r, r]or (r, +∞)respectively, where ris the limit of the region of practical equivalence (rope) as a linear combination of the observations with the weights wi. Then, θl, θe, θr∼Dir(nl+sI(−∞,−r)(z0), ne+sI[−r,r](z0), nr+sI(r,+∞)(z0)), where nl, ne, nrare the number of observations that fall in each interval. For a version of signed rank test the computation of the probabilities θl, θe, θr is similar to that of the sign test, although it involves pairs of observations in the modification of the probabilities. This time there is not a simple distribution for the probabilities, although we can estimate it using Monte Carlo sampling (w0, w1, . . . , wn)∼Dir(s, 1,...,1). 5.3 Imprecise Dirichlet Process In the Bayesian paradigm we must fix the prior strength sand the prior measure α∗according to the available prior information. If we do not have any information, we should specify a non-informative Dirichlet Process. In our problem we could use the limiting DP when s→0, but it brings out numerical problems like instability in the inversion of Bayesian Friedman covariance matrix. A solution then is assuming s > 0and α∗=δX1=···=Xm, so we get E[E[Xj−Xi]] = 0 for each pair i, j and E[E[Rj]] = m(m−1)/2. Then we are presuming that all the algorithms’ performances are equal, which is not non-informative. The alternative proposed in [47] is to use a prior near-ignorance DP (IDP). This involves fixing s > 0and letting α∗vary in the set of all probability measures, i. e. considering all the possible prior ranks. Posterior inference is then computed taking α∗into account, obtaining lower and upper bounds for the hypothesis probability. 5.4 Bayesian Friedman Test Bayesian sign test is generalised for the comparison among m≥3algorithms by the Bayesian Friedman test [51]. In this section γrepresents the level of 18 significance, as αdenotes the measure. We consider the function R(Xi) = P i6=k=1 I{Xi>Xk}+ 1, the i-th rank. The main goal here is to test if the point µ0= [(m+ 1)/2,...,(m+ 1)/2], i. e. the point of null hypothesis where all the algorithms have the same rank, is in (1 −γ)%SCR(E[R(X1), . . . , R(Xm)]), where the SCR is the symmetric credible region for E[R(X1), . . . , R(Xm)]. If the inclusion does not happen, there is a difference between algorithms with probability 1−γ. For a large number of nobservations, we can suppose that the mean and the covariance matrix tends to the sample mean µand covariance Σ. Then, we define ρ=Finv(1 −γ, m −1, n −m+ 1)(n−1)(m−1) n−m+1 }, where Finv is the inverse of the F distribution. So we assume that µ0∈(1 −γ)%SCR if (µ−µ0)TΣ−1(µ−µ0)|m−1≤ρ, where the notation |m−1means that we take the first m−1components. For a small number of observations, we can compute the SCR by sampling from the posterior DP, resulting in the linear combination of the weights and the Dirac delta functions centered in the observations P=w0δX0+ n P j=1 wjδXj, where X0 is the pseudo-observation. 6 Multiple Measures Tests A relevant concern in the optimisation community is the comparison between multi-objective algorithms. This issue can be addressed by considering a weighted sum of the scores of the different objectives or by calculating the Pareto frontier, selecting the algorithms that are not worse than others in all the criteria or objectives. The statistical approach for this scenario consists in the examination of the relation between two studied elements (algorithms in our case) through multiple observations (benchmarks) and measuring different properties. As in the tests oriented to the comparison of single-objective optimisation algorithms, there exist statistical tests from the different paradigms that address this issue: 6.1 Hotelling’s T2 The parametric approach to the comparison of multivariate groups is the Hotelling’s T2statistics [77]. For this test, we would start with two matrices, gathering the results of two algorithms in mmeasures. This test assumes that the difference from the samples comes from a multivariate Gaussian distribution Nm(µ, Σ) and µand Σare unknown. This assumption can be checked with the generalised version of Shapiro-Wilk’s test [78]. We have not found any proposal of the use of the Hotelling’s T2test to the comparison of algorithms, what can be motivated by the mentioned assumption required. Then, the null hypothesis is that the mean of the mmeasures is the same for the two matrices. This means 19 that if the null hypothesis is rejected, there is at least a measure where the algorithms obtain different values. However, this test has some drawbacks, as the prerequisite of the normality of the results. Besides, in multiple objective optimisation field, the researchers are interested in finding Pareto optimal solutions, and it is not enough knowing that there are differences between the compared algorithms. 6.2 Non-Parametric Multiple Measures Test A relevant proposal to multi-objective algorithms comparison comes from de Campos and Benavoli [79]. In this study, they propose a new test for the comparison of the results of two algorithms in multiple problems and measures. They make a non-parametric proposal for the comparison of two algorithms and its Bayesian version. In this comparison, we have two matrices with the results of two algorithms in nrows representing the problems or benchmark functions measured in m quality criteria or objectives. Then, let M1, . . . , Mmbe a set of mperformance measures, and, in the comparison of the algorithms Aand B, we call a dominance statement D(B,A)= [,≺,...,≺], where the comparison in the i-th entry of D(B,A)means that the algorithm Bis better than Afor Mi. Then we want to decide which (from 2npossibilities) D(B,A)is the most appropriate vector given the results matrices from each algorithm. We denote θ= [θ0, . . . , θ2m−1]the set of probabilities for each possible dominance statement and i∗the index of the most observed configuration. Then, the null hypothesis H0is defined as θi∗≤max(θ\θi∗), i. e. rejecting the null hypothesis would mean rejecting the fact that the probability assigned to the most observed configuration is less or equal to the second greatest probability. The computation of the statistic is detailed in [79]. 6.3 Bayesian Multiple Measures Test The Bayesian version of the Multiple Measures Test follows a Bayesian estimation approach and estimates the posterior probability of the vector of probabilities θ. A Dirichlet distribution is considered to be the prior distribution. Then the weights are updated according to the observations. The posterior probabilities of the dominance statements are computed by Monte Carlo sampling from the space of θand counting the fraction of times for each ithat θiis the maximum of θ. 7 Experimental Framework In this section we describe the framework that will be used in the experiments in Section 8. This way the use of the previously defined test in the scenario of a statistical comparison of the CEC’17 Special Session and Competition on Single Objective Real Parameter Numerical Optimisation can be illustrated [13]. 20 The results have been obtained from the organiser’s GitHub repository 2. The mean final results for the LSHADE-cnEpSin algorithm, whose results have not been correctly included in the organisation data, have been extracted from the original paper. 7.1 Benchmarks Functions The competition goal is finding the minimum of the test functions f(x), where x∈RD, D ∈ {10,30,50,100}. All the benchmark functions are shifted to a global optimum oand scalable and rotated according to Mia rotation matrix. The search range is [−100,100]Dfor all functions. Below we have included a simple summary of the test functions: •Unimodal functions: –Bent Cigar Function –Sum of Different Power Function. The results of this function have been discarded in the experiments because they presented some unstable behavior for the same algorithm presented in different languages. –Zakharov Function •Simple Multimodal Functions: –Disembark Function –Expanded Scaffer’s F6 Function –Lunacek Bi-Rastrigin Function –Non-Continuous Rastrigin’s Function –Levy Function –Schwefel’s Function •Ten Hybrid Functions formed as the sum of different basic functions. •Ten Composition Functions formed as a weighted sum of basic functions plus a bias according to which component optimum is the global one. 7.2 Contestant Algorithms The contestant algorithms are briefly described in the following list, and are ordered according to the ranking obtained in the competition. A short-name has been assigned to each algorithm in order to present clearer tables and plots. 1. EBOwithCMAR (EBO) [80]: Effective Butterfly Optimiser with a new phase which uses Covariace Matrix (CMAR). 2https://github.com/P-N-Suganthan/CEC2017-BoundContrained 21 2. jSO [81]: Improved variant of iL-SHADE algorithm based on a new weighted version of mutation strategy. 3. LSHADE-cnEpSin (LSCNE) [82]: New version of LSHADE-EpSin that uses an ensemble of sinusoidal approaches based on the current adaptation and a modification of the crossover operator with a covariance matrix. 4. LSHADE-SPACMA (LSSPA) [83]: Hybrid version of a proposed algorithm LSHADE-SPA and CMA-ES. 5. DES [84]: An evolutionary algorithm that generates new individuals using a non-elitist truncation selection and an enriched differential mutation. 6. MM-OED (MM)[85]: Multi-method based evolutionary algorithm with orthogonal experiment design (OED) and factor analysis to select the best strategies and crossover operators. 7. IDE-bestNsize (IDEN) [86]: Variant of individual-dependent differential evolution with a new mutation strategy in the last phase. 8. RB-IPOP-CMA-ES (RBI) [87]: New version of IPOP-CMA-ES with a restart trigger according to the midpoint fitness. 9. MOS (MOS11, MOS12, MOS13) [88]: Three large-scale global optimiser used in these scenarios. 10. PPSO [89]: Self tuning Particle Swarm Optimisation relying on Fuzzy Logic. 11. DYYPO [90]: Version of Yin-Yang Pair Optimisation that converts a static archive updating interval into a dynamic one. 12. TLBO-FL (TFL) [91]: Variant to the Teaching Learning Based Optimisation algorithm that includes focused learning of students. 8 Experiments and Results In this section we perform the previously described tests on the competition results in order to provide clear examples of their use. Setup considerations •Here we use the self-developed shiny application shinytests3, which makes use of our Rpackage rNPBST for the analysis of the results of the competition. The associated blocks of code and scripts are available as a vignette in the rNPBST package4. 3https://github.com/JacintoCC/shinytests 4https://jacintocc.github.io/rNPBST/articles/StatisticalAnalysis.html 22 •We will mainly use two data sets, one with all the results (all iterations of the execution of all the algorithms in all benchmark functions for all dimensions), cec17.extended.final, and the mean data set (which aggregates the results among the runs), cec17.final. •In most of the pairwise comparisons we have involved EBOwithCMAR and jSO algorithms, as they are the best-classified algorithms in the competition. •Table 3 shows the mean among different runs of the results obtained at the end of all of the steps of each algorithm on each benchmark function for the 10 dimension scenario. •The tables included in this section are obtained with the function AdjustFormatTable of the package used, which is helpful to highlight the rejected hypotheses by emboldening the associated p-values. 8.1 Parametric Analysis As we have described before, traditionally the statistical tests applied to the comparison of different algorithms belonged to the parametric family of tests. We start the statistical analysis of the results with these kinds of tests and the study of the prerequisites in order to use them safely. The traditional parametric test used in the context of a comparison of multiple algorithms over multiple problems (benchmarks) is the ANOVA test, as we have seen in subsection 3.1. This test makes some assumptions that should be checked before it is performed: 1. The distribution of the results for each algorithm among different benchmarks follows a Gaussian distribution. 2. The standard deviation of results among groups is equal. In Table 4 we gather the p-values associated with the normality of each group of mean results for an algorithm in a dimension scenario. All the null hypotheses are rejected because the p-values are less than 0.05, which means that we reject that the distribution of the mean results for each benchmark function follow a normal distribution. This conclusion could be expected because of the different difficulty of the benchmark functions in higher dimension scenarios. This is marked with boldface in subsequent tables. In some circumstances like the Multi-Objective Optimisation we need to include different measures in the comparison. We will now consider the results of the different dimensions as if they were different measures in the same benchmark function. Then, in order to perform the Hotelling’s T2test we first check the normality of the population with the multivariate generalisation of Shapiro-Wilk’s test. Table 5 shows that the normality hypothesis is rejected for every algorithm. Therefore, we stop the parametric analysis of the results here because the assumptions of parametric tests are not satisfied. 23 Benchmark DES DYYPO EBO IDEN jSO LSSPA MM MOS11 MOS12 MOS13 PPSO RBI TFL LSCNE 10.000 2855.620 0.000 0.000 0.000 0.000 0.000 691.550 2891.970 3916.210 239.250 0.000 2022.580 0.000 30.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 40.000 2.070 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 1.200 0.000 3.030 0.000 51.540 11.200 0.000 3.280 1.760 1.760 1.110 6.940 64.750 16.420 18.080 1.580 8.750 1.690 60.120 0.000 0.000 0.000 0.000 0.000 0.000 0.000 50.740 0.000 0.230 0.000 0.000 0.000 711.930 21.800 10.550 12.890 11.790 10.930 11.520 18.960 48.510 27.730 16.910 10.110 27.640 11.980 81.560 13.250 0.000 2.890 1.950 0.840 1.110 6.970 65.030 14.220 9.950 1.970 12.280 1.800 90.000 0.020 0.000 0.000 0.000 0.000 0.000 0.000 2622.080 0.000 0.000 0.000 0.010 0.000 10 5.660 367.030 37.210 190.420 35.900 21.830 17.890 360.260 1335.450 610.390 503.090 435.430 954.620 43.030 11 0.120 9.280 0.000 0.000 0.000 0.000 0.000 2.540 32.530 7.180 16.890 0.170 4.120 0.000 12 440.840 13 492.010 90.150 2.440 2.660 119.440 101.600 11 049.040 15 052.370 12 283.250 4551.420 110.460 65 562.150 101.280 13 3.310 5079.250 2.170 0.840 2.960 4.370 4.190 3293.930 11 951.200 4367.380 1391.140 4.170 2447.820 3.660 14 12.250 20.670 0.060 0.000 0.060 0.160 0.090 158.280 7602.910 275.820 37.310 15.930 67.280 0.080 15 3.250 43.640 0.110 0.010 0.220 0.410 0.070 89.500 5916.640 408.540 53.270 0.490 125.900 0.320 16 6.090 43.840 0.420 0.490 0.570 0.740 0.250 34.470 496.950 13.510 82.980 97.100 8.910 0.540 17 21.040 14.090 0.150 0.790 0.500 0.160 0.060 2.890 210.460 22.220 24.590 52.460 38.300 0.310 18 29.130 8757.310 0.700 0.050 0.310 4.350 0.970 1401.920 9489.180 6393.550 878.260 19.720 6150.510 3.860 19 2.520 92.670 0.020 0.010 0.010 0.230 0.000 9.450 5658.600 1093.150 22.480 1.820 60.590 0.040 20 12.170 8.010 0.150 0.000 0.340 0.310 0.070 0.030 281.400 7.700 27.820 106.050 14.610 0.260 21 201.580 100.370 114.020 149.230 132.380 100.710 104.070 173.060 278.590 147.270 104.330 137.350 142.400 146.360 22 100.000 97.470 98.460 96.080 100.000 100.040 100.020 87.960 1260.220 99.570 96.730 99.260 93.310 100.010 23 301.460 309.440 300.170 301.910 301.210 302.680 298.420 311.920 368.510 318.160 342.150 275.020 306.870 302.000 24 303.040 116.850 166.210 292.970 296.600 274.600 103.930 276.240 345.060 301.330 226.630 197.580 310.380 315.830 25 407.650 422.990 412.350 413.910 405.960 428.370 414.050 415.020 433.670 430.510 404.100 402.270 425.520 425.560 26 296.080 303.380 265.400 300.000 300.000 300.000 294.120 254.780 1149.080 256.020 266.670 272.690 300.890 300.000 27 395.720 396.490 391.570 392.800 389.390 389.650 389.500 392.670 452.080 396.180 426.520 394.860 392.510 389.500 28 526.070 300.980 307.140 322.810 339.080 317.240 336.810 368.570 515.800 366.630 294.120 402.260 446.980 384.880 29 236.080 260.310 231.190 236.650 234.200 231.440 235.680 269.450 559.670 280.450 277.870 265.830 274.500 228.410 30 154 002.280 6701.910 406.680 403.790 394.520 430.180 56 938.190 21 723.410 457 211.160 60 963.140 2993.250 2045.990 278 998.940 17 618.430 Table 3: Mean final results across runs in 10 dimension scenario 24 Algorithm Dim. 10 Dim. 30 Dim. 50 Dim. 100 DES 1.40 ·10−11 6.37 ·10−08 1.36 ·10−11 3.85 ·10−08 DYYPO 7.62 ·10−09 2.35 ·10−11 3.67 ·10−11 1.68 ·10−11 EBO 2.88 ·10−06 1.22 ·10−07 1.39 ·10−11 2.71 ·10−08 IDEN 3.15 ·10−06 4.24 ·10−08 1.59 ·10−11 1.21 ·10−10 jSO 1.28 ·10−06 1.42 ·10−07 1.40 ·10−11 9.00 ·10−09 LSSPA 2.84 ·10−06 3.76 ·10−07 1.39 ·10−11 3.71 ·10−08 MM 1.51 ·10−11 1.68 ·10−07 1.40 ·10−11 7.61 ·10−07 MOS11 3.38 ·10−10 3.75 ·10−10 1.51 ·10−10 1.89 ·10−09 MOS12 2.16 ·10−11 2.42 ·10−08 3.64 ·10−11 2.79 ·10−10 MOS13 1.09 ·10−10 2.82 ·10−09 3.03 ·10−11 9.56 ·10−11 PPSO 7.57 ·10−09 1.77 ·10−10 3.32 ·10−10 4.49 ·10−11 RBI 4.75 ·10−09 2.12 ·10−07 5.75 ·10−11 2.07 ·10−10 TFL 4.47 ·10−11 7.50 ·10−11 2.41 ·10−09 1.58 ·10−11 LSCNE 2.17 ·10−11 2.20 ·10−07 1.39 ·10−11 1.95 ·10−08 Table 4: p-values for Shapiro tests for the normality of the mean results 8.2 Non-parametric Tests In this subsection we perform the most popular tests in the field of the comparison of optimisation algorithms. We continue using the aggregated results across the different runs at the end of the iterations, except for the Page test for the study of the convergence. 8.2.1 Classic tests Pairwise comparisons First, we perform the non-parametric pairwise comparisons with the Sign, Wilcoxon and Wilcoxon Rank-Sum tests described in subsubsection 3.2.2 for the 10 dimension scenario. The hypotheses of the equality of the medians is only rejected by the Wilcoxon Rank-Sum test, and we can see in Table 6 how for example Wilcoxon’s statistics R+and R−are both high numbers, which means that there is no significant difference between the ranking of the observations where one algorithm outperforms the other. Then, the next step requires all the different algorithms in the competition to be involved in the comparison. Multiple comparisons We can see in Table 7 how the tests that involve multiple algorithms, described in subsubsection 3.2.3 reject the null hypotheses, that is, the equivalence of the medians of the results of the different benchmarks. We must keep in mind that a comparison between thirteen algorithms is not the recommended procedure if we want to compare our proposal. We should only include the state-of-the-art algorithms in the comparison, because the inclusion of an algorithm with lower performance could lead to the rejection of the null hypothesis, not due to the differences between our algorithm and the comparison 25 8.3.1 Bayesian Friedman test We start with the Bayesian version of the Friedman test, mentioned in subsection 5.4. In this test we do not obtain a single p-value, but the accepted hypothesis. Due to the high number of contestant algorithms and the memory needed to allocate the covariance matrix of all the possible permutations, here we will perform the imprecise version of the test. The null hypothesis of the equivalence of the mean ranks for the 10 dimension scenario is rejected. The mean ranks of the algorithms are shown in Table 14. Algorithm Mean Rank DES 7.98 DYYPO 9.62 EBO 3.55 IDEN 4.9 jSO 4.7 LSSPA 5.6 MM 4.28 MOS-11 8.35 MOS-12 13.32 MOS-13 10.41 PPSO 8.81 RBI 6.01 TFL 10.58 LSCNE 5.9 Table 14: Mean Rank of Bayesian Friedman Test 8.3.2 Bayesian Sign and Signed-Rank test The original proposal of the use of the Bayesian Sign and Signed-Rank tests included in subsection 5.2 is the comparison of classification algorithms and the proposed rope is [−0.01,0.01] for a measure in the range [0,1]. In the scenario of optimisation problems, we should be concerned that the possible outcomes are lower-bounded by 0 but in many functions, there is not an upper bound or the maximum is very high, so we must follow another approach. As the difference in the 100 dimension comparison is between 0 and 15000, we state that the region of practical equivalence is [−10,10]. The tests compute the probability of the true location of EBO-CMAR−jSO with respect to 0, so both tests’ results shows that there is a similar probability for the three hypotheses. In the Bayesian Sign test results, the hypothesis with a greater probability is the left region (i.e. the true location is less than 0 and then jSO obtain worse results), although the results are not significant enough to state that EBO is the winner algorithm. Following the results of the Bayesian Signed-Rank test, left region is also the most probable, although this 32 is not significant. Rope probability is also high, so the equivalence cannot be discarded, which means that there are several ties in the benchmark results. We can see the posterior probability density of the parameter in Figure 4, where each point represents an estimation of the probability of the parameter which belongs to each region of interest. The vertexes of the triangle represent the points where there is probability 1 of the true location being in this region. The proportion of the location of the points is shown in Table 15. This means that we have repeatedly obtained the triplets of the probability of each region to be the true location of the difference between the two samples, and then we have plotted these triplets to obtain the posterior distribution of the parameter. If we compare these results with a paired Wilcoxon test, we see that the null hypothesis of the equivalence of the means is rejected, although there is no information about if one algorithm outperforms the other. However, using the Bayesian paradigm we can see that this is not the situation, as we cannot establish the dominance of one algorithm over the other either. Following the frequentist paradigm we could be tempted to (erroneously) establish that EBO is better, according to a single statistic and the result of the Wilcoxon test. Wilcoxon Signed Ranks Statistic V 655. p-value 0.002 Bayesian Sign jSO 0.2777823 Posterior probability rope 0.3112033 EBO 0.4110144 Bayesian Signed-Rank jSO 0.2905800 Posterior probability rope 0.3409561 EBO 0.3684639 Table 15: Bayesian Sign and Signed Ranks tests 20 40 60 80 100 20 40 60 80 100 20 40 60 80 100 rope jSO EBO 20 40 60 80 100 20 40 60 80 100 20 40 60 80 100 rope jSO EBO Figure 4: Bayesian Sign and Bayesian Signed-Rank tests 33 8.3.3 Imprecise Dirichlet Process Test The Imprecise Dirichlet Process, as we have seen in subsection 5.3, consists of a more complex idea of the previous tests although the implications of the use of the Bayesian Tests could be clearer. In this test, we try to not introduce any prior distribution, not even the prior distribution where both algorithms have the same performance, but all the possible probability measures α, and then obtain an upper and a lower bounds for the probability in which we are interested. The input consists of the aggregated results among the different runs for all the benchmark functions for a single dimension. The other parameters of the function are the strong parameter sand the pseudo-observation c. With these data, we obtain two bounds for the posterior probability of the first algorithm outperforming the second one, i.e. the probability of P(X≤Y)≥0.5. These are the possible scenarios: •Both bounds are greater than 0.95: Then we can say that the first algorithm outperforms the second algorithm with 95% probability. •Both bounds are lower than 0.05: This is the inverse case. In this situation, the second algorithm outperforms the first algorithm with 95% probability. •Both bounds are between 0.05 and 0.95: Then we can say that the probability of one algorithm outperforming the other is lower than the desired probability of 0.95. •Finally, if only one of the bounds is greater than 0.95 or lower than 0.05, the situation is undetermined and we cannot decide. IDP - Wilcoxon test Posterior Distribution Upper Bound 0.591 Lower Bound 0.452 Table 16: Imprecise Dirichlet Process of Wilcoxon test According to the results of the Bayesian Imprecise Dirichlet Process (see Table 16), the probability under the Dirichlet Process of P(EBO ≤jSO)≥0.5, that is the probability of EBO-CMAR outperforming jSO, is between 0.59 and 0.45, so there is not a probability greater than 0.95 of EBO-CMAR outperforming jSO. These numbers represent the area under the curve of the upper and lower distributions when P(X≤Y)≥0.5. In Figure 5 we can see both posterior distributions. 34 0 2 4 0.2 0.4 0.6 0.8 P(X <= Y) Density Posterior distribution area Lower: 0.437 Upper: 0.591 JSO vs. EBO−CMAR Figure 5: Imprecise Dirichlet Process - Wilcoxon test 8.4 Multi-Objective Comparison In some circumstances like the Multi-Objective Optimisation, we need to include different measures in the comparison. We include in this section the illustration of the use of the tests presented in Section 6. 8.4.1 Multiple measures test - GLRT As we have mentioned in subsection 8.1 concerning the Hotelling’s T2test, we can be interested in the simultaneous comparison of multiple measures. This is the scenario of application of the Non-Parametric Multiple Measures test, described in subsubsection 8.4.1. We select the means of the executions of the two best algorithms and reshape them into a matrix with the results of each benchmark in the rows and the different dimensions in the columns. Then we use the test to see which hypothesis of dominance is the most probable and if we can state that the probability of this dominance statement is significant. According to the results shown in Table 17, we obtain that the most observed dominance statement is the configuration [<, <, <, <], it is, EBO-CMAR obtains a better result for all the dimensions. However, the associated null hypothesis which states that the mentioned configuration is no more probable than the following one, obtain a p-value of 0.39, showed in Table 17, so this hypothesis cannot be rejected. The second most probable configuration is [>, >, >, >]which means that EBO-CMAR obtains worse results in all the dimensions, so we cannot be certain which is the most probable situation in any scenario. It is relevant to note that the number of observations can be a real value as the weight of an observation is divided between the possible configuration when there is a tie for any measure. 35 GLRT Multiple Measures λ0.6917945 Configuration Number of observations < < < < 8.25 < < < > 2.62 < < > < 1.25 < < > > 0.12 < > < < 1.25 < > < > 0.12 < > > < 1.25 < > > > 3.12 > < < < 0.75 > < < > 1.62 > < > < 1.25 > < > > 0.12 > > < < 0.75 > > < > 0.12 > > > < 1.25 > > > > 5.12 p-value 0.3906 Table 17: Posterior configuration probability 8.4.2 Bayesian Multiple Measures Test In the Bayesian version of the Multiple Measures test the results are analogous to the frequentist version shown in the previous section. We use the same matrices with the results of each algorithm arranged by dimensions. Here we obtain that the most probable configuration is also the dominance of EBO in all the dimensions according to this test, but the posterior probability is 0.75, as is shown in Table 18, so we cannot say that the difference with respect to the remaining configurations is determinant. 9 Summary of results in CEC’2017 In this section, we include a summary of the statistical analysis performed within the context of the CEC’2017. Table 19 shows the official results of the algorithms in their scoring system and the scores computed following the indications of the report of the problem definition for the competition [13] with the available raw results of the algorithms. The Score1is defined using a weighted sum of the errors of the algorithms in all benchmark on different dimensions. The weights are 0.1,0.2,0.3,0.4, with the higher weights corresponding with the higher dimension scenarios. Then, if we call SE the summed error for an algorithm and 36 Configuration Probability < < < < 0.750 < < < > 0.020 < < > < 0.000 < < > > 0.000 < > < < 0.000 < > < > 0.000 < > > < 0.000 < > > > 0.030 > < < < 0.000 > < < > 0.010 > < > < 0.000 > < > > 0.000 > > < < 0.000 > > < > 0.000 > > > < 0.000 > > > > 0.170 Table 18: Posterior configuration probability SEmin the minimum sum for an algorithm, Score1is defined as Score1 = 0.5∗(1 −SE−SEmin SE ). Analogously, Score2is defined as a weighted sum of the ranks of the algorithms instead of using the error: Score2=0.5∗(1−SR−SRmin SR ), where SR is the weighted sum of the ranks. The difference between the official score and our computation may reside in the aggregation method for the results of the different runs or the programming language used for the computation of the scores. Different versions of the results have been used without major impact in the final ranking, which proves the robustness of the ranking and the algorithms. However, these differences do not affect the final ranking of the first classified algorithms. In the official CEC’17 summary, the results of the MOS12 algorithm are not included, so we have excluded them in the analyses of this section. The classic statistical analysis that should be made in the context of a competition would involve a non-parametric test with post-hoc test for a nversus nscenario, as we do not have a preference for the results of comparison of any specific algorithm. In order to preserve the relative importance of the results in the different dimension scenarios, we show the plots of the critical difference explained in subsection 3.3 for the four scenarios. As we can see in Figure 6, summary scores presented in Table 19 are consistent with the Critical Differences plots made with the mean values, as the first classified algorithms are also in the first positions of the graphical representation. However, the statistical tests make it possible to address the fact that there is a group of algorithms in the lead group in every dimension scenario whose associated hypothesis of equivalence cannot be discarded. These algorithms are LSSPA, DES, LSCNE, MM, jSO and EBO. Moreover, there is not a 37 Official CEC’17 Results Score Computation Algorithm Score1 Score2 Score Algorithm Score1 Score2 Score EBO 50.00 48.01 98.01 EBO 50.00 50.00 100.00 jSO 49.69 47.08 96.77 jSO 49.69 43.01 92.70 LSCNE 46.82 49.74 96.56 LSCNE 46.82 44.75 91.56 LSSPA 46.44 50.00 96.44 LSSPA 46.44 44.73 91.17 DES 45.94 43.20 89.14 DES 45.94 40.65 86.59 MM 45.96 40.12 86.07 MM 45.96 36.16 82.12 IDEN 29.85 27.68 57.53 IDEN 29.85 26.15 56.00 RBI 3.79 33.61 37.40 MOS13 18.94 17.33 36.27 MOS13 18.94 17.34 36.28 RBI 3.79 32.00 35.79 MOS11 11.09 19.30 30.39 MOS11 11.09 19.17 30.25 PPSO 3.93 17.36 21.28 PPSO 3.93 17.26 21.19 DYYPO 0.59 17.03 17.62 DYYPO 0.59 17.06 17.65 TFL 0.03 16.25 16.27 TFL 0.03 16.31 16.34 Table 19: CEC17 Results Scores with mean results single scenario where the winner equivalences with this lead group algorithms can be discarded. 345678 9 10 11 CD EBO MM jSO IDEN LSSPA LSCNE RBI DES MOS11 PPSO DYYPO MOS13 TFL (a) CD plot Dim10 345678 9 10 11 12 CD EBO jSO LSCNE LSSPA MM DES IDEN RBI MOS11 PPSO TFL MOS13 DYYPO (b) CD plot Dim30 345678 9 10 11 12 CD LSSPA EBO jSO LSCNE DES MM RBI IDEN MOS11 DYYPO TFL MOS13 PPSO (c) CD plot Dim50 345678 9 10 11 12 CD DES LSSPA LSCNE EBO jSO MM RBI IDEN MOS11 MOS13 PPSO DYYPO TFL (d) CD plot Dim100 Figure 6: Plots of Critical Differences In the Bayesian paradigm, after having rejected the equivalence of all the mean ranks of the algorithms with the Friedman tests, we repeatedly perform the Bayesian Signed-Rank for every pair of algorithms in each dimension scenario. The results are summarised in Figures 7-10. Especially in lower dimensions, there is less certainty than there was in the non-parametric analysis concerning the dominance of an algorithm over the other in each comparison, although 38 from the Bayesian perspective we can state where there is a tie between a pair of algorithms and the direction of the difference, while the equivalence with NHST cannot be assured. The tiles for each column and row represent the comparison between the two algorithms. The colour depends on the result of the comparison, indicating if the greater probability belongs to the region of an algorithm or the rope. The probability of this hypothesis is written in the tile as well as represented in the opacity of the colour, to highlight the greater probabilities. 0.47 0.53 0.61 0.61 0.59 0.61 0.44 0.53 0.55 0.40 0.580.61 0.63 0.56 0.59 0.59 0.57 0.39 0.46 0.38 0.48 0.430.50 0.72 0.76 0.73 0.76 0.65 0.75 0.73 0.50 0.740.73 0.92 0.75 0.71 0.60 0.68 0.66 0.44 0.660.71 0.68 0.75 0.60 0.72 0.66 0.44 0.680.79 0.85 0.61 0.73 0.67 0.43 0.710.71 0.58 0.75 0.66 0.44 0.730.71 0.58 0.50 0.50 0.540.55 0.63 0.63 0.420.62 0.57 0.640.58 0.670.48 0.63 0.47 0.53 0.61 0.61 0.59 0.61 0.44 0.53 0.55 0.40 0.58 0.61 0.63 0.56 0.59 0.59 0.57 0.39 0.46 0.38 0.48 0.43 0.50 0.72 0.76 0.73 0.76 0.65 0.75 0.73 0.50 0.74 0.73 0.92 0.75 0.71 0.60 0.68 0.66 0.44 0.66 0.71 0.68 0.75 0.60 0.72 0.66 0.44 0.68 0.79 0.85 0.61 0.73 0.67 0.43 0.71 0.71 0.58 0.75 0.66 0.44 0.73 0.71 0.58 0.50 0.50 0.54 0.55 0.63 0.63 0.42 0.62 0.57 0.64 0.58 0.67 0.48 0.63 TFL RBI PPSO MOS13 MOS11 MM LSSPA LSCNE jSO IDEN EBO DYYPO DES DES DYYPO EBO IDEN jSO LSCNE LSSPA MM MOS11MOS13 PPSO RBI TFL Algorithm 2 Algorithm 1 Winner Alg. 1 Alg. 2 rope Dimension 10 Figure 7: Bayesian Signed-Rank Dim10 As the error obtained in the competition increases, more comparisons are marked as significant and less ties between algorithms are detected. In Figure 7 we see how in the comparison of the first classified algorithms the most probable situation is a tie. This group starts to win with a greater probability in the 30 Dimension scenario (Figure 8), while the ties persist within the lead group. 39 0.95 0.52 0.47 0.50 0.56 0.43 0.93 0.92 0.92 0.59 0.950.49 0.95 0.95 0.95 0.95 0.95 0.59 0.45 0.53 0.90 0.560.95 0.54 0.76 0.75 0.59 0.94 0.92 0.92 0.63 0.950.75 0.49 0.46 0.54 0.95 0.94 0.93 0.55 0.930.52 0.65 0.52 0.94 0.92 0.92 0.62 0.950.76 0.60 0.94 0.92 0.92 0.62 0.950.76 0.94 0.93 0.93 0.68 0.950.61 0.47 0.53 0.91 0.560.94 0.60 0.90 0.520.92 0.89 0.570.92 0.870.65 0.95 0.95 0.52 0.47 0.50 0.56 0.43 0.93 0.92 0.92 0.59 0.95 0.49 0.95 0.95 0.95 0.95 0.95 0.59 0.45 0.53 0.90 0.56 0.95 0.54 0.76 0.75 0.59 0.94 0.92 0.92 0.63 0.95 0.75 0.49 0.46 0.54 0.95 0.94 0.93 0.55 0.93 0.52 0.65 0.52 0.94 0.92 0.92 0.62 0.95 0.76 0.60 0.94 0.92 0.92 0.62 0.95 0.76 0.94 0.93 0.93 0.68 0.95 0.61 0.47 0.53 0.91 0.56 0.94 0.60 0.90 0.52 0.92 0.89 0.57 0.92 0.87 0.65 0.95 TFL RBI PPSO MOS13 MOS11 MM LSSPA LSCNE jSO IDEN EBO DYYPO DES DES DYYPO EBO IDEN jSO LSCNE LSSPA MM MOS11MOS13 PPSO RBI TFL Algorithm 2 Algorithm 1 Winner Alg. 1 Alg. 2 rope Dimension 30 Figure 8: Bayesian Signed-Rank Dim30 The results of the 50 dimension scenario, shown in Figure 9, coincide with the conclusions obtained in the non-parametric analysis. In this scenario the lead group is reduced to EBO, jSO, LSCNE and DES, and the probabilities of the ties are lower than in previous scenarios. Finally in 100 dimension scenario, DES wins in the comparisons with all the remaining algorithms, with probabilities between 0.59 in the comparison versus LSCNE and 0.97 versus PPSO. 0.96 0.41 0.74 0.35 0.47 0.47 0.95 0.95 0.96 0.71 0.960.39 0.96 0.92 0.96 0.96 0.96 0.52 0.48 0.50 0.89 0.530.96 0.82 0.58 0.60 0.52 0.95 0.95 0.97 0.66 0.960.47 0.88 0.81 0.66 0.94 0.91 0.95 0.49 0.960.82 0.53 0.55 0.95 0.95 0.97 0.65 0.960.52 0.45 0.96 0.95 0.97 0.64 0.960.42 0.95 0.94 0.96 0.60 0.960.49 0.50 0.53 0.82 0.600.95 0.52 0.82 0.580.95 0.79 0.630.97 0.830.64 0.96 0.96 0.41 0.74 0.35 0.47 0.47 0.95 0.95 0.96 0.71 0.96 0.39 0.96 0.92 0.96 0.96 0.96 0.52 0.48 0.50 0.89 0.53 0.96 0.82 0.58 0.60 0.52 0.95 0.95 0.97 0.66 0.96 0.47 0.88 0.81 0.66 0.94 0.91 0.95 0.49 0.96 0.82 0.53 0.55 0.95 0.95 0.97 0.65 0.96 0.52 0.45 0.96 0.95 0.97 0.64 0.96 0.42 0.95 0.94 0.96 0.60 0.96 0.49 0.50 0.53 0.82 0.60 0.95 0.52 0.82 0.58 0.95 0.79 0.63 0.97 0.83 0.64 0.96 TFL RBI PPSO MOS13 MOS11 MM LSSPA LSCNE jSO IDEN EBO DYYPO DES DES DYYPO EBO IDEN jSO LSCNE LSSPA MM MOS11MOS13 PPSO RBI TFL Algorithm 2 Algorithm 1 Winner Alg. 1 Alg. 2 rope Dimension 50 Figure 9: Bayesian Signed-Rank Dim50 40 0.96 0.65 0.94 0.70 0.62 0.73 0.94 0.94 0.97 0.80 0.960.59 0.96 0.95 0.96 0.96 0.96 0.71 0.59 0.59 0.96 0.660.96 0.94 0.37 0.47 0.44 0.94 0.94 0.97 0.60 0.960.53 0.94 0.94 0.95 0.87 0.87 0.89 0.64 0.960.94 0.56 0.46 0.94 0.94 0.97 0.56 0.960.55 0.55 0.94 0.94 0.97 0.60 0.960.48 0.94 0.94 0.97 0.61 0.960.56 0.56 0.58 0.92 0.730.95 0.54 0.86 0.810.94 0.84 0.750.97 0.960.60 0.96 0.96 0.65 0.94 0.70 0.62 0.73 0.94 0.94 0.97 0.80 0.96 0.59 0.96 0.95 0.96 0.96 0.96 0.71 0.59 0.59 0.96 0.66 0.96 0.94 0.37 0.47 0.44 0.94 0.94 0.97 0.60 0.96 0.53 0.94 0.94 0.95 0.87 0.87 0.89 0.64 0.96 0.94 0.56 0.46 0.94 0.94 0.97 0.56 0.96 0.55 0.55 0.94 0.94 0.97 0.60 0.96 0.48 0.94 0.94 0.97 0.61 0.96 0.56 0.56 0.58 0.92 0.73 0.95 0.54 0.86 0.81 0.94 0.84 0.75 0.97 0.96 0.60 0.96 TFL RBI PPSO MOS13 MOS11 MM LSSPA LSCNE jSO IDEN EBO DYYPO DES DES DYYPO EBO IDEN jSO LSCNE LSSPA MM MOS11MOS13 PPSO RBI TFL Algorithm 2 Algorithm 1 Winner Alg. 1 Alg. 2 rope Dimension 100 Figure 10: Bayesian Signed-Rank Dim100 10 Discussion and Lessons Learnt In this section we include some considerations about the use of the tests described in previous sections and other proposed tests. Criticisms on the p-value The criticisms made regarding the p-value and NHST are not limited to the Evolutionary Optimisation field or even Computer Science [92], but occur more frequently in other research fields, like psychology [93] or neuroscience [94]. A recent Nature paper [95] warns about the common mistakes in the interpretation about the meaning of the p-value, especially the statements about the alleged “no difference” between the groups when the null hypothesis is not rejected. This is one of the most direct and powerful arguments for the promotion of the use of the Bayesian tests, as the posterior distribution reflects the behaviour of the parameter of interest, and it is more difficult to conclude that the algorithms’ performance is the same if it is not reflected in the plots of the distribution. The other main practical criticism is related to the researcher’s intention and the effect size, as an elevated number of samples could derive in the rejection of the null hypothesis even with a tiny effect size. The Bayesian approach is not affected in the same way, as increasing the number of samples should be reflected in a posterior distribution closer to the underlying one. Another controversial aspect of the NHST is their performance in a dichotomous way in order to determine the result of the experiment. However, these criticisms are not restricted to the frequentist paradigm and also affects the Bayesian perspective. Similar opinions appeared in ASA’s statement [96], which represents a major setback to the use of NHST. They do recognise that to obtain 41 [49] G. Corani, A. Benavoli, A Bayesian approach for comparing cross-validated algorithms on multiple data sets, Machine Learning 100 (2-3) (2015) 285– 304. doi:10.1007/s10994-015-5486-z. [50] A. Benavoli, C. P. de Campos, Statistical Tests for Joint Analysis of Performance Measures, in: Advanced Methodologies for Bayesian Networks - Second International Workshop, AMBN 2015, Yokohama, Japan, November 16-18, 2015. Proceedings, 2015, pp. 76–92. doi:10.1007/ 978-3-319-28379-1_6. [51] A. Benavoli, G. Corani, F. Mangili, M. Zaffalon, A Bayesian nonparametric procedure for comparing algorithms, in: Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015, 2015, pp. 1264–1272. [52] Y. Wang, J. Li, Y. Li, R. Wang, X. Yang, Confidence Interval for F_1 Measure of Algorithm Performance Based on Blocked 3x2 Cross-Validation, IEEE Transactions on Knowledge and Data Engineering 27 (3) (2015) 651– 659. doi:10.1109/TKDE.2014.2359667. [53] J. Perolat, I. Couso, K. Loquin, O. Strauss, Generalizing the Wilcoxon rank-sum test for interval data, International Journal of Approximate Reasoning 56 (2015) 108–121. doi:10.1016/j.ijar.2014.08.001. [54] P. K. Singh, R. Sarkar, M. Nasipuri, Statistical validation of multiple classifiers over multiple datasets in the field of pattern recognition, International Journal of Applied Pattern Recognition 2 (1) (2015) 1–23. [55] L. Gondara, Classifier comparison using precision, arXiv preprint arXiv:1609.09471 (2016). [56] G. Corani, A. Benavoli, J. Demšar, F. Mangili, M. Zaffalon, Statistical comparison of classifiers through Bayesian hierarchical modelling, Machine Learning (2016) 1–21. [57] R. Eisinga, T. Heskes, B. Pelzer, M. Te Grotenhuis, Exact p-values for pairwise comparison of Friedman rank sums, with application to comparing classifiers, BMC Bioinformatics 18 (1) (Dec. 2017). doi:10.1186/ s12859-017-1486-2. [58] Z. Yu, Z. Wang, J. You, J. Zhang, J. Liu, H. Wong, G. Han, A New Kind of Nonparametric Test for Statistical Comparison of Multiple Classifiers Over Multiple Datasets, IEEE Transactions on Cybernetics 47 (12) (2017) 4418–4431. doi:10.1109/TCYB.2016.2611020. [59] T. Eftimov, P. Korošec, B. K. Seljak, Comparing multi-objective optimization algorithms using an ensemble of quality indicators with deep statistical comparison approach, in: 2017 IEEE Symposium Series on Computational Intelligence (SSCI), 2017, pp. 1–8. doi:10.1109/SSCI.2017.8280910. 48 [60] B. Calvo, J. Ceberio, J. A. Lozano, Bayesian Inference for Algorithm Ranking Analysis, in: Proceedings of the Genetic and Evolutionary Computation Conference Companion, GECCO ’18, ACM, New York, NY, USA, 2018, pp. 324–325. doi:10.1145/3205651.3205658. [61] F. Campelo, F. Takahashi, Sample size estimation for power and accuracy in the experimental comparison of algorithms, Journal of Heuristics 25 (2) (2019) 305–338, 00002. doi:10.1007/s10732-018-9396-7. [62] T. Eftimov, P. Korošec, A novel statistical approach for comparing metaheuristic stochastic optimization algorithms according to the distribution of solutions in the search space, Information Sciences 489 (2019) 255–273, 00000. doi:10.1016/j.ins.2019.03.049. [63] F. Pesarin, L. Salmaso, Permutation Tests for Complex Data: Theory, Applications and Software, John Wiley & Sons, 2010. [64] D. W. Nordstokke, B. D. Zumbo, A new nonparametric Levene test for equal variances, Psicológica 31 (2) (2010). [65] E. Kasuya, Wilcoxon signed-ranks test: Symmetry should be confirmed before the test, Animal Behaviour 79 (3) (2010) 765–767. doi:10.1016/ j.anbehav.2009.11.019. [66] W. J. Dixon, A. M. Mood, The statistical sign test, Journal of the American Statistical Association 41 (236) (1946) 557–566. [67] F. Wilcoxon, Individual Comparisons by Ranking Methods, Biometrics Bulletin 1 (6) (1945) 80–83. doi:10.2307/3001968. [68] W. J. Conover, R. L. Iman, Rank transformations as a bridge between parametric and nonparametric statistics, The American Statistician 35 (3) (1981) 124–129. [69] A. L. Rhyne, R. G. D. Steel, Tables for a Treatments Versus Control Multiple Comparisons Sign Test, Technometrics 7 (3) (1965) 293–306. [70] R. L. Iman, J. M. Davenport, Approximations of the critical region of the fbietkan statistic, Communications in Statistics - Theory and Methods 9 (6) (1980) 571–595, 00709. doi:10.1080/03610928008827904. [71] J. L. Hodges, E. L. Lehmann, Rank Methods for Combination of Independent Experiments in Analysis of Variance, The Annals of Mathematical Statistics 33 (2) (1962) 482–497. doi:10.1214/aoms/1177704575. [72] D. Quade, Rank analysis of covariance, Journal of the American Statistical Association 62 (320) (1967) 1187–1200. 49 [73] D. Berrar, J. A. Lozano, Significance tests or confidence intervals: Which are preferable for the comparison of classifiers?, Journal of Experimental & Theoretical Artificial Intelligence 25 (2) (2013) 189–206. doi:10.1080/ 0952813X.2012.680252. [74] J. Seldrup, C. Lentner, K. Diem, Geigy Scientific Tables: Introduction to Statistics, Statistical Tables, Mathematical Formulae, Ciba-Geigy, 1982. [75] P. I. Good, J. W. Hardin, Common Errors in Statistics (and How to Avoid Them), John Wiley & Sons, 2012. [76] J. K. Kruschke, Bayesian data analysis, Wiley Interdisciplinary Reviews: Cognitive Science 1 (5) (2010) 658–676. [77] G. Willems, G. Pison, P. J. Rousseeuw, S. Van Aelst, A robust Hotelling test, Metrika 55 (1) (2002) 125–138, 00058. doi:10.1007/s001840200192. [78] J. A. Villasenor Alva, E. G. Estrada, A generalization of Shapiro–Wilk’s test for multivariate normality, Communications in Statistics—Theory and Methods 38 (11) (2009) 1870–1883, 00126. [79] C. P. de Campos, A. Benavoli, Joint Analysis of Multiple Algorithms and Performance Measures, New Generation Computing 35 (1) (2016) 69–86. doi:10.1007/s00354-016-0005-8. [80] A. Kumar, R. K. Misra, D. Singh, Improving the local search capability of Effective Butterfly Optimizer using Covariance Matrix Adapted Retreat Phase, in: 2017 IEEE Congress on Evolutionary Computation (CEC), 2017, pp. 1835–1842. doi:10.1109/CEC.2017.7969524. [81] J. Brest, M. S. Maučec, B. Bošković, Single objective real-parameter optimization: Algorithm jSO, in: Evolutionary Computation (CEC), 2017 IEEE Congress On, IEEE, 2017, pp. 1311–1318. [82] N. H. Awad, M. Z. Ali, P. N. Suganthan, Ensemble sinusoidal differential covariance matrix adaptation with Euclidean neighborhood for solving CEC2017 benchmark problems, in: Evolutionary Computation (CEC), 2017 IEEE Congress On, IEEE, 2017, pp. 372–379. [83] A. W. Mohamed, A. A. Hadi, A. M. Fattouh, K. M. Jambi, LSHADE with semi-parameter adaptation hybrid with CMA-ES for solving CEC 2017 benchmark problems, in: Evolutionary Computation (CEC), 2017 IEEE Congress On, IEEE, 2017, pp. 145–152. [84] D. Jagodziński, J. Arabas, A differential evolution strategy, in: Evolutionary Computation (CEC), 2017 IEEE Congress On, IEEE, 2017, pp. 1872–1876. 50 [85] K. M. Sallam, S. M. Elsayed, R. A. Sarker, D. L. Essam, Multi-method based orthogonal experimental design algorithm for solving CEC2017 competition problems, in: Evolutionary Computation (CEC), 2017 IEEE Congress On, IEEE, 2017, pp. 1350–1357. [86] P. Bujok, J. Tvrdík, Enhanced individual-dependent differential evolution with population size adaptation, in: Evolutionary Computation (CEC), 2017 IEEE Congress On, IEEE, 2017, pp. 1358–1365. [87] R. Biedrzycki, A version of IPOP-CMA-ES algorithm with midpoint for CEC 2017 single objective bound constrained problems, in: Evolutionary Computation (CEC), 2017 IEEE Congress On, IEEE, 2017, pp. 1489–1494. [88] A. LaTorre, J.-M. Peña, A comparison of three large-scale global optimizers on the CEC 2017 single objective real parameter numerical optimization benchmark, in: Evolutionary Computation (CEC), 2017 IEEE Congress On, IEEE, 2017, pp. 1063–1070. [89] A. Tangherloni, L. Rundo, M. S. Nobile, Proactive Particles in Swarm Optimization: A settings-free algorithm for real-parameter single objective optimization problems, in: Evolutionary Computation (CEC), 2017 IEEE Congress On, IEEE, 2017, pp. 1940–1947. [90] D. Maharana, R. Kommadath, P. Kotecha, Dynamic Yin-Yang Pair Optimization and its performance on single objective real parameter problems of CEC 2017, in: Evolutionary Computation (CEC), 2017 IEEE Congress On, IEEE, 2017, pp. 2390–2396. [91] R. Kommadath, P. Kotecha, Teaching Learning Based Optimization with focused learning and its performance on CEC2017 functions, in: Evolutionary Computation (CEC), 2017 IEEE Congress On, IEEE, 2017, pp. 2397–2403. [92] D. Berrar, W. Dubitzky, On the Jeffreys-Lindley Paradox and the Looming Reproducibility Crisis in Machine Learning, in: Data Science and Advanced Analytics (DSAA), 2017 IEEE International Conference On, IEEE, 2017, pp. 334–340. [93] S. L. Chow, Précis of statistical significance: Rationale, validity, and utility, Behavioral and brain sciences 21 (2) (1998) 169–194. [94] F. Melinscak, L. Montesano, Beyond p-values in the evaluation of brain–computer interfaces: A Bayesian estimation approach, Journal of neuroscience methods 270 (2016) 30–45. [95] V. Amrhein, S. Greenland, B. McShane, Scientists rise up against statistical significance, Nature 567 (7748) (2019) 305, 00002. doi:10.1038/ d41586-019-00857-9. 51 [96] R. L. Wasserstein, N. A. Lazar, The ASA’s Statement on p-Values: Context, Process, and Purpose, The American Statistician 70 (2) (2016) 129– 133. doi:10.1080/00031305.2016.1154108. [97] R. L. Wasserstein, A. L. Schirm, N. A. Lazar, Moving to a World Beyond “p <0.05”, The American Statistician 73 (sup1) (2019) 1–19, 00003. doi: 10.1080/00031305.2019.1583913. [98] I. R. Silva, On the correspondence between frequentist and Bayesian tests, Communications in Statistics - Theory and Methods 47 (14) (2018) 3477– 3487. doi:10.1080/03610926.2017.1359296. [99] I. Couso, A. Álvarez-Caballero, L. Sánchez, Reconciling Bayesian and Frequentist Tests: The Imprecise Counterpart, in: A. Antonucci, G. Corani, I. Couso, S. Destercke (Eds.), Proceedings of the Tenth International Symposium on Imprecise Probability: Theories and Applications, Vol. 62 of Proceedings of Machine Learning Research, PMLR, 2017, pp. 97–108. 52