scieee AI-readable full text Open interactive document viewer

Comparison of tree-based methods used in survival data

Yabaci, Aysegul,Sigirli, Deniz

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Yabaci, Aysegul; Sigirli, Deniz Article Comparison of tree-based methods used in survival data Statistics in Transition new series (SiTns) Provided in Cooperation with: Polish Statistical Association Suggested Citation: Yabaci, Aysegul; Sigirli, Deniz (2022) : Comparison of tree-based methods used in survival data, Statistics in Transition new series (SiTns), ISSN 2450-0291, Sciendo, Warsaw, Vol. 23, Iss. 1, pp. 21-38, https://doi.org/10.2478/stattrans-2022-0002 This Version is available at: https://hdl.handle.net/10419/266293 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-sa/4.0/ STATISTICS IN TRANSITION new series, March 2022 Vol. 23 No. 1, pp. 21–38, DOI 10.2478/stattrans-2022-0002 Received – 03.05.2021; accepted – 19.07.2021 Comparison of tree-based methods used in survival data Aysegul Yabaci1, Deniz Sigirli2 ABSTRACT Survival trees and forests are popular non-parametric alternatives to parametric and semiparametric survival models. Conditional inference trees (Ctree) form a non-parametric class of regression trees embedding tree-structured regression models into a well-defined theory of conditional inference procedures. The Ctree is applicable in a varietyof regression-related issues, involving nominal, ordinal, numeric, censored, as well as multivariate response variables and arbitrary measurement scales of covariates. Conditional inference forests (Cforest) consitute a survival forest method which combines a large number of Ctrees. The Cforest provides a unified and flexible framework for ensemble learning in the presence of censoring. The random survival forests (RSF) methodology extends the random forests method enabling the approximation of rich classes of functions while maintaining generalisation errors low. In the present study, the Ctree, Cforest and RSF methods are discussed in detail and the performances of the survival forest methods, namely the Cforest and RSF have been compared with a simulation study. The results of the simulation demonstrate that the RSF method with a log-rank score distinction criteria outperforms the Cforest and the RSF with log-rank distinction criteria. Key words: tree-based methods, conditional inference trees, conditional inference forests, random survival forests. 1. Introduction Tree-based methods constitute classification and regression models in the form of a tree structure according to data sets. Understanding the decision rules used in the creation of tree structures makes the use of the method common. Decision trees perform decision making with a multi-stage and sequential approach in solving the classification and regression problem (Safavian et al. 1991). The Classification and Regression Trees (C&RT) provide a visual representation of the effect of independent variables on dependent variables and the interaction between 1 Department of Biostatistics and Medical Informatics, Bezmialem Vakif University, Faculty of Medicine, Istanbul, Turkey. E-mail: [email protected]. ORCID: https://orcid.org/0000-0002-5813-3397. 2 Department of Biostatistics, Faculty of Medicine, Uludag University, Bursa, Turkey. E-mail: [email protected]. ORCID: https://orcid.org/ 0000-0002-4006-3263. 22 A. Yabaci, D. Sigirli: Comparison of tree-based methods… them, which is used to estimate the class membership of a discrete or continuous dependent variable without pre-requisite presentation of the independent variable. In general, if the dependent variable is categorical, the name of the method is the classification tree, and if it is continuous, the method is called the regression tree (Breiman et al. 1984). Survival trees and forests are popular non-parametric alternatives to parametric and semi-parametric survival models. A single tree set can be classified according to survival characteristics by taking into consideration independent variables, while a very powerful estimating tool can be obtained through tree sets created by the combination of trees. The aim of this study is to evaluate the performances of random survival forests (RSF) and conditional inference forest (Cforest) methods as tree-based methods used in survival data analysis, for different conditional censored survival function estimators, for different sample sizes and for cases where the proportional hazard assumption is provided and not provided. 2. Methods of comparison 2.1. Conditional inference trees  Ctree Let 𝑇 show the actual time of death and 𝐶 be the time of censoring, 𝑇𝑚𝑖𝑛𝑇 ,𝐶 is the dependent variable and ∆𝐼𝑇𝐶 is the state variable. Let 𝑿𝑋,…,𝑋 be the vector of p dimensional covariate from 𝒳𝒳 …𝒳 sample space. The situation in which covariates are measured on any scale is discussed. Given the covariate of 𝑿, 𝑇 is the conditional distribution of the dependent variable, presented in the form of ℱ𝑻|𝑿 to be a function of the common variables of ℱ𝑻|𝑿 in the Eq. (1) that is bound to suppose. ℱ|𝑿 = ℱ󰇛𝑇 󰇻 𝑓𝑋,…𝑋󰇢 . (1) Let ℒ be given as in Eq. (2), where some 𝑋 (j=1,..,p ; i=1,..,n) covariate values are missing, and independent and identically distributed observation values are n units of the random sample ‘learning sample’. ℒ󰇝󰇛𝑻,∆, 𝑿󰇜 𝑖1, … , 𝑛 󰇞. (2) For each node in the tree, there is a unit weights vector. Let the unit weight vector be shown as 𝒘󰇟𝑤,…𝑤󰇠. If the observation values of the relevant variable are located on this node, the corresponding value in the weight vector is 1, and if not 0 (Hothorn et al. 2006b). STATISTICS IN TRANSITION new series, March 2022 23 The following steps are taken to create conditional inference trees: Step 1: For 𝒘 unit weights, the general null hypothesis that there is independence between any of the p covariates and the dependent variable is tested. If this hypothesis is not rejected, then it is stopped. In other cases, 𝑿∗ is chosen as the 𝑗∗nth covariate, which has the strongest relationship with T. Step 2: The 𝑋∗ variable, which divides the 𝐴∗⊂ 𝑋∗ set into two discrete sets 𝐴∗ and 𝑋∗/𝐴∗, is selected. Step 3: Step-1 and step-2, 𝒘 and 𝒘 unit weights are modified and repeated. In Step 1, the absence hypothesis is as follows: 𝐻󰇊𝐻   Here, p partial hypotheses are defined as follows: 𝐻:ℱ|𝑿𝒋ℱ ; j1, … , p When the 𝐻 hypothesis cannot be rejected at the specified α level of significance, the division stops. The relationship between the T and each 𝑋 , 𝑗1, … , 𝑝 covariate is tested by the 𝐻 hypotheses, which are partial hypotheses. For this hypothesis, the test statistics or p values are used to select the covariate that is the most associated with T. Weights of 𝑤 can be set to 0 or 1. The symmetric group of all permutations of elements corresponding to the unit weight 𝑤1 is shown with 𝑆󰇛ℒ,𝑤󰇜. In this case, the relationship between T and 𝑋 , 󰇛 𝑗1, … , 𝑝󰇜 is measured by the linear test statistic given below (Hothorn et al. 2006b). 𝑻󰇛ℒ,𝑤󰇜𝑣𝑒𝑐󰇛∑𝑤𝑔󰇛𝑋󰇜ℎ󰇛󰇛𝑇,󰇛𝑇,….𝑇󰇜󰇜󰆒󰇜  󰇜 ∈ ℝ (3) Where 𝑔∶ 𝒳→ℝ is the non-random transformation of the covariate 𝑋. For continuous covariate, 𝑔󰇛𝑥󰇜𝑥 unit transformation can also be applied. Also, it is possible to rank or nonlinear transformations. The effect function ℎ∶ 𝒯𝒯→ℝ is based on response variables in symmetric permutation and is obtained as in Eq. (4). In survival data, ℎ can be selected as log-rank score. ℎ󰇛𝑇,󰇛𝑇,….𝑇󰇜󰇜∑𝑤 𝐼󰇛𝑇𝑇󰇜 𝑖1, . . 𝑛  (4) To divide the covariate selected in Step-1 into two, the permutation test is used in Step-2. The test statistic, which is a special case of 𝑻󰇛ℒ,𝑤󰇜 the test statistic, is calculated as in Eq.(5). 𝑻∗ 󰇛ℒ,𝑤󰇜𝑣𝑒𝑐∑𝑤 𝐼𝑋∗∈𝐴ℎ󰇛  𝑇,󰇛𝑇,….𝑇󰇜󰆒∈ℝ (5) 24 A. Yabaci, D. Sigirli: Comparison of tree-based methods… This linear statistic gives two sampling test statistics that measure the discordance between samples {𝑇𝑤 0 𝑣𝑒 𝑋∈𝐴;𝑖1,..,𝑛 and {𝑇𝑤 0 𝑣𝑒 𝑋∉𝐴;𝑖 1, . . 𝑛. Conditional expected value µ∗  and covariance 𝛴∗  are calculated as in Eq. (6) and Eq. (7) respectively. 𝜇𝔼󰇡𝑻󰇛ℒ,𝑤󰇜󰇻𝑆󰇛ℒ,𝑤󰇜󰇢𝑣𝑒𝑐󰇛󰇛∑𝑤𝑔󰇛𝑋󰇜󰇜 𝔼 󰇛ℎ | 𝑆󰇛ℒ,𝑤󰇜󰇜󰆒󰇜  (6) ∑𝕍󰇛𝑻󰇛ℒ,𝑤󰇜 | 𝑆󰇛ℒ,𝑤󰇜 ). (7) Using this expected value and covariance, 𝑻∗ 󰇛ℒ,𝑤󰇜's standardized test statistic is obtained from c(𝑡∗ ,µ ∗ ,𝛴∗ 󰇜. The distinction that corresponds to the maximum of this test statistic is indicated by "𝐴∗ ". The test statistic that is maximized over all possible subsets of A is as in Eq. (8) 𝐴∗𝑎𝑟𝑔𝑚𝑎𝑥 c(𝑡∗ ,µ ∗ ,𝛴∗ 󰇜. (8) Then, as stated in Step 2 of the algorithm, "𝑤" and "𝑤" unit weights are determined by the functions 𝒘,𝑤 Ι󰇛 𝑿∗∈ 𝐴∗󰇜 and 𝒘,𝑤 Ι󰇛 𝑿∗∉ 𝐴∗󰇜 and the weights are modified and repeat Step-1 and Step-2. 2.2. Conditional inference forest method  Cforest Assume that the conditional distribution function of T the response variable is dependent on random variable X with the function 𝑓:𝒳→ℝ. In this case, ℱ| ℱ|󰇛󰇜. The conditional censoring survival function is given in the form of 𝐺󰇛𝑇∣𝐗󰇜 ℙ󰇛𝐶𝑡∣𝐗x󰇜. Let 𝜓 be the function space of all candidate estimators 𝜓:𝒳→ℝ. Estimation of the regression function f, as defined by full data loss function L is found by minimizing the expected value of the risk function. However, the full data function cannot be calculated because all data cannot be reached in the presence of censored observation. Therefore, instead of the full data loss function, the observed data loss function 𝐿󰇛𝑇,𝜓󰇛𝐗󰇜∣𝜂󰇜 is used. In this case, the expected value of the observed data loss function is obtained as given in Eq. (9). Here, the expected value of the full loss data function is intended to be minimized according to the candidate estimators 𝜓 𝜖 Ψ (Hothorn et al. 2006a). 𝔼,𝑿𝐿𝑇,𝜓󰇛𝐗󰇜 𝐿󰇛𝑇,𝜓󰇛𝐗󰇜 ∣ ∣ 𝜂󰇜𝑑ℱ,,𝐗𝔼,,𝐗 𝐿󰇛𝑇,𝜓󰇛𝐗󰇜 ∣ ∣ 𝜂󰇜. (9) In Eq. (9), η is the nuisance parameter and can be defined as a conditional censored survival function. The observed loss data function can be defined as Eq. (10) by using 𝐺󰇛𝑇∣𝑿󰇜. 𝐿󰇛𝑇,𝜓󰇛𝑿󰇜 ∣ ∣ 𝐺󰇜𝐿𝑇,𝜓󰇛𝑿󰇜∆ 𝑇∣ ∣ 𝑿 . (10) STATISTICS IN TRANSITION new series, March 2022 25 The full data loss function is weighted by the inverse of the probability of censored after T time. In this case, the expected value of the observed data loss function is obtained as in Eq. (11). 𝔼,∆,𝑿 𝐿󰇛𝑇,𝜓󰇛𝑿󰇜 ∣ ∣ 𝐺󰇜 𝑛𝐿𝑇,𝜓󰇛𝑿𝒊󰇜 ∣ ∣ 𝐺   𝑛𝐿𝑇,𝜓󰇛𝑿𝒊󰇜 ∣ ∣ 𝐺   ∆ 𝐺󰇛𝑇 ∣ ∣ 𝑿𝒊󰇜 . (11) The regression function predictor 𝑓󰆹 is obtained by minimizing this equation according to the candidate predictors 𝜓∈Ψ. Here the 𝐺 conditional censored survival function is unknown and its estimator is used instead. As a 𝐺 estimator, the nonparametric estimator, Cox estimator or the cumulative Aalen estimator can be used. In the case of 𝑤∆ 𝐺󰇛𝑇∣𝑿𝒊󰇜, 𝐰󰇛𝑤,𝑤,..,𝑤󰇜 is called IPC ( the inverse probability of censored weights). The conditional inference forest (cforest) algorithm has been proposed by Hothorn et al to find the values of 𝜓 that minimize the expected value of the observed data loss function. 𝐰 weight vector is calculated by using the observed learning sample ℒ 󰇝󰇛𝑇,∆,𝑿𝒊󰇜;𝑖1, … , 𝑛󰇞 and 𝑤∆ 𝐺󰇛𝑇∣𝑿𝒊󰇜. If the learning sample contains a censored observation value, it is 𝑤0 because it is ∆0. The steps of the algorithm are as follows (Hothorn et al. 2006a): Step 1: Set 𝑚1 and 𝑀1. Step 2: From the multinomial distribution with parameter 𝑛 and 󰇛∑𝑤  󰇜𝒘, a random vector of the unit numbers 𝐯󰇛𝑣,…,𝑣󰇜 is drawn. Step 3: With a regression tree, the sample space 𝒳 is divided into 𝐾󰇛𝑚󰇜 cells and created 𝜋𝑅,…,𝑅󰇛󰇜 pieces are created. This regression tree is created using the learning sample ℒ with case counts 𝐯. In the permutations of the ℒ learning sample, i th observation takes place once. Step 4: Increase 𝑚 by one, repeat Step 2 and step 3 until 𝑚𝑀. In Step 3, using the learning sample obtained in Step 2, a survival tree is obtained with a conditional inference trees algorithm. Let 𝒯 denote 𝑚 th survival tree and 𝒯󰇛𝒙󰇜 denote terminal node with 𝒙 covariate value in the 𝑚 th tree. Each 𝒙 value will take place on a single terminal node. 𝑁󰇛𝑠󰇜𝐼󰇛𝑇𝑠,∆1󰇜 and 𝑍󰇛𝑠󰇜𝐼󰇛𝑇𝑠󰇜 𝑁 ∗󰇛𝑠,𝑥󰇜∑𝑣𝐼󰇛𝑋∈𝒯󰇛𝑥󰇜󰇜𝑁󰇛𝑠󰇜  (12) 𝑍 ∗󰇛𝑠,𝑥󰇜∑𝑣𝐼𝑋∈𝒯󰇛𝑥󰇜𝑍󰇛𝑠󰇜  (13) 26 A. Yabaci, D. Sigirli: Comparison of tree-based methods… Where 𝑁 ∗󰇛𝑠,𝑥󰇜 and 𝑍 ∗󰇛𝑠,𝑥󰇜 are respectively the number of uncensored events in the terminal node up to the time of 𝑠 corresponding to the 𝒙 covariate value, and the number of units at risk at 𝑠 time. In this case, when 𝒙 is given a covariate, the ensemble survival function for t time is equal to that of Eq. (14). 𝑆󰆹󰇛𝑡∣𝑥󰇜∏󰇡1∑ ∗󰇛,󰇜   ∑ ∗󰇛,󰇜   󰇢.  (14) 2.2. Random survival forest method  RSF The algorithm steps of the RSF method are as follows: Step 1: Extract M bootstrap sample from the original data. Each bootstrap sample should exclude average 37% of the original data. The data that is excluded is called outof-bag data (OOB). Step 2: Create a survival tree for each bootstrap samples. On each node of the tree, randomly 𝑝 candidate variable is selected. The node is separated by using candidate variables that maximise the survival difference between child nodes. Step 3: continue the split until at least one observed case remains on each terminal node. Step 4: Cumulative hazard function (CHF) is calculated for each tree. Average to obtain the ensemble CHF. Step 5: Using OOB data, estimation error is calculated for the ensemble cumulative hazard function (Ishwaran et al. 2008a). Logrank test is being used to compare two groups survival, by putting equal weights to each individual (Mantel N. 1966; Karadeniz et al. 2018). Two methods can be used as separation criteria in the algorithm. The first is the log-rank distinction and the second is the log-rank score distinction (Segal 1988; Ciampi et al.1986; Hothorn and Lausanne 2003). i. Log-rank distinction criteria Let 𝑇 ; 𝑖1,..,𝑛 denote the survival time of 𝑖 th unit and 𝑋 covariate for the distinction on a node, 𝑋𝑐 and 𝑋𝑐 according to the cut point of c. Let 𝑠𝑠 ⋯𝑠 denote discrete time of death on a node for 𝑧1, … , 𝑁. For the m th tree, 𝑁 ∗󰇛𝑠,𝑥󰇜 show the number of people dying in 𝑠 time on child nodes d=1,2. 𝑁 ∗󰇛𝑠,𝑥󰇜 =𝑁 ∗󰇛𝑠,𝑥󰇜𝑁  ∗󰇛𝑠,𝑥󰇜 is in format. For the m th tree, , 𝑍 ∗󰇛𝑠,𝑥󰇜 indicates the number of units at risk at 𝑠 time on child nodes d=1,2. In this case, 𝑍 ∗󰇛𝑠,𝑥󰇜 𝑍 ∗󰇛𝑠,𝑥󰇜𝑍 ∗󰇛𝑠,𝑥󰇜 and 𝑍 ∗󰇛𝑠,𝑥󰇜#󰇝𝑇𝑠 , 𝑥𝑐󰇞, 𝑍 ∗󰇛𝑠,𝑥󰇜#󰇝𝑇𝑠 , 𝑥𝑐󰇞. Where 𝑥, is the value that the 𝑋 covariate takes for unit i th. 𝑛 is the total number of units observed in the d th child node. Thus, 𝑛 #󰇝𝑖:𝑥𝑐󰇞 and 𝑛#󰇝𝑖:𝑥𝑐󰇞 are equal to 𝑛𝑛𝑛. STATISTICS IN TRANSITION new series, March 2022 27 The log-rank test statistic for the c cut-off value of the 𝑋 covariate is as in Eq. (15). 𝐿𝑜𝑔𝑅𝑎𝑛𝑘𝑋,𝑐 ∑ ∗, ∗,  ∗󰇡,󰇢  ∗󰇡,󰇢   ∑ ∗󰇡,󰇢  ∗󰇡,󰇢󰇭󰇛 ∗󰇡,󰇢  ∗󰇡,󰇢󰇜󰇛 ∗󰇡,󰇢  ∗󰇡,󰇢  ∗󰇡,󰇢 󰇜  ∗,󰇮   . (15) 𝐿𝑜𝑔𝑅𝑎𝑛𝑘𝑋,𝑐 provides a measure for node distinction. The distinction occurs between the two terminal nodes that has the highest 𝐿𝑜𝑔𝑅𝑎𝑛𝑘𝑋,𝑐 value. The best distinction value of 𝐿𝑜𝑔𝑅𝑎𝑛𝑘𝑋∗,𝑐∗𝐿𝑜𝑔𝑅𝑎𝑛𝑘𝑋,𝑐 is determined by the value of the 𝑋 covariate and c cut-off value (Segal 1988; Hothorn and Lausen 2003). ii. Log-rank score distinction criteria Another distinction rule is the log-rank score distinction rule proposed by Hothorn and Lusen (2003). Assume that the values of the 𝑋 covariate are sorted as 𝑥𝑥 ⋯𝑥. For each 𝑇 survival time, ranks are obtained as in Eq. (16). 𝛼∆∑∆    . (16) Where,𝛤#󰇝𝑠: 𝑇𝑇󰇞. In this case, the log-rank score statistic is obtained as in Eq. (17). 𝐿𝑜𝑔𝑅𝑎𝑛𝑘𝑠𝑘𝑜𝑟𝑋,𝑐∑  󰇛 󰇜 . (17) In Eq. (17), 𝛼 and 𝑠 is defined as the sample mean and sample variance of ranks, respectively. 𝐿𝑜𝑔𝑅𝑎𝑛𝑘𝑠𝑐𝑜𝑟𝑒𝑋,𝑐 provides log-rank score for node distriction. 2.4. Estimators used in estimating G conditional censored survival function i. Nonparametric estimator Let 𝐺󰇛𝑇∣𝑋󰇜𝑃󰇛𝐶𝑡∣𝑋𝑥󰇜 and 𝐾󰇛𝑡󰇜 denote respectively conditional survival function of the censoring time and any kernel function. The nonparametric estimator used by Graf et al. is given in Eq. (18). (Gerds and Schumacher 2007). 𝐺󰇥𝐺: 𝑠𝑢𝑝𝑇∣ ∣ 𝑋󰇛󰆓󰇜∣ |󰆓|𝐾󰇛𝑡󰇜0󰇦. (18) ii. Cox estimator Let α and 𝐻󰇛𝑡󰇜 denote respectively regression coefficient and initial cumulative hazard function. Cox regression estimator is given in Eq. (19) (Gerds and Schumacher 2007). 𝐺𝐺,󰇛󰇜:𝐺󰇛𝑇∣𝑋󰇜𝑒𝑥𝑝󰇝exp 󰇛α󰆒𝑋󰇜𝐻󰇛𝑡󰇜󰇞;α∈ℝ . (19) 28 A. Yabaci, D. Sigirli: Comparison of tree-based methods… iii. Aalen estimator Let 𝛼󰇛𝑡󰇜 denote time-dependent regression coefficient. Cumulative Aalen regression estimator are given as in Eq. (20) (Gerds and Schumacher 2007). 𝐺𝐺:𝐺󰇛𝑇∣𝑋󰇜𝑒𝑥𝑝󰇥𝑋󰆒𝛼󰇛𝑠󰇜.𝑑𝑠   󰇦. (20) 2.5. Criteria used to evaluate model performance 2.5.1. Brier Score  BS The prediction error defined as the time dependent expected Brier score is one of the measures for assessing the predictive performances of rival survival modeling strategies. If the score is close to zero, the class estimates are accepted to be reliable. Let ∆𝐼𝑇 𝑡 be state of i th unit for t time. When X is given, the probability of survival predicted at t time for the i th unit is shown as 𝑆󰆹󰇛𝑡∣ ∣ 𝑋󰇜 . In this case, the Brier score is the same as the Eq. (21) 𝐵𝑆𝑡,𝑆󰆹𝐸𝐼󰇛𝑇 𝑡󰇜𝑆󰆹󰇛𝑡 ∣ ∣ 𝑋󰇜  . (21) The expected value is calculated based on the data of the i th unit which is not included in the learning set. The first critical value for the Brier score is 33%. This corresponds to the risk predicted by the random number drawn from the U[0,1] distribution. The second critical value is 25% and corresponds to 50% risk estimation for each unit. Another criterion is the Brier score value obtained from the model from which all independent variables are extracted (Ishwaran et al. 2008a). Residual squares are weighted using the inverse probabilities of the censored weights given in Eq. (22). 𝑊 󰇛𝑡󰇜󰇛󰇜∆ 󰇛∣󰇜󰇛󰇜 󰇛∣󰇜 . (22) Here 𝐺󰇛𝑡∣𝑥󰇜𝑃󰇛𝐶𝑡∣𝑋𝑥󰇜 is the estimate of the conditional survival function for the i th unit of censoring time. If an independent set of data 𝐷 is available, the expected Brier score is the same as in Eq. (23). 𝐵𝑆 𝑡,𝑆󰆹∑𝑊 󰇛𝑡󰇜󰇝 ∈𝐼󰇛𝑇𝑡󰇜𝑆󰆹󰇛𝑡∣𝑋󰇜󰇞. (23) Where n is the number of units in 𝐷 (i=1,.,n) and calculated from the learning data 𝑆󰆹. 2.5.2. Integrated Brier Score  IBS Prediction errors can be summed up with IBS as follows: 𝐼𝐵𝑆󰇛𝑇𝐻,𝜏󰇜𝑇𝐻󰇛𝑡,𝑆󰆹   󰇜𝑑𝑡 (24) STATISTICS IN TRANSITION new series, March 2022 35 Table 10. The mean and standard error values according to the IBS criteria of RSF and Cforest method for different survival times in cases where n=100 and the proportional hazard assumption is not provided IBS (scenario 2, n=100) AppErr BootCvErr NoInfErr Boot632plusErr 𝓍  𝑠𝓍  𝓍  𝑠𝓍  𝓍  𝑠𝓍  𝓍  𝑠𝓍  RSF(logrank-nonparametric) 0.0296 0.0056 0.1658 0.0350 0.2840 0.0054 0.1379 0.0269 RSF (logrank-Cox) 0.0290 0.0042 0.1642 0.0243 0.2804 0.0044 0.1374 0.0248 RSF (logrank-Aalen) 0.0285 0.0021 0.1627 0.0146 0.2777 0.0032 0.1362 0.0158 RSF (logrankscorenonparametric) 0.0297 0.0070 0.1743 0.0370 0.3040 0.0116 0.1429 0.0279 RSF (logrankscore-Cox) 0.0292 0.0068 0.1732 0.0283 0.3004 0.0087 0.1414 .0270 RSF(logrankscore-Aalen) 0.0295 0.0035 0.1718 0.0156 0.2977 0.0054 0.1402 0.0257 Cforest(nonparametric) 0.1007 0.0200 0.1531 0.0200 0.2566 0.0102 0.1376 0.0215 Cforest (Cox) 0.1005 0.0191 0.1508 0.0188 0.2389 0.0097 0.1374 0.0200 Cforest (Aalen) 0.0964 0.0120 0.1479 0.0117 0.2332 0.0087 0.1353 0.0123 *AppErr: apparent prediction, BootCvErr: Boostrap Cross-Validation prediction, noinferr: ignorance prediction error, Boot632plusErr: 0.632+ prediction Table 11. The mean and standard error values according to the IBS criteria of RSF and Cforest method for different survival times in cases where n=200 and the proportional hazard assumption is not provided IBS (scenario 2, n=200) AppErr BootCvErr NoInfErr Boot632plusErr 𝓍  𝑠𝓍  𝓍  𝑠𝓍  𝓍  𝑠𝓍  𝓍  𝑠𝓍  RSF(logrank-nonparametric) 0.0278 0.0026 0.1630 0.0325 0.2835 0.0040 0.1369 0.0249 RSF (logrank-Cox) 0.0272 0.0012 0.1616 0.0223 0.2784 0.0035 0.1363 0.0228 RSF (logrank-Aalen) 0.0266 0.0009 0.1605 0.0135 0.2757 0.0018 0.1352 0.0138 RSF (logrankscorenonparametric) 0.0280 0.0060 0.1635 0.0357 0.3020 0.0100 0.1408 0.0249 RSF (logrankscore-Cox) 0.0275 0.0058 0.1626 0.0281 0.2904 0.0075 0.1405 0.0260 RSF (logrankscore-Aalen) 0.0267 0.0025 0.1610 0.0154 0.2857 0.0040 0.1402 0.0157 Cforest(nonparametric) 0.1007 0.0187 0.1504 0.0177 0.2446 0.0084 0.1366 0.0205 Cforest (Cox) 0.0977 0.0162 0.1506 0.0156 0.2287 0.0056 0.1358 0.0195 Cforest (Aalen) 0.0951 0.0100 0.1459 0.0105 0.2300 0.0043 0.1328 0.0108 *AppErr: apparent prediction, BootCvErr: Boostrap Cross-Validation prediction, noinferr: ignorance prediction error, Boot632plusErr: 0.632+ prediction 36 A. Yabaci, D. Sigirli: Comparison of tree-based methods… Table 12. The mean and standard error values according to the IBS criteria of RSF and Cforest method for different survival times in cases where n=300 and the proportional hazard assumption is not provided IBS (scenario 2, n=300) AppErr BootCvErr NoInfErr Boot632plusErr 𝓍  𝑠𝓍  𝓍  𝑠𝓍  𝓍  𝑠𝓍  𝓍  𝑠𝓍  RSF(logrank-nonparametric) 0.0265 0.0022 0.1618 0.0315 0.2817 0.0035 0.1358 0.0237 RSF (logrank-Cox) 0.0261 0.0010 0.1606 0.0218 0.2773 0.0027 0.1352 0.0222 RSF (logrank-Aalen) 0.0255 0.0002 0.1608 0.0126 0.2737 0.0008 0.1347 0.0126 RSF (logrankscorenonparametric) 0.0275 0.0053 0.1625 0.0347 0.3010 0.0098 0.1407 0.0239 RSF (logrankscore-Cox) 0.0273 0.0054 0.1616 0.0261 0.2902 0.0065 0.1400 0.0250 RSF (logrankscore-Aalen) 0.0266 0.0017 0.1607 0.0144 0.2847 0.0030 0.1399 0.0137 Cforest(nonparametric) 0.1005 0.0167 0.1503 0.0157 0.2444 0.0084 0.1356 0.0205 Cforest (Cox) 0.0967 0.0152 0.1501 0.0136 0.2305 0.0056 0.1348 0.0197 Cforest (Aalen) 0.0941 0.0104 0.1449 0.0102 0.2284 0.0040 0.1318 0.0103 *AppErr: apparent prediction, BootCvErr: Boostrap Cross-Validation prediction, noinferr: ignorance prediction error, Boot632plusErr: 0.632+ prediction 5. Results and discussion In this study, Cforest method (Hothorn and et al. 2006a), which aims to minimize the proposed empirical risk function for right-censored data and a community with a low correlation structure by creating different trees, and RSF method (Ishwaran and et al. 2008a), which is an extension of Brieman's random forest method for rightcensored data, are compared according to C-Index and IBS criteria. According to the C-Index criterion; in all cases, RSF method has higher mean C - Index values and lower standard error values than Cforest method. When we examined the sample size, it was observed that the mean C-Index values for both scenarios and both methods were increased and standard error values were decreased with the increase in sample size. It is observed gave the best results for the RSF method and the non parametric estimator has lower mean C-Index values than the Aalen estimator and Cox estimator. In the Cforest method, it was observed that the nonparametric estimator had lower C-Index mean values and similar results were obtained in Cox and Aalen estimator. When the RSF method was examined in terms of two different separation criteria, it was determined that the logrank distinction had higher mean C-Index values and lower standard error values. Compared to the situation in which the proportional hazard assumption is provided and not provided, it has been observed that both methods perform better in the absence of the proportional hazard assumption. STATISTICS IN TRANSITION new series, March 2022 37 However, when the proportional hazard assumption provided, there has been a further decrease in the mean C-Index values for the RSF method compared to the Cforest method. According to the IBS criterion; for all cases, in both scenarios, and for all 𝐺 estimation methods (Cox, Aalen and Nonparametric), the RSF method has lower mean and standard error values than the Cforest method. With the increase in sample size, model performance was observed to increase in all cases according to IBS criteria. For all methods and for both scenarios, Aalen estimator has a lower error value than nonparametric estimator and Cox estimator. When examined according to RSF separation criteria, it was determined that logrank distinction criteria had lower IBS mean values and standard error values. In this study, it was observed that all methods performed better in the case that the proportional hazard assumption is not provided, compared to the case that the proportional hazard assumption is provided. Mogensen and et al. (2012) examined the performance of the RSF, Cforest and Cox regression models using the "cost" data set included in the PEC package. As a result, while some cross-validation methods found the performance of the methods to be similar, some cross-validation methods found the performance of the RSF method to be higher. Gerds and Schumacher (2007) used marginal Kaplan-Meier, Cox, Aalen and nonparametric estimators for calculating IBS values. However, if the censored mechanism of the Kaplan-Meier estimator is dependent on the common variables, it gives error. For this reason, they recommended the use of three other predictors for the case where the censored survival function is dependent on the common variables. In their simulation study, they stated that the Aalen estimator was better than the Cox estimator. The results of our simulation study showed that the Aalen estimator has better performance in both methods. Ciampi (1986) proposed the use of logrank test statistic to compare two child nodes in decision trees. Ishwaran and et al.(2008a) stated that the model obtained by using logrank criteria is higher than C-Index value when they apply the RSF method on 11 sets of data according to different separation rules. According to the results of our simulation study, it was determined that the logrank distinction criteria showed higher performance than the logrank score distinction criteria in the case where the proportional hazard assumption is provided and not provided in the RSF method. As a result, it has been shown that the RSF method performs better than the Cforest. For both methods, it can be said that the Aalen estimator performs better than the other estimators. The performance of both methods was better if the proportional hazard assumption was not provided. In addition, the RSF method shows that the logrank distinction criteria, which is one of two different separation criteria, performs better than the logrank score distinction criteria. 38 A. Yabaci, D. Sigirli: Comparison of tree-based methods… References Breiman, L., Friedman, J., Olshen, R. and Stone, C., (1984). Classification and regression trees. Wadsworth Int. Group, 37(15), pp. 237–251. Ciampi, A., Thiffault, J., Nakache, J. P. and Asselain, B., (1986). Stratification by stepwise regression, correspondence analysis and recursive partition: a comparison of three methods of analysis for survival data with covariates. Computational statistics & data analysis, 4(3), pp. 185–204. Gerds, T. A., Schumacher, M., (2007). Consistent estimation of the expected Brier score in general survival models with right‐censored event times. Biometrical Journal, 48(6), pp. 1029–1040. Hothorn, T., Bühlmann, P., Dudoit, S., Molinaro, A. and Van Der Laan, M. J., (2006a). Survival ensembles. Biostatistics, 7(3), pp. 355–373. Hothorn, T., Hornik, K. and Zeileis, A., (2006b). Unbiased recursive partitioning: A conditional inference framework. Journal of Computational and Graphical statistics, 15(3), pp. 651–674. Hothorn, T., Hornik, K. and Zeileis, A., (2005). party: A laboratory for recursive part (y) itioning. R package version 0.3-2. Ishwaran, H., Kogalur, U. B., Blackstone, E. H. and Lauer, M. S., (2008). Random survival forests. The annals of applied statistics, 2(3), pp. 841–860. Ishwaran, H., Kogalur, U. B., (2019). randomForestSRC: Fast unified random forests for survival, regression, and classification (RF-SRC). R Package Version, 2(1). Mogensen, U. B., Ishwaran, H. and Gerds, T. A., (2012). Evaluating random forests for survival analysis using prediction error curves. Journal of statistical software, 50(11), 1. Safavian, S. R., Landgrebe, D., (1991). A survey of decision tree classifier methodology. IEEE transactions on systems, man, and cybernetics, 21(3), pp. 660– 674. Segal, M. R., (1988). Regression trees for censored data. Biometrics, pp. 35–47. Team, R. C., (2013). R: A language and environment for statistical computing.