scieee AI-readable full text Open interactive document viewer

Influence diagnostics in exponentiated-Weibull regression models with censored data

Ortega, Edwin M. M.; Cancho, Vicente G.; Bolfarine, Heleno

Abstract

Diagnostic methods have been an important tool in regression analysis to detect anomalies, such as departures from the error assumptions and the presence of outliers and influential observations with the fitted models. The literature provides plenty of approaches for detecting outlying or influential observations in data sets. In this paper, we follow the local influence approach (Cook 1986) in detecting influential observations with exponentiated-Weibull regression models. The relevance of the approach is illustrated with a real data set, where it is shown that by removing the most influential observations, there is a change in the decision about which model fits the data better.

Full text

Statistics & Operations Research Transactions SORT 30 (2) July-December 2006, 171-192 Statistics & Operations Research Transactions Influence diagnostics in exponentiated-Weibull regression models with censored data c Institut d’Estad´ ıstica de Catalunya [email protected] ISSN: 1696-2281 www.idescat.net/sort Edwin M. M. Ortega1, Vicente G. Cancho2& Heleno Bolfarine3 1ESALQ-USP, 2ICMC-USP, 3IME-USP Abstract Diagnostic methods have been an important tool in regression analysis to detect anomalies, such as departures from the error assumptions and the presence of outliers and influential observations with the fitted models. The literature provides plenty of approaches for detecting outlying or influential observations in data sets. In this paper, we follow the local influence approach (Cook 1986) in detecting influential observations with exponentiated-Weibull regression models. The relevance of the approach is illustrated with a real data set, where it is shown that by removing the most influential observations, there is a change in the decision about which model fits the data better. MSC: 62H10, 62J20, 62N01 Keywords: Exponentiated-Weibull distribution; censored data; local influence; influence diagnostic; survival data. 1 Introduction In this paper we consider data sets representing the elapsed time until the occurrence of an event of interest such as the recurrence of a disease, death of a patient, failure of equipment, performance of a task, and so on. Address for correspondence: Departamento de Ciˆ encias Exatas, USP Av. P´ adua Dias 11 - Caixa Postal 9, 13418900 Piracicaba - S˜ ao Paulo - Brasil. [email protected] Edwin M. M. Ortega. ESALQ, University of S˜ ao Paulo, Piracicaba, Brazil. [email protected] Vicente G. Cancho. ICMC, University of S˜ ao Paulo, S˜ ao Carlos, S˜ ao Paulo, Brazil. [email protected] Heleno Bolfarine. IME, University of S˜ ao Paulo, S˜ ao Paulo, Brazil. [email protected] Received: March 2006 Accepted: October 2006 172 Influence diagnostics in exponentiated-Weibull regression models with censored data This elapsed time is generally termed survival time or life time. To simplify notation, the units under study are called individuals, and the data set is called survival data. It is common, in such kind of data, to have censored observations, that is, for some individuals the exact time of death is not known. It is only known that it lies beyond a certain value (censoring time). In this paper, we consider that the censoring times are random and noninformative. Moreover, in most situations, survival times can be affected by covariates (explanatory variables), such as the age at disease onset, blood pressure, cholesterol level, treatment type and many other important factors. In this paper, we consider the exponentiated Weibull model, which includes as special cases the Weibull and exponential models. As considered by Mudhokar et al. (1995), it can be used in adjusting survival data with bathtub-type risk functions. Cancho et al. (1999) conducted a Bayesian study for exponentiated-Weibull regression models and Bolfarine and Cancho (2001) considered an exponentiated-Weibull survival model with a survival fraction. An important step in regression analysis is to conduct a robustness study to detect influential or extreme observations that can cause important distortions on the results of the analysis. Numerous approaches have been proposed in the literature with a view to detect influential or outlying observations that can seriously affect parameter estimates. Studies of case deletion have were started with Cook (1977). Important reviews on the main approaches to detect influential observations are considered in Cook and Weisberg (1982) and Chaterjee and Hadi (1988). A general framework to detect influence of observations was proposed by Cook (1986) and has often been applied with regression models. The method basically indicates how sensitive the analysis is when small perturbations are made to the data or the model. For instance, under the normal error, Lawrance (1988) investigated local influence applications in linear models with a response transformation parameter, Beckman et al. (1987) presented influence influence studies in mixed effects analysis of variance, Tsai and Wu (1992) considered first-order autoregressive models with nonconstant variances, and Paula (1993) used local influence methods with linear regression models when there are inequality constraints on the parameters. Moving away from normal models, Petit and Bin Daud (1989) investigated local influence with proportional hazard regression models, Escobar and Meeker (1992) adapted local influence methods to regression analysis with censoring, and O’Hara et al. (1992) and Kim (1995) applied local influence methods with multivariate regression. More recently, Galea et al. (1997) and Liu (2000) used local influence with elliptical linear regression models; Kwan and Fung (1998) applied the methodology to factor analysis and Gu and Fung (1998) discussed local influence in canonical correlation analysis. An interesting discussion and comparison with other influence measures is considered in Fung and Kwan (1997). An important extension of the method to assess the local influence of observations on the predictions from the fitted model was proposed by Thomas and Cook (1990). Edwin M. M. Ortega, Vicente G. Cancho & Heleno Bolfarine 173 In Sections 2 and 3, we review the exponentiated-Weibull regression model considered in Bolfarine et al. (2001). In Sections 4 and 5, we discuss the local influence method and local influence on predictions. Likelihood displacement is used to evaluate the influence of observations on the maximum likelihood estimators. Section 6 presents the results of an analysis with a real data set, including a residual analysis. 2 The exponentiated-Weibull distribution The Weibull family of distributions has been widely used in the analysis of survival data specially in medical and engineering application. This family is suitable in situations where the risk function is constant or monotone. It is not, however, suitable in situations where the risk function is unimodal or presents a bathtub shape. Many parametric families have been considered for modeling survival data with a more general shape for the risk function. For example, Prentice (1974) considered the generalized F distribution; Stacy (1962) proposed the generalized gamma distribution while Mudhokar et al. (1995) presented an extension of the Weibull distribution, which is called the exponentiated Weibull family of distributions, and can adequately fit data sets presenting unimodal, monotone and bathtub shaped risk functions. The exponentiated-Weibull distribution considered in Mudhokar et al. (1995) with parameters α,θand σconsiders that life time Thas a density function given by f(t;α, θ, σ)=αθ σ1−exp −t σαexp −t σαt σα−1 ,∀t>0 (1) where α>0, θ>0 are shape parameters and σ>0 is a scale parameter. As a special cases, there is the Weibull distribution when θ=1 and the exponential distribution when α=1, θ=1. The survival function corresponding to random variable Twith exponentiated-Weibull density is given by S(t;α, θ, σ)=P(T≥t)=1−1−exp −ti σαθ .(2) The great flexibility of this model in fitting survival data can be depicted from its risk function, which can be monotonically decreasing if α≤1andαθ ≤1, monotonically increasing if α≥1andαθ ≥1 and present a bathtub shape if α>1andαθ < 1. Let t1,t2,...,tnbe a random sample of random variate Twith exponentiated-Weibull distribution. The likelihood function corresponding to the observed sample is given by 174 Influence diagnostics in exponentiated-Weibull regression models with censored data L(t;α, θ, σ)=αrθrσ−rαexp ⎡⎢⎢⎢⎢⎢⎣− i∈Fti σσ⎤⎥⎥⎥⎥⎥⎦ i∈F tα−1 i1−exp −ti σαθ−1 (3)  i∈C1−1−exp −ti σαθ , where ris the observed number of failures, Fdenotes the set of uncensored observations and Cdenotes the set of censored observations. No explicit expressions are available for the maximum likelihood estimators of α,σand θ, which are obtained by maximizing the log-likelihood numerically. One approach that can be used is the Newton-Raphson algorithm. 3 Exponentiated-Weibull Regression models In many practical applications, lifetimes are affected by covariates such as cholesterol level, blood pressure and many others. The covariate vector is denoted by x= (x1,x2,...,xp)Twhich is related to responses Y=log(T) through a regression model. It is also considered that the scale parameter σof the exponentiated-Weibull model depends on the matrix of explanatory variables X. Considering the transformation σ=exp(μ)andα=1/δ, it follows that the density function of Ycan be written as f(y)=θ δ1−exp −exp y−μ δθ−1 exp y−μ δ−exp y−μ δ (4) y>0, where α>0, θ>0, and −∞ <μ<∞. Using (4), we can write the above model as a log-linear model Y=μ+δZ(5) where variable Zfollows the density f(z)=θ1−exp[−exp(z)]θ−1exp[z −exp(z)],∀−∞<z<∞(6) with survival function given by S(y)=1−1−exp −expy−μ δθ .(7) We consider now the regression model based on the log-exponentiated-Weibull given in (5), relating response Yand covariate vector x, so that the conditional distribution Y|x can be represents as Edwin M. M. Ortega, Vicente G. Cancho & Heleno Bolfarine 175 Yi=xT iβ+δZi,i=1,...,n,(8) where β=(β1,...,βp)T,δ>0andθ>0 are unknown parameters, xT i= (xi1,xi2,...,xip) is the explanatory vector and Zfollows the distribution in (7). In this case, the survival function of Y|xis given by S(y)=1−1−exp −expy−xTβ δθ .(9) Moreover, corresponding to sample (y1,x1),(y2,x2),...,(yn,xn)ofnobservations from distribution (4), where yirepresents the logarithm of the survival time and xithe covariate vector associated with the i-th individual, the log-likelihood function can be written as l(γ)=rlog(θ)−rlog(δ)+(θ−1)  i∈F log1−exp[−exp(zi)]+(10) + i∈Fzi−exp(zi)+ i∈C log1−[1 −exp(−exp(zi))]θ, where ris the number of uncensored observations (failures) and zi=yi−xT iβ δ.Maximum likelihood estimates for the parameter vector γ=(θ, δ, βT)Tcan be obtained by maximizing the likelihood function while Bayesian estimation is discussed by Cancho et al. (1999). In this paper, software Ox (MAXBFGS subroutine) (see Doornik, 1996) was used to compute maximum likelihood estimates (MLE). Covariance estimates for the maximum likelihood estimators γcan also be obtained using the Hessian matrix. Confidence intervals and hypothesis testing can be conducted by using the large sample distribution of MLE which is a normal distribution with the covariance matrix as the inverse of the Fisher information as long as regularity conditions are satisfied. More specifically, the asymptotic covariance matrix is given by I−1(γ) with I(γ)=−E[¨ L(γ)] such that ¨ L(γ)=∂2l(γ) ∂γ∂γT. Since it is not possible to compute the Fisher information matrix I(γ) due to the censored observations (censoring is random and noninformative), it is possible to use in its place the matrix of second derivatives of the log likelihood, −¨ L(γ), evaluated at the MLE γ=γ, which is consistent. Then ¨ L(γ)=⎛⎜⎜⎜⎜⎜⎜⎜⎜⎝ Lθθ Lθδ Lθβ .Lδδ Lδβ ..Lββ ⎞⎟⎟⎟⎟⎟⎟⎟⎟⎠ with the submatrices in appendix A. 176 Influence diagnostics in exponentiated-Weibull regression models with censored data 4Influence diagnostics Let l(γ) denote the log-likelihood function from the postulated model, where γ= (θ, δ, βT)T,andletωbe a n×1 vector of perturbations restricted to some open subset Ω⊂Rn. The perturbations are made on the log-likelihood function. We will assume, in particular, the case-weights perturbation scheme such that the log-likelihood function takes the form l(γ|ω)= i∈F ωilog f(yi;γ)+ i∈C ωilogS(yi;γ), where 0 ≤ωi≤1andω0=(1,1,...,1)Tis the vector of no perturbation. Note that l(γ|ω0)=l(γ). To assess the influence of the perturbations on the maximum likelihood estimate ˆγ, we consider the likelihood displacement LD(ω)=2{l(ˆγ)−l(ˆγω)}, where ˆγωdenotes the maximum likelihood estimate under model l(γ|ω). The LD(ω) measures distance between ˆγand ˆγωin terms of the log-likelihood difference. It is a nonnegative function with a global minimum at ω0. The idea of local influence (Cook, 1986) is concerned about characterizing the behaviour of LD(ω) around ω0. The procedure consists in selecting a unit direction d, d=1, and then to consider the plot of LD(ω0+ad)againsta, where a∈R.Thisplot is called lifted line. Note that, since LD(ω0)=0, LD(ω0+ad) has a local minimum at a=0. Each lifted line can be characterized by considering the normal curvature Cd(γ) around a=0. This curvature is interpreted as the inverse radius of the best fitting circle at a=0. The suggestion is to consider direction dmax corresponding to the largest curvature Cdmax (γ). The index plot of dmax may reveal those observations that, under small perturbations, exercise notable influence on LD(ω). Cook(1986) showed that normal curvature at direction dtakes the form Cd(γ)=2|dTΔT(¨ L)−1Δd|where −¨ L is the observed Fisher information matrix for the postulated model (ω=ω0)andΔis the (p+1) ×nmatrix with elements Δji =∂2L(γ|ω)/∂θi∂ωj, evaluated at γ=ˆγand ω=ω0, j=1,...,p+2andi=1,...,n. Then, Cdmax is the largest eigenvalue of the matrix B=ΔT(¨ L)−1Δ,anddmax is the corresponding eigenvector. The index plot of dmax for matrix ΔT(¨ L)−1Δmay show how to perturb the log-likelihood function to obtain larger changes in the estimate of γ.Wefind, after some algebraic manipulation, the following expressions for the weighted log-likelihood function and for the elements of matrix Δ: In this case the log-likelihood function takes the form l(γ|ω)=rlog(θ)−rlog(δ) i∈F wi+(θ−1)  i∈F wilog1−exp −exp yi−xT iβ δ (11) Edwin M. M. Ortega, Vicente G. Cancho & Heleno Bolfarine 177 + i∈F wiyi−xT iβ δ− i∈F wiexp yi−xT iβ δ+ + i∈C wilog1−1−exp −exp yi−xT iβ δ θ Let us denote Δ=(Δ1,...,Δp+2)T. Then the elements of vector Δ1take the form Δ1i=⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ r  θ +log(gi)ifi∈F −(gi) θ 1−(gi) θlog(gi) θif i∈C On the other hand, the elements of vector Δ2can be shown to be given by Δ2i= ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ 1  δ−r−[ θ−1] hizi gi −zi%1−exp{  zi}&if i∈F  θzi hi(gi) θ−1  δ%1−(gi)& θif i∈C The elements of vector Δj, for j=3,...,p+2, may be expressed as Δji = ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ xij  δ−( θ−1) h−i gi +exp{  zi}−1if i∈F xij θ hi(gi) θ−1  δ1−(gi) θif i∈C where  hi=exp  zi−exp{  zi}, gi=1−exp −exp{  zi}e zi=yi−xT i β  δ However, if the interest is only in vector β, the normal curvature in direction dis given by Cd(β)=2|dTΔT(¨ L−1−B22)Δd|(see Cook, 1986), where B22 ='00 0¨ L−1 22 ( 178 Influence diagnostics in exponentiated-Weibull regression models with censored data with ¨ L22 denoting the submatrix of ¨ Lobtained according to partition ¨ L(γ)='L11 L12 L21 L22 ( The index plot of the largest eigenvector of ΔT(¨ L−1−B22)Δmay reveal those observations most influential on ˆ β. On the other hand, considering the direction for the i-th individual the total local influence in that direction is given by Ci=2|ΔT i(¨ L−1−B22)Δi|.(12) 5 Local influence on predictions Let zap×1 be a vector of values of the explanatory variables, for which we do not have necessarily an observed response. Then, the prediction at zis ˆμ(z)=)p j=1zjˆ βj. Analogously, the point prediction at zbased on the perturbed model becomes ˆμ(z,ω)= )p j=1zjˆ βjω, where ˆ βω=(ˆ β1ω,...,ˆ βpω)Tdenotes the maximum likelihood estimate from the perturbed model. Thomas and Cook (1990) have investigated the effect of small perturbations on predictions at some particular point zin continuous generalized linear models and by assuming φknown or estimated separately from ˆ β.φ−1is defined as a dispersion parameter. For more details, see McCullagh and Nelder (1989). They defined three objective functions based on different residuals. Because the diagnostic calculations were identical for the proposed functions, they concentrated the application of the methodology on the objective function f(z,ω)={ˆμ(z)−ˆμ(z,ω)}2. Similarly, we will concentrate our study on investigating the normal curvature of the surface formed by vector ωand function f(z,ω), around ω0. The normal curvature at unit direction dtakes, in this case, form Cd(z)=2|dT¨ fd |, where ¨ f=∂2f/∂ω∂ωTis evaluated at ω0and ˆ β. From Thomas and Cook (1990) one has that ¨ f=ΔT(¨ L−1 ββ zzT¨ L−1 ββ )Δ, where Δ=∂2l(γ|ω)/∂β∂ωT. Consequently dmax(z)∝−ΔT¨ L−1 ββ z. In the sequel, we discuss the calculation of dmax(z) under additive perturbations for the response and for each continuous explanatory variable. Edwin M. M. Ortega, Vicente G. Cancho & Heleno Bolfarine 179 5.1 Response perturbation Consider the regression model (8) by assuming now that each yiis perturbed as yi→yi+σωi=y∗ i,i=1,...,n, with σplaying a role of scale parameter. Below, we give the expressions for the log-likelihood function and for the elements of matrix Δ, with z∗ i=(y∗ i−xT iβ)/δ,i=1,...,n. Here, the perturbed log-likelihood function becomes expressed as l(γ|ω)=rlog(θ)−rlog(δ)+(θ−1)  i∈F log1−exp −exp y∗ i−xT iβ δ(13) + i∈Fy∗ i−xT iβ δ− i∈F expy∗ i−xT iβ δ+ + i∈C log1−1−exp−exp y∗ i−xT iβ δ θ where y∗ i=yi+σωi. Matrix Δ=(Δ1,...,Δp+2)Tis given in appendix B. Vector dmax(z) is constructed by taking z=xi, which corresponds to the n×1 vector dmax(xi)∝−ΔT¨ L−1 ββ xi.(14) A large value for the ith component of (15), dmaxi(xi), indicates that the ith observation should have substantial local influence on ˆyi. Then, the suggestion is to take the index plot of the n×1 vector (dmax1(x1),...,dmaxn(xn))Tin order to identify those observations with high influence on its own fitted value. 5.2 Explanatory variable perturbation Consider now an additive perturbation on a particular continuous explanatory variable, namely Xt, by making xitω=xit +ωiSt, where Stis a scaled factor. This perturbation scheme leads to the following expressions for the log-likelihood function and for the elements of matrix Δ. The perturbed log-likelihood function is, in this case, expressed as l(γ|ω)=rlog(θ)−rlog(δ)+(θ−1)  i∈F log1−exp −exp yi−x∗T iβ δ(15) + i∈Fyi−x∗T iβ δ− i∈F expyi−x∗T iβ δ+ 186 Influence diagnostics in exponentiated-Weibull regression models with censored data -4 -3 -2 -1 0 1 2 3 4 Index Deviance Deviance Residual Residual 5 5 34 34 101 101 -4 -3 -2 -1 0 1 2 3 4 Index Deviance Deviance Residual Residual 5 5 34 34 101 101 Figure 8:Index plot of the deviance residual rDi. 0 0,1 0,2 0,3 0,4 0,5 0,6 0,7 0,8 0,9 1 1,1 0 3 6 9 12 15 18 21 24 time Exponentiated-Weibull Regression Models Kaplan-Meier Method Survival Survival Distribution Distribution Function Function 0 0,1 0,2 0,3 0,4 0,5 0,6 0,7 0,8 0,9 1 1,1 0 3 6 9 12 15 18 21 24 time Exponentiated-Weibull Regression Models Kaplan-Meier Method Survival Survival Distribution Distribution Function Function Figure 9:Plot of the Survivor Function. residual (see, for instance, definition in McCullagh and Nelder, 1989, section 2.4) that has been largely applied in generalized linear models (GLMs). Various authors have investigated the use of deviance residuals in GLMs (see, for instance, Williams, 1987; Hinkley et al., 1991; Paula 1995) as well as in other regression models (see, for example, Farhrmeir and Tutz, 1994). In the exponentiated-Weibull regression model, the residual Edwin M. M. Ortega, Vicente G. Cancho & Heleno Bolfarine 187 deviance can be expressed here as rDi=sign(rMi)−2*rMi+δilog(δi−rMi)+1 2 where rMiis the residual martingale corresponding to the exponentiated-Weibull regression model. Analyzing the residual deviances, obtained after computing the residual martingales, it follows that individual 5 presented residual deviances greater than 3 (Figure 8). 6.4 Impact of the detected influential observations To reveal the impact of the detected influential observations, we estimate the parameters again without the influential observations. Let ˆγand ˆγ0be the maximum likelihood estimates of the models that are obtained from the data sets with and without the influential observations, respectively. Lee, Lu and Song (2006) define the following two quantities to measure the difference between ˆγand ˆγ0: TRC = np  i=1 |ˆγi−ˆγ0 i| ˆγi and MRC =maxi|ˆγi−ˆγ0 i| ˆγi where TRC is total relative changes, MRC maximum relative changes and npis the number of parameters. We find that TRC =5.490 and MRC =0.415. In order to compare the impact of the non-influential observations, we repeat the analysis after removing the same number randomly selected from non-influential observations. We find that TRC =1.786 and MRC =0.123. Hence, the ML results are more sensitive to the influential observations. Table 3:Maximum likelihood estimates for the complete data set. Parameter Estimate SE p-value θ5.556 24.761 — δ3.561 1.6481 — β00.28453 14.625 0.4705 β12.216 0.24192 <0.0001 β20.10296 0.0012952 0.0021 β3−0.12659 0.00083604 <0.0001 β40.038063 0.0000811 <0.0001 β50.0021795 0.00026622 0.4468 β60.2631 0.038086 0.0888 β70.028371 0.022559 0.4251 188 Influence diagnostics in exponentiated-Weibull regression models with censored data 6.5 A reanalysis of golden shiner data The model was estimated one more time, but without observation 5. Next, we present the results of the model fitting We can observe from Table 3 that the variable x2became significant. The survival function was also fitted again for the exponentiated-Weibull regression model (see Figure 8) in which we can observe a good model fitting. 7 Concluding remarks In this work, we have discussed applications of influence diagnostics in exponentiatedWeibull regression models with censored data. Appropriate matrices for assessing local influence as well as predictions on the fitted models under different perturbation schemes are obtained. Model fitting is also considered by using deviance residuals and graphs of the survival function. The approach was applied to simulated and real data sets, which clearly indicates the usefulness of the approach. Acknowledgements The authors acknowledge the partial financial support from Fundac¸˜ ao de Amparo ` a Pesquisa do Estado de S˜ ao Paulo (FAPESP) and CNPq-Brasil. The authors are thankful to two anonymous reviewers for valuable comments that substantially improved the paper. Appendix A: Matrix of second derivatives ¨ L(γ) (γ) (γ) Here, we derive the necessary formulas to obtain the second order partial derivatives of the log-likelihood function. After some algebraic manipulations, we obtain Lθθ =−r θ2+ i∈Cgθ i[log(gi)]2 (1 −gθ i)2; Lθδ =−1 δ i∈F zihi gi +θ i∈Czihigθ−1 ilog(gi)−gθ i+1 (1 −gθ i)2; Lθβ =−1 δ i∈F xijhi gi + i∈Cxijhigθ−1 i−gθ i+θlog(gi)+1 (1 −gθ i)2; Edwin M. M. Ortega, Vicente G. Cancho & Heleno Bolfarine 189 Lδδ =r δ2+(θ−1) δ2 i∈Fgizihi(2 +hi−hiexp{zi})−(zihi)2 g2 i+ +1 δ2 i∈F2zi(1 −exp{zi})−z2 iexp{zi}+ +θ δ2 i∈Czihigθ−1 i (1 −gθ i)2%zig−1 ihi(θ−1) +zi(1 −exp{zi})−zigθ i+ zigθ i(g−1 ihi+exp{zi})&; Lδβ =−1 δ2 i∈F (θ−1)xijhi g2 i−gi(1 +zi−ziexp{zi})+zihi− xijg2 i1−exp{zi}−ziexp{zi} −θ δ2 i∈C xijgθ−1 ihi (1 −gθ i)2(1 −gθ−1 i)1−zig−1 ihi+zi(1 −exp{zi})+ θzig−1 ihi; Lββ =−(θ−1) δ2 i∈F xijxikhigi(−1+exp{zi})+hi g2 i −1 δ2 i∈F xijxik exp{zi}+ +θ δ2 i∈C xijxikhigθ−1 i(1 −gθ−1 i)−1+exp{zi}−(θ−1)hi−θhigθ−1 i (1 −gθ i)2, where hi=exp zi−exp{zi},gi=1−exp −exp{zi}and zi=yi−xT iβ δ. Appendix B: Local influence on predictions: Response perturbation Here, we provide the derivatives of elements Δij of matrix Δconsidering the response variables perturbation scheme. The elements of vector Δ1take the form Δ1i=⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩  h∗ is g∗ i δif i∈F −(g∗ i) θ−1 h∗ is  δ%1−(g∗ i) θ& θlog(g∗ i)(g∗ i) θ 1−(g∗ i) θ+1+1if i∈C 190 Influence diagnostics in exponentiated-Weibull regression models with censored data On the other hand, the elements of the vector Δ2are expressed as Δ2i= ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ −( θ−1) s h∗ i  δ2g∗ i−z∗ i h∗ i g∗ i +z∗ i(1 −exp{  z∗ i})+s  δ2exp{  z∗ i}( z∗ i+1) −1if i∈F s θ h∗ i(g∗ i) θ−1  δ2[1 −(g∗ i)θ]z∗ i h∗ i g∗ i θ(g∗ i) θ 1−(g∗ i) θ+ θ−1−exp{  z∗ i}+1+1if i∈C, while the elements of the vector Δj,j=3,...,p+2 are expressed as Δji = ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ sxij  δ2,−( θ−1) h∗ i g∗ i1− h∗ i g∗ i −exp{  z∗ i}+exp{  z∗ i}-if i∈F s θxij h∗ i(g∗ i) θ−1 1−(g∗ i) θ h∗ i g∗ i θ(g∗ i) θ 1−(g∗ i) θ+ θ−1+1−exp{  z∗ i}if i∈C, where  h∗ i=exp  z∗ i−exp{  z∗ i}, g∗ i=1−exp −exp{  z∗ i}and  z∗ i=y∗ i−xT i β  δ. Appendix C: Local influence on predictions: Explanatory variable perturbation In this appendix we provide the derivatives of elements Δij of matrix Δ, considering the explanatory variables perturbation scheme. The elements of vector Δ1are expressed as Δ1i= ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ −s h∗ i βt  δg∗ i if i∈F s h∗ i βt(g∗ i) θ−1  δ[1 −(g∗ i) θ]1+ θlog(g∗ i)1+(g∗ i) θ 1−(g∗ i) θif i∈C, the elements of vector Δ2are expressed as Δ2i= ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ s βt  δ2( θ−1) h∗ i g∗ iz∗ i− h∗ i g∗ i −exp{  z∗ i}+1+1−exp{  z∗ i}(1 + z∗ i)+1if i∈F −s βt θ h∗ i(g∗ i) θ−1  δ2[1 −(g∗ i) θ]z∗ i h∗ i g∗ i θ(g∗ i) θ 1−(g∗ i) θ+ θ−1−exp{  z∗ i}+1+1if i∈C the elements of vector Δj, for j=1,...,pand jt, take the forms Edwin M. M. Ortega, Vicente G. Cancho & Heleno Bolfarine 191 Δji =⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ xij sβt δ2(θ−1)h∗ i g∗ ih∗ i g∗ i +exp{z∗ i}−1+exp{z∗ i}if i∈F sβtθh∗ ixij(g∗ i)θ−1 δ2[1 −(g∗ i)θ]h∗ i g∗ i−θ(g∗ i)θ 1−(g∗ i)θ−θ+1if i∈C the elements of vector Δtare given by Δti =⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ s δ−(θ−1) h∗ i g∗ ixit βt δh∗ i g∗ i +exp{z∗ i}−1+exp{z∗ i}+1−βtxit δ−1se i∈F sθh∗ i(g∗ i)θ−1 δ[1 −(g∗ i)θ]−βtxit δh∗ i g∗ iθ(g∗ i)θ 1−(g∗ i)θ+θ−1−exp{z∗ i}+1+1se i∈C where  h∗ i=exp  z∗ i−exp{  z∗ i}, g∗ i=1−exp −exp{  z∗ i}and  z∗ i=yi−x∗T i  δ References Barlow, W. E., and Prentice, R. L.(1988). Residual for relative risk regression. Biometrika, 75, 65-74. Beckman, R. J., Nachtsheim, C. J. and Cook, R. D. (1987). Diagnostics for mixed-model analysis of variance. Technometrics, 29, 413-426. Bolfarine, H. and Cancho, V. (2001). Modelling the presence of immunes by using the exponentiatedWeibull model. Journal of Applied Statistics, 28, 659-671. Cancho, V.; Bolfarine, H. and Achcar, J. A. (1999). A Bayesian analysis for the exponentiated-Weibull distribution. Journal Applied Statistical Science, 8, 227-242. Chatterjee, S. and Hadi, A. S. (1988). Sensitivity Analysis in Linear Regression. New York: John Wiley. Cox, D. R. and Snell, E. J. (1968). A general definition of residuals. Journal of the Royal Statistical Society B, 30, 248-275. Collet, D. (1994). Modelling Survival Data in Medical Research. Chapman and Hall: London. Cook, R. D. (1977). Detection of influential observations in linear regression. Technometrics, 19 15-18. Cook, R. D. (1986). Assessment of local influence (with discussion). Journal of the Royal Statistical Society, 48, 133-169. Cook, R. D. and Weisberg, S. (1982). Residuals and Influence in Regression. New York: Chapman and Hill. Davison, A. C. and Gigli, A. (1989). Deviance residuals and normal scores plots. Biometrika, 76, 211-221. Doornik, J. (1996). Ox: An Object-Oriented Matrix Programming Language. International Thomson Business Press. Escobar, L. A. and Meeker, W. Q. (1992). Assessing influence in regression analysis with censored data. Biometrics, 48, 507-528. Fahrmeir, L. and Tutz, G. (1994). Multivariate Statistical Modelling Based on Generalized Linear Models. Springer-Verlag: New York. Fleming, T. R. and Harrington, D. P. (1991). Counting Process and Survival Analysis. New York: John Wiley. 192 Influence diagnostics in exponentiated-Weibull regression models with censored data Fung, W. K. and Kwan, C. W. (1997). A note on local influence based on normal curvature. Journal of the Royal Statistical Society B, 59, 839-843. Galea, M., Paula, G. A. and Bolfarine, H. (1997). Local influence in elliptical linear regression models. The Statistician, 46, 71-79. Gu, H. and Fung, W. K. (1998). Assessing local influence in canonical analysis. Annals of the Institute of Statistical Mathematics, 50, 755-772. Hinklei, D. V., Reid, N. and Snell, E. J.(1991). Statistical Theory and Modelling-In honor of Sir David Cox. London: Chapman and Hall. Kim, M. G. (1995). Local influence in multivariate regression. Communications in Statistics Theory and Methods, 20, 1271-1278. Kwan, C. W. and Fung, W. K. (1998). Assesing local influence for specific restricted likelihood: Applications to factor analysis. Psychometrika, 63, 35-46. Lawrence, A. J. (1988). Regression transformation diagnostics using local influence. Journal of the American Statistical Association, 83, 1067-1072. Lawless, J. F. (1982). Statistical Models and Methods for lifetime data. New York: John Wiley. Lee, S. Y., Lu, B. and Song. X. Y. (2006). Assessing local influence for nonlinear structural equation models with ignorable missing data. Computational Statistics and Data Analysis, 50, 1356-1377. Lesaffre, E. and Verbeke, G. (1998). Local influence in linear mixed models. Biometrics, 54, 570-582. Liu, S. Z. (2000). On local influence for elliptical linear models. Statistical Papers, 41, 211-224. McCullagh, P. and Nelder, J. A. (1989). Generalized Linear Models, 2nd Edition. London: Chapman and Hall. Mudholkar, G. S.; Srivastava, D. K. and Friemer, M. (1995). The exponentiated Weibull family: A reanalysis of the bus-motor-failure data. Technometrics, 37, 436-445. Nelson, W. B. (1990). Accelerated Testing; Statistical Models, Test Plans and Data Analysis.NewYork: John Wiley. O’Hara, R. J., Lawless, J. F. and Carter, E. M. (1992). Diagnostics for a cumulative multinomial generalized linear model with application to grouped toxicological mortality data. Journal of the American Statistical Association, 87, 1059-1069. Ortega, E. M. M., Bolfarine, H. and Paula G. A. (2003). Influence diagnostics in generalized log-gamma regression models. Computational Statistics and Data Analysis, 42, 165-186. Paula, G. A. (1993). Assessing local influence in restricted regressions models. Computational Statistics and Data Analysis, 16, 63-79. Paula, G. A. (1995). Influence residuals in restricted generalized linear models. Journal of Statistical Computation and Simulation, 51, 63-79. Pettitt, A. N. and Bin Daud, I. (1989). Case-weight measures of influence for proportional hazards regression. Applied Statistics, 38, 51-67. Prentice, R. L. (1974). A log-gamma model and its maximum likelihood estimation. Biometrica, 61, 539544. Stacy, E. W. (1962). A generalization of the gamma distribution. Annals of Mathematical Statistics,33 1187-1192. Thomas, W. and Cook, R. D. (1990). Assessing influence on predictions from generalized linear models. Technometrics, 32, 59-65. Tsai, C. and Wu, X. (1992). Transformation-Model diagnostics. Technometrics, 34, 197-202. Therneau, T. M., Grambsch, P. M. and Fleming, T. R. (1990). Martingale-based residuals for survival models. Biometrika, 77, 147-60. Williams, D. A. (1987). Generalized linear model diagnostic using the deviance and single case deletion. Applied Statistics, 36, 181-191.