Comparison of parameter estimators for cause-specific hazards in competing risks data with missing cause of failure
Full text
Comparison of parameter estimators for cause-specific hazards in competing risks data with missing cause of failure Jimin Lee1and Kedai Cheng2 1Department of Mathematics and Statistics University of North Carolina Asheville, Asheville, NC 28804, USA Email: [email protected] 2Department of Mathematics and Statistics University of North Carolina Asheville, Asheville, NC 28804, USA Email: [email protected] ABSTRACT We consider the semiparametric proportional hazards model for the cause-specific hazard function in analysis of competing risks data with missing cause of failure. The inverse probability weighting equation is well used for estimating the regression parameters under the missing at random (MAR), provided the model for the nonmissing cause case is correctly specified. In this study, we consider a parametric logistic regression, discriminant analysis, and neural network for the model of missingness probability. Simulation studies demonstrate that the inverse probability weighted (IPW) estimators by a correctly specified logistic regression model and by discriminant analysis perform well and are appropriate for practical use. The IPW estimator by neural network works considerably well for a sufficient large sample. The application of the proposed methods are illustrated with data from a bone marrow transplant study and Women’s Interagency HIV study. Keywords: Missing cause of failure, Inverse probability weighted estimator, Logistic regression, Neural network, Discriminant analysis. Mathematics Subject Classification: 62D10, 62F12, 62N01, 62N02 1 Introduction In medical studies, the responses to a treatment can be classified in terms of failure from disease of interest or from other causes. Hence, in the competing risks framework, each individual is exposed to Kdistinct types of risks and the eventual failure is classified to precisely one of the risks. Let T∗denote the time to failure, Δ∗the cause of failure, and Zap-dimensional vector of possibly time-dependent covariates. Then, an estimable quantity in competing risks International Journal of Applied Mathematics and Statistics, Int. J. Appl. Math. Stat.; Vol. 64; Year 2025, Issue:2 (Published on 03/10/2025 ) ISSN 0973-1377 (Print), ISSN 0973-7545 (Online) Copyright © 2025 by International Journal of Applied Mathematics and Statistics www.ceser.in/ceserp
data is the cause-specific hazard function of cause k∈{1,...,K}defined by λ∗ k(t|Z) = lim h→0 1 hP(t≤T∗<t+h, Δ∗=k|T∗≥t, Z), which is the instantaneous rate of experiencing the event of type kat time t, having not experienced any of the Kcompeting events until time t. Without loss of generality in our study, we consider only two causes of failure, the cause of interest as cause 1 and the other as cause 2 (i.e. Δ∗=1or 2). In many applications involving follow-up studies, however, individuals may be subject to censoring. Let Cbe a censoring time and T∗=min(T∗ 1,T∗ 2), where T∗ 1and T∗ 2denote the latent failure times from causes 1 and 2, respectively. Then the observed data consist of observations of (T,Δ,Z), where T=min(T∗,C)and Δ=Δ ∗I(T∗≤C). If the failure time T∗is observed, Δis the cause of failure and Δ=0otherwise. The observable cause-specific hazard function of cause kin the presence of censoring is given by λk(t|Z) = lim h→0 1 hP(t≤T<t+h, Δ=k|T≥t, Z),k=1,2. Throughout the paper, we assume that Zis an external covariate process (Kalbfleish and Prentice, 2002) and that the censoring time Cis conditionally independent of (T∗,Δ∗)given Z. Under this assumption, it was shown that λ∗ k(t|Z)=λk(t|Z)if the distribution of Cis continuous at t(Lu and Tsiatis, 2001; Gao and Tsiatis, 2005; Lu and Liang, 2008). A number of statistical models for the relationship between the cause-specific hazard function of interest and regression covariates have been studied, among others, by (Prentice et al., 1978; Benichou and Gail, 1990; Cheng et al., 1998; Shen and Cheng, 1999; Scheike and Zhang, 2003). In this article we study the proportional hazards model for describing the relationship λ∗ 1(t|Z)=λ0(t)eβT 0Z(t),(1.1) where λ0(·)is a nonnegative but otherwise unspecified baseline hazard function and β0is a pdimensional vector of regression parameters. The parameter β0can be consistently estimated by treating all the failure times with Δ=1as censored observations and using the partial likelihood score equation proposed by (Cox, 1972; Cox, 1975). The estimator will be called the full-case estimator, denoted by βFin the paper. In practice, however, the information needed for the cause of failure may be lost, or it may be difficult to determine the cause of death for some individuals (Andersen et al., 1996). When we have missing causes in data, a naive method for estimating the regression parameter β0is to simply ignore the missing data and use the partial likelihood score equation to the complete data only. The so-called complete-case analysis and its estimator, denoted by βC, is clearly inefficient and can lead to serious bias. Thus, analysis of competing risks data with missing cause of failure has received considerable attention and a number of methods have been proposed. (Lu and Tsiatis, 2001) proposed a parametric model to model the probability that the missing cause is the cause of interest while allowing the inclusion of additional auxiliary covariates and then estimated the regression parameters by using a multiple imputation method. (Gao and Tsiatis, 2005) considered linear transformation models and (Lu and Liang, 2008) considered the additive hazards model for the relationship in analysis of competing risks data with missing cause of failure. International Journal of Applied Mathematics and Statistics 2
In this study, we derive an estimator of the regression parameters in model (1.1), namely the inverse probability weighted estimator and establish its theoretical properties. The approach, following the idea of (Horvitz and Thompson, 2012), uses the inverse probability weighted complete-case technique to estimate the regression parameter. The fundamental idea of IPW is to weight observed data points where the cause is known by the inverse of the probability that their cause of failure would be observed. This approach uses only the complete cases and relies on the missing at random (MAR) assumption, and the model for the missing cause case must be correctly specified for the IPW estimator to be valid (Scharfstein et al., 1999; Lu and Tsiatis, 2001; Gao and Tsiatis, 2005; Lu and Liang, 2008). Thus, we consider several models for the probability of missing causes; logistic regression model, discriminant analysis, and neural network for the model. The paper is structured as follows. In Section 2, the inverse probability weighted estimator is developed. The asymptotic properties of the corresponding estimator is established in Section 2. We review logistic regression, discriminant analysis, and neural network to estimate the probability of missing causes in Section 3. In Section 4, we investigate and compare the finite sample properties of the proposed estimators through simulations. A bone marrow transplant data set and Women’s Interagency HIV study data set are analyzed in Section 4. Some conclusions and discussions are given in Section 5. Technical notations are detailed in Appendix. 2 Estimating equation and asymptotic properties Since the cause of failure may not be observed for some individuals, we define the misssingness indicator Ras follows. If an individual’s failure is observed, then R=1when the cause of failure information Δis observed and R=0otherwise. If an individual is censored, we always define R=1. We also introduce auxiliary covariates Awhich are not of interest for modeling the cause-specific hazard function but may be used to describe the misssingness mechanism. The utilization of auxiliary information has been considered by (Lu and Tsiatis, 2001; Gao and Tsiatis, 2005; Lu and Liang, 2008; Gilbert et al., 2008), among others. Then the observed data will consist of Oi={Ti,Z i,A i,Δi,R i}if R1=1, and Oi={Ti,Z i,A i,R i}if R1=0 for i=1,...,n. We assume that {Oi,i=1,...,n}are independent identically distributed. We also assume that the cause of failure is missing at random (Rubin, 1976); that is, the probability that the cause of failure is non-missing given Δ>0and W=(T,Z,A)depends only on the observed W, but not on the unobserved Δ, P(R=1|Δ,Δ>0,W)=P(R=1|Δ>0,W).(2.1) The assumption implies that P(R=1|Δ>0,W)=P(R=1|Δ=1,Δ>0,W)=P(R=1|Δ=2,Δ>0,W). See (Lu and Tsiatis, 2001; Gao and Tsiatis, 2005; Lu and Liang, 2008). International Journal of Applied Mathematics and Statistics 3
2.1 Inverse probability weighted estimator Following the inverse selection probability idea of (Horvitz and Thompson, 2012), the method of inversely weighting the probability of complete-case has been commonly used in missing data problems. To do that, we need to estimate the probability of a complete case, π(Q)≡ P(R=1|Q), where Q=(W, Δ). By the MAR assumption and R=1when Δ=0,wehave π(Q)=P(R=1|W, Δ) =P(R=1|W, Δ,Δ>0) ·I(Δ >0) + P(R=1|W, Δ,Δ=0)·I(Δ = 0) =P(R=1|W, Δ,Δ>0) ·I(Δ >0) + 1 ·I(Δ = 0) =P(R=1|W, Δ>0) ·I(Δ >0) + I(Δ = 0) =r(W)·I(Δ >0) + I(Δ = 0),(2.2) where r(W)=P(R=1|W, Δ>0). The probability of complete-case r(Wi)may be specified by a parametric model r(Wi,ψ 0), in terms of a few unknown parameters ψ0. Accordingly, let π(Qi,ψ 0)=r(Wi,ψ 0)I(Δ >0) + I(Δ = 0). By (2.1) and (2.2), the likelihood Lregarding to π(Q, ψ0)is L(π)= n i=1 π(Qi)I(Ri=1)(1 −π(Qi))I(Ri=0) = n i=1 π(Qi)I(Ri=1)I(Δi>0) ·π(Qi)I(Ri=1)I(Δi=0) ·(1 −π(Qi))I(Ri=0)I(Δi>0) = n i=1 r(Wi)I(Ri=1)I(Δi>0) ·1·(1 −r(Wi))I(Ri=0)I(Δi>0), because π(Qi)=r(Wi)when Δi>0,and π(Qi)=1when Δi=0 = n i=1 r(Wi)RiI(Δi>0) ·(1 −r(Wi))(1−Ri)I(Δi>0) . This implies that the maximum likelihood estimator ψof ψcan be estimated by maximizing the likelihood based on only uncensored data n i=1{r(Wi,ψ)}RiI(Δi>0){1−r(Wi,ψ)}(1−Ri)I(Δi>0). We define the counting process Ni(t)=I(Δi=1)I(Ti≤t)and at-risk process Yi(t)=I(Ti≥ t). Let a⊗0=1,a⊗1=a, and a⊗2=aaTfor a vector a. Let S(m)(t, β, ψ)= 1 n n i=1 Ri π(Qi,ψ)Yi(t)eβTZi(t)Zi(t)⊗m,m=0,1,2, Z(t, β, ψ)= S(1)(t, β, ψ) S(0)(t, β, ψ), V(t, β, ψ)= S(2)(t, β, ψ) S(0)(t, β, ψ)− Z(t, β, ψ)⊗2. Then we consider the following inverse probability weighted estimating equation for β0, UI(β, ψ)= n i=1 τ 0 Ri π(Qi, ψ)Zi(t)− Z(t, β, ψ)dNi(t),(2.3) International Journal of Applied Mathematics and Statistics 4
where τ>0is the end of follow-up time. The inverse probability weighted estimator of βsolves the above equation and is denoted by βI. When there is no missing cause, the equation (2.3) consequently becomes the partial likelihood score equation proposed by (Cox, 1972). The cumulative baseline hazard function Λ0(t)=t 0λ0(u)du can be estimated by ΛI0(t)= n i=1 t 0 Ri π(Qi, ψ) dNi(u) n S(0)(u, βI, ψ). 2.2 Asymptotic results Conditions C. (C.1) λ0(t)is continuous on [0,τ]. The distribution of Cis continuous on [0,τ]and P(C>τ)> 0. The covariate processes Zi(t)have paths that are left continuous and of bounded variation, and satisfy the moment condition EZi(t)4exp(2MZi(t))<∞, where M is a constant such that β∈[−M,M]pand A=max k,l |akl|for a matrix A=(akl). (C.2) Each component of s(j)(t, β),j =0,1,2is continuous on [0,τ]×[−M,M]pfor some M>0,s(0)(t, β)>0on [0,τ]×[−M,M]p, and sup t∈[0,τ],β∈[−M,M]p S(j)(t, β)−s(j)(t, β)=Op(n−1/2). (C.3) The matrix Σ=τ 0v(t, β0)λ0(t)s(0)(t, β0)dt is positive definite. (C.4) There is a σ>0such that r(Wi)≥σfor all iwith Δi>0.r(Wi,ψ)is twice continuously differentiable with respect to ψ. There exists ψ∗satisfying the equation ES∗ ψi =0, where S∗ ψi is the corresponding score functions for r(Wi,ψ)given in (5.2) in the Appendix. The information matrix I∗ ψalso given in (5.2) in the Appendix is positive definite. We let ψ0be the true value of ψsuch that r(Wi)=r(Wi,ψ 0). In general, under Condition (C.4), there exists ψ∗such that ψconverges in probability to ψ∗and ψconsistently estimates ψ0(Haberman, 1974; Haberman, 1977; Gourieroux and Monfort, 1981; White, 1982). Thus, if the model for r(Wi)is correctly specified, we have ψ∗=ψ0and ψconverges in probability to ψ0. Theorem 2.1. Assume Conditions C are satisfied. If r(Wi,ψ)is correctly specified for r(Wi), then βIconverges in probability to β0and √n βI−β0converges in distribution to a zeromean Gaussian random vector with covariance matrix Σ−1Eω1ωT 1Σ−1, where Σ=τ 0 v(t, β0)λ0(t)s(0)(t, β0)dt, ωi=τ 0{Zi(t)−z(t, β0)}Ri π(Qi,ψ 0)dMi(t)−VψI−1 ψSψi, Mi(t)=Ni(t)−t 0Yi(u)eβT 0Zi(u)dΛ0(u),Vψis given in (5.1), Iψand Sψi are given in (5.3) in the Appendix. International Journal of Applied Mathematics and Statistics 5
The asymptotic covariance matrix Σ−1Eω1ωT 1Σ−1can be consistently estimated by Σ−1n−1 n i=1 ωiωT i Σ−1, where Σ= 1 n n i=1 τ 0 Ri π(Qi, ψ) V(t, βI, ψ)dNi(t), ωi=τ 0 Ri π(Qi, ψ){Zi(t)− Z(t, βI, ψ)}d Mi(t)− Vψ I−1 ψ Sψi, and Mi(t)=Ni(t)−t 0Yi(u)e βT IZi(u)d ΛI0(u).Here Vψ, Iψand Sψi are obtained by replacing with their respective sample estimators and substituting ( βI, ψ)for (β0,ψ 0)in Vψ,Iψ, and Sψi. See (Hyun et al., 2012) for the proof of Theorem 2.1. 3 Several models for r(W) Logistic regression Logistic regression is often used to model the probability that an experimental unit falls into one of the two groups based on predictor variables measured on the experimental unit. Since Ris binary, one can posit a parametric logistic model, logit{r(Wi,ψ)}=WT iψwhich was often used in this study area (Lu and Tsiatis, 2001; Gao and Tsiatis, 2005; Lu and Liang, 2008; Hyun et al., 2012), among others. However, model misspecification occurs, leading to biased parameter estimates and inaccurate inferences. Discriminant analysis Discriminant analysis is a method used to classify observations into nonoverlapping groups based on predictor variables. For example, a doctor might used discriminant analysis to identify patients at low or high risk for a disease based on factors like cholesterol levels, sex, blood pressure, and other variables, There are different types of discriminant analysis, including linear discriminant analysis which is a common approach that assumes a linear relationship between the predictor variables and the groups, and quadratic discriminant analysis that allows non-linear relationship between the predictor variables and the groups. Discriminant analysis relies on certain assumptions, including independence of observations, normality of predictor variables, and equal variance-covariance matrices across groups for linear discriminant analysis. The only difference from quadratic discriminant analysis is that we do not assume that the variance-covariance matrix is identical for different groups. We consider IPW estimator under linear and quadratic discriminant analysis for the probability of non-missingness, r(W)=P(R=1|W, Δ>0). International Journal of Applied Mathematics and Statistics 6
Neural network A (artificial) neural network is a computational model inspired by the structure of biological neural networks. A neural network typically consists of an input layer, one of more hidden layers, and an output layer. Each node in the hidden layer receives input, processes it, and send an output to other connected nodes. See Figure 1 for neural network diagram. Neural networks are being used in the area of prediction and classification of outcomes in medicine areas where regression models have traditionally been used. However, logistic regression assumes a linear relationship between input variables and the probability of the outcome, but neural networks can capture complex, non-linear relationship between inputs and the output. We consider IPW estimator under neural network model for the probability of non-missingness, r(W)=P(R=1|W, Δ>0). Figure 1: A neural network is an interconnected group of nodes. Each circular node represents an artificial neuron and an arrow represents a connection from the output of one neuron to the input of another. x0 x1 x2 x3 Input layer h(1) 1 h(1) 2 h(1) 3 h(1) 4 Hidden layer 1 h(2) 1 h(2) 2 h(2) 3 Hidden layer 2 ˆy1 Output layer 4 Numerical results In this study, we estimate the parameter β0under these methods and evaluate the estimates based on Mean Squared Error (MSE) and Mean Absolute Error (MAE), where MSE calculates the average of the squared differences between IPW estimates and β0, and MAE is the average of the absolute differences between IPW estimates and β0. 4.1 Simulation studies We present simulation studies conducted to evaluate the performance of the proposed methods. We consider a univariate covariate Z, where Zfollows a normal distribution N(0.5,1/12). Given Z, the latent failure time T∗ 1of interest is generated from the proportional hazards model λ∗ 1(t|Z)=λeβ0Z, where λ=1and β0=−0.5. The other latent failure time T∗ 2is generated from a Gompertz distribution with a hazard function λ∗ 2(t|Z)=eθ+νt, where θ=−0.5and ν=0.2. The censoring time Cis generated from an exponential distribution which yields about 25% censoring level. We also consider a single auxiliary covariate Awhich follows the standard normal distribution N(0,1) or a Bernoulli distribution with success probability of 0.5. For missing International Journal of Applied Mathematics and Statistics 7
cause of failure, we consider a logistic regression model logit{r(W, ψ)}=ψ1+ψ2Tb+ψ3Z+ψ4A, where Tb=Tλ−1 λ, the Box–Cox transformation with λ=0.25 and Ais the standard normal distribution (Model 1). We also consider logit{r(W, ψ)}=ψ1+ψ2T+ψ3Z+ψ4A, where Ais the Bernoulli distribution (Model 2). With ψ=(0.5,0.5,0.5,0.5), we have about 32% missingness (Model 1) and about 17% missingness (Model 2). To study the performance of the estimators when r(W)is misspecified, we posit two different parametric models of r(W, ψ), where one is a correctly specified logistic model and the other is a misspecified model logit{r(W, ψ)}=ψ1+ψ4A. The simulation studies consist of 1000 runs with the sample size n= 400 and n= 800. Neural network consists of 3 nodes in one hidden layer. We consider IPW estimators from linear and quadratic discriminant analysis. Model 1 has all normally distributed predictors which are required for discriminant analysis. The results from Tables 1 and 2 show that as expected, the full-case estimator βFshows the smallest biases, SE, MSE, and MAE in all the settings, but the complete-case estimator βC shows larger biases, SE, MSE, and MAE in almost all the settings. When the parametric logistic model for r(W)is correctly specified, the IPW estimator βIc shows small biases, but the IPW estimator βIm tends to have larger bias, MSE, and MAE when logit{r(W, ψ)}is misspecified and βIm is clearly sensitive to the misspecification. Correctly specifying a statistical model in real-world data analysis is incredibly difficult in most cases. The IPW estimator under linear discriminant analysis, βIL performs quite well and shows similar result as βIc. Since neural networks generally require a large number of observations to achieve good performance, βIN shows poor performance when n= 400 and as sample size increase its performance improves. Model 2 has that only the auxiliary predictor Ais normally distributed and the other two predictors are not normally distributed. Thus, it does not meet the normally condition for discriminant analysis. However, the results from Table 2 show that βIL and βIQ perform well and all the other estimators show similar results and patterns as in Table 1. In conclusion, βIL and βIQ under discriminant analysis and the correctly specified IPW estimator βIc show similar performance and have better performance than the other estimators, except the full-case estimator. 4.2 Bone marrow transplant (Sierra et al., 2002) described the characteristics and outcomes of 452 patients with primary myelodysplastic syndrome (MDS) who received transplants from HLA-identical siblings and were registered with the International Bone Marrow Transplant Registry (IBMTR). The study has two competing risks; treatment related death defined as death in complete remission and relapse defined as recurrence of myelodysplasia. In this example we consider 408 patients with complete covariate information obtained from the timereg R-package. Among these 408 patients, 161 patients died in complete remission, 87 patients relapsed, and 160 patients were censored. The covariates considered in our study are age of patient standardized at mean of 35 years old and platelet before transplantation (1 for more than 100 ×109per L, or 0 for less). We apply the semiparametric proportional hazards model for the cause-specific hazard function to the data set with the two covariates: platelet and age. We also consider tcell (T-cell International Journal of Applied Mathematics and Statistics 8
depleted BMT; 1 for yes, 0 for no) for auxiliary variable. In the data set, the causes of failure are all known. We can obtain βF. For illustration purposes, we delete some failure causes by the three following missing mechanisms; missing completely at random (MCAR), missing at random (MAR), and not missing at random (NMAR). For the MCAR, the causes of failure are randomly selected to yield about 23% missing causes. For the MAR, the logistic model is chosen as logit{r(W)}=0.1+0.5T−1.0age which yields about 23% missing causes, where Tis the failure time. For the NMAR, the logistic model is chosen as logit{r(W)}=0.1+1.0T−1.0age −0.5I(Δ = 1) which yields about 22% missing causes, where Δ=1corresponds to the death in complete remission, Δ=2to relapse, and Δ=0does to being censored. Because correctly specifying a statistical model is often difficult, if not impossible, in real-world data analysis, we posit two logistic models for r(W) with logit{r(W, ψ)}=ψ1+ψ2Tb+ψ3platelet +ψ4age +ψ5tcell to obtain the estimators for β, where where Tb=Tλ−1 λwith λ=0.06 (Model 1) and logit{r(W, ψ)}=ψ1+ψ2T+ψ3platelet + ψ4age +ψ5tcell (Model 2). The results of the estimation of βare summarized in Tables 3 and 4. For comparison they also includes the full-case estimator, βF, based on the original data without artificial missing. The results are close under all the missing mechanisms and almost all estimators are closer to the full-case estimator than the IPW estimator from misspecified logistic regression. The analyses are consistent with the findings from the earlier study that patients with high platelet counts have a lower risk of treatment related mortality than those with low platelet counts and a higher risk rate is seen among the older patients. 4.3 Women’s Interagency HIV Study (Lau et al., 2009) analyzed the Women’s Interagency HIV Study (WIHS) data. The WIHS was established in August 1993 to investigate the impact of HIV infection on US women. In 1994–1995, 2,054 HIV-positive and 569 HIV-negative women were enrolled. The study sample consisted of 1,164 women enrolled in WIHS, who were infected with HIV, alive, and free of clinical AIDS on December 6, 1995. Women were followed until the first of the following: time to initiation of therapy prior to AIDS or death (679 women, cause=1), time to the occurrence of AIDS or death prior to initiation therapy (359 women, cause=2), or administrative censoring on September 28, 2006 (126 women, cause=0). The data set also includes history of injection drug use at WIHS enrollment (drug, 1=yes, 0=no), whether an individual was African American (black, 1=yes, 0=no), age on December 6, 1995, standardized at mean 36.4 years, and CD4 nadir prior to December 6, 1995 (CD4). We apply the proportional hazards model for the cause-specific hazard function to the data set with the three covariates: the history of injection drug use, age, African American, and CD4 as auxiliary variable. Because the cause of failure are all known, we delete some causes of failure by missing at random (MAR) mechanism by logit{r(W)}=0.1+0.5T+0.5drug −1.5age +0.5black which yields about 20% missing causes. However, we posit two logistic models for r(W)with logit{r(W, ψ)}=ψ1+ψ2Tb+ψ3drug +ψ4age +ψ5CD4 to obtain the estimators for β, where where Tb=Tλ−1 λwith λ=0.06 (Model 1) and logit{r(W, ψ)}=ψ1+ψ2T+ψ3drug +ψ4age + ψ5CD4 (Model 2). The study found that women with an injection drug use history were less likely to initiate therapy International Journal of Applied Mathematics and Statistics 9
Table 4: Estimation of the effects of platelet and age in the bone marrow transplant data, Model 2. Platelet Age Missingness Estimator Est. SE p-value Est. SE p-value None βF−0.543 0.186 0.0035 0.377 0.087 <0.0001 βI−0.422 0.217 0.0522 0.393 0.107 0.0002 MAR βIL −0.428 0.214 0.0405 0.372 0.103 0.0003 βIQ −0.558 0.287 0.0522 0.392 0.109 0.0003 βIN −0.636 0.219 0.0038 0.375 0.097 0.0001 βI−0.775 0.209 0.0002 0.362 0.087 <0.001 MCAR βIL −0.774 0.209 0.0002 0.362 0.088 <0.001 βIQ −0.412 0.229 0.0725 0.443 0.144 0.0012 βIN −0.830 0.221 0.0002 0.366 0.086 <0.001 βI−0.568 0.228 0.0128 0.392 0.128 0.0021 NMAR βIL −0.563 0.222 0.0113 0.330 0.109 0.0025 βIQ −0.735 0.230 0.0014 0.378 0.105 0.0003 βIN −0.683 0.235 0.0037 0.453 0.141 0.0013 Est., the estimate; SE, the standard error estimate; p-value pertaining to testing no covariate effect; βF, the full-case estimator with no missing causes; βI, the IPW estimator under logistic regression; βIL and βIQ, the IPW estimators under linear and quadratic discriminant analysis, respectively; βIN, the IPW estimator under neural network. International Journal of Applied Mathematics and Statistics 16
Table 5: Estimation of the effects of drug use, age, and African American in the WIHIV study. Drug Use Age African American Missingness Estimator Est. SE p-value Est. SE p-value Est. SE p-value None βF−0.383 0.187 <0.0001 0.132 0.038 0.0005 −0.394 0.077 <0.0001 (Model 1) βI−0.494 0.117 <0.0001 0.126 0.069 0.0659 −0.321 0.090 0.0004 MAR βIL −0.513 0.120 <0.0001 0.129 0.072 0.0720 −0.317 0.091 0.0005 βIQ −0.943 0.244 0.0001 0.246 0.096 0.0103 −0.181 0.123 0.1418 βIN −0.472 0.108 <0.0001 0.127 0.056 0.0225 −0.320 0.089 0.0003 (Model 2) βI−0.511 0.117 <0.0001 0.165 0.069 0.0173 −0.308 0.090 0.0006 MAR βIL −0.511 0.121 <0.0001 0.126 0.075 0.0801 −0.309 0.090 0.0006 βIQ −0.960 0.201 <0.0001 0.342 0.070 <0.0001 −0.102 0.115 0.3745 βIN −0.490 0.107 <0.0001 0.149 0.058 0.0102 −0.320 0.088 0.0003 Est., the estimate; SE, the standard error estimate; p-value pertaining to testing no covariate effect; βF, the full-case estimator with no missing causes; βI, the IPW estimator under logistic regression; βIL and βIQ, the IPW estimators under linear and quadratic discriminant analysis, respectively; βIN, the IPW estimator under neural network. International Journal of Applied Mathematics and Statistics 17