scieee AI-readable full text Open interactive document viewer

Unconditional and Conditional Quantile Treatment Effect: Identification Strategies and Interpretations

Fort, Margherita

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Fort, Margherita Working Paper Unconditional and Conditional Quantile Treatment Effect: Identification Strategies and Interpretations Quaderni - Working Paper DSE, No. 857 Provided in Cooperation with: University of Bologna, Department of Economics Suggested Citation: Fort, Margherita (2012) : Unconditional and Conditional Quantile Treatment Effect: Identification Strategies and Interpretations, Quaderni - Working Paper DSE, No. 857, Alma Mater Studiorum - Università di Bologna, Dipartimento di Scienze Economiche (DSE), Bologna, https://doi.org/10.6092/unibo/amsacta/3720 This Version is available at: https://hdl.handle.net/10419/159696 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/3.0/ Unconditional and Conditional Quantile Treatment Effect: Identification Strategies and Interpretations Margherita Fort Quaderni - Working Paper DSE N° 857 Unconditional and Conditional Quantile Treatment Effect: Identification Strategies and Interpretations Margherita Fort, University of Bologna) This version: May 2012 Abstract This paper reviews strategies that allow one to identify the effects of policy interventions on the unconditional or conditional distribution of the outcome of interest. This distiction is irrelevant when one focuses on average treatment effects since identifying assumptions typically do not affect the parameter’s interpretation. Conversely, finding the appropriate answer to a research question on the effects over the distribution requires particular attention in the choice of the identification strategy. Indeed, quantiles of the conditional and unconditional distribution of a random variable carry a different meaning even if identification of both these set of parameters may require conditioning on observed covariates. Keywords: impact heterogeneity, quantile treatment effects, rank invariance. JEL codes: C18 1 Introduction In the recent years there has been a growing interest in the evaluation literature for models that allow essential heterogeneity in the treatment parameters and more generally for models that are informative on the impact distribution. The recent increase in the attention devoted to the identification and estimation of quantile treatment effects (QTEs) is due to their intrinsic ability to characterize the heterogenous impact of the treatment on various points of the outcome ∗Department of Economics, University of Bologna, IZA and CHILD; Piazza Scaravilli 2 Bologna, [email protected]. Paper prepared for the 46th Italian Statistical Society meeting (invited). This paper benefited from comments by E. Rettore, B. Pacini and F. Mealli and participants to the 46th Italian Statistical Society meeting. Financial support of MIURFIRB 2008 project RBFR089QQC-003-J31J10000060001 grant is gratefully acknowledged. The usual disclaimer applies. 2 distribution. QTEs are informative about the impact distribution when the potential outcomes observed under various levels of the treatment are comonotonic random variables. The variable describing the relative position of an individual in the outcome distribution thus plays a special role in this setting, representing at the same time the main dimension along which treatment effects are allowed to vary as well as a key ingredient to relate potential outcomes. Several identification approaches currently used in the literature for the assessment of mean effects have thus been extended to quantiles. Most of these strategies require to condition on a set of variables to achieve identification. While conditioning on a set of observed regressors does not affect the interpretation of the parameters in a mean regression, this is not the case for quantiles. The law of iterated expectations guarantees that the parameters of a mean regression have both a conditional and an unconditional mean interpretation. This does not carry over to quantiles, where conditioning on covariates affects the interpretation of the residual disturbance term. Indeed, since quantile regression allows one to characterize the heterogeneity of the treatment response only along this latter dimension, conditioning on covariates in quantile regression generally affects the interpretation of the results. This paper reviews strategies aimed at identifying quantile treatment effects, covering strategies that deal with the identification of conditional and unconditional quantile treatment effects with particular attetion to cross-sectional data applications in which the treatment is endogenous without conditioning on additional covariates. The aim of the paper is to provide useful guidance for users of quantile regression methods in choosing the most appropriate approach while addressing a specific research question. The remainder of the paper is organized as follows. After introducing the basic notation and the key parameters of interest in Section 2, Section 3 reviews solutions to the identification of quantile treatment effects. The review covers strategies that are appropriate only when the outcome of interest is a continuous variable, i.e. in cases where the quantiles of the outcome distribution are unambiguosly defined. It concludes illustrating some of the methods through two illustrative examples aimed at assessing the distributional impacts of training on earnings and of education on wages. Section 4 concludes. 3 2 What Are We After: Notation and Parameters of Interest In this section I first introduce the notation used throughout the paper and then define the objects whose identification is sought. Ydenotes the observed outcome, Dthe intensity of the treatment received and Wa set of observable individual characteristics. Wmay include exogenous variables Xand instruments Z.1Yis restricted to be continuous while D, W can be either continuous or discrete random variables. Both Yand Dcan be decomposed in two components: one of which is deterministic and one of which is stochastic. These two components need not be additively separable. The stochastic components account for differences in the distribution of Dand Yacross otherwise identical individuals. The econometric models reviewed in Section 3 place restriction on : i) the scale of D; ii) the number of independent sources of stochastic variation in the model; iii) the distribution (joint, marginal, conditional) of these stochastic components and Dor W≡(X, Z); iv) the scale of Z.Yd idenotes the potential outcome for individual iif the value of the treatment is d: it represents the outcome that would be observed had the individual ibeen exposed to level dof the treatment. FYd(·), fyd(·) and F−1 Yd(·) = q(d, ·) denote the corresponding cumulative distribution and density function and the quantile function. The conditional distribution and conditional quantile are denoted by FYd(·|x) and F−1 Yd(·|x) = q(d, x, ·). We are interested in characterizing the dependence structure between Yand Deventually conditioning on a set of covariates Win the presence of essential heterogeneity and in the absence of general equilibrium effects. Knowledge of the joint distribution (Yd)d∈D or the conditional joint distribution (Yd|x)d∈D would allow to characterize a distribution for the outcome for any possible level of the treatment. When potential outcomes are comonotonic, they can be described as different functions of the same (single) random variable and quantile treatment effects (QTEs) are informative on the impact distribution. The potential outcome could be written as yd=q(d, u), u eU(0,1), q(d, u) is increasing in uas is refereed in the literature as the structural quantile function. If the potential outcomes are not comonotonic, QTEs are informative on the distance between potential outcomes distributions, which may be interesting per se, but not on the impact distribution. We thus concentrate on strategies that focus on 1Capital letters denote random variables and lower case letters denote realizations. 4 QTEs.2In the binary case, QTEs (see equation (2)) are defined as the horizontal distance between the distribution function in the presence and in the absence of the treatment ([9]; [15]) .3 δ(τ) = F−1 Y1(τ)−F−1 Y0(τ) 0 < τ < 1 (2) We can distinguish conditional and unconditional quantile treatment effects by characterizing the uniformly distributed random variable that describes the quantile of the outcome variable. This distinction becomes clearer if we think about a specific empirical example. Motivating Example: Returns to Education or Training There is a large literature that studies the returns to education. Key questions in this literature (e.g. does additional education cause wage increase? does additional schooling increase wages more for the more able than for the less able? does additional schooling increase or decrease wage inequality?) can be addressed using quantile regression methods. In this applications, the treatment is likely endogenous in the outcome equation without conditioning on additional covariates: typically researchers seek instruments that allow to isolate exogenous variation in education in the wage equation. Suppose we could measure the individual ability aithat drives the endogeneity of education in the wage equation. Now, consider the alternative specifications for the wage model presented in equation (3), (4) where Ddenotes schooling (the treatment). Yi=α0(f(εi, ai)) + α1(f(εi, ai))D(3) Yi=β0(εi)ai+β1(εi)D(4) These specifications differ because they impose different structures of the variables governing the heterogeneity in the returns to education. In equation (3) the relative position of an individual in the wage distribution is determined by (εi, ai), i.e. by both an unobserved uniformly distributed error component εi and by the observed individual ability level while in equation (4) the relative position of the individual is only determined by εi. In both cases, we can think 2The review will not cover strategies that focus on other objects and may deliver QTEs as byproduct such as [8], for instance. 3In the continuous case δ(τ) represents the change in Yinduced by a change in Dfrom dto d+ when is small. δ(τ) = ∂QY(τ|d) ∂d 0< τ < 1 (1) 5 Table 1: Moment conditions under assumptions in [2] and [11] Quantile conditional unconditional Y1E[{1(Y <q(1,x)) −τ} · wy,d,x·D|X] = 0 E[{1(Y <q(1)) −τ} · wy,d·D] = 0 Y0E[{1(Y <q(0,x)) −τ} · wy,d,x·(1 −D)|X] = 0 E[{1(Y <q(0)) −τ} · wy,d·(1 −D)] = 0 weight 1−D[1−P(Z=1|Y,D,X)] 1−P(Z=1|X)−(1−D)P(Z=1|Y,D,X)] P(Z=1|X)E[Z−P(Z=1|X) P(Z=1|X)[1−P(Z=1|X)] |Y, D](2D−1) Note: Positive weights are reported. See [2] and [11] for other definitions of weights. about the relative position of an individual in the wage distribution as his/her proneness ([9]) to earn a high wage for a given level of schooling D. However, in model (3) we would refer to the total proneness/ability while in model (4) we would be speaking only about unobserved proneness/ability.4Using model (3) we can explore whether the returns to education vary depending on the individuals’ total ability levels while using model (4) we can study how the returns to education vary for given observed ability levels. Individuals who earn high wages conditional on some specific level of ability may not be the same individuals who earn high wages in the sample. However conditioning on observed ability maybe important to be able to isolate the causal effect of schooling Don the distribution of wages Y. Equation (5) and Equation (6) represent the structural quantile function corresponding to model (3) and (4) respectively5: equation (5) is an example of an unconditional quantile regression model while equation (6) is an example of a conditional quantile regression model. This distinction might be empirically relevant since, in general, for a given τ∈(0,1), α1(τ)6=β1(τ). f(ε, a)≡ε∗, ε∗eU(0,1) QY(τ|d) = α0(τ) + α1(τ)d(5) εeU(0,1) QY(τ|d) = β0(τ)ai+β1(τ)d(6) 3 Identification Strategies and Estimation In cross-sectional applications, two main identification approaches have been extended to QTEs: strategies based on the unconfoundedness assumption and strategies based on the availability of an instrumental variable. In the first case, the researcher must be willing to assume that the joint distribution of the potential outcomes is independent of the treatment conditional on a set of exogenous 4To the best of my knowledge, [17] is the first to distinguish between total and observed proneness. 5Under comonotonicity of potential outcomes, the structural quantile function describes the link between potential outcomes. 6 covariates. Under this assumptions, conditional QTEs can be estimated as originally proposed by [14] and unconditional QTEs can be estimated as proposed by [10]. [2] and [6], [7] propose identifying assumptions for conditional quantiles when an instrumental variable is available. The assumptions of [2] guarantee identification of conditional and unconditional QTEs when the treatment is binary and endogenous and a binary instrument is available. They lead to the moment conditions described in Table 1: in both cases, identification relies on previous results ([1], [13]) that guarantee that in the subpopulation of compliers comparisons by treatment D, conditional on X, have a causal interpretation. Recall that compliers are individuals whose treatment status is affected by the instrument Zbut that this sub-population cannot be identified directly from the data, because it is defined by means of potential outcomes. The moment conditions highlight that is possible to construct weights that ’find compliers in the population in an average sense’ ([1]). The weights will differ when one is interested in the conditional or in the unconditional quantiles. Only the weights considered in the second case ’simultaneously balance the distribution of the covariates between treated and non-treated compliers’ ([12]). In both cases weights are functions of P(Z= 1|X) and observed variables. Estimation thus proceeds in two steps: 1) weights are estimated; 2) weighted quantile regressions are run.6. Estimation requires two steps also under the identification strategy proposed by [6], [7] and [17], [18] but does not involve re-weighting. The crucial assumption for identification in the approach by [6] is rank invariance or rank similarity, i.e. we require that the individual’s rank in the potential outcome distribution, conditional on exogenous covariates, is not systematically affected by the treatment. The assumptions by [6] lead to the moment condition in equation (7). Equation (7) suggests an estimation procedure that first requires to compute the conditional quantiles of the random variable Y−q(d, x, τ) given X and Z; then, choose as estimate of q(d, x, τ) the one that minimizes the absolute value of the coefficient associated with Zin the first step.7 P r[Y−q(d, x, τ)≤0|X, Z] = τ. (7) 6When identification is achieved relying on uncounfoundedness, the moment conditions are similar but the weights are identically 1 for conditional quantiles ([14]) and are D P(D=1|X)+1−D 1−P(D=1|X)for unconditional quantiles ([10]). 7This approach can be used when the treatment and instrument are binary, discrete as well as continuous. 7 The instrumental variable approach for the identification of unconditional QTEs proposed by [17] delivers the moment condtions in equation (8) E[Z{1(Y ≤q(d, τ)) −τX}]=0, τX≡P[Y ≤q(d, τ)|X].E[1(Y ≤q(d, τ)) −τ]=0. (8) These moment conditions reflect the idea that, first, the instrument Zdoes not affect the distribution of the disturbance once Xis controlled for and, second, the joint distribution of Xand the disturbance is unrestricted. Estimation involves first an estimation of the quantiles of Y−q(d, τ) given X and Z and τX; then, a second step choose as estimate of q(d, τ) the value that minimizes the coefficient of Zaveraging over all possible values of X. We now apply these strategies to two illustrative examples taken from the literature. Table 2 reports estimates of the effect of training (or education) on the conditional and unconditional distribution of earnins (or log wages) using data of males from [2] and data from [5], respectively.8Column (1) and (2) reports results delivered when training or education are treated as exogenous in the estimation of conditional and unconditional quantiles respectively. Column (3) and (4) report estimates that address the endogeneity of training or education in the outcome equation relying on [2]. These estimates apply to the sub-population of compliers. Column (5)-(8) report estimates based on [6] or [17]. These approaches guarantee global identification of conditional and unconditional QTEs. We discuss the top-panel estimates first: in the example from [2] the treatment assignment is randomized thus covariates are not needed for identification. Indeed, under both the identification approaches considered, training effects on the conditional and unconditional quantiles do not exhibit substantial differences in magnitude and all suggest that the effect of training is larger at the top of the earnings distribution.9In addition, both the identification strategies deliver similar results, suggesting that key assumptions are unlikely to be violated in both cases. Let’s now turn to the estimates in the bottom part of 8In the second example, only reforms that increased compulsory schooling for 3 years are considered (i.e. only Greece, Italy and Finland) and the original treatment (years of education) and instrument (years of compulsory schooling) were recoded to binary. Estimates of column (1), (2), (3), (4) have been computed by the author using the STATA package ivqte by [12], except column (3) for the first example (taken from the article). Estimates in column (1) replicate original results in the papers except that standard errors are now robust to heteroskedasticity; estimates of columns (5)-(8) are taken from [18] for the AAI02 example and obtained using the STATA package ivqreg by Do Wan Kwack available from Christian Hansen’s research page. 9When endogeneity of training is addressed, point estimates of the returns to training are generally lower in the unconditional distribution with respect to the returns observed holding race, age, education and marital status fixed. 8