Measuring quality for use in incentive schemes: The case of "shrinkage" estimators
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Mehta, Nirav Article Measuring quality for use in incentive schemes: The case of "shrinkage" estimators Quantitative Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Mehta, Nirav (2019) : Measuring quality for use in incentive schemes: The case of "shrinkage" estimators, Quantitative Economics, ISSN 1759-7331, The Econometric Society, New Haven, CT, Vol. 10, Iss. 4, pp. 1537-1577, https://doi.org/10.3982/QE950 This Version is available at: https://hdl.handle.net/10419/217174 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/
Supplementary Material Supplement to “Measuring quality for use in incentive schemes: The case of “shrinkage” estimators” (Quantitative Economics, Vol. 10, No. 4, November 2019, 1537–1577) Nirav Mehta Department of Economics, University of Western Ontario Appendix B: Cutoff model proofs and extensions B.1 Direct conditioning on class size The difference in administrator’s value from using different teacher-quality estimators derives from the assumption that the administrator chooses a cutoff policy based on only test score information. Such a one-dimensional policy is quite simple and, therefore, is of considerable clear policy relevance; this demonstrated by Table A.1, which documents existing incentive schemes and shows that none condition on class size. Moreover, when compared with a policy that may also explicitly condition on class sizes, a test-score-based cutoff may attenuate issues of class size manipulation for the sake of affecting the administrator’s posterior about the quality of a particular teacher. However, allowing the administrator to explicitly take into account class size may still be of interest. This section shows how the theoretical results in Section 3 would be affected. Now suppose the administrator, instead of only indirectly taking it into account when maximizing her utility, could instead explicitly condition on class size ni.Ifniwas a strictly monotonic function of teacher quality θ, then the administrator could achieve a perfect classification of teachers by inverting n(θ)—even if she ignored all teachers’ test scores. A more realistic case would allow for multiple teacher qualities for at least one class size. Suppose that the distribution of teacher qualities for each class size was normally distributed. Because the administrator can explicitly condition on class size she can hold a separate cutoff-based classification problem for each class size level; denote the administrator’s value from using the fixed effects and empirical Bayes estimators as vFE CPn(κ) and vEB CPn(κ), respectively. Then by Proposition 1, the administrator would obtain the same value for either estimator given the desired cutoff κ,thatis, vFE CPn(κ) =vEB CPn(κ) for all (n κ). Therefore, we can without loss of generality consider only the fixed-effects estimator, with optimal cutoff policy c∗FE n. Further note that the administrator’s expected objective would be at least as high if she is allowed to split her original objective into one objective for each class size; if the cutoff for c∗FE n1=c∗FE n2for all class sizes n1n2, then her value under the separate class size scheme would be the same as that from her original objective. Nirav Mehta: [email protected] ©2019 The Author. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at http://qeconomics.org.https://doi.org/10.3982/QE950
2Nirav Mehta Supplementary Material B.2 Administrator’s problem with infinite precision We want to prove that as the variance of the measurement error tends to 0 (which implies σ→0) all teachers will be correctly categorized, giving vFE CP(κ) =vEB CP(κ) =1for all desired κ. First, consider the fixed effects estimator. The administrator’s utility for a teacher with true quality θunder the fixed effects estimator and cutoff policy cis uCPθ ˆ θFE;cκ=α1ˆ θFE ≥c|θ≥κ+(1−α)1ˆ θFE <c|θ<κ p →α1{θ≥c|θ≥κ}+(1−α)1{θ<c|θ<κ}(S.1) which is maximized at c=κ. The administrator’s utility from using cutoff policy cunder the empirical Bayes estimator, for the same teacher, is plim σ→0 uCPθ ˆ θEB;cκ =α1ˆ θEB ≥c|θ≥κ+(1−α)1ˆ θEB <c|θ<κ =α1λ(θ)θ ≥c|θ≥κ+(1−α)1λ(θ)θ < c|θ<κ (S.2) which is maximized at c=κ/λ(F−1(κ)). The probabilities of the events in both (S.1) and (S.2)areall1, giving an expected utility of 1 for all teacher qualities, which then integrates to a value of 1for each estimator. B.3 Proof of Proposition 2 Recall that we are considering first the case where κ c∗EB <0. Differentiating the administrator’s value with respect to β−and evaluating at β−=0, our goal is to show when ∂vEB CP ∂β−β−=0 =(1−α)κ −∞ −c∗EBθ (δ−)2σ φθ−c∗EB δ− σφ(θ/σθ) σθ(κ/σθ)dθ +α0 κ c∗EBθ (δ−)2σ φθ−c∗EB δ− σφ(θ/σθ) σθ1−(κ/σθ)dθ < 0(S.3) The conjugate nature of the normal distribution (Bromiley (2003)) allows us to combine the above densities into one Gaussian density,1letting us write 1 σφ( c∗EB δ−−θ σ)×1 σθφ( θ σθ)= 1If X1∼F1=N(μ1σ2 1),withdensityf1and X2∼F2=N(μ2σ2 2),withdensityf2,thenf1(θ) ×f2(θ) = fP(θ) SP,where fP(θ) =1 σP φθ−μP σP
Supplementary Material Measuring quality for use in incentive schemes 3 fP(θ) SP,whereSPis a positive constant and fP(θ) =1 σP φθ−μP σP μP= c∗EB δ− σ2 θ+0×σ2 σ2+σ2 θ= c∗EB δ− σ2 θ σ2+σ2 θ=c∗EB σ2 P=σ2σ2 θ σ2+σ2 θ where the last equality on the second line follows because δ−=λ(θ)|β−=0=σ2 θ σ2+σ2 θ .Dividing through by SP,wehave∂vEB CP ∂β−|β−=0<0if κ −∞ c∗EBθ (δ−)2fP(θ) dθ 0 κ c∗EBθ (δ−)2fP(θ) dθ >α 1−α (κ/σθ) 1−(κ/σθ) ⇒κ −∞ θfP(θ)dθ 0 κ θfP(θ)dθ >α 1−α (κ/σθ) 1−(κ/σθ)(S.4) The top and bottom terms on the left side of (S.4) are expectations of truncated normal random variables, scaled by truncation probabilities, that is, κ −∞ θfP(θ)dθ =FP(κ) −FP(−∞)EFP[θ|θ<κ] =FP(κ) −FP(−∞)⎛ ⎜ ⎜ ⎝μP+σP φ−∞−μP σP−φκ−μP σP κ−μP σP−−∞−μP σP⎞ ⎟ ⎟ ⎠ =κ−μP σPμP+σP −φκ−μP σP κ−μP σP μP=μ1σ2 2+μ2σ2 1 σ2 1+σ2 2 σ2 P=σ2 1σ2 2 σ2 1+σ2 2 and SPis a positive and constant scaling factor that depends on μ1μ2σ1σ2according to Sp=φ( μ1−μ2 σ2 1+σ2 2 ).
4Nirav Mehta Supplementary Material and 0 κ θfP(θ)dθ =FP(0)−FP(κ)EFP[θ|κ<θ<0] =FP(0)−FP(κ)μP+σP φκ−μP σP−φ−μP σP −μP σP−κ−μP σP =−μP σP−κ−μP σPμP+σP φκ−μP σP−φ−μP σP −μP σP−κ−μP σP Putting the above expressions back into the comparison (S.4) and rearranging, we have μP+σP −φκ−μP σP κ−μP σP μP+σP φκ−μP σP−φ−μP σP −μP σP−κ−μP σP #1 ×1−κ σθ κ σθ #2 × κ−μP σP −μP σP−κ−μP σP #3 >α 1−α(S.5) Recall that we are considering the case where κ<0and c∗EB <0.2Therefore, both the numerator and denominator of expression #1 are negative, with the numerator greater in absolute value than the denominator, which implies that expression #1 >1.Moreover, κ<0implies that (κ/σθ)<1/2, which implies that expression #2 >1as well. For example, consider α≤1/2. In these cases, the expressions #1 and #2 would satisfy the above inequality. However, the expression #3 could potentially be less than 1. Due to our well-known lack of a closed-form expression for the standard normal CDF , it is impossible to sign the above inequality in a purely analytical manner for all possible cases. Moreover, the inequality would be violated for extreme parameterizations, such as where the administrator only placed value on one type of error (making α go to either the corner of zero or one). However, it is still possible to show why Condition (S.5) would likely hold for a wide range of reasonable values of (ακσθσc∗EB). To start, assume α=1/2, in which case what is left to be shown is that the left side of the inequality (S.5) is greater than one (in Appendix B.4, I use numerical methods to verify that estimator rankings consistent with (S.5) hold for a wide range of α). For 2I show below that similar conditions obtain when considering κ>0and c∗EB >0.
Supplementary Material Measuring quality for use in incentive schemes 5 simplicity, express the cutoff policy as a fraction of the desired cutoff κ,thatis,cEB =γcκ, where γc∈[01](i.e., cEB willbethesamesignasκand no larger in absolute value than κ) and write μP=c∗EB =γcκ. Because #1 is always greater than one, it is sufficient to satisfy the following condition: 1< 1−κ σθ κ σθ× κ−μP σP −μP σP−κ−μP σP ⇒ −γc √γσ κσ (1−γc) √γσ κσ<1 (κσ)(S.6) where κσ≡κ/σθexpresses the desired cutoff in terms of standard deviations of teacher quality and γσ≡σ2 σ2+σ2 θ∈(01). With the above simplifications, whether Condition (S.6) will be satisfied (which is sufficient for satisfying Condition (S.5)whenα=1/2) depends on (γσγcκσ)∈(01)× [01]×(−∞0]. The following cases analyze when (γσγcκσ)satisfy Condition (S.6): Case a: κσ→0For all (γσγc)∈(01)×[01], Condition (S.6) will also be satisfied as κσ→0, as the left side tends to one and the right side tends to two. Case b: κσ<0There are two relevant subcases: Case b.i: 1−γc≤√γσ Lemma S.1. Condition (S.6)will be satisfied if 1−γc≤√γσ. Proof. The numerator of the left side of (S.6) will always be smaller than the numerator of the right for any finite κσ.If1−γc≤√γσ,thenwehave(1−γc) √γσκσ≥κσ⇒ ((1−γc) √γσκσ)≥(κσ), that is, the denominator of the left will be at least as large as the denominator of the right. Case b.ii: 1−γc>√γσ Lemma S.2. There exists a threshold ˆγσ(κσγc)∈(01)such that,given (κσγc),Condition (S.6)is met for γσ>ˆγσ(κσγc). Proof. Note that (γσγc)only enter the left side of (S.6). The limit of the left side of Condition (S.6)asγσ→0is (∞)/(−∞)=1/0=∞, which is greater than the right side for any finite κσ.Asγσ→1, the above sufficient condition 1−γc≤√γσis satisfied for γc∈(01)because 1−γc≤1⇔γc≥0, meaning the left side is less than the right side. Because the left side is continuous in γσ,thereexistsa ˆγσ(κσγc)by the Intermediate Value Theorem such that G(κσγcˆγσ(κσγc)) ≡ ( −γc √ˆγσ(κσγc)κσ) ( (1−γc) √ˆγσ(κσγc)κσ)−1 (κσ)=0.The
6Nirav Mehta Supplementary Material solution to Gwill be unique if the left side of (S.6) is monotonically decreasing in γσ. The derivative of the left side with respect to γσis ∂−γc √γσ κσ (1−γc) √γσ κσ ∂γσ =2κσ √γσ ⎡ ⎢ ⎢ ⎢ ⎣ γcφ−γc √γσ κσ(1−γc) √γσ κσ+(1−γc)φ(1−γc) √γσ κσ−γc √γσ κσ (1−γc) √γσ κσ2⎤ ⎥ ⎥ ⎥ ⎦ (S.7) which is negative because the term in brackets is positive and κσ<0. Lemma S.3. The threshold ˆγσ(κσγc)is decreasing in γc. Proof. Having established that there exists a cutoff ˆγσsolving Gin Lemma S.2,wecan implicitly differentiate Garound the solution to obtain ∂ˆγσ ∂γc=− ∂G ∂γc ∂G ∂γσ =− ∂−γc √γσ κσ (1−γc) √γσ κσ ∂γc ∂−γc √γσ κσ (1−γc) √γσ κσ ∂γσ The derivative in the denominator, (S.7), has been shown to be negative, which means that ∂ˆγσ ∂γcwill have the same sign as the derivative in the numerator (because the entire expression is multiplied by negative one): ∂−γc √γσ κσ (1−γc) √γσ κσ ∂γc =κσ √γσ ⎡ ⎢ ⎢ ⎢ ⎣ φ(1−γc) √γσ κσ−γc √γσ κσ−φ−γc √γσ κσ(1−γc) √γσ κσ (1−γc) √γσ κσ2⎤ ⎥ ⎥ ⎥ ⎦
Supplementary Material Measuring quality for use in incentive schemes 7 Table B.1. Summary of cases. Case Parameterization Will (S.5) be Satisfied? Proof aκσ→0Yes Direct (see page 5of text) b.i κσ<0and 1−γc≤√γσYes Lemma S.1 b.ii κσ<0and 1−γc>√γσDepends on (γσγcκσ)Lemma S.2 The denominator of the term in the brackets is positive iff φ(1−γc) √γσ κσ−γc √γσ κσ>φ −γc √γσ κσ(1−γc) √γσ κσ ⇔ φ(1−γc) √γσ κσ (1−γc) √γσ κσ> φ−γc √γσ κσ −γc √γσ κσ which is satisfied because each side represents the reverse hazard function for the normal distribution r(x) =φ(x) (x), which is decreasing (see Theorem 17 of Chechile (2011)), and we have (1−γc) √γσκσ<0<−γc √γσκσ. Putting the above together, although it is not possible to obtain explicit conditions that exactly characterize the set of parameters satisfying the desired inequality, it is possible to show that the inequality would be satisfied for a wide range of parameters. Table B.1 summarizes the cases considered above. First, all (γcγσ)∈(01)×[01]satisfy the inequality as κσ→0(case a). Intuitively, as κσ→0, the countervailing decrease in Type I errors vanishes, meaning the increase in Type II errors will necessarily dominate. For κσ<0(case b), we know that there exists a cutoff value of γσfor each (κσγc)such that the inequality will be satisfied, and that this cutoff is decreasing in γc(making it more likely to hold). If γcis large enough (case b.i) then κσdoes not need to be considered to verify whether the inequality will be satisfied. Otherwise (case b.ii), whether the inequality will be satisfied depends on (γσγcκσ).Intuitively,as1−γcincreases (relative to √γσ) the weights φ(θ−c∗EB δ− σ)in (S.3) increase more for larger values of |θ|, causing the increase in Type II errors to be larger than the decrease in Type I errors. Therefore, it is instructive to compute the “worst-case” cutoff where γc=0. Setting γc=0and rearranging (S.6), we obtain (0) κσ √γσ<1 (κσ)⇔γσ>κσ −1(κσ) 22 where the sign reversal in the last inequality is due to −1((κσ) 2)<0. Take, for example, κσ=−1, meaning the administrator wishes to identify the bottom 16% teachers. In this case, ˆγσ(κσ=−1γc=0)≈050. Several empirical studies find that test scores are
8Nirav Mehta Supplementary Material Figure B.1. Cutoff ˆγσwhen κσ=−1and κσ=−2. comprised of roughly equal parts signal and noise (see, e.g., Staiger and Rockoff (2010), which means we can use γσ=1/2as a rough approximation. This, combined with the fact that ˆγσwould only decrease as γcincreased from its lower bound of zero, makes it very likely that Condition (S.6) would be satisfied. Figure B.1 plots the sufficient condition γσ≥(1−γc)2(dotted curve) and the κσdependent ˆγσwhen κσis −1(solidcurve)and−2 (dot-dashed curve). Any value of γσ above a scenario-specific curve would satisfy Condition (S.6) for that scenario. For example, the result that all values of γσwould satisfy the inequality as γc→1can be seen as the sufficient condition goes down to zero when γc→1.Asκσincreases in absolute value the cutoff ˆγσincreases, where ˆγσapproaches the sufficient condition (1−γc)2as |κσ|→∞. Across a wide variety of class size and teacher quality scenarios and desired cutoffs (κ), the typical value of γcis in the range of 0275 to 04. For example, the 25th and 75th percentiles of γcare, respectively, 0286 and 030 when using the calibrated relationship between class size and teacher quality for Reading in the LAUSD. The red segment shows the range of typical values of γc, coupled with the typical value of γσ≈1/2.Muchofthe red segment lies in the “sufficiency” part of the plot, meaning Condition (S.6)wouldbe satisfied for any nonpositive κσ. We can see that the typical values of (γcγσ)would also easily satisfy Condition (S.6)whenκσ=−1and κσ=−2. It is important to remember that Condition (S.6)isonlyasufficientconditionforthe derivative of the value under empirical Bayes with respect to β−being negative, as it ignores the first term in (S.5). For example, consider γc=0275 and γσ=05. Setting γc to the lowest value in the typical range is conservative as it makes (S.6) harder to satisfy.
Supplementary Material Measuring quality for use in incentive schemes 15 where Prˆ θEB ≥q=∞ −∞ Prˆ θEB ≥q|θf(θ)dθ=∞ −∞ Pr!≥q λn(θ)−θ"f(θ)dθ =∞ −∞ 1− q λn(θ)−θ σ(θ) f(θ)dθ We then have Prˆ θEB <q =1−Prˆ θEB ≥q=∞ −∞ q λn(θ)−θ σ(θ) f(θ)dθ As before, these quantities when using fixed effects are obtained by setting λ(·)=1: Eθ|ˆ θFE ≥q =1 Prˆ θFE ≥q∞ −∞∞ q−θ θ1−q−θ σ(θ)f(θ)dθ Prˆ θFE ≥q=∞ −∞ 1−q−θ σ(θ)f(θ)dθ Prˆ θFE <q =∞ −∞ q−θ σ(θ)f(θ)dθ The administrator’s value from using the empirical Bayes estimator, using the result that she will use a cutoff signal policy, is then vEB HT2(χ) =max q−Prˆ θEB <q χ+Prˆ θEB ≥qEθ|ˆ θEB ≥q =max q!−Prˆ θEB <q χ+Prˆ θEB ≥q1 Prˆ θEB ≥q ×∞ −∞∞ q/λ(n(θ))−θ θ1−q/λn(θ)−θ σ(θ) f(θ)dθ" =max q!−Prˆ θEB <q χ+∞ −∞ θ1−q/λn(θ)−θ σ(θ) f(θ)dθ" =max q!∞ −∞ (−χ) q λn(θ)−θ σ(θ) f(θ)dθ +∞ −∞ θ1−q/λn(θ)−θ σ(θ) f(θ)dθ"(S.8) and the administrator’s value from using fixed effects is vFE HT2(χ) =max q!∞ −∞ (−χ)q−θ σ(θ)f(θ)dθ+∞ −∞ θ1−q−θ σ(θ)f(θ)dθ"(S.9)
16 Nirav Mehta Supplementary Material Intuitively, under either estimator the administrator will either replace a teacher with some probability, which in expectation reduces her objective by an expected value of χ, or retains the teacher with the complementary probability, in which case her objective is increased by that teacher’s quality θ. Because n(θ) is not constant (as it was in HT-0), the reliability of signals varies by teacher and the analytical characterization of the administrator’s reservation signal from Model HT-0 no longer obtains. The estimator-specific reservation signals, q∗EB and q∗FE, are respectively obtained by numerically solving (S.8)and(S.9). The ranking of the administrator’s utility from HT-2, by class size scenario n(θ),is the same as her ranking under the cutoff-based model. Proposition S.2. In Model HT2, the administrator’s preferred estimator depends on the relationship between teacher quality and class size.In particular,when the relationship between teacher quality and class size is negative-(positive-)quadratic,the administrator will prefer the fixed effects (empirical Bayes)estimator. Proof. I will focus on teachers with θ<0, since below-average teachers will be most affected by the reservation policy, which will generally be negative.5As in the cutoff model, parameterize the empirical Bayes weights using λ(θ) =max{λδ−+β−θ}if θ<0 max{λδ++β+θ}if θ≥0 where λ>0, and assume homoskedastic errors, that is, σ(θ) =σfor all θ.6Differentiating the administrator’s value with respect to β−and evaluating the derivative at β−=0 (as will be made clear below, this is convenient because the environment in which the derivative is evaluated is consistent with the one studied in HT-0, the constant class size scenario studied by Proposition 4), we obtain ∂vEB HT2 ∂β−β−=0=0 −∞ q∗EB δ2 − θ[χ+θ]1 σ φ q∗EB δ−−θ σ1 σθ φθ σθdθ (S.10) because ∂vEB HT2 ∂q∗EB ×∂q∗EB ∂β−=0due to the envelope theorem. As in the proof of Proposition 2, combine the normal densities into one density, fP(θ)/SP,whereSPis a positive constant (see footnote 1for details). The expression (S.10) becomes ∂vEB HT2 ∂β−β−=0=1 SP0 −∞ q∗EB δ2 − θ[χ+θ] m(θ) fP(θ) dθ 5Figure C.2 shows that estimator rankings are similar for the negative-quadratic and increasing scenarios, and between the positive-quadratic and decreasing scenarios. 6The numerical solutions presented later in this section and the quantitative results in Section 5.3 allow for mechanical heteroskedasticity in σ(θ).
Supplementary Material Measuring quality for use in incentive schemes 17 Figure C.1. |m(θ)|and fP(θ). where fP(θ) =1 σPφ(θ−μP σP),μP=q∗EB δ− σ2 θ σ2+σ2 θ=q∗EB,σ2 P=σ2 θγσ,andγσ=σ2 σ2+σ2 θ .The function m(θ) is strictly concave (in particular, negative quadratic) in θand has zeros at θ=−χand θ=0. Therefore, we have ∂vEB HT2 ∂β−β−=0 <0⇔−χ −∞ −m(θ)fP(θ) dθ > 0 −χ m(θ)fP(θ) dθ Observe that −m(−χ−a) > m(−χ+a) for 0<a<χand −m(−χ−a) ≥0for a>0. Therefore, if fPwere centered around a value no greater than −χ,thatis,μP≤−χ, then by the symmetry of the normal distribution we would have #−χ −∞−m(θ)fP(θ) dθ > #0 −χm(θ)fP(θ) dθ because −m(θ)fP(−χ−θ) > m(θ)fP(−χ+θ) for θ>0,andthus, ∂vEB HT2 ∂β−|β−=0<0. Note that the measurement errors have been assumed to be homoskedastic and that λ(·)is constant because we are evaluating the derivative at β−=0, which means the optimal fixed effects policy from Model HT-0 (see (8)) would obtain when using fixed effects: q∗FE =−χ ρ,whereρ=σ2 θ σ2 θ+σ2 <1. Then, by Proposition 4 we have q∗EB =ρq∗FE =−χ, which satisfies μP≤−χ, and the desired inequality is obtained.7 Figure C.1 illustrates this scenario, where the solid curve is |m(θ)|,whichis−m(θ) for θ≤−χand m(θ) for θ∈[−χ 0], and the dot-dashed curve is fP, drawn with μP=−χ. 7Although not as straightforward to solve for explicitly, q∗EB is very close to −χacross a wide variety of n(θ), even when allowing for both mechanical heteroskedasticity in σ(θ) and the effect of class size in λ(θ).
18 Nirav Mehta Supplementary Material Figure C.2. Ratio of values of FE over EB, across class size scenarios. To illustrate this theoretical result, Figure C.2 plots the ratio of value from the FE over EB estimators for a wide range of replacement costs χ,8across different class size scenarios: constant, increasing, decreasing, negative quadratic, and positive quadratic. The constant class size scenario (dotted black line) represents a special case of HT-2 where n(θ) =n, which is simply model HT-0. Unsurprisingly, then, we obtain the same value for all replacement costs χ. Under the negative quadratic scenario (dot-short-dashed blue curve) the administrator would obtain higher value from using fixed effects for every χ. This is the same result as was obtained for a wide range of parameterizations of the cutoff-based model. Also, as in the cutoff-based model, the estimator ranking is reversed under the positive-quadratic class size scenario (long-dashed brown curve); that is, she would prefer to use empirical Bayes instead of fixed effects. The rankings of increasing (short-dashed red) and negative-quadratic (dot-short-dashed blue) scenarios are similar. As argued above, most of the difference in value comes from the part of the teacher quality distribution at highest risk of being replaced—those with below-average qualities. By the same reasoning, the rankings of decreasing (dot-long-dashed green) and positive-quadratic (long-dashed brown) scenarios are also similar.9 As with model HT-1, an environment with multiple periods could be modeled by suitably adjusting the desired threshold quality. For example, adding more periods could 8Wiswall (2013) reported that teachers with 30 years of experience have value-added that is one standard deviation higher than new teachers and 075 standard deviations higher than teachers with 5years of experience; this implies a 025 sd difference acquired in the first 5years of experience. Therefore, I set χ=025σθ=0054 for the baseline quantitative results and for this figure use a range for the replacement cost running from zero to 030, over five times this value. 9This figure takes into account the mechanical heteroskedasticity caused by the variation in n(θ).
Supplementary Material Measuring quality for use in incentive schemes 19 be accommodated by decreasing the replacement cost, as the administrator would have a relatively higher gain from replacing when there are more periods of output. Because they range from a cost of zero to several times the estimated difference in value-added between a teacher with 5years experience and no experience, Figure C.2 then likely also characterizes estimator rankings for multiperiod environments. The takeaway from this section is that (i) the administrator’s preferred estimator depends on the class size scenario n(θ), (ii) though the difference in values from using either estimator depends on other model parameters (T χ), the preferred estimator does not, and (iii) the administrator would prefer the same estimator in HT-2 as she would in the cutoff model. Appendix D: Details for quantitative exercises D.1 Calibrated error variances I calibrate σ2 θand σ2 from Table B-2 of Schochet and Chiang (2012) normalizing the total variance to one. To most closely match a policy where an administrator would like to rank teachers across a school district, I calibrate σ2 θ=0046 by summing the average of schooland teacher-level variances in random effects. To most closely approximate an environment where both student and aggregate-level shocks may affect student test scores, I calibrate σ2 =0953 by summing the average of classand student-level variances in random effects. Note that, due to the much greater student-level error variance, the approximate sizes of σ2 θand σ2 are approximately the same if school-level variances are excluded from σ2 θor class-level variances are excluded from σ2 , lending robustness to the quantitative findings. D.2 Heteroskedasticity correction for relationship between class size and teacher quality The advantage of the indirect inference approach is that it can be implemented using a vector of auxiliary moments which do not necessarily correspond to structural econometric parameters. This is useful in the current context, where the microdata to directly correct for heteroskedasticity are not available.10 Indirect inference algorithm The following is done separately for Reading and Math. 0. Estimate the relationship between class size (ni) and teacher i’s estimated quality in the subject ( ˆ θi) by running the regression ni=βdata 0+βdata 1ˆ θi+βdata 2(ˆ θi)2+ ei. The regression coefficients (ˆ βdata 0ˆ βdata 1ˆ βdata 2)and residual standard error ˆσdata e form the first set auxiliary parameters to fit. Compute the 25th, 50th, and 75th percentiles of the empirical distribution of class sizes, (ndata p25 ndata p50 ndata p75 ). These are the remaining auxiliary parameters. The target vector of auxiliary parameters is then (ˆ βdata 0ˆ βdata 1ˆ βdata 2ˆσedatandata p25 ndata p50 ndata p75 ). 10If microdata had been available, then one could in principle use an approach like the one in Lockwood and McCaffrey (2014) to account for the nonlinearities produced by heteroskedastic errors.
20 Nirav Mehta Supplementary Material 1. Given σ2 θ, simulate teacher quality θsim ionce for each teacher in the sample. (Recall the population mean has been normalized to 0.) 2. Simulate the random component of class sizes nsim iiid, which is distributed normal with mean zero and standard deviation σnsim . As described below, this algorithm chooses the parameter σnsim . Note these are independent from teacher quality to get an idea of the role heteroskedasticity plays. 3. Assign incremental class sizes according to ninc(θsim i)=a0+a1θsim i+a2(θsim i)2.As described below, this algorithm chooses the parameters (a0a1a2). The final simulated class size for teacher iis then nsim i=round{nsim iiid+ninc(θsim i)}, that is, class sizes are integer-valued. 4. Given σ2 and nsim isimulate an average shock for each teacher, sim i;formsimulated estimated teacher quality according to ˆ θsim i=θsim i+sim i. 5. Regress nsim i=βsim 0+βsim 1ˆ θsim i+βsim 2(ˆ θsim i)2, estimating the auxiliary coefficients (ˆ βsim 0ˆ βsim 1ˆ βsim 2)and auxiliary residual standard error ˆσsim e. Compute the 25th, 50th, and 75th percentiles of the simulated distribution of class sizes, (nsim p25nsim p50nsim p75).The simulated vector of auxiliary parameters is then (ˆ βsim 0ˆ βsim 1ˆ βsim 2ˆσsim ensim p25nsim p50nsim p75). 6. Compute the Euclidean distance between target auxiliary parameters and simulated auxiliary parameters (e.g., ˆ βdata 0and ˆ βsim 0, resp.) as a function of the parameters governing class size, d(a0a1a2σnsim ). Repeat steps 1–6 for the vector (a0a1a2σnsim ), until the distance between data and simulated auxiliary moments is minimized. D.3 Details for quantitative illustration for hidden action model Output in the hidden action model depends on several parameters, including the variance of measurement error on output, σ2 η. I adjust the error variance in several steps, using Reading test scores as the measure: 1. Simulate teacher quality, class sizes, and measurement errors using the parameters from Section 5.1, for 30,000 teachers. Each simulated teacher then has a simulated quality θs iand a simulated fixed-effect estimate ˆ θsFE i. 2. Use the empirical Bayes weights λ(·)to generate simulated EB measures of teacher quality according to ˆ θsEB i=λ(n(θs i)) ˆ θsFE i. 3. Standardize θs i,ˆ θsFE i,and ˆ θsEB ito have variances of 1, to make the residual variances comparable. 4. Finally, I estimate the residual variance from a regression of standardized ˆ θsFE ion the standardized true (simulated) quality θs iand the residual variance from a regression of standardized empirical Bayes measure ˆ θsEB ion standardized true (simulated) quality. The ratio of residual variances, or amount unexplained in each regression, tells us how much more (or less) the fixed effects estimator would inform the administrator about teacher quality.
Supplementary Material Measuring quality for use in incentive schemes 21 Table D.1. Regressions of simulated teacher quality on FE and EB estimates. Dependent Variable: θs(Standardized) (1) (2) ˆ θsFE (standardized) 0718 (0004) ˆ θsEB (standardized) 0707 (0004) Constant 0002 −0001 (0004)(0004) Observations 30,000 30,000 R20516 0500 Residual Std. Error (df =29,998)0696 0707 Note: Standard errors are reported in parenthesis. The regression results, shown in Table D.1, indicate that the fixed-effects estimator explains about 32% more variation in teacher quality than the empirical Bayes estimator (1−069562/070702=0032). That is, the fact that the EB estimator makes it more difficult to separate highand low-performing teachers when the class size function is negative quadratic, as it is in the data, can be modeled as increasing the measurement error variance on teacher output, σ2 η, by this amount. References Boyd, D., H. Lankford, S. Loeb, and J. Wyckoff (2013), “Measuring test measurement error: A general approach.” Journal of Educational and Behavioral Statistics, 38 (6), 629– 663. [11] Bromiley, P. (2003), “Products and convolutions of Gaussian probability density functions.” Tina-Vision Memo, 3 (4). [2] Chechile, R. A. (2011), “Properties of reverse hazard functions.” Journal of Mathematical Psychology, 55 (3), 203–222. [7] Lockwood, J. and D. McCaffrey (2014), “Should nonlinear functions of test scores be used as covariates in a regression model?” In Value-Added Modeling and Growth Modeling With Particular Application to Teacher and School Effectiveness, Chapter 1 (R. Lissitz and H. Jiang, eds.), 1–36, Information Age, Charlotte, NC. [19] Schochet, P. and H. Chiang (2012), “What are error rates for classifying teacher and school performance using value-added models?” Journal of Educational and Behavioral Statistics.[19] Staiger, D. and J. Rockoff (2010), “Searching for effective teachers with imperfect information.” Journal of Economic Perspectives, 24 (3), 97–117. [8] Wiswall, M. (2013), “The dynamics of teacher quality.” Journal of Public Economics, 100, 61–78. [14,18]
22 Nirav Mehta Supplementary Material Co-editor Peter Arcidiacono handled this manuscript. Manuscript received 17 August, 2017; final version accepted 18 July, 2018; available online 10 April, 2019.