scieee AI-readable full text Open interactive document viewer

Small area estimation: its evolution in five decades

Ghosh, Malay

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Ghosh, Malay Article Small area estimation: its evolution in five decades Statistics in Transition New Series Provided in Cooperation with: Polish Statistical Association Suggested Citation: Ghosh, Malay (2020) : Small area estimation: its evolution in five decades, Statistics in Transition New Series, ISSN 2450-0291, Exeley, New York, Vol. 21, Iss. 4, pp. 1-22, https://doi.org/10.21307/stattrans-2020-022 This Version is available at: https://hdl.handle.net/10419/236772 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc-nd/4.0/ STATISTICS IN TRANSITION new series, Special Issue, August 2020 Vol. 21, No. 4, pp. 1–22, DOI 10.21307/stattrans-2020-022 Received – 31.01.2020; accepted – 30.06.2020 Small area estimation: its evolution in five decades Malay Ghosh1 ABSTRACT The paper is an attempt to trace some of the early developments of small area estimation. The basic papers such as the ones by Fay and Herriott (1979) and Battese, Harter and Fuller (1988) and their follow-ups are discussed in some details. Some of the current topics are also discussed. Key words: template, article, journal. 1. Prologue Small area estimation is witnessing phenomenal growth in recent years. The vastness of the area makes it near impossible to cover each and every emerging topic. The review articles of Ghosh and Rao (1994), Pfeffermann (2002, 2013) and the classic text of Rao (2003) captured the contemporary research of that time very successfully. But the literature continued growing at a very rapid pace. The more recent treatise of Rao and Molina (2015) picked up many of the later developments. But then there came many other challenging issues, particularly with the advent of “big data”, which started moving the small area estimation machine faster and faster. It seems real difficult to cope up with this super-fast development. In this article, I take a very modest view towards the subject. I have tried to trace the early history of the subject up to some of the current research with which I am familiar. It is needless to say that the topics not covered in this article far outnumber those that are covered. Keeping in mind this limitation, I will make a feeble attempt to trace the evolution of small area estimation in the past five decades. 2. Introduction The first and foremost question that one may ask is “what is small area estimation”? Small area estimation is any of several statistical techniques involving estimation of parameters in small ‘sub-populations’ of interest included in a larger ‘survey’. The term ‘small area’ in this context generally refers to a small geographical area such as a county, census tract or a school district. It can also refer to a ‘small domain’ cross-classified by 1Department of Statistics, University of Florida, Gainesville, FL, USA. E-mail: [email protected]fl.edu. ORCID: https://orcid.org/0000-0002-8776-7713. 2 M. Ghosh: Small area estimation: its evolution... several demographic characteristics, such as age, sex, ethnicity, etc. I want to emphasize that it is not just the area, but the ‘smallness’ of the targeted population within an area that constitutes the basis for small area estimation. For example, if a survey is targeted towards a population of interest with prescribed accuracy, the sample size in a particular subpopulation may not be adequate to generate similar accuracy. This is because if a survey is conducted with sample size determined to attain prescribed accuracy in a large area, one may not have the resources available to conduct a second survey to achieve similar accuracy for smaller areas. A domain (area) specific estimator is ‘direct’ if it is based only on the domain-specific sample data. A domain is regarded as ‘small’ if domain-specific sample size is not large enough to produce estimates of desired precision. Domain sample size often increases with population size of the domain, but that need not always be the case. This requires use of ‘additional’ data, be it either administrative data not used in the original survey, or data from other related areas. The resulting estimates are called ‘indirect’ estimates that ‘borrow strength’ for the variable of interest from related areas and/or time periods to increase the ‘effective’ sample size. This is usually done through the use of models, mostly ‘explicit’, or at least ‘implicit’ that links the related areas and/or time periods. Historically, small area statistics have long been used, albeit without the name “small area” attached to it. For example, such statistics existed in eleventh century England and seventeenth century Canada based on either census or on administrative records. Demographers have long been using a variety of indirect methods for small area estimation of population and other characteristics of interest in postcensal years. I may point out here that the eminent role of administrative records for small area estimation cannot but be underscored even today. A very comprehensive review article in this regard is due to Erciulescu, Franco and Lahiri (2020). In recent years, the demand for small area statistics has greatly increased worldwide. The need is felt for formulating policies and programs, in the allocation of government funds and in regional planning. For instance, legislative acts by national governments have created a need for small area statistics. A good example is SAIPE (Small Area Income and Poverty Estimation) mandated by the US Legislature. Demand from the private sector has also increased because business decisions, particularly those related to small businesses, rely heavily on local socio-economic conditions. Small area estimation is of particular interest for the transition economics in central and eastern European countries and the former Soviet Union countries. In the 1990’s these countries have moved away from centralized decision making. As a result, sample surveys are now used to produce estimates for large areas as well as small areas. 3. Examples Before tracing this early history, let me cite a few examples that illustrate the ever increasing current day importance of small area estimation. One important ongoing small STATISTICS IN TRANSITION new series, Special Issue, August 2020 3 area estimation problem at the U.S. Bureau of the Census is the small area income and poverty estimation (SAIPE) project. This is a result of a Bill passed by the US House of Representatives requiring the Secretary of Commerce to produce and publish at least every two years beginning in 1996, current data related to the incidence of poverty in the United States. Specifically, the legislation states that “to the extent feasible”, the secretary shall produce estimates of poverty for states, counties and local jurisdictions of government and school districts. For school districts, estimates are to be made of the number of poor children aged 5-17 years. It also specifies production of state and county estimates of the number of poor persons aged 65 and over. These small area statistics are used by a broad range of customers including policy makers at the state and local levels as well as the private sector. This includes allocation of Federal and state funds. Earlier the decennial census was the only source of income distribution and poverty data for households, families and persons for such small geographic areas. Use of the recent decennial census data pertaining to the economic situation is unreliable especially as one moves further away from the census year. The first SAIPE estimates were issued in 1995 for states, 1997 for counties and 1999 for school districts. The SAIPE state and county estimates include median household income number of poor people, poor children under age 5 (for states only), poor children aged 5-17, and poor people under age 18. Also starting 1999, estimates of the number of poor school-aged children are provided for the 14,000 school districts in the US (Bell, Basel and Maples, 2016). Another example is the Federal-State Co-Operative Program (FSCP). It started in 1967. The goal was to provide high-quality consistent series of post-censal county population estimates with comparability from area to area. In addition to the county estimates, several members of FSCP now produce subcounty estimates as well. Also, the US Census Bureau used to provide the Treasury Department with Per Capita Income (PCI) estimates and other statistics for state and local governments receiving funds under the general revenue sharing program. Treasury Department used these statistics to determine allocations to local governments within the different states by dividing the corresponding state allocations. The total allocation by the Treasury Dept. was $675 billion in 2017. United States Department of Agriculture (USDA) has long been interested in prediction of areas under corn and soybeans. Battese, Harter and Fuller (JASA, 1988) considered the problem of predicting areas under corn and soybeans for 12 counties in North-Central Iowa based on the 1978 June enumerative survey data as well as Landsat Satellite Data. The USDA statistical reporting Service field staff determined the area of corn and soybeans in 37 sample segments of 12 counties in North Central Iowa by interviewing farm operators. In conjunction with LANDSAT readings obtained during August and September 1978, USDA procedures were used to classify the crop cover for all pixels in the 12 counties. 4 M. Ghosh: Small area estimation: its evolution... There are many more examples. An important current day example is small area “poverty mapping” initiated by Elbers, Lanjouw and Lanjouw (2003). This was extended as well as substantially refined by Molina and Rao (2010) and many others. 4. Synthetic Estimation An estimator is called ‘Synthetic’ if a direct estimator for a large area covering a small area is used as an indirect estimator for that area. The terminology was first used by the U.S. National Center for Health Statistics. These estimators are based on a strong underlying assumption is that the small area bears the same characteristic for the large area. For example, if y1,···,ymare the direct estimates of average income for mareas with population sizes N1,···,Nm, we may use the overall estimate ¯ys=∑m j=1Njyj/Nfor a particular area, say, i,where N=∑m j=1Nj. The idea is that this synthetic estimator has less mean squared error (MSE) compared to the direct estimator yiif the bias ¯ys−yiis not too strong. On the other hand, a heavily biased estimator can affect the MSE as well. One of the early use of synthetic estimation appears in Hansen, Hurwitz and Madow (1953, pp 483-486). They applied synthetic regression estimation in the context of radio listening. The objective was to estimate the median number of radio stations heard during the day in each of more than 500 counties in the US. The direct estimate yiof the true (unknown) median Miwas obtained from a radio listening survey based on personal interviews for 85 county areas. The selection was made by first stratifying the population county areas into 85 strata based on geographical region and available radio service type. Then one county was selected from each stratum with probability proportional to the estimated number of families in the counties. A subsample of area segments was selected from each of the sampled county areas and families within the selected area segments were interviewed. In addition to the direct estimates, an estimate xiof Mi, obtained from a mail survey was used as a single covariate in the linear regression of yion xi. The mail survey was first conducted by sampling 1,000 families from each county area and mailing questionnaires. The xiwere biased due to nonresponse (about 20% response rate) and incomplete coverage, but were anticipated to have high correlation with the Mi. Indeed, it turned out that Corr(yi,xi)=.70. For nonsampled counties, regression synthetic estimates were ˆ Mi=.52 +.74xi. Another example of Synthetic Estimation is due to Gonzalez and Hoza (JASA, 1978, pp 7-15). Their objective was to develop intercensal estimates of various population characteristics for small areas. They discussed synthetic estimates of unemployment where the larger area is a geographic division and the small area is a county. Specifically, let pij denote the proportion of labor force in county ithat corresponds to cell j(j=1,···,G). Let ujdenote the corresponding unemployment rate for cell j STATISTICS IN TRANSITION new series, Special Issue, August 2020 5 based on the geographic division where county ibelongs. Then, the synthetic estimate of the unemployment rate for county iis given by u∗ i=∑G j=1pijuj. These authors also suggested synthetic regression estimate for unemployment rates. While direct estimators suffer from large variances and coefficients of variation for small areas, synthetic estimators suffer from bias, which often can be very severe. This led to the development of composite estimators, which are weighted averages of direct and synthetic estimators. The motivation is to balance the design bias of synthetic estimators and the large variability of direct estimators in a small area. Let yij denote the characteristic of interest for the jth unit in the ith area; j=1,···,Ni;i= 1,···,m. Let xij denote some auxiliary characteristic for the jth unit in the ith local area. Note that the population means are ¯ Yi=∑Ni j=1yij/Niand ¯ Xi=∑Ni j=1xij/Ni. We denote the sampled observations as yij,j=1,···,niwith corresponding auxiliary variables xij, j=1,···,ni. Let ¯xi=∑ni j=1xij/ni.¯xiis obtained from the sample. In addition, one needs to know ¯ Xi, the population average of auxiliary variables. A Direct Estimator (Ratio Estimator) of ¯ Yiis ¯yR i=(¯yi/¯xi)¯ Xi. The corresponding Ratio Synthetic Estimator of ¯ Yiis (¯ys/¯xs)¯ Xi, where ¯ys=∑m i=1Ni¯yi/∑m i=1Niand ¯xs=∑m i=1Ni¯xi/∑m i=1Ni. A Composite Estimator of ¯ Yiis (ni/Ni)¯yi+(1−ni/Ni)(¯ys/¯xs)¯ X i, where ¯ X i=(Ni−ni)−1∑Ni j=ni+1xij/(Ni−ni).NoteNi¯ Xi=ni¯xi+(Ni−ni)¯ X i. All one needs to know is the population average ¯ Xiin addition to the already known sample average ¯xi to find ¯ X i. Several other weights in forming a linear combination of direct and synthetic estimators have also been proposed in the literature. The Composite Estimator proposed in the previous paragraph can be given a modelbased justification as well. Consider the model yij ind ∼(bxij,σ2xij). Best linear unbised estimator of bis obtained by minimizing ∑m i=1∑ni j=1(yij −bxij)2/xij. The solution is ˆ b=¯ys/¯xs. Now estimate ¯ Yi=( ∑ni j=1yij+∑Ni j=ni+1yij)/Niby ∑ni j=1yij/Ni+ˆ b∑Ni j=ni+1xij/Ni. This simplifies to the expression given in the previous paragraph. Holt, Smith and Tomberlin (1979) provided more general model-based estimators of this type. 5. Model-Based Small Area Estimation Small area models link explicitly the sampling model with random area specific effects. The latter accounts for between area variation beyond that is explained by auxiliary variables. We classify small area models into two broad types. First, the “area level” models that relate small area direct estimators to area-specific covariates. Such models are necessary if unit (or element) level data are not available. Second, the “unit level” models that relate the unit values of a study variable to unit-specific covariates. Indirect 6 M. Ghosh: Small area estimation: its evolution... estimators based on small area models will be called “model-based estimators”. The model-based approach to small area estimation offers several advantages. First, “optimal” estimators can be derived under the assumed model. Second, area specific measures of variability can be associated with each estimator unlike global measures (averaged over small areas) often used with traditional indirect estimators. Third, models can be validated from the sample data. Fourth, one can entertain a variety of models depending on the nature of the response variables and the complexity of data structures. Fifth, the use of models permits optimal prediction for areas with no samples, areas where prediction is of utmost importance. In spite of the above advantages, there should be a cautionary note regarding potential model failure. We will address this issue to a certain extent in Section 7 when we discuss benchmarking. Another important issue that has emerged in recent years, is design-based evaluation of small area predictors. In particular, design-based mean squared errors (MSE’s) is of great appeal to practitioners and users of small area predictors, because of their long-standing familiarity with the latter. Two recent articles addressing this issue are Pfeffermann and Ben-Hur (2018) and Lahiri and Pramanik (2019). The classic small area model is due to Fay and Herriot (JASA, 1979) with Sampling Model: yi=θi+ei,i=1,...,mand Linking Model: θi=xT ib+ui,i=1,...,m. The target is estimation of the θi,i=1,...,m. It is assumed that eiare independent (0,Di), where the Diare known and the uiare iid (0,A), where Ais unknown. The assumption of known Dican be put to question because they are, in fact, sample estimates. But the assumption is needed to avoid nonidentifiablity in the absence of microdata. This is evident when one writes yi=xT ib+ui+ei. In the presence of microdata, it is possible to estimate the Dias well. An example appears in Ghosh, Myung and Moura (2018). A few notations are needed to describe the Fay-Herriot procedure. Let y=(y1,...,ym)T; θ=(θ1,...,θm)T:e=(e1,...,em)T;u=(u1,...,um)T;XT=(x1,...,xm);b=(b1,...,bp)T. We assume Xhas rank p(<m). In vector notations, we write y=θ+eand θ=Xb+u. For known A, the best linear unbiased predictor (BLUP) of θiis (1−Bi)yi+BixT i˜ bwhere ˜ b=(XTV−1X)−1XTV−1y,V=Diag(D1+A,···,Dm+A)and Bi=Di/(A+Di). The BLUP is also the best unbiased predictor under assumed normality of yand θ. It is possible to give an alternative Bayesian formulation of the Fay-Herriott model. Let yi|θiind ∼N(θi,Di);θi|bind ∼N(xT ib,A). Then the Bayes estimator of θiis (1−Bi)yi+BixT ib, where Bi=Di/(A+Di). If instead we put a uniform(Rp)prior for b, the Bayes estimator of θiis the same as its BLUP. Thus, there is a duality between the BLUP and the Bayes estimator. STATISTICS IN TRANSITION new series, Special Issue, August 2020 7 However, in practice, Ais unknown. A hierarchical prior joint for both band Ais π(b,A)=1. (Morris, 1983, JASA). Otherwise, estimate Ato get the resulting empirical Bayes or empirical BLUP. We now describe the latter. There are several methods for estimation of A. Fay and Herriot (1979) suggested solving iteratively the two equations (i) ˜ b=(XTV−1X)−1XTV−1yand (ii) ∑m i=1(yi−xT i˜ b)2= m−p. The motivation for (i) comes from the fact that ˜ bis the best linear unbiased estimator (BLUE) of bwhen Ais known. The second is a method of moments equation noting that the expectation of the left hand side equals m−p. The Fay-Herriot method does not provide an explicit expression for A. Prasad and Rao (1990, JASA) suggested instead a unweighted least squares approach, which provides an exact expression for A. Specifically, they proposed the estimator ˆ bL=(XTX)−1XTy. Then E||y−Xˆ bL||2=(m−p)A+∑m i=1Di(1−ri),ri=xT i(XTX)−1xi,i=1,···,m. This leads to ˆ AL=max0,||y−Xˆ bL||2−∑m i=1Di(1−ri) m−pand accordingly ˆ BL i=Di/(ˆ AL+Di). The corresponding estimator of θis ˆ θEB i=(1−ˆ BL i)yi+ˆ BL ixT i˜ b(ˆ AL), where ˜ b(ˆ AL)=[XTV−1(ˆ AL)X]−1XTV−1(ˆ AL)y. Prasad and Rao also found an approximation to the mean squared eror (Bayes risk) of their EBLUP or EB estimators. Under the subjective prior θiind ∼N(xT ib,A), the Bayes estimator of θiis ˆ θB i=(1−Bi)yi+BixT ib,Bi=Di/(A+Di). Also, write ˜ θEB i(A)=(1−Bi)yi+ BixT i˜ b(A). Then E(ˆ θEB i−θi)2=E(ˆ θB i−θi)2+E(˜ θEB i(A)−ˆ θi B)2+E(ˆ θEB i−˜ θEB i(A))2. The cross-product terms vanish due to their method of estimation of A, by a result of Kackar and Harville (1984). The first term is the Bayes risk if both band Awere known. The second term is the additional uncertainty due to estimation of bwhen Ais known. The third term accounts for further uncertainty due to estimation of A. One can get exact expressions E(θi−ˆ θB i)2=Di(1−Bi)=g1i(A),say and E(ˆ θEB i(A)− ˆ θB i)2=B2 ixT i(XTV−1X)−1xi=g2i(A),say. However, the third term, E(ˆ θEB i−ˆ θEB i(A))2 needs an approximation. An approximate expression correct up to O(m−1), i.e. the remainder term is of o(m−1), as given in Prasad and Rao, is 2B2 i(Di+A)−1A2∑m i=1(1− Bi)2/m2=g3i(A),say. Further, an estimator of this MSE correct up to O(m−1)is g1i(ˆ A)+g2i(ˆ A)+2g3i(ˆ A). This approximation is justified by noticing E[g1i(ˆ A)] = g1i(A)− g3i(A)+o(m−1). A well-known example where this method has been applied is estimation of median income of four-person families for the 50 states and the District of Columbia in the United States. The U.S. Department of Health and Human Services (HHS) has a direct need for such data at the state level in formulating its energy assistance program for low-income families. The basic source of data is the annual demographic supplement to the March sample of the Current Population Survey (CPS), which provides the median income of 8 M. Ghosh: Small area estimation: its evolution... four-person families for the preceding year. Direct use of CPS estimates is usually undesirable because of large CV’s associated with them. More reliable results are obtained these days by using empirical and hierarchical Bayesian methods. Here sample estimates of the state medians for the current year (c) as obtained from the Current Population Survey (CPS) were used as dependent variables. Adjusted census median (c) defined as the base year (the recent most decennial census) census median (b) times the ratio of the BEA PCI (per capita income as provided by the Bureau of Economic Analysis of the United States Bureau of the Census) in year (c) to year (b) was used as an independent variable. Following the suggestion of Fay (1987), Datta, Ghosh, Nangia and Natarajan (1996) used the census median from the recent most decennial census as a second independent variable. The resulting estimates were compared against a different regression model employed earlier by the US Census Bureau. The comparison was based on four criteria recommended by the panel on small area estimates of population and income set up by the US committee on National Statistics. In the following, we use eias a generic notation for the ith small area estimate, and ei.TR the “truth”, i.e. the figure available from the recent most decennial census. The panel recommended the following four criteria for comparison. Average Relative Absolute Bias =(51)−1∑51 i=1|ei−ei,TR|/ei,TR. Average Squared Relative Bias =(51)−1∑51 i=1(ei−ei,TR)2/e2 i,TR. Average Absolute Bias =(51)−1∑51 i=1|ei−ei,TR|. Average Squared Deviation =(51)−1∑51 i=1(ei−ei,TR)2. Table 1 compares the Sample Median, the Bureau Estimate and the Empirical BLUP according to the four criteria as mentioned above. Table 1. Average Relative Absolute Bias, Average Squared Relative Bias, Average Absolute Bias and Average Squared Deviation (in 100,000) of the Estimates. Bureau Estimate Sample Median EB Aver. rel. bias 0.325 0.498 0.204 Aver. sq. rel bias 0.002 0.003 0.001 Aver. abs. bias 722.8 1090.4 450.6 Aver. sq. dev. 8.36 16.31 3.34 There are other options for estimation of A. One due to Datta and Lahiri (2000) uses the MLE or the residual MLE (RMLE). With this estimator, gDL 3iis approximated by 2D2 i(A+Di)−3[∑m i=1(A+Di)−2]−1, while g1iand g2iremain unchanged. Finally, Datta, Rao and Smith (2005), went back to the original Fay-Herriot method of estimation of A, and obtained gDRS 3i=2D2 i(A+Di)−3m[∑m i=1(A+Di)−2]−1. The string of inequalities m−1m ∑ i=1 (A+Di)2≥[m−1m ∑ i=1 (A+Di)]2≥m2[ m ∑ i=1 (A+Di)−1]2 STATISTICS IN TRANSITION new series, Special Issue, August 2020 15 0.1 0.2 0.3 0.4 Figure 1: Map of posterior means of θ’s. (ii) In Mississippi, Georgia, Alabama and New Mexico, 55%+ counties have poverty rates >the third quartile (18.9%). (iii) In New Hampshire, Connecticut, Rhode Island, Wyoming, Hawaii and New Jersey, 70%+ counties have poverty rates <the first quartile (11.1%). (iv) Examples of counties with high poverty ratios are Shannon, SD; Holmes, MS; East Carroll, LA; Owsley, KY; Sioux, IA. (v) Examples of counties with large random effects are Madison, ID; Whitman, WA; Harrisonburg, VA; Clarke, GA; Brazos, TX. Dr. Pfeffermann suggested splitting the counties, whenever possible, into a few smaller groups, and then use the same global-local priors for estimating the random effects separately for the different groups. From a pragmatic point of view, this may sometimes be necessary for faster implementation. It seems though that the MCMC implementation even for such a large number of counties was quite easy since all the conditionals were standard disributions, and samples could be generated easily from these distributions at each iteration. 9. Variable Transformation Often the normality assumption can be justified only after transformation of the original data. Then one performs the analysis based on the transformed data, but transform back properly to the original scale to arrive at the final predictors. One common example is transformation of skewed positive data, for example, income data where log transfor- 16 M. Ghosh: Small area estimation: its evolution... mation gets a closer normal approximation. Slud and Maiti (2006) and Ghosh and Kubokawa (2015) took this approach, providing final results for the back-transformed original data. For example, consider a multiplicative model yi=φiηiwith zi=log(yi),θi=log(φi) and ei=log(ηi). Consider the Fay-Herriott (1979) model (i) zi|θiind ∼N(θi,Di)and (ii) θiind ∼N(xT iβ,A).θihas the N( ˆ θB i,Di(1−Bi)) posterior with ˆ θB i=(1−Bi)zi+BixT iβ, Bi=Di/(A+Di).NowE(φi|zi)=E[exp(θi)|zi]=exp[ˆ θB i+(1/2)Di(1−Bi)]. Another interesting example is the variance stabilizing transformation. For example, suppose yiind ∼Bin(ni,pi). The arcsine transformation is given by pi=sin−1(2pi−1). The back transformation is pi=(1/2)[1+sin(θi)]. A third example is the Poisson model for count data. There yiind ∼Poisson(λi). Then one models zi=y1/2 ias independent N(θi,1/4)where where θi=λ1/2 i. An added advantage in the last two examples is that the assumption of known sampling variance, which is really untrue, can be avoided. 10. Final Remarks As acknowledged earlier, the present article leaves out a large number of useful current day topics in small area estimation. I list below a few such topics which are not covered at all here. But there are are many more. People interested in one or more of the topics listed below and beyond should consult the book of Rao and Molina (2015) for their detailed coverage of small area estimation and an excellent set of references for these topics. • Design consistency of small area estimators. • Time series models. • Spatial and space-time models. • Variable Selection. • Measurement errors in the covariates. • Poverty counts for small areas. • Empirical Bayes confidence intervals. • Robust small area estimation. • Misspecification of linking models. STATISTICS IN TRANSITION new series, Special Issue, August 2020 17 • Informative sampling. • Constrained small area estimation. • Record Linkage. • Disease Mapping. • Etc, Etc., Etc. Acknowledgements I am indebted to Danny Pfeffermann for his line by line reading of the manuscript and making many helpful suggestions, which improved an earlier version of the paper. Partha Lahiri read the original and the revised versions of this paper very carefully, and caught many typos. A comment by J.N.K. Rao was helpful. The present article is based on the Morris Hansen Lecture delivered by Malay Ghosh before the Washington Statistical Society on October 30, 2019. The author gratefully acknowledges the Hansen Lecture Committee for their selection. REFERENCES ARMAGAN, A., CLYDE, M., and DUNSON, D. B., (2013). Generalized double pareto shrinkage. Statistica Sinica, 23, pp. 119–143. ARMAGAN, A., DUNSON, D. B., LEE, J., and BAJWA, W. U., (2013). Posterior consistency in linear models under shrinkage priors. Biometrika, 100, pp. 1011–1018. BATTESE, G. E., HARTER, R. M., and FULLER, W. A., (1988). An error components model for prediction of county crop area using survey and satellite data. Journal of the American Statistical Association, 83, pp. 28–36. BELL, W. R., DATTA, G, S., and GHOSH, M., (2013). Benchmarking small area estimators. Biometrika, 100, pp. 189–202. BELL, W. R., BASEL, W. W., and MAPLES, J. J., (2016). An overview of U.S. Census Bureau’s Small Area Income and Poverty Estimation Program. In Analysis of Poverty Data by Small Area Estimation. Ed. M. Pratesi. Wiley, UK, pp. 349–378. BERG, E., CECERE, W., and GHOSH, M., (2014). Small area estimation of county level farmland cash rental rates. Journal of Survey Statistics and Methodology, 2, pp. 1–37. Bivariate hierarchical Bayesian model for estimating cropland cash rental rates at the county level. Survey Methodology, in press. 18 M. Ghosh: Small area estimation: its evolution... BOOTH, J. G., HOBERT, J., (1998). Standard errors of prediction in generalized linear mixed models. Journal of the American Statistical Association, 93, pp. 262–272. BUTAR, F. B., LAHIRI, P., (2003). On measures of uncertainty of empirical Bayes small area estimators. Journal of Statistical Planning and Inference, 112, pp. 63–76. CARVALHO, C. M., POLSON, N. G., SCOTT, J. G., (2010). The horseshoe estimator for sparse signals. Biometrika, 97, pp. 465–480. CHEN, S., LAHIRI, P., (2003).A comparison of different MPSE estimators of EBLUP for the Fay-Herriott model. In Proceedings of the Section on Survey Research Methods. Washington, D.C. American Statistical Association, pp. 903–911. DAS, K., JIANG, J., RAO, J. N. K., ((2004). Mean squared error of empirical predictor. Annals of Statistics, 32, pp. 818–840. DATTA, G. S., GHOSH, M., (1991). Bayesian prediction in linear models: applications to small area estimation. The Annals of Statistics, 19, pp. 1748–1770. DATTA, G., GHOSH, M., NANGIA, N., and NATARAJAN, K., (1996). Estimation of median income of four-person families: a Bayesian approach. In Bayesian Statistics and Econometrics: Essays in Honor of Arnold Zellner. Eds. D. Berry, K. Chaloner and J. Geweke. North Holland, pp. 129–140. DATTA, G. S., LAHIRI. P., (2000). A unified measure of uncertainty of estimated best linear unbiased predictors in small area estimation problems. Statistica Sinica,10, pp. 613–627. DATTA, G. S., RAO, J. N. K., and SMITH, D. D., (2005). On measuring the variability of small area estimators under a basic area level model. Biometrika, 92, pp. 183–196. DATTA, G. S., GHOSH, M., STEORTS, R., and MAPLES, J. J., (2011). Bayesian benchmarking with applications to small area estimation. TEST, 20, pp. 574–588. DATTA, G. S., HALL, P., and MANDAL, A., (2011). Model selection and testing for the presence of small area effects and application to area level data. Journal of the American Statistical Association, 106, pp. 362–374. DATTA, G. S., MANDAL, A., (2015). Small area estimation with uncertain random effects. Journal of the American Statistical Association, 110, pp. 1735–1744. STATISTICS IN TRANSITION new series, Special Issue, August 2020 19 ELBERS, C., LANJOUW, J. O., and LANJOUW, P., (2003). Micro-level estimation of poverty and inequality. Econometrica, 71. pp. 355–364. ERCIULESCU, A. L., FRANCO, C., and LAHIRI, P., (2020). Use of administrative records in small area estimation. To appear in Administrative Records for Survey Methodology. Eds. P. Chun and M. Larson. Wiley, New York. FAY, R. E., (1987). Application of multivariate regression to small domain estimation. In Small Area Statistics. Eds. R. Platek, J.N.K. Rao, C-E Sarndal and M.P. Singh. Wiley New York, pp. 91–102. FAY, R. E., HERRIOT, R. A., (1979). Estimates of income for small places: an application of James-Stein procedure to census data. Journal of the American Statistical Association, 74, pp. 269–277. GHOSH, M., (1992). Constrained Bayes estimation with applications. Journal of the American Statistical Association, 87, pp. 533–540. GHOSH, M., RAO, J. N. K., (1994). Small area estimation: an appraisal. Statistical Science, pp. 55–93. GHOSH, M., NATARAJAN, K., STROUD, T. M. F., and CARLIN, B. P., (1998). Generalized linear models for small area estimation. Journal of the American Statistical Association, 93, pp. 273–282. GHOSH, M., STEORTS, R., (2013). Two-stage Bayesian benchmarking as applied to small area estimation. TEST, 22, pp. 670–687. GHOSH, M., KUBOKAWA, T., and KAWAKUBO, Y., (2015). Benchmarked empirical Bayes methods in multiplicative area-level models with risk evaluation. Biomerika, 102, pp. 647–659. GHOSH, M., MYUNG, J., and MOURA, F. A. S., (2018). Robust Bayesian small area estimation. Survey Methodology, 44, pp. 101–115. GONZALEZ, M. E., HOZA, C., (1978). Small area estimation with application to unemployment and housing estimates. Journal of the American Statistical Association, 73, pp. 7–15. GRIFFIN, J. E., BROWN, P. J., (2010). Inference with normal-gamma prior distributions in regression problems. Bayesian Analysis, 5, pp. 171–188. HANSEN, M. H., HURWITZ, W. N., and MADOW, W. G., (1953). Sample Survey Methods and Theory. Wiley, New York. 20 M. Ghosh: Small area estimation: its evolution... HOLT, D., SMITH, T. M. F., and TOMBERLIN, T. J., (1979). A model-based approach for small subgroups of a population. Journal of the American Statistical Association, 74, pp. 405–410. JIANG, J., LAHIRI, P., (2001). Empirical best prediction of small area inference with binary data. Annals of the Institute of Statistical Mathematics, 53, pp. 217–243. JIANG, J., LAHIRI, P., and WAN, S-M., (2002). A unified jackknife theory. The Annals of Statistics, 30, pp. 1782–1810. JIANG, J., LAHIRI, P., (2006). Mixed model prediction and small area estimation (with discussion). TEST, 15, pp. 1–96. JIANG, J., NGUYEN, T., and RAO, J. S., (2011). Best predictive small area estimation. Journal of the American Statistical Association, 106, pp. 732–745. KACKAR, R. N., HARVILLE, D. A., (1984). Approximations for standard errors of estimators of fixed and random effects in mixed linear models. Journal of the American Statistical Association, 79, pp. 853–862. LAHIRI, P., RAO, J. N. K., (1995). Robust estimation of mean squared error of small area estimators. Journal of the American Statistical Association, 90, pp. 758–766. LAHIRI, P., PRAMANIK, S., (2019). Evaluation of synthetic small area estimators using design-based methods. Austrian Journal of Statistics, 48, pp. 43–57. LOUIS, T. A., (1984). Estimating a population of parameter values using Bayes and empirical Bayes methods. Journal of the American Statistical Association, 79, pp. 393–398. MALEC, D., DAVIS, W. W., and CAO, X., (1999). Model-based small area estimates of overweight prevalence using sample selection adjustment. Statistics and Medicine, 18, pp. 3189–3200. MOLINA, I., RAO, J. N. K., (2010). Small area estimation of poverty indicators. Canadian Journal of Statistics, 38, pp. 369–385. MOLINA, I., RAO, J. N. K., and DATTA, G. S., (2015). Small area estimation under a Fay-Herriot model with preliminary testing for the presence of random effects. Survey Methodology. PFEFFERMANN, D., TILLER, R. B., (2005). Bootstrap approximation of prediction MSE for state-space models with estimated parameters. Journal of Time Series Analysis, 26, pp. 893–916. STATISTICS IN TRANSITION new series, Special Issue, August 2020 21 MORRIS, C. N., (1983). Parametric empirical Bayes inference: theory and applications. Journal of the American Statistical Association, 78, pp. 47–55. POLSON, N. G., SCOTT, J. G., (2010). Shrink globally, act locally: Sparse Bayesian regularization and prediction. Bayesian Statistics, 9, pp. 501–538. PFEFFERMANN, D., (2002). Small area estimation: new developments and direction. International Statistical Review, 70, pp. 125–143. PFEFFERMANN, D., (2013). New important developments in small area estimation. Statistical Science, 28, pp. 40–68. PRASAD, N. G. N., RAO, J. N. K., (1990). The estimation of mean squared error of small area estimators. Journal of the American Statistical Association, 85, pp. 163–171. RAGHUNATHAN, T. E., (1993). A quasi-empirical Bayes method for small area estimation. Journal of the American Statistical Association, 88, pp. 1444–1448. RAO, J. N. K., (2003). Some new developments in small area estimation. Journal of the Iranian Statistical Society, 2, pp. 145–169. RAO, J. N. K., (2006). Inferential issues in small area estimation: some new developments. Statistics in Transition, 7, pp. 523–526. RAO, J. N. K., Molina, I., (2015). Small Area Estimation, 2nd Edition. Wiley, New Jersey. SCOTT, J. G., BERGER, J. O., (2010). Bayes and empirical-Bayes multiplicity adjustment in the variable-selection problem. The Annals of Statistics, 38, pp. 2587–2619. SLUD, E. V., MAITI, T., (2006). Mean squared error estimation in transformed FayHerriot models. Journal of the Royal Statistical Society, B, 68, pp. 239–257. TANG, X., GHOSH, M., Ha, N-S., and SEDRANSK, J., (2018). Modeling random effects using global-local shrinkage priors in small area estimation. Journal of the American Statistical Association, 113, pp. 1476–1489. WANG, J., FULLER, W. A., and QU, Y., (2008). Small area estimation under restriction. Survey Methodology, 34, pp. 29–36. 22 M. Ghosh: Small area estimation: its evolution... YOSHIMORI, M., LAHIRI, P., (2014). A new adjusted maximum likelihood method for the Fay-Herriott small area model. Journal of Multivariate Analysis, 124, pp. 281–294. YOU, Y., RAO, J. N. K., and HIDIROGLOU, M. A., (2013). On the performance of self-benchmarked small area estimators under the Fay-Herriott area level model. Survey Methodology, 39, pp. 217–229.