Panel stochastic frontier analysis with dependent error terms
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
El Mehdi, Rachida; Hafner, Christian M. Article Panel stochastic frontier analysis with dependent error terms International Econometric Review (IER) Provided in Cooperation with: Econometric Research Association (ERA), Ankara Suggested Citation: El Mehdi, Rachida; Hafner, Christian M. (2021) : Panel stochastic frontier analysis with dependent error terms, International Econometric Review (IER), ISSN 1308-8815, Econometric Research Association (ERA), Ankara, Vol. 13, Iss. 2, pp. 24-40, https://doi.org/10.33818/ier.1033722 This Version is available at: https://hdl.handle.net/10419/290160 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc-nd/4.0/
International Econometric Review (IER) 24 Panel Stochastic Frontier Analysis with Dependent Error Terms Rachida El Mehdi and Christian M. Hafner Receiving Time: May 21, 2021 Acceptance Time: October 21, 2021 ABSTRACT In presence of panel data, technical efficiency is used to compare the performances of Decision-Making Units (DMUs). The novelty of this paper is the consideration of the dependence between the two error terms in the case of panel data and the introduction of time effect models in the Stochastic Frontier Analysis (SFA). Hence, our SFA model considers the balanced panel case, several models describing the evolution of the inefficiency over time and the dependence between the two error terms. The inefficiency and noise terms being dependent, a copula function which reflects the dependence between them is included in their joint density. The model is estimated by maximum likelihood and the Akaike Information Criterion (AIC) is used for model selection. Moreover, a likelihood ratio test is performed for the nested models. A bootstrap algorithm is proposed for statistical inference on the Technical Efficiency (TE) measures. Results for Moroccan policy of the production and sales of drinking water from 2001 to 2007 identify the most and least efficient provinces, and a generally positive trend of estimated TE measures. Key words: Bootstrap, Copulas, Efficiency, Panel data, Stochastic frontier analysis JEL Classifications : C13, C18, D24 1. INTRODUCTION When panel data are available, it is recommended to use the structure of the data to estimate technical efficiencies in a Stochastic Frontier Analysis (SFA) because a panel contains more information than a single cross section. Furthermore, as noted in Schmidt and Sickles (1984) and Kumbhakar and Lovell (2000) some strong distributional assumptions used in the crosssectional data case can be relaxed with the panel data and the technical efficiency can be estimated consistently when 𝑇, the number of time observations for each Decision-Making Unit (𝐷𝑀𝑈), is large. Hence, repeated observations can be considered as a substitute for some strong distributional assumptions. They can also constitute a weakening of the independence assumption between the technical inefficiency term and the regressors. An overview of the research on panel SFA models reveals that Jondrow et al. (1982) generalized the cross-sectional model to the panel data model and used the conditional Rachida El Mehdi, SmartICT Lab, National School of Applied Sciences, Mohammed First University, (email: [email protected]), Tel: +212666146503 Christian M. Hafner, Louvain Institute of Data Analysis and Modelling in Economics and Statistics, and ISBA, Universit´e catholique de Louvain, (email: [email protected])
International Econometric Review (IER) 25 expectation of the inefficiency term 𝑢 given the realized value of the error 𝜖 to estimate efficiency. Schmidt and Sickles (1984) and Kumbhakar and Lovell (2000) have proposed for the balanced panel case, models where they supposed that technical efficiency varies across producers but is either constant or varies through time for each producer. Battese et.al. (2000) adopted an unbalanced panel to investigate efficiency of labour in the Swedish banking industry. Kim and Lee (2006) assumed a time varying pattern of technical efficiency movements to analyze the productivity growth of several East Asian countries over a period of twenty years. However, all studies handled the panel SFA model with independence between noise and inefficiency terms. Recently, Smith (2008) handled the panel data model with dependent error components using a simulated example but without making inference on the estimated efficiency. Allowing for dependence is generally a desirable feature, because the efficiency of a given 𝐷𝑀𝑈 at a given period of time might depend on whether or not the 𝐷𝑀𝑈 was ‘lucky’, expressed by the random noise term. For example, if the 𝐷𝑀𝑈 was unlucky in a particular period, it might attempt to compensate this by an increased efficiency, generating a dependence between noise and efficiency. Furthermore, several studies have assessed the performance of water services such as Faria et al. (2005) and Tupper and Resende (2004), which compare the technical efficiency of Brazilian public and private companies in water supply; Sampaio et.al. (2005) which deals with the cost efficiency of the public water service in Portugal, and Vishwakarma and Kulshrestha (2010) which analyses the water supply utility of urban cities in India using the stochastic production frontier analysis. However, most of these studies use cross-sectional data. In the absence of such study in the water domain based on Moroccan data, the performance of the entities responsible of the water management in all regions will be measured by estimating the efficiencies in case of panel data and a nonparametric confidence interval will be proposed. This research is an extension of El Mehdi and Hafner (2014b) to the panel data case. Our objective is to deal with the dependence of the error terms in the panel SFA approach using models for the time variation of efficiencies. We evaluate the efficiency and compare the 𝐷𝑀𝑈 performances through an empirical data set on the water management in Morocco. Thus, in this work the production frontiers and panel data are considered to estimate technical efficiency when the two components of the error term are dependent. Efficiency being estimated, statistical inference is needed to draw reliable conclusions. Hence, this work presents also an associated procedure to build confidence intervals on the efficiencies in this considered case. The remainder of the paper is organized as follows: Section 2 describes the model with the copula function, Section 3 presents the procedure of statistical inferences on the Technical Efficiency (TE) measure in the case of panel data with dependent error terms, and the last section presents results of an empirical analysis of the water area in all Moroccan regions with the numerical procedure estimation of technical efficiency in order to compare the 𝐷𝑀𝑈𝑠. Finally we conclude by a summary of the results with some remarks and open issues.
International Econometric Review (IER) 26 2. EFFICIENCY MEASURES FOR PANEL DATA The principle of the efficiency measure estimation for panel data is the same as that for crosssectional data. We need however to make additional assumptions about the temporal pattern of inefficiency. There are also differences between the two procedures in terms of the simulated likelihood function definition. When information about all is 𝐷𝑀𝑈𝑠 available at 𝑇 different time periods, it is preferable to use a stochastic frontier model which is adequate for panel data. In frontier analysis this model is more appropriate because even if it is not fundamentally different from the cross-sectional model, it has several advantages as it increases the degree of freedom to estimate parameters, provides consistent efficiency estimates when 𝑇 is increasing and does not require that the inefficiencies are independent of the regressors. 2.1. The panel data production frontier model The panel stochastic frontier model, when the inefficiencies are assumed to vary systematically with time, is specified as follows: 𝑦𝑖𝑡 =𝑓(𝑥𝑖𝑡 ,𝛽)+𝜖𝑖𝑡 =𝑓(𝑥𝑖𝑡 ,𝛽)+𝑣𝑖𝑡 −𝑢𝑖𝑡 , 𝑖 =1,2,…,𝑛 ; 𝑡=1,2,…,𝑇 (2.1) where 𝑦𝑖𝑡 =𝑙𝑜𝑔(𝑌𝑖𝑡); 𝑌𝑖𝑡 : the observed output for 𝐷𝑀𝑈𝑖 at the tth time period (one output), so 𝑌𝑖𝑡 ∈ℝ+; 𝑥𝑖𝑡 =𝑙𝑜𝑔(𝑋𝑖𝑡); 𝑋𝑖𝑡 : a vector of length 𝑝 that describes the observed inputs for observation 𝑖 at time 𝑡, so 𝑋𝑖𝑡 ∈ℝ+ 𝑝 where 𝑝 is the number of the inputs; 𝛽 : a vector of unknown parameters to be estimated, 𝛽 ∈ ℝ𝑙+(1×𝑇) where 𝑙 is the number of parameters excluding the time-varying intercepts. If intercepts are constant over time, then 𝛽∈ ℝ𝑙+1. Moreover, 𝜖𝑖𝑡 is the error term for observation 𝑖 at time 𝑡, 𝑓(𝑥𝑖𝑡 ,𝛽) the production frontier, 𝑛 is the number of DMUs under study and 𝑇 is the number of periods or the number of observations for each DMU. The two components of the error term are motivated by the idea that deviations from the frontier might not be entirely under the control of the DMU and that the performance of a DMU is affected by these two components. Hence, the term 𝜖𝑖𝑡 is divided into two parts, the inefficiency term 𝑢𝑖𝑡 which is constrained to be non-negative (𝑢𝑖𝑡 ≥0) and the statistical noise term 𝑣𝑖𝑡 which is usually a normal with zero as mean and 𝜎𝑉 as standard deviation (𝑣𝑖𝑡 ∼ N(0 , 𝜎𝑉 2) ). Furthermore, distributional assumptions will be imposed on both terms 𝑢𝑖𝑡 and 𝑣𝑖𝑡. In particular, it is assumed that the components of the first are independently and identically positively distributed and the components of the second are independently and identically normally distributed. In addition, both terms are assumed continuous and independent of 𝑥𝑖𝑡. At first, it is supposed that the two terms are mutually independent and the model is estimated by the Maximum Likelihood (ML) method. At a second stage, they will be allowed to be dependent and the ML estimates of the first stage will be considered as initial values in the numerical optimization.
International Econometric Review (IER) 27 As proposed in the literature on panel SFA, the frontier model considers either the time-constant or the time-varying efficiency (see Schmidt and Sickles (1984) and Kumbhakar and Lovell (2000)). In our study we consider the frontier model with time-varying efficiency which is, in our opinion, more realistic and reflects the inefficiency variability over time. Nevertheless, we shall limit our analysis to the fixed intercept over time and to a comparison between some timevarying models in order to select one of them. 2.2. Time-varying efficiency When 𝑇 is large, the assumption of a time-constant inefficiency is typically not appealing, as one would expect that inefficient DMUs are forced to improve over time. So, a time varying inefficiency is needed and a random-effects model should be used. To define a random-effects model, one has developed an extension of the fixed-effects model to a more general model to get consistent estimators of 𝑢𝑖 when 𝑇 is large. Among these we refer to Jondrow et al. (1982), which derived panel generalizations of the conditional inefficiency predictors, Battese and Coelli (1988) where the term 𝑢𝑖 has a more general truncated-normal distribution, and Battese et.al. (1989) which extend the model to allow unbalanced data. The frontier model is called a random-effects model when it is described by 𝑦𝑖𝑡 =𝛽0𝑡 + ∑𝛽𝑗 𝑙 𝑗=1 𝑥𝑖𝑗𝑡 +𝑣𝑖𝑡 −𝑢𝑖𝑡 = 𝛽𝑖𝑡 + ∑𝛽𝑗 𝑙 𝑗=1 𝑥𝑖𝑗𝑡 +𝑣𝑖𝑡 (2.2) where 𝛽𝑖𝑡 = 𝛽0𝑡 −𝑢𝑖𝑡 is the intercept for DMU𝑖 in the time period 𝑡 and where 𝛽0𝑡 is the intercept common to all DMUs in the time period 𝑡, see e.g. Kumbhakar (1990) and Kumbhakar and Lovel (2000). Of course, 𝑛×𝑇 parameters 𝛽𝑖𝑡 should be estimated but Cornwell et.al. (1990) reduce this number to 3×𝑛. In the same way, Kumbhakar (1990) suggested a model in which the 𝑢𝑖𝑡 are specified by the expression (2.5). He suggested estimating the model with the maximum likelihood method but does not provide an empirical application. Battese and Coelli (1992) suggested a time-varying model for unbalanced panel data with the exponential function of time for 𝑢𝑖𝑡 in (2.3) bellow. They also proposed in their later work, Battese and Coelli (1995), a model where 𝑢𝑖𝑡 follows a normal distribution truncated at zero. The ML estimation and the efficiency calculations of these cases have been included in the FRONTIER programs implemented by Coelli (1996). Schmidt and Sickles (1984) suggested not to specify an implicit distribution for the inefficiency when the panel data are available and to estimate the fixedeffects model with the traditional panel data methods. In extension of this approach, Cornwell et.al. (1990) and Lee and Schmidt (1993) have developed an approach in which they introduce the variation of the effect of inefficiencies over time. Both of the latter approaches propose a variation of inefficiencies more flexible than proposed in (2.3) and (2.5). In this research, the Kumbhakar (1990) and the Cornwell et.al. (1990) random-effects models will not be considered given the large number of parameters to be estimated, and so just one fixed intercept will be estimated. The Kumbhakar (1990) time effect expressed here by the
International Econometric Review (IER) 28 formula (2.5), Battese and Coelli (1992) and a variety of time effects models are considered with a function 𝜂(𝑡)≥0 that describes the evolution of inefficiency over time such that 𝑢𝑖𝑡 = 𝜂(𝑡)𝑢𝑖 . These models are denoted as 𝑀1 : 𝑢𝑖𝑡 = [𝑒𝑥𝑝{−𝜂1(𝑡−𝑇)}] 𝑢𝑖 , 𝑢𝑖 ~ 𝑁+(0 ,𝜎𝑈 2) ; (2.3) 𝑀2 : 𝑢𝑖𝑡 = [𝑒𝑥𝑝{−𝜂1(𝑡−𝑇)}] 𝑢𝑖 ; (2.4) 𝑀3 : 𝑢𝑖𝑡 = [1+𝑒𝑥𝑝{𝜂1𝑡+ 𝜂2𝑡2}]−1 𝑢𝑖 ; (2.5) 𝑀4: 𝑢𝑖𝑡 = [1+𝜂1(𝑡−𝑇)𝑠𝑖𝑛(𝑡−𝑇)] 𝑢𝑖 ; (2.6) 𝑀5: 𝑢𝑖𝑡 = [1+𝜂1𝑠𝑖𝑛(𝜂2𝑡)] 𝑢𝑖 ; (2.7) 𝑀6: 𝑢𝑖𝑡 = [1+𝜂1𝑠𝑖𝑛(𝜂2(𝑡−𝑇))] 𝑢𝑖 ; (2.8) 𝑀7: 𝑢𝑖𝑡 = [1+𝜂1(𝑡−𝑇)𝑠𝑖𝑛(𝜂2(𝑡−𝑇))] 𝑢𝑖 ; (2.9) 𝑀8: 𝑢𝑖𝑡 = [𝜂0+𝜂1𝑡+1 2𝜂2𝑡2+2∑(𝑎ℎ𝑠𝑖𝑛(ℎ𝑡)−𝑏ℎ𝑐𝑜𝑠(ℎ𝑡)) 𝐻 ℎ=1 ] 𝑢𝑖; (2.10) where, for all models and except for M1, 𝑢𝑖 ~ N+(μ ,σU 2) and 𝜇 is the mean of the original normal distribution. That indicates that the inefficiency term 𝑢𝑖 is a normal truncated at zero with mean 𝜇. Furthermore, M1 and M2 are the Battese and Coelli (1992, 1995) models and M3 is the Kumbhakar (1990) model. The proposed M4 − M7 models include sinusoidal functions to allow for possible periodicity effects in the inefficiency. For example, M7 models a timevarying amplitude of the sine function depending on the parameter 𝜂1. The last considered model M8 is the Fourier Flexible Form of Gallant (1984) which can closely approximate any smooth function 𝜂(𝑡) for sufficiently large H. In our study the Akaike Information Criterion (AIC) is used to select a model among M1 to M8. Moreover, a likelihood ratio test is performed for the nested models. 2.3. Model estimation Considering the models described by (2.2) and (2.3)-(2.10), the methods used to estimate all models depend on the distributional assumptions. When 𝑣𝑖𝑡 is i.i.d. normal, 𝑢𝑖𝑡 is i.i.d. with positive support, 𝑣𝑖𝑡 and 𝑢𝑖𝑡 are mutually independent and independent from the regressors, the Maximum Likelihood Estimation (MLE) is feasible. Schmidt and Sickles (1984) conjectures that given suitable regularity conditions the ML estimates of (2.2) are consistent and asymptotically efficient as 𝑛 → ∞ regardless of 𝑇. In the particular, when 𝑣𝑖𝑡∼𝑖𝑖𝑑 𝑁(0 , 𝜎𝑉 2) and 𝑢𝑖𝑡 ~ 𝑖𝑖𝑑𝑁+(0 ,𝜎𝑈 2), the MLE leads for 𝜖𝑖= (𝜖𝑖1,…,𝜖𝑖𝑡,…,𝜖𝑖𝑇)′, to the log-likelihood function, ignoring an additive constant, 𝑙 =ln(𝐿)= −𝑛 2𝑙𝑛 𝜎∗2−1 2∑𝑎∗𝑖 𝑛 𝑖=1 −𝑛.𝑇 2ln𝜎𝑉 2−𝑛 2𝑙𝑛 𝜎𝑈 2 +∑𝑙𝑛[1−Φ(−𝜇∗𝑖 𝜎∗)] 𝑛 𝑖=1 , (2.11) which leads to the technical efficiency (TE) estimate for all 𝑖 =1,…,𝑛 𝑇𝐸𝑖𝑡 =𝐸(𝑒𝑥𝑝{−𝑢𝑖𝑡}|𝜖𝑖) , = 1−Φ(𝜂(𝑡)𝜎∗− 𝜇∗𝑖 𝜎∗) 1−Φ(− 𝜇∗𝑖 𝜎∗)𝑒𝑥𝑝{−𝜂(𝑡)𝜇∗𝑖 +1 2𝜂2(𝑡)𝜎∗2}, (2.12)
International Econometric Review (IER) 29 where 𝐿 is the likelihood function, 𝜂(𝑡) is the time function, 𝜎∗2= 𝜎𝑉 2𝜎𝑈 2 𝜎𝑉 2+𝜎𝑈 2∑𝜂2(𝑡) 𝑡, 𝜇∗𝑖 = (∑ 𝜂(𝑡)𝜖𝑖𝑡 𝑡)𝜎𝑉 2 𝜎𝑉 2+𝜎𝑈 2∑𝜂2(𝑡) 𝑡, 𝑎∗𝑖 =1 𝜎𝑉 2[∑𝜖𝑖𝑡 2 𝑡−𝜎𝑈 2(∑ 𝜂(𝑡)𝜖𝑖𝑡 𝑡)2 𝜎𝑉 2+𝜎𝑈 2∑𝜂2(𝑡) 𝑡] and Φ is the standard normal cdf. In comparison with the cross-section data, calculation of the log-likelihood function for panel data is similar to El Mehdi and Hafner (2014b), but it changes at the level of computing the 𝜖𝑖 density which is 𝑔(𝜖𝑖) where 𝜖𝑖= (𝜖𝑖1,…,𝜖𝑖𝑡,…,𝜖𝑖𝑇)′ . In the expression of 𝑔(𝜖𝑖), the joint density 𝑓(𝑢𝑖 ,𝑣𝑖 ) of 𝑢𝑖 and 𝑣𝑖 is replaced by 𝑓1(𝑢𝑖) 𝑓2(𝑣𝑖)=𝑓1(𝑢𝑖)∏ 𝑓2(𝜖𝑖𝑡 + 𝜂(𝑡)𝑢𝑖) 𝑡 . Of course, this last expression is integrated by 𝑢𝑖 to get 𝑔(𝜖𝑖). When 𝑢𝑖 and 𝑣𝑖 are dependent, their joint density when panel data is available becomes 𝑓1(𝑢𝑖)𝑓2( 𝑣𝑖) 𝑐𝜃(𝐹1(𝑢𝑖),𝐹2(𝑣𝑖))= 𝑓1(𝑢𝑖)∏ 𝑓2(𝜖𝑖𝑡 + 𝜂(𝑡)𝑢𝑖) 𝑡∏ 𝑐𝜃(𝐹1(𝑢𝑖),𝐹2(𝜖𝑖𝑡 + 𝜂(𝑡)𝑢𝑖)) 𝑡 (2.13) where 𝑐 is a bivariate copula density which expresses the dependence between the two variables 𝑢𝑖 and 𝑣𝑖, and 𝐹1(𝑢𝑖) and 𝐹2(𝑣𝑖) are two uniform variables which are the cdf of 𝑓1(𝑢𝑖) and 𝑓2(𝑣𝑖) respectively and called the margins. The independence case is a special case of this model when the copula is the product copula, for which 𝑐(.,.)=1. But for general copula functions, the ML estimation will become more complicated. Given that the 𝑣𝑖𝑡 are supposed independent and identically distributed, the density of 𝜖𝑖 becomes 𝑔(𝜖𝑖)=∫𝑓(𝜖𝑖,𝑢𝑖) 𝑑𝑢𝑖 +∞ 0=∫𝑓1(𝑢𝑖) ∏𝐴𝑖𝑡 𝑡 𝑑𝑢𝑖 +∞ 0=𝐸(∏ 𝐴𝑖𝑡 𝑡) , (2.14) where Ait = f2(𝜖it +η(t)ui) cθ(F1(ui),F2(ϵit + η(t)ui)) . See the Appendix A for more details. Therefore, assuming the independence across DMUs, the log-likelihood function can be written as 𝑙(𝜗)=log𝐿(𝜗)= log𝐿(𝜎𝑈,𝜎𝑉,𝜃,𝛽0,𝛽,𝜂𝑘) = ∑𝑙𝑜𝑔 𝑛 𝑖=1 𝑔𝜗(𝜖𝑖)=∑𝑙𝑜𝑔 𝑛 𝑖=1 𝑔𝜗(𝑦𝑖−(𝛽0+∑𝛽𝑗𝑥𝑖𝑗 𝑙 𝑗=1 )), (2.15) where 𝛽0= (𝛽01,…,𝛽0𝑡,…,𝛽0𝑇)′, 𝛽 = (𝛽1,…,𝛽𝑗,…,𝛽𝑙)′ are vectors with a length equal respectively to the time periods 𝑇 and the number of inputs 𝑙, 𝜂𝑘 is a vector of 𝑘 parameters in the time-varying function and where 𝑥𝑖𝑗 = (𝑥𝑖𝑗1,…,𝑥𝑖𝑗𝑡,…,𝑥𝑖𝑗𝑇)′ and 𝑦𝑖= (𝑦𝑖1,…,𝑦𝑖𝑡,…,𝑦𝑖𝑇)′. For simplicity, all intercepts 𝛽0𝑡, 𝑡=1,…,𝑇 are considered the same and denoted by 𝛽0 in the empirical analysis. Generally, the expression of the function 𝑙(𝜗) is complex in the dependence case and to obtain analytical derivatives becomes a tedious or even impossible task in several cases. So, the loglikelihood is optimized numerically using the mle function in the R software and using the simplex numerical method called the Nelder-Mead method.
International Econometric Review (IER) 30 Once the parameters (𝜎𝑈,𝜎𝑉,𝜃,𝛽0,𝛽,𝜂𝑘) are estimated, the technical efficiency can be estimated using the expected value of (𝑒𝑥𝑝{−𝑢𝑖𝑡}|𝜖𝑖) as 𝑇𝐸𝑖𝑡 =𝐸[𝑒𝑥𝑝{−𝑢𝑖𝑡}|𝜖𝑖] = 𝐸(𝑒𝑥𝑝{−𝜂(𝑡)𝑢𝑖}∏ 𝐴𝑖𝑡 𝑡) / 𝐸(∏ 𝐴𝑖𝑡 𝑡) . (2.16) See the Appendix A for more details. Given again the complexity of the 𝑇𝐸𝑖𝑡 expression, the expectation will be estimated for a large number 𝑚 of Monte Carlo draws by 𝑇𝐸 𝑖𝑡 ≅ [1 𝑚∑(𝑒𝑥𝑝{−𝜂(𝑡)𝑢𝑗}∏𝐴𝑖𝑗𝑡 𝑡 ) 𝑚 𝑗=1 ]/[1 𝑚∑(∏𝐴𝑖𝑗𝑡 𝑡) 𝑚 𝑗=1 ]. (2.17) The following section develops an algorithm based on the bootstrap to construct confidence intervals for the estimated technical efficiencies. 3. INFERENCE FOR THE TECHNICAL EFFICIENCY MEASURE Since the technical efficiencies of each DMU at each time 𝑡 are unknown and estimated by 𝑇𝐸 𝑖𝑡, an inference about them is required. To build the confidence interval at a level 𝛼, given that the true sampling distribution is not available, we see that a modified algorithm of the parametric bootstrap Algorithm#3 of Simar and Wilson (2010) adapted to the dependence case and to the panel framework is more appropriate. Hence, we developed a procedure to estimate the associated confidence bounds when 𝑣 is normal and 𝑢 is half-normal which can be generalized for any positive distribution of 𝑢 such as the truncated-normal. The method is easy to apply but it is quite computationally intensive. The steps are the following: 1. Estimate 𝜗=(𝜎𝑈,𝜎𝑉,𝜃,𝛽0,𝛽,𝜂𝑘) according to (2.15), using the observed (𝑥𝑖𝑡 ,𝑦𝑖𝑡), 𝑖= 1,2,…,𝑛 and 𝑡=1,2,…,𝑇 and using a numerical optimization procedure to get 𝜗 = (𝜎𝑈,𝜎𝑉,𝜃 ,𝛽 0,𝛽 ,𝜂𝑘) and to compute the point estimates 𝑇𝐸 as described before 2. For 𝑖=1,2,…,𝑛, draw 𝑢𝑖∗ ~ 𝑁+(0 ,𝜎𝑈 2) and 𝑣𝑖𝑡 ∗ ~ 𝑁(0 ,𝜎𝑉 2), 𝑡=1,2,…,𝑇 such that 𝑢𝑖∗ and 𝑣𝑖𝑡 ∗ are dependent with dependence characterized by the Clayton copula. Then compute 𝑦𝑖𝑡 ∗= 𝛽 0+ ∑𝛽 𝑗𝑥𝑖𝑗𝑡 𝑙 𝑗=1 +𝑣𝑖𝑡 ∗− 𝜂(𝑡)𝑢𝑖∗. There are several procedures to generate the pair (𝑢𝑖 ∗,𝑣𝑖𝑡 ∗) according to the Clayton copula, we use the one described in Nelsen (1999), page 41. The four steps of this procedure are a. Draw 𝑇+1 independent uniform random variables 𝑤1𝑖, ℎ2𝑖1 ,…,ℎ2𝑖𝑡,…,ℎ2𝑖𝑇, such that 𝑤1𝑖~ 𝑈(0,1) and ℎ2𝑖𝑡 ~ 𝑈(0,1) for 𝑡=1,2,…,𝑇. b. Set 𝑤2𝑖𝑡 = [𝑤1𝑖 −𝜃 (ℎ2𝑖𝑡 −𝜃 /(1+𝜃 )−1)+1]−1/𝜃 for all 𝑡=1,2,…,𝑇. c. Set 𝑢𝑖∗= 𝐹1−1(𝑤1𝑖) and 𝑣𝑖𝑡 ∗= 𝐹2 −1(𝑤2𝑖𝑡) for all 𝑡=1,2,…,𝑇 and where 𝐹1 and 𝐹2 are the cdf of the 𝑁+(0 ,𝜎𝑈 2) and 𝑁(0 ,𝜎𝑉 2) respectively.
International Econometric Review (IER) 31 d. Repeat steps a to c 𝑛 times to generate 𝑛×𝑇 pairs (𝑢𝑖 ∗,𝑣𝑖𝑡 ∗). 3. Using the pseudo-data 𝒫𝑏,𝑛 ∗={(𝑥𝑖𝑡 ,𝑦𝑖𝑡 ∗)}𝑖=1 𝑛, compute bootstrap estimates 𝜗 𝑏 ∗= 𝑎𝑟𝑔𝑀𝑎𝑥𝜗∈Θ 𝑙(𝜗|𝒫𝑏,𝑛 ∗) after replacing 𝑦𝑖𝑡 by 𝑦𝑖𝑡 ∗ in (2.15) and then compute the bootstrap estimates 𝑇𝐸 𝑏 ∗ using (A.4) after replacing 𝜖 by 𝜖𝑏 ∗=𝑦−𝛽 0 ∗−𝛽 ∗.𝑥, where 𝑥𝑖𝑡 and 𝑦𝑖𝑡 represent the observed data. 4. Repeat steps 2 and 3, 𝐵 times to obtain estimates ℬ∗= {𝜗 𝑏 ∗}𝑏=1 𝐵. Therefore, use ℬ∗ to get 𝜉∗= {𝑇𝐸 𝑏 ∗}𝑏=1 𝐵 . Each individual 𝑖 is described by a sub-matrix of 𝜉∗ denoted 𝜉𝑖∗, it has 𝑇 rows and 𝐵 columns. For each individual 𝑖 at time period 𝑡 (row 𝑡 of the 𝜉𝑖∗ matrix, denoted 𝜉𝑖𝑡 ∗, 𝑖 =1,2,…,𝑛, compute the (𝛼 2 ⁄ ) and the (1−𝛼 2 ⁄ ) quantiles for 𝜉𝑖𝑡 ∗ by considering its 𝐵 components. The 100×(1−𝛼) percentile bootstrap confidence interval of the statistic of interest 𝑇𝐸 is obtained by the probability 𝑃((𝜉𝑖𝑡 ∗)𝛼 2 ⁄< 𝑇𝐸𝑖𝑡 < (𝜉𝑖𝑡 ∗)1−𝛼 2 ⁄)=1−𝛼. Hence using the 100×(𝛼 2) and 100×(1−𝛼 2) percentiles, we define the lower and the upper bounds of the confidence interval as 𝑇𝐸𝑖𝑡 ∈ [(𝜉𝑖𝑡 ∗)𝛼 2 ⁄ ,(𝜉𝑖𝑡 ∗)1−𝛼 2 ⁄]. We note that the estimation procedure presented in Section 2.3 leads sometimes to a positive skewness of the composite error term which consequently leads to biased parameter estimates and to biased technical efficiencies estimates because all of these latter will be close to one. If this is the case, the procedure presented in this section allows us to overcome this problem. To perform our procedure, a simulation example is proposed. The model describing data is supposed to be log-linear where there are one input and one output such that for all 𝑖= 1,2,…,𝑛 and for all 𝑡=1,2,…,𝑇, we have 𝑙𝑜𝑔(𝑌𝑖𝑡)=𝛽0+𝛽1𝑙𝑜𝑔(10(1+𝑋𝑖𝑡)) where 𝑋𝑖𝑡 ~ 𝑈(0,1) and parameters will be set to 𝛽0=𝑙𝑜𝑔(10) and 𝛽1=1. As for the noise term and the inefficiency term, they are supposed to be normal as usual for the first such that 𝑣𝑖𝑡∼𝑁(0 , 𝜎𝑉 2) with 𝜎𝑉=0.5 and half-normal for the second such that 𝑢𝑖𝑡 = 𝜂(𝑡)𝑢𝑖 and 𝑢𝑖∼𝑁+(0 , 𝜎𝑈 2) with 𝜎𝑈=1 and the two components are dependent using the Clayton copula with dependence parameter 𝜃=1. The time varying function is supposed to be 𝜂(𝑡)= 𝑒𝑥𝑝{−𝜂1(𝑡−𝑇)} with 𝜂1=−0.1. We suppose that 𝑛=50, 𝑇=10 and 𝐵=500. To compute the true efficiencies, the number of simulations to approximate numerically the integral is set to 𝑚=10000 which is large enough to have a good approximation of the expectation in equation (A.4) evaluated at the true values. The bootstrap procedure shows that all estimated efficiencies are covered by their confidence intervals. The percentage for the true efficiencies is evaluated at 87%, 91.2% and 100% for a significance level of 10%, 5% and 1% respectively, so that the bootstrap coverage ratio is reasonably close to the nominal level given our moderate number of bootstrap replications.
38 International Econometric Review (IER) APPENDIX 1. Density of the error 𝝐𝒊 Density of 𝜖𝑖 in the panel SFA when the dependence between the two error terms is considered and when 𝑢𝑖 ∼ 𝑁+(0 ,𝜎𝑈 2) becomes 𝑔(𝜖𝑖)=∫𝑓(𝜖𝑖,𝑢𝑖) 𝑑𝑢𝑖 +∞ 0 (A.1) = ∫𝑓(𝜖𝑖1,…,𝜖𝑖𝑡,…,𝜖𝑖𝑇,𝑢𝑖) 𝑑𝑢𝑖 +∞ 0 = ∫𝑓1(𝑢𝑖) ∏𝑓2(𝜖𝑖𝑡 +𝜂(𝑡)𝑢𝑖) 𝑡 +∞ 0 . 𝑐𝜃(𝐹1(𝑢𝑖),𝐹2(𝜖𝑖1 + 𝜂(1)𝑢𝑖), …,𝐹2(𝜖𝑖𝑇 + 𝜂(𝑇)𝑢𝑖)) 𝑑𝑢𝑖 = ∫𝑓1(𝑢𝑖) ∏𝑓2(𝜖𝑖𝑡 +𝜂(𝑡)𝑢𝑖) 𝑡 +∞ 0. ∏ 𝑐𝜃(𝐹1(𝑢𝑖),𝐹2(𝜖𝑖𝑡 + 𝜂(𝑡)𝑢𝑖)) 𝑑𝑢𝑖𝑡 = ∫𝑓1(𝑢𝑖) +∞ 0. ∏[𝑓2(𝜖𝑖𝑡 +𝜂(𝑡)𝑢𝑖) .𝑐𝜃(𝐹1(𝑢𝑖),𝐹2(𝜖𝑖𝑡 + 𝜂(𝑡)𝑢𝑖))] 𝑡𝑑𝑢𝑖 = ∫𝑓1(𝑢𝑖) ∏𝐴𝑖𝑡 𝑡 𝑑𝑢𝑖 +∞ 0=𝐸(∏ 𝐴𝑖𝑡 𝑡) (A.2) where Ait = f2(𝜖it +η(t)𝑢i) cθ(F1(𝑢i),F2(𝜖it + η(t)𝑢i)). 2. Technical efficiency for each DMU at time t The associated technical efficiency of 𝐃𝐌𝐔𝐢𝐭 is expressed as 𝑇𝐸𝑖𝑡 =𝐸[𝑒𝑥𝑝{−𝑢𝑖𝑡}|𝜖𝑖] (A.3) =𝐸[𝑒𝑥𝑝{−𝜂(𝑡)𝑢𝑖}|(𝜖𝑖1,…,𝜖𝑖𝑡,…,𝜖𝑖𝑇)] =∫𝑒𝑥𝑝{−𝜂(𝑡)𝑢𝑖} 𝑓1(𝑢𝑖|(𝜖𝑖1,…,𝜖𝑖𝑡,…,𝜖𝑖𝑇)) 𝑑𝑢𝑖 +∞ 0 = ∫𝑒𝑥𝑝{−𝜂(𝑡)𝑢𝑖} 𝑓( 𝑢𝑖 , 𝜖𝑖) 𝑔(𝜖𝑖) 𝑑𝑢𝑖 +∞ 0 = 1 𝑔(𝜖𝑖)∫𝑒𝑥𝑝{−𝜂(𝑡)𝑢𝑖} 𝑓1(𝑢𝑖)∏𝐴𝑖𝑡 𝑡 𝑑𝑢𝑖 +∞ 0 = 𝐸(𝑒𝑥𝑝{−𝜂(𝑡)𝑢𝑖}∏ 𝐴𝑖𝑡 𝑡 ) 𝐸(∏ 𝐴𝑖𝑡 𝑡) (A.4) REFERENCES Battese, G.E. and T.J. Coelli (1988). Prediction of firm level technical efficiencies with a generalized frontier production function and panel data. Journal of Econometrics, 38(3), 387-399. Battese, G.E. and T.J. Coelli (1992). Frontier Production Functions, Technical Efficiency and Panel Data: With Application to Paddy Farmers in India. The Journal of Productivity Analysis, 3(1), 153-169. Battese, G.E. and T.J. Coelli (1995). A Model for Technical Inefficiency Effects in a Stochastic Frontier Production Function for Panel Data. Empirical Economics, 20, 325332. Battese, G.E., A. Heshmati, and L. Hjalmarsson (2000). Efficiency of labour use in Swedish
39 International Econometric Review (IER) banking industry: a stochastic frontier approach. Empirical Economics, 25(4), 623640. Battese, G.E., T.J. Coelli, and T.C. Colby (1989). Estimation of Frontier Production Functions and the Efficiencies of Indian Farms Using Panel Data from ICRISAT’s Village Level Studies. Journal of Quantitative Economics, 5(2), 327-348. Bhat, C. and N. Eluru (2009). A Copula-Based Approach to Accommodate Residential SelfSelection Effects in Travel Behavior Modeling. Transportation Research, Part B, 43(7), 749-765. Coelli, T.J. (1995). Estimators and hypothesis tests for a stochastic frontier function: A Monte Carlo analysis. Journal of Productivity Analysis, 6(3), 247-268. Coelli, T.J. (1996). A Guide to FRONTIER, Version 4.1: A Computer Program for Stochastic Frontier Production and Cost Function Estimation. Centre for efficiency and Productivity Analysis, CEPA Working Paper 96/07, Department of Econometrics, University of New England. Cornwell, C., P. Schmidt, and R.C. Sickles (1990). Production frontiers with cross-sectional and time-series variation in efficiency levels. Journal of Econometrics, 46(1-2), 185200. Daraio, C. and L. Simar (2007). Advanced Robust and Nonparametric Methods in Efficiency Analysis: Methodology and Applications. Springer. De Witte, K. and R.C. Marques (2008). Designing incentives in local public utilities, an international comparison of the drinking water sector. Social Science Research Network SSRN 1084807. Efron, B. (1982). The jackknife, the bootstrap, and other resampling plans. CBMS-NSF Regional Conference Series in Applied Mathematics, #38. Philadelphia: SIAM. El Mehdi, R. and C.M. Hafner (2014a). Local government efficiency: The case of Moroccan municipalities. African Development Review, 26(1), 88-101. El Mehdi, R. and C.M. Hafner (2014b). Inference in stochastic frontier analysis with dependent error terms. Mathematics and Computers in Simulation (MATCOM), 102(C), 104-116. Faria, R.C., G.S. Souza, and T.B. Moreira (2005). Public versus private water utilities: Empirical evidence for brazilian companies. Economics Bulletin, 8(2), 1-7. Gallant, A. Ronald (1984). The fourier flexible form. American Journal of Agricultural
40 International Econometric Review (IER) Economics, 66(2), 204-208. Jondrow, J., C. A. Knox Lovell, I. S. Materov, and P. Schmidt (1982). On the estimation of technical inefficiency in the stochastic frontier production function model. Journal of Econometrics, 19(2-3), 233-238. Kim, S. and Y.H. Lee (2006). The productivity debate of East Asia revisited: a stochastic frontier approach. Applied Economics, Taylor and Francis Journals, 38(14), 1697-1706. Kumbhakar, S.C. (1990). Production frontiers, panel data, and time-varying technical inefficiency. Journal of Econometrics, 46(1-2), 201211. Kumbhakar, S.C. and C.A. Knox Lovell (2000). Stochastic Frontier Analysis. First edition. Cambridge University Press, United Kingdom. Lee, Y.H. and P. Schmidt (1993). A Production Frontier Model with Flexible Temporal Variation in Technical Inefficiency. The Measurement of Productive Efficiency: Techniques and Applications. Edited by H. Fried, C.A.K. Lovell and S. Schmidt, Oxford University Press, pp. 237-255. Nelsen, R. B. (1999). An Introduction to Copulas. First edition. Springer, New York. Sampaio, A., C. Barros, and J. Ramajo (2005). Technical Inefficiency in Municipal Water Distribution Service: A Case Study for Portugal. Anales de Economia Aplicada, XIX Reuni´on Anual. Edi¸c˜oes Asepelt (Associa¸c˜ao de Economia Aplicada) Espanha Badajoz. Schmidt, P. and R.C. Sickles (1984). Production frontiers and panel data. Journal of Business & Economic Statistics, 2(4), 367-374. Simar, L. and P.W. Wilson (2010). Inferences from cross-sectional, stochastic frontier models. Econometric Reviews, 29(1), 62-98. Smith, M.D. (2008). Stochastic frontier models with dependent error components. The Econometrics Journal, 11(1), 172-192. Tupper, H.C. and M. Resende (2004). Efficiency and regulatory issues in the Brazilian water and sewage sector: an empirical study. Utilities Policy, 12(1), 29-40. Vishwakarma, A. and M. Kulshrestha (2010). Stochastic Production Frontier Analysis of Water Supply Utility of Urban Cities in the State of Madhya Pradesh, India. International Journal of Environmental Sciences, 1(3), 357-367.