scieee AI-readable full text Open interactive document viewer

Data envelopment analysis efficiency in the public sector using provider and customer opinion: An application to the Spanish health system

Tapia García, Jesús Alberto,Salvador González, Bonifacio

Abstract

Producción Científica

Full text

Health Care Management Science https://doi.org/10.1007/s10729-021-09589-7 Data envelopment analysis efficiency in the public sector using provider and customer opinion: An application to the Spanish health system Jes´ us A. Tapia1Bonifacio Salvador1 Received: 18 March 2020 / Accepted: 23 December 2021 ©The Author(s) 2022 Abstract Measuring the relative efficiency of a finite fixed set of service-producing units (hospitals, state services, libraries, banks,...) is an important purpose of Data Envelopment Analysis (DEA). We illustrate an innovative way to measure this efficiency using stochastic indexes of the quality from these services. The indexes obtained from the opinion-satisfaction of the customers are estimators, from the statistical view point, of the quality of the service received (outputs); while, the quality of the offered service is estimated with opinion-satisfaction indexes of service providers (inputs). The estimation of these indicators is only possible by asking a customer and provider sample, in each service, through surveys. The technical efficiency score, obtained using the classic DEA models and estimated quality indicators, is an estimator of the unknown population efficiency that would be obtained if in each one of the services, interviews from all their customers and all their providers were available. With the object of achieving the best precision in the estimate, we propose results to determine the sample size of customers and providers needed so that with their answers can achieve a fixed accuracy in the estimation of the population efficiency of these service-producing units through the use of a novel one bootstrap confidence interval. Using this bootstrap methodology and quality opinion indexes obtained from two surveys, one of doctors and another of patients, we analyze the efficiency in the health care system of Spain. Keywords Data envelopment analysis Sampling survey research Public sector Sample size Bootstrap Efficiency confidence interval 1 Introduction Data Envelopment Analysis (DEA) has become a widely used technique to compare the efficiency of serviceproducing units because it easily handles the multiple outputs characteristic of public sector production, is nonparametric and does not require input price data [34]. The popular application is with deterministic information and classic DEA models [3,10], for instance, in health care Jes´us A. Tapia [email protected] Bonifacio Salvador [email protected]a.es 1Department of Statistic and Operative Research, University of Valladolid, Campus Miguel, Delibes, Paseo Bel´ en, no ¯7,47011 Valladolid, Spain [1,37]or[25]), universities and research institutes [26], government [17], public libraries [22,23], schools [35], public transport services [15]orbanks[28]. In recent years, some approaches consider the opinion of the customer to be crucial for measuring DEA efficiency in public services ([5,16]). Consumer satisfaction-opinion surveys are a common tool for building opinion indexes, which measure the quality of the service, and to be used as output variables in DEA ( [20,27,31,32,35,40,41]and [42]). Besides customer opinion-satisfaction (output), provider opinion-satisfaction (input) is also a fundamental protagonist in public service-producing units (SPUs), our decisionmaking units. The provider’s positive opinion-satisfaction concerning the service offered may result in a better customer opinion concerning the service received. For instance, a city’s public bus drivers with a good opinion of their salary, timetable, partners, driven vehicles, etc., may influence a better satisfaction-opinion of the travellers using the J. A. Tapia and B. Salvador service. One way to measure the service quality offered is with satisfaction-opinion indexes obtained through survey samples carried out with the providers. In Tapia et al. [40– 42] the protagonism of the opinion of the providers is not considered, that is, the inputs are deterministic and only the outputs are estimated using a customer sample. Introducing the opinion-satisfaction of the providers as estimated indexes increases the field of application of this work, where opinion-satisfaction indexes, estimated from a sample of providers (inputs) and a sample of customers (outputs), are used as information to measure the DEA efficiency of the public services. To do so, it is necessary to determine the sampling design, the sample size and the estimators of the satisfaction-opinion indexes of both the providers and the customers in each SPU. In practice, this stochastic input and output information is available in many services, as it is increasingly common to conduct opinion-satisfaction surveys for both providers and customers. For instance, the sample data used in the Spanish health system application of Section 5. The efficiency obtained with DEA models, using the opinion indexes estimated with the survey answer of samples of providers and consumers as data, will be an estimation of the population DEA efficiency. This population efficiency is an unknown non-evaluable parameter, since it would be necessary to use the indexes obtained with the opinion of all the customers and all the providers of all the services as data and this is a census of the entire population [40]. This statistical analysis of the DEA efficiency gives rise to the problem of determining the sample size of customers and providers needed to guarantee an a priori fixed accuracy in the estimation, with confidence interval, of the DEA efficiency of each public service, which is the object of our investigation. Liu et al. [30] considered the statistical analysis and sampling process as an important direction for handling the DEA. Ceyhan and Benneyan [7] investigated the impact of the sample size on the measures of efficiency when the DEA problem was carried out on values that include such estimated proportions as defect, satisfaction, mortality, or adverse event rates estimated from samples. Nevertheless, they do not propose a solution to the necessary sample size in order to control the error in the measures of efficiency. The problem of calculating the sample size of customers necessary when the objective is to estimate, with a fixed precision in the estimate error, the population efficiency in a finite set of public services using stochastic data output (customer opinionsatisfaction indexes) and known (non-stochastic) data input was resolved in Tapia et al. [40,42]. In this paper, we propose a solution to determine the customer and provider sample sizes needed to estimate the DEA efficiency in public services with bootstrap confidence intervals and a fixed accuracy. These intervals capture the random variations introduced in the DEA analysis by using outputs and inputs estimated with a sample. So far, Simar and Wilson’s [38,39] methodology has been the most common for measuring efficiency with bootstrap confidence intervals in such public services as health care ([11]), universities and research institutes [4], government [6], public libraries [29], schools [19], tourism [2], banks [28] or public transport services [21]. The problem posed by Simar and Wilson is different from ours because the stochastic character of the inputs and outputs is different. In Simar and Wilson, the stochastic character of the input and output information comes from considering the set of available SPUs, , as a sample from an infinite population and the sample observations in are realizations of identically, independently distributed (iid) random variables with a probability density function with support over can produce ,[14]. In this work, the set of available SPUs are the only units whose efficiency we wish to evaluate. Therefore, we do not consider them a sample as in the classical DEA models, [9]. The stochastic character of our output and input information comes from the fact that these data are opinion-satisfaction indexes that it is necessary to estimate by taking independent random samples in each SPU, and the probability, in the statistical model, depends on the sample design used in each SPU. In Shwartz [37], a random sample of patients arriving for health care services were bootstrap resampled to obtain data input and interval estimates of the DEA efficiency. These interval estimates are conservative (large) and the problem of the patient sample size necessary to obtain the desired accuracy in the estimation of the DEA efficiency is not resolved. In other approaches, where samples in public services are used to estimate the data (Charles2014), the efficiency is estimated with stochastic DEA models. These models, which use LP problems subject to constraints defined in terms of probability, are also called chance-constrained problems. The deterministic characterization of “efficient” is then changed by the probabilistic characterization “probably efficient”. A vast number of papers show a wide range of uses of chance-constrained programming (CCP), including [8,9,12,33]or[43]. One of the main advantages of the technique presented in this study is its simplicity, as we use the original DEA models with constant (CCR) and variable returns-to-scale (BCC), that is, linear programming (LP) problems subject to deterministic constraints [3,10]. In this paper, we therefore examine the implications of input and output data estimated with a provider and a customer sample, respectively, for the performance analysis of public services using DEA analysis techniques. In Section 2, we examine the nature of the problem. Data envelopment analysis efficiency in the public sector using provider... Section 3presents a theoretical method to determine the customer and provider sample size needed to estimate the population DEA efficiency with bootstrap confidence interval, which is examined in Section 4. Section 5includes the empirical application in the Spanish health system we have undertaken. Section 6contains our conclusions. All the software is our own elaboration using MATLAB and it is included in Appendix II. 2 Nature of the problem We consider a finite fixed set of SPUs, service-producing units, one provider interview to estimate inputs and one customer interview to estimate outputs. For example, in each hospital of a homogeneous set of , it is possible to estimate the general satisfaction of the personnel, or their annual time of formation, or other personal opinion indexes (input data) by interviewing a sample of the personnel. It is also possible to estimate the satisfaction of the patients attention they have received, the human and material resources of the hospital, etc., by interviewing to a patient sample (output data). The main problem approached in this paper is to estimate the unknown parameter population DEA efficiency. The population efficiency, or census efficiency, is unknown because, in order to know it, it would be necessary to interview all the providers and all the customers (census) of each one of the public services and, with this population data, to obtain the opinion indexes and use them as input/output data in the classic LP model CCR or BCC of Table 1; in real applications, these censuses are completely non-viable. The error made in estimating the input/output data with samples is transferred to the DEA efficiency estimation. In this study, we propose a methodology that guarantees a reasonable quality of the DEA efficiency estimation. Formally, in the jth service-producer unit (SPU for short), we consider the finite provider population Uj U1j UNxjjof size and the customer population WjW1j WNyjjof size . Each provider is a quantitative vector Ukj 1,where is the answer of the kth provider of the jth SPU to the ith opinion provider item, i.e., each customer Whj 1is a quantitative vector where is the answer of the hth customer of the jth SPU to the rth customer opinion item. The population input and output data are opinion indexes obtained as a function of the provider and customer population answers, in general: Xj11 1 1 1(5) Yj11 1 1 1(6) These functions and can be of any type, with or without weights, whenever they admit an interpretation of population opinion-satisfaction indexes as input and population opinion-satisfaction indexes as output. Using the data XjYj1in the Table 1 models, we obtain the population DEA efficiency scores 1... . If it were possible to carry out a census and to know these population indexes, the information would be fixed or deterministic and the character of the problem would not be stochastic. A lack of knowledge concerning the population information input and output makes the taking of samples to estimate 1... necessary. In the SPU ,let U1j UnxjjUjand W1j WnyjjWjbe random samples of size and of the provider and customer populations, respectively. To obtain the estimators of the input indexes Eq. 5and the output indexes Eq. 6,the same functions and are used with the random samples: Xj11 1 1 1(7) Yj11 1 1 1(8) Having observed the provider and customer sample answers ukj 11and whj 11, respectively, Eqs. 9 and 10 are the estimates of the input and output indexes, respectively; that is, the values that the estimators Eqs. 7 and 8take with the sample answers of the customers and providers: x11 1 1 11(9) y11 1 1 11(10) Using XjYj1in the Table 1models, we obtain the estimators 1of the population efficiency scores 1, on the understanding that the DEA model is maximized, or minimized, with the data xjyj1in order to obtain the estimation of the estimator . Therefore, our statistical model corresponds to independent, random samples in each SPU, that is, the sample space is 1,where samples ukj of size and samples whj of size in SPU , and the probability depends on the sample design used. The first objective of this paper, having fixed 0 1 and 0 1 , is to obtain the provider and customer sampling size, and , respectively, to estimate and J. A. Tapia and B. Salvador Table 1 DEA models with variable (BCC) and constant returns-to-scale (CCR) and the estimator such that: 1 1 . (11) The second objective is to determine a confidence interval for the populational efficiency, in each SPU, with a fixed accuracy. 3 How many providers and customers need to be interviewed? The provider and customer sample size problem Eq. 11 can only be analytically resolved in the case of one input and one output and the CCR model. A rigorous proof of all the results are provided in the Appendix I. 3.1 CCR model with one provider and one customer opinion index Let us consider these assumptions: C1 Fixed SPUs. C2 One provider population opinion index 1 as input and one customer population opinion index 1as output. We consider 1and 1to be the corresponding estimators. C3 CCR model Eq. 3with output orientation (CCR-O). Let 1and 1 .Inthis situation and with the given notation, the population efficiency obtained with the CCR-O model and its estimator are: 1 (12) 1 . (13) Having fixed 0 1 , we thus consider the sets of : 1(14) The following Lemmas 1, 2 and 3 are used to prove Theorem 1 which, in this particular case, establishes the relation between the accuracy of the provider and customer opinion estimation indexes and that of the population CCRO efficiency estimation. Lemma 1 Under assumptions C1, C2 and C3, with the given notation, fixed 0 1 and such that 1, then: 1 1 1 1 1 1 1 1 1 1 1 1. Lemma 2 Under assumptions C1, C2 and C3, with the given notation, fixed 0 1 ,and defined as in Lemma 1 and such that : 1 Consider the set of 1 1 1 1(15) Data envelopment analysis efficiency in the public sector using provider... then: implies implies Lemma 3 Under assumptions C1, C2 and C3, with the given notation, fixed 0 1 and for any 1 we have that 12 12 1 1 1 1 12 12 Theorem 1 Let us consider the assumptions C1, C2 and C3. Having fixed 0 1 and 0 1 ,if 1 1 ,then 12 12 12 1213. (16) Lemma 4 is an instrumental result, used to prove Theorem 2 and Corollary 1, which establishes how to determine the provider and customer sample size in each SPU, so the CCR-O efficiency estimator has the precision fixed in Eq. 11. Lemma 4 Under the hypotheses of Theorem 1, if 12 12112 12 then 4 1214 12 Theorem 2 Let us consider the assumptions C1, C2 and C3. Having fixed 0 1 and 0 1 , for every 1,let be the sampling size in the SPU , such that 61(17) and be the sampling size such that 61(18) then 4 121 1 . (19) Corollary 1 Let us consider the assumptions C1, C2 and C3. Having fixed 0 1 and 0 1 , for every 1,let be the sampling size in the SPU ,such that 2 2 1 61 (20) and be the sampling size such that 2 2 161 (21) then 1 1 (22) Remark 1 gives the explicit formula to obtain the sample size under the usual simple random sample without replacement sample design. Remark 1 If the design in each SPU is simple random sampling without replacement, and the output (i.e. input) is the mean of all the answers of the population to a survey item, then the sampling size that it verifies 11or 10 1 is ( [36]) 1 (23) with 2 112 2 2and 1121112 , where 2is the population variance and the normal standard distribution function. 3.2 CCR or BCC model with two or more provider and/or customer opinion indexes In this section, we report on our simulation study to check that Theorem 2 and Corollary 1 also work in the BCC model with two or more estimated provider and/or customer opinion indexes. If we consider items (the same in all SPUs) to estimate the provider opinion indexes (inputs) 1 with 1, Remark 1 calculates the sample size necessary to achieve 61 1 ... 1... . (24) J. A. Tapia and B. Salvador We propose to determine the provider sample size in the SPU as 1... (25) i.e., if we consider items (the same in all SPUs) to estimate the customer opinion indexes (outputs) the provider sample size in the SPU is determined as 1... (26) where is the sample size necessary to achieve 61 1 ... 1... . (27) 3.2.1 Simulation study We use the [13] health center data (Table 2)tosimulate a population model: in the jth health center a population size of providers and customers, and ,are generated from a random uniform distribution, according to the intervals [10000 50000]and [30000 80000], respectively. For each provider, we generate two item answers 1 2 , i.e., for each customer 1 2 , from a bivariate normal distribution as 1 22 240 0241.... 1... 12 1 22 240 0241.... 1... 12 where , and ,are the original value doctor, nurse and outpatient, inpatient of the jth health center, columns 2, 3, 4 and 5 of Table 2, respectively. We consider the population mean to the simulated answers to the provider and customer items in the jth center, 1... 12, to simulate the population inputs and outputs, columns3,4,5and6inTable3: 1 2 1112 1 2 1112 The last two columns in Table 3show the population efficiency scores CCR and BCC with output orientation. To check the relation between sample size, estimation of the input/output indexes and estimation of the DEA efficiency, using Theorem 2 and Corollary 1 and these simulated population data, we follow the next steps: i. In the jth health center, 1 12, a previous simple random sample without replacement of 25 providers 025 is taken to estimate the two inputs 0 1 0 2, and their variances 2 0 1 2 0 2 using the sample means and the sample quasivariances, respectively, i.e., we estimate the two outputs 0 1 0 2and their variances, 2 0 1 2 0 2 with a previous simple random sample without replacement of 25 customers 025 . ii. Fixed 0.1or0.2and1 =0.9 as in Corollary 1, and with the estimates of step i., the sample sizes, and , are determined using Eqs. 23,25 and 26. iii. In the jth health center, the simple random samples without replacement of size and are taken and the inputs 1 2 and outputs 1 2 are estimated. With the data 1 2 1 2 1... 12, the estimated efficiencies 1... 12 are obtained, maximizing the LP model Eqs. 3or 4with output orientation. iv. One thousand iterations of step iii. are carried out obtaining, for the jth health center, 1000 estimated efficiency scores 1... 1000 and 1000 intervals 1 1 1000 (28) v. The probability is approximated by calculating 1 1000 1000 1 (29) Table 4shows the sampling sizes, and , obtained for the jth health center in the last iteration, for the two values of and 0.1. In health center 3 or 6, the customer sample size increases up to 6 times when fixing a maximum 0.2to0.1. The probabilities approximated with Eq. 29 take the value one for all the health centers, the two values of , the CCR-O or BCC-O model and with 0.1. Therefore, the confidence intervals for the population Data envelopment analysis efficiency in the public sector using provider... Table 2 Number of doctors, nurses, outpatients and inpatients in 12 health centers Health center Doctor Nurse Outpatient Inpatient Efficiency score CCR-O Efficiency score BCC-O 1 2.0 15.1 10 9 1 1 2 1.9 13.1 15 5 1 1 3 2.5 16 16 5.5 0.883 0.925 4 2.7 16.8 18 7.2 1 1 5 2.2 15.8 9.4 6.6 0.763 0.767 6 5.5 25.5 23 9 0.835 0.955 7 3.3 23.5 22 8.8 0.902 1 8 3.1 20.6 15.2 8 0.796 0.826 9 3 24.4 19 10 0.960 0.990 10 5 26.8 25 10 0.871 1 11 5.3 30.6 26 14.7 0.955 1 12 3.8 28.4 25 12 0.958 1 Font: Table 1.5 [13] efficiency score , obtained with the samples of the size of Table 4and Corollary 1, are very conservative. However, in the next section, we will see that these same sample sizes allow less conservative bootstrap DEA efficiency confidence intervals to be obtained. 4 Description of the bootstrap efficiency confidence interval technology Bootstrap uses resampling to estimate the value of a parameter of a population ([18]). For the problem suggested in Section 2, we propose bootstrap resampling of the samples of the provider and customer answers to the opinion item to obtain confidence intervals for the population efficiencies, following these steps: i. Having fixed and a probability 1 , we determine the sample sizes and , in the SPU ,using Corollary 1, Remark 1 and Eqs. 25 and 26. 1. In the SPU , we take a provider and a customer simple random samples without replacement ukj 1... 1... and whj 1... 1... , respectively, to estimate the provider and customer opinion indexes 1and 1, respectively, for example, with the sample means: 11... 1... (30) 11... 1... . (31) ii. In the SPU , we take a bootstrap sample with replacement ukj 1... from ukj 1... , i.e., whj 1... of size from whj 1... , with which we obtain the bootstrap version of the inputs, xj1... ,and outputs, yj1... . With the data 1... 1... 1and the DEA model of Table 1, we obtain the bootstrap version of the estimated DEA efficiency, 1... . iii. The step iii. is repeated Btimes and the Bbootstrap versions of the estimated DEA efficiency for the SPU , 1... are =1,..., . 2. In the SPU , the observed percentile bootstrap confidence interval for the population efficiency score is obtained, having fixed a coverage intention of level 1- ,as 2 1 2 (32) where is the -percentile of the values =1,..., . 4.1 Simulation study To illustrate the bootstrap efficiency confidence interval methodology and to check the estimation quality of the DEA population efficiency obtained, a simulation was performed using the population model simulated from Table 3. First, we fixed 0.2 and the confidence 1 0.9 to calculate the provider and customer sample size to J. A. Tapia and B. Salvador Table 3 Simulated population model Health center Provider population size Nxj X1 X2 Customer Population size Nyj Y1 Y2 Population efficiency score CCR-O Population efficiency score BCC-O 1 16684 2.04 15.52 57128 10.29 9.25 1 1 2 26950 1.95 13.40 36283 15.36 5.15 1 1 3 26583 2.56 16.41 65876 16.41 5.66 0.882 0.926 4 18974 2.77 17.22 33569 18.44 7.37 1 1 5 43106 2.26 16.26 62522 9.63 6.79 0.761 0.764 6 21305 5.65 26.27 68349 23.68 9.24 0.834 0.952 7 10599 3.40 23.98 36417 22.61 9.07 0.901 1 8 27432 3.18 21.13 73596 15.60 8.22 0.798 0.828 9 11113 3.09 24.94 38962 19.54 10.29 0.954 0.987 10 31717 5.15 27.44 76479 25.70 10.28 0.875 1 11 39802 5.44 31.44 36356 26.63 15.12 0.955 1 12 32282 3.91 29.26 55847 25.67 12.33 0.955 1 estimate the input and output data; supposing a simple random sample without replacement, these sample sizes are columns 2 and 3 in Table 4. The steps ii.-v. are iterated 1000 times, fixing the confidence of the bootstrap efficiency interval 1 0.9 or 0.95 and 2000 resampling Bootstraps ( 2000). The confidence of the bootstrap interval for the DEA efficiency in the SPU is approximated with: 1 1000 1000 =1 1... 12 (33) where 1... 1000 are the 1000 bootstrap efficiency confidence intervals obtained in step v. Table 5shows the approximate confidence of the bootstrap intervals, output orientation DEA models. We observe that the control of the coverage level 1 leads to the achievement of the confidence of the bootstrap efficiency interval required by the experimenter. The amplitude of the bootstrap efficiency confidence intervals is analysed with the approximation of the expected value of the bounds 1000 =1 1000 and 1000 =1 1000 . (34) Table 6shows the approximation of the expected values of the bounds of the bootstrap efficiency confidence interval, considering the BCC model with output orientation. If Table 4 Customer and provider sample size, taking two values of , 0.2 or 0.1, and a confidence 1 0.9 0.2 0.1 Health center Provider sample size Customer sample size Provider sample size Customer sample size 1 463 560 1747 2654 2 681 358 1858 3627 3 398 380 2481 1268 4 544 487 1047 1749 5 495 404 2440 1807 6 441 403 1650 2397 7 399 551 1533 2131 8 492 418 2010 2519 9 542 371 1495 2045 10 394 515 1731 2900 11 457 612 2872 1456 12 464 257 1265 2015 Data envelopment analysis efficiency in the public sector using provider... Table 5 Simulated confidence of bootstrap efficiency intervals, having fixed 0.2 10.90 1 0.95 SPU CCR-O BCC-O CCR-O BCC-O 1100 100 100 100 2100 100 100 100 392.5 91.9 96.5 95.8 498.4 99.9 99.5 99.9 592.2 91 96.5 95.5 689.7 89.9 95.2 94.6 789.3 98.3 93.7 99.2 891.1 92.4 96.3 97.3 991.9 91.4 96.4 95.8 10 90.6 98.9 94.8 99.4 11 89.4 100 95 100 12 92 100 95 100 Output-oriented CCR and BCC model we look at the SPUs in which the expected efficiency confidence interval contains the one, 1, 2, 4, 7, 10, 11, 12 , these SPUs coincide with the efficient population units (value 1 in column 9 from Table 3). As expected, the increase in the trust 1 of the interval Bootstrap leads to an increase in the amplitude. In conclusion, the bootstrap efficiency confidence interval methodology has the advantage that, after determining the provider and customer sample size using Corollary 1, Remark 1 and Eqs. 25 and 26, the experimenter can achieve the confidence required, 1 , to estimate the population efficiency. 5 Application to the Spanish health system This section provides an empirical analysis of health production for Spain’s 18 Autonomous Communities (CCAA). Spain’s Health Ministry has, for some time, been compiling the statistic “Health Barometer” (HB), where a group of individuals is selected in each CCAA, and a questionnaire is carried out to test the health system. One of the survey question blocks take the opinion of the individual concerning the attention provided by the doctors of primary attention and pediatrics in Spain with the following items: Table 6 Approximation of the expected values of the bounds of the bootstrap efficiency confidence intervals, 0.2 0.1 and 1 0.9 or 0.95 10.9 1 0.95 SPU 11 1 1 1 21 1 1 1 3 0.880 0.966 0.872 0.974 4 0.979 1 0.972 1 5 0.737 0.821 0.731 0.835 6 0.894 0.985 0.885 0.990 7 0.956 1 0.948 1 8 0.796 0.854 0.791 0.860 9 0.941 0.998 0.932 0.999 10 0.954 1 0.943 1 111111 121111 Output-oriented BCC model