Smoothing the Catalan tourism micro-data time series
Abstract
In this paper we propose a method for smoothing the Catalan tourism time series between 1997 and 2000. These time series, built upon a micro database drawn from a survey conducted by the Statistical Institute of Catalonia, are somewhat volatile due, it would seem, to the incomplete nature of the information. The application of a smoothing procedure based on the combination of classical techniques and weighted moving averages allows us to overcome the problems caused by this lack of information and to obtain time series that evolve smoothly over time.
Full text
Q¨ UESTII´ O,vol. 26, 1-2, p. 197-211, 2002 SMOOTHING THE CATALAN TOURISM MICRO-DATA TIME SERIES M. ART´ IS ORTU ˜ NO J. L. CARRION I SILVESTRE ` A. COSTA S ´ AENZ DE SAN PEDRO J. SURI˜ NACH CARALT In this paper we propose a method for smoothing the Catalan tourism time series between 1997 and 2000. These time series, built upon a micro database drawn from a survey conducted by the Statistical Institute of Catalonia, are somewhat volatile due, it would seem, to the incomplete nature of the information. The application of a smoothing procedure based on the combination of classical techniques and weighted moving averages allows us to overcome the problems caused by this lack of information and to obtain time series that evolve smoothly over time. Keywords: Smothing micro-data, tourism time series AMS Classification (MSC 2000): 62P20 This paper is a joint undertaking between the Statistical Institute of Catalonia (Idescat) and the An`alisi Quantitativa Regional Research Group of the University of Barcelona. An`alisi Quantitativa Regional (AQR) Research Group Departament d’Econometria, Estad´ıstica i Economia Espanyola Universitat de Barcelona. Av. Diagonal, 690, 08034 Barcelona. Institut d’Estad´ıstica de Catalunya, Via Laietana, 58, 08003 Barcelona. – Received October 2001. – Accepted January 2002. 197
1. INTRODUCTION Studies of the tourism sector in Catalonia have traditionally drawn on macro aggregates corresponding to each tourist season, such as private consumption and gross domestic product. In so doing they have tended to rely on one of the main data sources for this sector i.e. the survey of the supply of hotel accommodation in each of the Spanish regions. This survey, undertaken by the Spanish Statistical Institute (INE), provides information about hotel occupancy, in other words, information provided by the supply side of the tourism market. In 1997, the Statistical Institute of Catalonia (Idescat) introduced a new survey of Catalan tourism, but in contrast to the survey described above this sought to obtain information from the demand side. This survey provides analysts with valuable information about visitors to Catalonia, whose point of origin is one of the other Spanish regions. Given its recent introduction, these statistics offer information for the most recent tourism seasons only and comparative data is only provided with the previous tourism seasons and, consequently, no time series is defined. It should, however, be noted that the time interval of each tourist season has changed since the introduction of the survey which hinders the definition of an appropriate time series. Thus, three tourism seasons were identified for 1997 and 1998: January to May, June to August and September to December; whilefor1999and 2000fourseasons wereidentified: Januaryto April, May to June, July to August and September to December. The recent introduction of the survey and the varying time intervals used in defining the tourist season hinder comparisons. Furthermore, although the survey was designed to embrace all the Spanish regions, only a few observations are eventuallyincluded within the database, and so the information describing individual characteristics tends to be highly heterogeneous. This heterogeneity becomes even more marked when the data are raised to the entire population. Consequently using this database to calculate growth rates gives highly volatile time series. Therefore, the aim of this paper is to present a methodology for computing time series from the micro data (the survey) but, in contrast with the original (populationraised) time series, with a smoothed temporal pattern. It is not, however, our aim to supply the analyst with a specific set of smoothed time series but rather to design a methodology that allows practitioners to obtain smoothed time series automatically whatever the concepts crossed in the database. The successful achievement of this goal depends on the application of simple smoothing methods that can be adequately employed in all cases. This paper is organised as follows. In Section 2 we describe the database provided by the survey carried out by Idescat. We describe some of the characteristics of this database and define the time series that constitute the focus of this paper. Section 3 outlines 198
the methodology that is applied in order to smooth these time series. This section contains three sub-sections that offer a detailed description of the transformations involved at each stage of our methodological proposal. Section 4 presents the results obtained. Finally, Section 5 concludes. 2. DESCRIPTION OF THE DATABASE The availability of a database built upon the conductingof a surveyat different points in time allows us to undertake the analysis at different time intervals. As a last resort, the information contained in the database can always be used to define a daily time series. However, problems arise owing to the absence of observations and distortions in the significance of the findings as the time frequency of the analysis increases. For these two reasons, in this paper, the temporal reference is fixed at monthly intervals and the monthly time series is the basic information to which our methodology is applied. This specification allows us to use classical smoothing techniques including exponential smoothing and Holt-Winters smoothing procedures. In addition, we can obtain time series of varying temporal frequency (quarterly and annual) by aggregating monthly time series. The micro database used here provides information about individual personal characteristics, including age, profession,marital status and regionof residence. It also contains details about their holidays: number and characteristics of the other group members, destination, type of accommodation, amount of expenditure, number of days spent in Catalonia and the numberof overnight stays, among others. However,here we focus on just two of these variables: (1) the numberof overnightstays and (2) the numberof tourists. Both variables are classified by type of accommodation(hotel, family and friends’ households, other types of accommodation and total) and by destination (Barcelona, Costa Daurada, other destinations in Catalonia and all destinations in Catalonia). The different combinations give rise to forty time series. The definition of these time series is strongly conditioned by the quality of the information comprising the tourism micro-database. Thus, firstly, although in aggregating the information we have tried to avoid missing or zero values, this has been unavoidable in certain periods for some time series. This might have a detrimental effect on the quality of the output following the application of the smoothing procedure. Secondly, graphic inspection of the time series indicates that there might be some outliers, the presence of which implies growth rates of doubtful validity. Finally, there would seem to be an Easter Week effect due to the fact it is a moveable feast and as such does not always occur in the same time period. 199
In this analysis these first two problems with the information are left for future consideration, particularly given that Idescat is planning to modify some of these anomalies. The third problem is discussed below. 3. METHODOLOGICAL PROPOSAL In this section we present the methodology adopted in smoothing the time series described in the previous section. One of the reasons why these time series are apparently so erratic is that the survey loses precision as the geographical and conceptual range is increased. Our proposal tries to compensate for this absence of observations by increasing the amount of information used when estimating the micro data of one particular month. The increase in the amount of information is achieved by the joint consideration of the micro-observations referring to the same month in two consecutive years. Thus, we compute the average number of tourists and overnight stays in the same month for two consecutive years and assign this mean value to these months. Hence, we take into account information that refers to two similar periods (month) and, as a consequence, we are able to reduce the volatility of the time series. Figure 1. Brief description of the methodological proposal. This simple methodallows us to obtain time series that have a smoother pattern throughout the time period under consideration. The main problem arises, however, when deciding the weightings that should be applied when computing this mean value. One possibility is the specification of equal weights for each time period. Yet, it might be argued that a weightingsystem that gives greater weighting to more recent information is preferable to a system that attaches the same importance to the two sets of information. 200
If this is the case, the analyst needs to select these weightings. This is not, however, a straightforward decision, given that different weightings will result in different numbers of tourists and overnight stays. We try to overcome this drawback by suggesting a method by which the weightings can be estimated. The method comprises three stages. 3.1. First stage: The computation of the original time series The approach relies on the definition of two sets of time series. The first set is the one defined by the original time series, that is, the time series that are derived from raising the data of the survey to the population. As mentioned in the previous section, the forty time series thus obtained are highly volatile over time, which is the problem that this paper seeks to rectify. The large number of observations available for the short period under analysis (forty-eightobservations in just four years)means that the application of the stochastic approach to the modelling of these processes is not the most appropriate and that the classical approach should be the one to be adopted. 3.2. Second stage: Definition of the time series of reference In the second stage of the analysis we obtain the set of forty time series following the application of a classical smoothing procedure to the original time series. This second set of smoothed time series serves as a referent for defining the system of weightings to be used when computing the average. Before applying the classical methodological approach to the modelling of the time series we need to know the type of time series that is being dealt with. Here, the characterisation of the time series was performed using two test statistics. In order to decide the consideration of a time trend we applied the Daniel test, while for the seasonal component we applied the Kruskal-Wallis test. These tests indicated that in most cases the patterns of the time series are given by both components. Nevertheless, it should be noted that these results are not entirely reliable since these statistical tools are more suited to moderate or large sample sizes. Table A.1 in the Appendix shows the results of the application of both tests. The most appropriate smoothing procedure for these data is that of Holt-Winters since there are trend and seasonal components in most of the forty time series. Graphical inspection indicates that the additive model can provide a good fit, although this conclusion might need to be revised as further information comes available. Before the Holt-Winters smoothing procedure can be applied to the time series under consideration, we need to analyse the effect of Easter Week on these time series. The only periodthat might have hadan influence on the time series was in 1997. In this year Easter fell in the month of March while for the remainingyears it fell in April. In order 201
to avoid distortions that might affect the output of the smoothing procedure we decided to test for the presence of a 1997 Easter Week effect and, if there was found to be such an effect to correct the time series to take it into account. This meant the estimation of a regression model that specifies the time series as a function of an independent term, a time trend, a seasonal dummies set, and an impulse dummy that captures the effect that can be assigned to March 1997. Only in three cases was this impulse dummy found to have a statistical significance of 10%, and the effect was corrected in each case. The three time series were the total numbers of overnight stays in Catalonia (CATPE T), overnightstays with family or in a friend’s household in Catalonia (CATPE F) and overnight stays with family or in a friends’ household using this minimisation criterion in Barcelona (BCNPE F). Table 1. Estimated coefficients for the Holt-Winters smoothing procedure. Overnight stays Tourists ALFA BETA GAMMA ALFA BETA GAMMA CAT T 00 000 0 CAT H0.02 0.03 0 0.03 0 0 CAT F0 0 0 0.01 0 0 CAT R00 000 0 CAT NH 0 0 0 0.01 0 0 BCN T0 0 0 0.02 0 0 BCN H0.01 0.09 0 0.26 0 0 BCN F00 000 0 BCN R0 0 0 0.01 0.18 0 BCN NH 0 0 0 0.01 0.07 0 CD T00 000 0 CD H0.01 0.08 0 0 0 0 CD F00 000 0 CD R00 000 0 CD NH 00 000 0 RD T00 000 0 RD H00 000 0 RD F00 000 0 RD R00 000 0 RD NH 0.1 0 0 0 0 0 Note: CAT T refers to all types of accommodation used in Catalonia (CAT). CAT H indicates those people staying in a hotel. CAT Fthose staying with family or in a friend’s household. CAT R denotes the other types of accommodation used. Finally, CAT NH denotesthosestaying in accommodation other than a hotel. This notation is repeated for the territorial division considered here: BCN-Barcelona, CD-Costa Daurada and RD-remaining destinations. 202
Note: CATPE T denotes the raw time series of the total number of overnight staysin Catalonia, CATPE TSMAE denotes the smoothedtime series using the Holt-Winters procedure with the estimated coefficients and CATPE TSMAM is the smoothed time series using the Holt-Winters procedure with a fixed value for the coefficients. GCATPE T, GCATPE TSMAE and GCATPE SMAM are the corresponding growth rates. CATTU T, CATTU TSMAE and CATTU TSMAM refer to the tourist series. Figure 2. Overnight stays and tourists in Catalonia. Levels and growth rates of the original and smoothed time series.
Once the time series affected by the Easter Week had been modified, we applied the Holt-Winters smoothing procedure to estimate the value of the parameters of the independent term (α , the slope (β and the seasonal parameter (γ using the criteria of the minimisation of the sum of squared residuals. The estimated coefficients for each time series are presented in Table 1. Note that although it is possible to fix the value of these parameters, the estimation provides a better fit. As can be seen from Table 1, in most cases the estimated parameters equal zero, indicating that the corresponding component —independent term, trend and seasonality— is stable, that is, it does not vary during the time period under analysis. The monthly time series for the level and rate of growth for the number of tourists visiting Catalonia and the number of overnight stays following the application of the Holt-Winters smoothing procedure are given in Figure 2. We denote these time series as the smoothed-HW time series. Each figure contains information about the original time series, the smoothed-HW time series with manual selection of the parameters and the smoothed-HW time series with the estimated parameters obtained using the minimisation of the squared sum of errors’ criteria. A number of comments should be made. First of all, it can be seen that the smoothed time series built on the use of the estimated coefficients show a smoother behaviour than those in which the value of such coefficients is imposed (in this case the parameter values were fixed at 0.1). Second, these results indicate that the estimation of the initial values used in obtaining the smoothed-HW time series influences the computation of the growth rates. Thus, for instance, we encounter a contradiction for the time series of tourists coming to Catalonia in which the type of accommodation is not specified (CATTU T). In this case the growth rates computed using the original time series are negative, while with the smoothed-HW time series they are positive. This is also the case for the time series of tourists coming to Catalonia and staying in other types of accommodation (CATTU RD) and overnight stays in hotel accommodation in Catalonia (CATPE T). Third, and in contrast to the smoothed-HW time series, for some time series and periods the original time series present null values which means the growth rates are discontinuous in these cases. 3.3. Third stage: Application of the Seasonal Weighted Moving Average (SWMA) smoothing procedure In the third stage of the analysis we select the weightings that best fit the time series smoothed in the previous stage. The estimation of these weightings is carried out by specifyingthe criteria of minimisationof the sum of squarederrors, wherethe errors are given by the difference between the smoothed-HW macro time series and the weighted average time series —hereafter smoothed-micro time series. This section describes the methodology adopted in this optimisation procedure. 204
Once the macro time series in question has been smoothed using the Holt-Winters procedure, employing an additive specification and by estimating the parameters of the model, we proceeded to select the set of weights used in the procedure applied in this paper to the micro series. This procedure can be understood as the computation of seasonal weighted moving averages (SWMA). For instance, in computing the smoothed time series of tourists for January 1998 using the SWMA procedure we need to take into account the information of the original time series of tourists that corresponds to January 1997 and January 1998. To compute the observation for February 1998 of the smoothed time series we need to look at the observations of the original time series referring to February 1997 and February 1998, and so on. The important aspect of our proposal is the system of weightings applied in computing the average. As mentioned above, different smoothed time series are obtained depending on the set of weights used. The greater the weight given to more recent values in the time series, the more the smoothed-micro time series tends to resemble the original time series. In the first stage we smoothed the original time series by applying five sets of weights: 50/50, 40/60, 30/70, 20/80 and 10/90. Yet, in order to avoid being subjective when selecting the system of weightings, we estimated this parameter through the minimisation of the sum of squared error of the difference between the smoothed time series using the SWMA procedure —the smoothed-micro time series— and the smoothed-HW time series. This estimation can be outlined as follows. If we denote the original time series by Y, the smoothed-HW time series by Y and the smoothed-micro time series using the SWMA procedure by ˆ Y, the target function to be minimised is the function given by: f p T ∑ i s j y i ˆyi 2 j 1 2 s , wheresdenotes the order of seasonality —here s 12. Thesmoothedmicro time series is computed from: ˆyi pyi 1 p yi s where pis the weight (parameter) to be estimated. Therefore, the optimisation program can be represented by: min T ∑ i s j y i pyi 1 p yi s 2 subject to 0 p 1 The necessary condition establishes that: ∂f p ∂p T ∑ i s j2 y i pyi 1 p yi s yi yi s 0 205