Full text
International Journal of Geo-Information Article The Value of OpenStreetMap Historical Contributions as a Source of Sampling Data for Multi-Temporal Land Use/Cover Maps Cláudia M. Viana * , Luis Encalada and Jorge Rocha Centre for Geographical Studies, Institute of Geography and Spatial Planning, Universidade de Lisboa, Rua Branca Edmée Marques, 1600-276 Lisboa, Portugal; [email protected] (L.E.); jorge.r[email protected] (J.R.) *Correspondence: [email protected]; Tel.: +351-210-442-938 Received: 30 November 2018; Accepted: 23 February 2019; Published: 28 February 2019 Abstract: OpenStreetMap (OSM) is a free, open-access Volunteered geographic information (VGI) platform that has been widely used over the last decade as a source for Land Use Land Cover (LULC) mapping and visualization. However, it is known that the spatial coverage and accuracy of OSM data are not evenly distributed across all regions, with urban areas being likelier to have promising contributions (in both quantity and quality) than rural areas. The present study used OSM data history to generate LULC datasets with one-year timeframes as a way to support regional and rural multi-temporal LULC mapping. We evaluated the degree to which the different OSM datasets agreed with two existing reference datasets (CORINE Land Cover and the official Portuguese Land Cover Map). We also evaluated whether our OSM dataset was of sufficiently high quality (in terms of both completeness accuracy and thematic accuracy) to be used as a sampling data source for multi-temporal LULC maps. In addition, we used the near boundary tag accuracy criterion to assesses the fitness of the OSM data for producing training samples, with promising results. For each annual dataset, the completeness ratio of the coverage area for the selected study area was low. Nevertheless, we found high thematic accuracy values (ranged from 77.3% to 91.9%). Additionally, the training samples thematic accuracy improved as they moved away from the features’ boundaries. Features with larger areas (>10 ha), e.g., Agriculture and Forest, had a steadily positive correlation between training samples accuracy and distance to feature boundaries. Keywords: OpenStreetMap (OSM); Volunteered Geographic Information (VGI); land use land cover; mapping; accuracy; sampling data 1. Introduction Since the end of the 20th century, land use and land cover (LULC) maps have been extensively generated at different spatial and temporal scales. The 2000 Global Land Cover map [1] and the 2000 CORINE Land Cover (CLC) [ 2 ] map are two examples. Multi-temporal LULC maps can be used to monitor land changes over time, enabling the creation of indicators that can measure changes and support land management. Mapping results have a significant impact on our understandings of LULC patterns and can affect monitoring, characterization, and quantification outcomes. Nevertheless, one of the main challenges regarding LULC map production is the difficulty of distinguishing and accurately mapping land attributes. Over the last decade, volunteered geographic information (VGI) platforms [ 3 ], as well as other data contributed by social network communities [ 4 , 5 ], have been widely used as sources for LULC mapping and visualizations [ 6 – 13 ]. Crowdsourced content from online platforms, accessed, and exchanged by ISPRS Int. J. Geo-Inf. 2019,8, 116; doi:10.3390/ijgi8030116 www.mdpi.com/journal/ijgi
ISPRS Int. J. Geo-Inf. 2019,8, 116 2 of 18 citizens, has emerged as a supplementary data source with significant implications for LULC database production [ 14 ]. This nontraditional data source is not necessarily a substitute for official data, but is considered to be complementary [15]. OpenStreetMap (OSM) is a free, open-access VGI platform to which volunteers from all over the world collaboratively contribute data, and it is a particularly promising source of information for LULC analysis [ 10 , 16 , 17 ]. Several factors have been essential to its success, including its availability of up-to-date data with global coverage, improvements in data quality (mainly driven by an increase in the number of contributors over time) [ 11 , 18 ], and its extensive volume and variety of thematic attribute data [ 11 ]. OSM therefore has the potential for use in the long-term mapping and monitoring of LULC changes [ 6 ] and could plausibly be used to improve the production, verification, and validation of LULC maps [ 6 – 8 , 11 , 19 – 21 ]. A multi-temporal trajectory can be achieved using OSM data [ 22 ], not only for LULC mapping but also for ground-validated data creation [19]. However, it is known that both the spatial coverage and the accuracy of OSM data are not homogeneous across all regions [ 16 , 23 ], with urban areas being likelier to have promising contributions (in both quantity and quality) than rural areas [ 24 , 25 ]. Several authors have recently highlighted quality issues on VGI platforms [4,26–32]. OSM data still lack formal standards, such as those established by the International Organization for Standardization (ISO) 19157: 2013 Geographic information-Data quality [ 11 , 33 ]. Five data quality criteria are commonly mentioned in the literature: (1) Completeness, (2) temporal accuracy, (3) logical consistency, (4) positional accuracy, and (5) thematic accuracy [ 34 – 36 ]. The thematic accuracy [ 37 ] of the OSM platform’s LULC mapping is currently a hot research topic [ 7 , 9 – 11 , 34 ], and a number of studies have compared OSM data with authoritative reference data [ 6 , 7 , 16 , 33 , 35 , 38 ]. For example, two studies of mainland Portugal [ 6 , 7 ] found a total of 76.7% agreement between the data from OSM and CLC maps, with artificial surfaces, forests, and water bodies presenting promising results. Similarly, a study of Vienna [ 21 ] compared OSM data to data from the Global Monitoring for Environment and Security Urban Atlas (GMESUA) and found an agreement rate of 76% to 91%. More recently, another comparison of OSM and GMESUA data [ 8 ] found an agreement rate of about 90%. Completeness accuracy [ 23 ] is also commonly used to assess OSM data since it allows for the evaluation of territorial coverage. It has been used mainly to evaluate the completeness of OSM’s data on road networks [22,39,40], buildings [23,41,42], and LULC features [6,9,34]. Related Works Recent studies [ 43 – 45 ] have also included OSM contributors’ update history in data quality assessments. OSM full history file stores extra information associated with data contributions, such as timestamps that provide the exact date and time of contributions [ 39 – 41 , 46 ]. For example, one study [ 40 ] that assessed OSM’s road network completeness and positional accuracy using OSM data history concluded that the use of historical information improved the quality of OSM data by up to 14%. OSM data history can moreover be used for a number of different purposes [ 40 ] and allow for new perspectives on the reliability of OSM data at different scales and timeframes [ 22 , 40 ]. OSM data history is available since 2005. In the present study, OSM data history has been accessed to generate different LULC datasets based on the contribution year (using the timestamps) in order to create regional and rural multi-temporal LULC maps. This had two primary purposes. First, we sought to evaluate the degree to which OSM datasets bounded by one-year timeframes agreed with extant authoritative datasets (i.e., CLC and the official Portuguese Land Cover Map—COS), using both completeness and thematic accuracy as quality parameters. Second, we sought to evaluate whether the OSM datasets were of sufficient quality to be used as sampling data sources for multi-temporal LULC maps. For this second assessment, we used near boundary tag accuracy (NBTA) to evaluate the fitness of the OSM data for producing training samples, by looking at the extent to which a feature’s proximity to a boundary influenced its attribute (tag) accuracy.
ISPRS Int. J. Geo-Inf. 2019,8, 116 3 of 18 This research was conducted in the largest Portuguese district, Beja. It is a predominantly rural region with high natural and economic value, and is characterized by a mixed agro-silvo-pastoral ecosystem [ 47 ]. In addition, the region has recently undergone rapid LULC changes [ 47 – 49 ], and the limited number of references LULC data for this region emphasizes the importance of finding supplementary LULC data to support the identification and monitoring of these changes. Our research is discussed in the rest of this paper. Section 2describes the study area and our data. Section 3details our methodology. Section 4presents our OSM quality results. Finally, in Section 5we discuss the implications and main conclusions of our results. 2. Study Area and Data 2.1. Study Area Beja is a district located in the southeast of Portugal, in the Alentejo region (Figure 1). It is the largest Portuguese district, with an area of 10,229.05 km 2 , and as of 2011 it had a population of 152,758 residents [ 50 ]. Urban density is low, and agricultural and forest areas dominate. The southeastern part of the district is flat, while in the northern and western parts the extensive plains are intersected by tiny hills. The valley of the River Guadiana, which traverses the eastern part of the district in a north-south direction, is the district’s main geographical feature. ISPRS Int. J. Geo-Inf. 2018, 7, x FOR PEER REVIEW 3 of 18 ecosystem [47]. In addition, the region has recently undergone rapid LULC changes [47–49], and the limited number of references LULC data for this region emphasizes the importance of finding supplementary LULC data to support the identification and monitoring of these changes. Our research is discussed in the rest of this paper. Section 2 describes the study area and our data. Sections 3 details our methodology. Section 4 presents our OSM quality results. Finally, in Section 5 we discuss the implications and main conclusions of our results. 2. Study Area and Data 2.1. Study Area Beja is a district located in the southeast of Portugal, in the Alentejo region (Figure 1). It is the largest Portuguese district, with an area of 10,229.05 km², and as of 2011 it had a population of 152,758 residents [50]. Urban density is low, and agricultural and forest areas dominate. The southeastern part of the district is flat, while in the northern and western parts the extensive plains are intersected by tiny hills. The valley of the River Guadiana, which traverses the eastern part of the district in a north-south direction, is the district’s main geographical feature. Figure 1. Location of Beja District and its Municipalities. 2.2. Datasets 2.2.1. The OSM Dataset and History File OSM data can be freely downloaded, for example from Geofabrik or Planet OSM, in the form of raw datasets covering different regions (countries, continents, or any other administrative level) around the globe. OSM data represent physical features (objects) on the ground, and their tags (i.e., labels) are used to describe the objects (i.e., class description). The OSM data include a variety of physical feature types, with land use, natural features, waterways, amenities, and highways being the most commonly represented [51]. Feature descriptions can be found on the OSM wiki page (https://wiki.openstreetmap.org/wiki/Map_Features). OSM history file includes records of all historical contributions, including recent ones, and are accessible as either XML or PBF formatted files. They provide a history of every modification made to a geographical feature’s shape or tag. This means that if, for example, a feature’s shape has been changed once, there will then be two entries for the same feature—the original feature and the modified one. A sample illustration and more information can be found in Nasiri et al. [40]. In the present study, we accessed the history file to see the timestamps (day/month/year and time) of each feature’s creation and modifications. Figure 1. Location of Beja District and its Municipalities. 2.2. Datasets 2.2.1. The OSM Dataset and History File OSM data can be freely downloaded, for example from Geofabrik or Planet OSM, in the form of raw datasets covering different regions (countries, continents, or any other administrative level) around the globe. OSM data represent physical features (objects) on the ground, and their tags (i.e., labels) are used to describe the objects (i.e., class description). The OSM data include a variety of physical feature types, with land use, natural features, waterways, amenities, and highways being the most commonly represented [ 51 ]. Feature descriptions can be found on the OSM wiki page (https://wiki.openstreetmap.org/wiki/Map_Features). OSM history file includes records of all historical contributions, including recent ones, and are accessible as either XML or PBF formatted files. They provide a history of every modification made to a geographical feature’s shape or tag. This means that if, for example, a feature’s shape has been
ISPRS Int. J. Geo-Inf. 2019,8, 116 4 of 18 changed once, there will then be two entries for the same feature—the original feature and the modified one. A sample illustration and more information can be found in Nasiri et al. [ 40 ]. In the present study, we accessed the history file to see the timestamps (day/month/year and time) of each feature’s creation and modifications. 2.2.2. Official Reference Datasets The two reference LULC datasets used in this study were the 2012 CLC and the 2015 official Portuguese COS. The CLC is produced by the Portuguese General Directorate for Territorial Development (DGT) in coordination with the European Environment Agency, while the COS is produced exclusively by the DGT. Both datasets are freely available for download from the DGT website (http://mapas.dgterritorio.pt/geoportal/catalogo.html), and both use hierarchical and a priori nomenclature systems. The COS nomenclature was produced to match the CLC one. Thus, despite the fact that COS has five disaggregation levels compared to the CLC’s three, the COS’s first three levels are similar to the three CLC levels, thereby enabling comparisons between them. The dataset characteristics and metadata are shown in Table 1, and their descriptive statistics are shown in Table 2. More details about their nomenclature characteristics can be seen in Estima and Painho [7]. Table 1. Characteristics of the CORINE Land Cover (CLC) and the official Portuguese Land Cover Map (COS) datasets. Characteristics CORINE Land Cover (CLC) Portuguese Land Cover Map (COS) Data model Vector Spatial representation Polygons Nomenclature Hierarchical (3 levels—44 classes) Hierarchical (5 levels—225 classes) Scale 1:100,000 1:25,000 Spatial resolution 20 m 0.5 m Minimum Mapping Unit (MMU) 25 ha 1 ha Minimum distance between lines 100 m 20 m Base data Satellite images Air-photo maps Production method Semi-automated production and visual interpretation Visual interpretation Table 2. Descriptive statistics of reference datasets (area in ha). Beja District CLC (2012) COS (2015) Total polygons 11,306 34,793 Minimum area of polygons 0.01 0.01 Maximum area of polygons 107,667.31 26,542.24 Mean area of polygons 90,78 29.50 Standard Deviation area 1,065.51 280.20 3. Methods Figure 2shows our main steps. Briefly, our first step was to download the OSM data history file and filter the contributions according to their timestamps. We obtained datasets for seven different years and resolved all logical inconsistencies, such as overlapping features. Second, we established a relationship between the OSM and CLC/COS first level of nomenclature. Third, we intersected the datasets to determine the area corresponding to the OSM dataset that matched the reference dataset. Fourth, we calculated the datasets’ completeness and thematic accuracies for 2012 and 2015. Fifth, random points were generated as training samples. Finally, in step six we calculated the NBTA for 2015.
ISPRS Int. J. Geo-Inf. 2019,8, 116 5 of 18 ISPRS Int. J. Geo-Inf. 2018, 7, x FOR PEER REVIEW 5 of 18 Figure 2. Workflow representing the proposed methodology. 3.1. Processing the OSM Datasets We began by downloading from Geofabrik the latest OSM history file available at the time (May 7, 2018) that covered Portugal. All OSM objects with the feature types of land use, natural features, airways, amenities, buildings, highways, historic features, leisure features, man-made features, power structures, public transportation, railroads, shops, sports, tourism, forests, and waterways were retrieved. We limited our results to features (polygons) found in the Beja district. By filtering the features by contribution year (using the timestamp data), we generated multi-temporal datasets that included only the features which were created or modified during any given year. We found contributions for seven different years (2011–2017). Since reference data for the study area only exists for 2012 (CLC) and 2015 (COS), we identified and selected all the OSM data contributions that had been created/modified in these two years (2012 and 2015) and used these two OSM datasets in our study for comparison with the reference datasets. A common problem when processing OSM data is the presence of logical inconsistencies, such as overlapping features [6,9,34,35]. This problem is even more common when history files are used, since they contain records of every single modification. It is therefore essential to dissolve or remove the overlapping features in order to not overestimate areas and to ensure that only one attribute (tag) is kept for analysis [6,9,35]. To ensure logical consistency between the annual datasets, when overlapped features had the same attribute we dissolved the overlapping areas, while when the overlapped features’ attributes did not match the conflict areas were removed. In addition, we examined the provenance of OSM data for our study area to confirm that there were no bulk imports from known sources (e.g., CLC; COS). In this study the polygons´ geometry for each reference dataset and both OSM datasets was compared by using the Feature Compare tool of ArcGIS 10.5 software. 3.2. Relationship Between Dataset Nomenclatures Several studies have stressed the difficulties of using datasets from different agencies due to the lack of direct relationships between their classes [6,9,35,52]. In Estima and Painho [6] , the authors attempted to reconcile the three nomenclature levels of the CLC, the OSM land use feature, and the OSM natural areas feature, based on the official descriptions of the CLC and OSM classes. Given their remarkable results, we decided to partially follow their nomenclature correspondence for the first level of the CLC and COS nomenclature, namely, (1) artificial surfaces, (2) agricultural areas, (3) forests, (4) wetlands, and (5) water (see Table 3). For the purposes of the study, the OSM features type airways, amenities, buildings, highways, historic features, leisure features, man-made features, power structures, public transportation, railroads, shops, sports, and tourism were considered to be Figure 2. Workflow representing the proposed methodology. 3.1. Processing the OSM Datasets We began by downloading from Geofabrik the latest OSM history file available at the time (7 May 2018) that covered Portugal. All OSM objects with the feature types of land use, natural features, airways, amenities, buildings, highways, historic features, leisure features, man-made features, power structures, public transportation, railroads, shops, sports, tourism, forests, and waterways were retrieved. We limited our results to features (polygons) found in the Beja district. By filtering the features by contribution year (using the timestamp data), we generated multi-temporal datasets that included only the features which were created or modified during any given year. We found contributions for seven different years (2011–2017). Since reference data for the study area only exists for 2012 (CLC) and 2015 (COS), we identified and selected all the OSM data contributions that had been created/modified in these two years (2012 and 2015) and used these two OSM datasets in our study for comparison with the reference datasets. A common problem when processing OSM data is the presence of logical inconsistencies, such as overlapping features [ 6 , 9 , 34 , 35 ]. This problem is even more common when history files are used, since they contain records of every single modification. It is therefore essential to dissolve or remove the overlapping features in order to not overestimate areas and to ensure that only one attribute (tag) is kept for analysis [ 6 , 9 , 35 ]. To ensure logical consistency between the annual datasets, when overlapped features had the same attribute we dissolved the overlapping areas, while when the overlapped features’ attributes did not match the conflict areas were removed. In addition, we examined the provenance of OSM data for our study area to confirm that there were no bulk imports from known sources (e.g., CLC; COS). In this study the polygons ´ geometry for each reference dataset and both OSM datasets was compared by using the Feature Compare tool of ArcGIS 10.5 software. 3.2. Relationship between Dataset Nomenclatures Several studies have stressed the difficulties of using datasets from different agencies due to the lack of direct relationships between their classes [ 6 , 9 , 35 , 52 ]. In Estima and Painho [ 6 ], the authors attempted to reconcile the three nomenclature levels of the CLC, the OSM land use feature, and the OSM natural areas feature, based on the official descriptions of the CLC and OSM classes. Given their remarkable results, we decided to partially follow their nomenclature correspondence for the first level of the CLC and COS nomenclature, namely, (1) artificial surfaces, (2) agricultural areas, (3) forests,
ISPRS Int. J. Geo-Inf. 2019,8, 116 6 of 18 (4) wetlands, and (5) water (see Table 3). For the purposes of the study, the OSM features type airways, amenities, buildings, highways, historic features, leisure features, man-made features, power structures, public transportation, railroads, shops, sports, and tourism were considered to be (1) artificial surfaces. The OSM feature type wood was classed as (3) forest, while the OSM feature type waterway was classed as (5) water. Table 3. Nomenclature correspondence for the first level of the CLC and COS nomenclature. Landuse feature type OSM tag CLC/COS ( Level 1) OSM tag CLC/COS ( Level 1) Abutters 1 Harbour 1 Allotments 2 Industrial 1 Basin 5 Landfill 1 Beach 3 Leisure 1 Brownfield 1 Meadow 2 Cemetery 1 Military ? Commercial 1 Museum 1 Conservation 3 Not_known ? Construction 1 Orchard 2 Farm 2 Park 1 Farmland 2 Public 1 Farmyard 2 Quarry 1 Garages 1 Railway 1 Garden 1 Recreation_groun 1 Grass 2 Reservoir 5 Greenfield 3 Retail 1 Greenhouse 2 Salt_pond 4 Greenhouse_horti 2 Scrub 3 Residential 1 Scrubs 3 University 1 Vineyard 2 Village_green 1 Waste_water_plan 1 Wood 3 Water 5 Natural feature type OSM tag CLC/COS ( Level 1) OSM tag CLC/COS ( Level 1) Grassland 2 Fell 3 Scrub 3 Bare_rock 3 Wood 3 Park 3 Scree 3 Forest 3 Beach 3 Wetland 4 Sand 3 Water 5 Rock 3 Riverbank 5 3.3. Accuracy Assessment Criteria 3.3.1. OSM Completeness Accuracy It can be difficult to draw definitive conclusions from OSM data due to the heterogeneity of contributions across different regions [ 23 ]. This challenge poses particular limitations when analyzing rural areas [ 16 , 53 ]. Accordingly, the completeness accuracy criterion is commonly used to assess the quality of OSM data [ 22 , 34 , 35 ]. Completeness accuracy is defined as the completeness of a dataset and measures the presence or absence of features. Typically, the total number of features is computed for point features, while for line features the total length is computed. These calculations are then compared to the reference dataset [ 23 ]. However, for polygon features, it is assumed that the overall area of the study area is the maximum area possible, and it is therefore not necessary to compare it to reference data [35]. In this study, completeness accuracy was calculated as a ratio of the OSM features’ overall area (A OSM ) and the total area of the study area (A Ref ). It is presented here as a percentage, as described
ISPRS Int. J. Geo-Inf. 2019,8, 116 7 of 18 in Equation 1. A ratio of 100 means that the OSM dataset provides full coverage of the study area. This criterion was measured for all OSM dataset features for 2012 and 2015 for each LULC class. Completeness = AOSM ARef ×100 (1) 3.3.2. OSM Thematic Accuracy Thematic accuracy is another common criterion used to evaluate the quality of OSM dataset features in LULC mapping [ 6 – 8 , 21 ]. Thematic accuracy describes the accuracy of the features’ attributes (tags) by computing the differences between OSM dataset features and those of the chosen reference datasets [35]. In this study, we used an overlap function to determine which features overlapped between the OSM datasets (2012 and 2015) and the reference datasets (CLC and COS, respectively). Following common statistical approaches [ 35 , 54 ], for each annual dataset the overlapping areas were computed in a confusion matrix, which presented the correct and incorrect mapped areas for each LULC class. The rows denote the occurrences of an actual class (OSM) and the columns denote the occurrences of reference data classes (CLC or COS). Several measures were obtained from this, including overall thematic accuracy, individual user accuracy, producer accuracy, and the Kappa index of agreement [ 34 , 35 , 55 ]. The overall thematic accuracy measure provides the overall percentage of the correctly mapped OSM features by dividing the correct mapped area by the total mapped area. User accuracy measures the probability that any given LULC class from an OSM dataset will actually match the reference dataset, while producer accuracy indicates the probability that a particular LULC class from the reference dataset is classified as such in OSM [ 35 ]. The Kappa index measures the degree of agreement between the OSM dataset and the reference dataset [55]. 3.3.3. Near Boundary Tag Accuracy (NBTA) We also used NBTA to measure the fitness of the OSM data as a source of sampling data to support regional and rural LULC mapping. NBTA measures the extent to which the proximity of an OSM feature’s boundary influences the accuracy of the attribute (tag) in the training sample. We computed the NBTA for the most recent OSM dataset (2015). First, it was necessary to create training samples by generating random points inside the OSM features. Since OSM feature areas are very heterogeneous, it would be inappropriate specify an exact number of random points to be generated inside each OSM feature; instead, the total number of random points generated was proportional to the area of each OSM feature. In addition, the shortest distance allowed between any two randomly placed points was 30 m, because this is the most common spatial resolution of satellite images, such as Landsat (Figure 3). We used point features instead of polygon features since this minimized the effects of comparing data with different mapping scales. Each random point generated took the attribute of the corresponding OSM polygon feature. ISPRS Int. J. Geo-Inf. 2018, 7, x FOR PEER REVIEW 7 of 18 3.3.2. OSM Thematic Accuracy Thematic accuracy is another common criterion used to evaluate the quality of OSM dataset features in LULC mapping [6–8,21]. Thematic accuracy describes the accuracy of the features’ attributes (tags) by computing the differences between OSM dataset features and those of the chosen reference datasets [35]. In this study, we used an overlap function to determine which features overlapped between the OSM datasets (2012 and 2015) and the reference datasets (CLC and COS, respectively). Following common statistical approaches [35,54], for each annual dataset the overlapping areas were computed in a confusion matrix, which presented the correct and incorrect mapped areas for each LULC class. The rows denote the occurrences of an actual class (OSM) and the columns denote the occurrences of reference data classes (CLC or COS). Several measures were obtained from this, including overall thematic accuracy, individual user accuracy, producer accuracy, and the Kappa index of agreement [34,35,55]. The overall thematic accuracy measure provides the overall percentage of the correctly mapped OSM features by dividing the correct mapped area by the total mapped area. User accuracy measures the probability that any given LULC class from an OSM dataset will actually match the reference dataset, while producer accuracy indicates the probability that a particular LULC class from the reference dataset is classified as such in OSM [35]. The Kappa index measures the degree of agreement between the OSM dataset and the reference dataset [55]. 3.3.3. Near Boundary Tag Accuracy (NBTA) We also used NBTA to measure the fitness of the OSM data as a source of sampling data to support regional and rural LULC mapping. NBTA measures the extent to which the proximity of an OSM feature’s boundary influences the accuracy of the attribute (tag) in the training sample. We computed the NBTA for the most recent OSM dataset (2015). First, it was necessary to create training samples by generating random points inside the OSM features. Since OSM feature areas are very heterogeneous, it would be inappropriate specify an exact number of random points to be generated inside each OSM feature; instead, the total number of random points generated was proportional to the area of each OSM feature. In addition, the shortest distance allowed between any two randomly placed points was 30 m, because this is the most common spatial resolution of satellite images, such as Landsat (Figure 3). We used point features instead of polygon features since this minimized the effects of comparing data with different mapping scales. Each random point generated took the attribute of the corresponding OSM polygon feature. Figure 3. Example of random points distribution. Second, thematic accuracy was assessed following the same procedure described in Section 3.3.2. Also, the Euclidean distance from each random point to the nearest segment of the OSM feature’s boundary was computed. The distance values (in ascending order) and the accumulated thematic accuracy were then plotted. The purpose of this process was to ascertain whether the training samples generated for each LULC class near the border of an OSM feature might be inherently less accurate. However, the relationship between the training samples accuracy and their proximity to the corresponding feature boundaries did not consider the influence of each feature area in explaining the degree of accuracy, since the distance to a feature boundary varies across feature areas. Are features with larger or smaller areas likely to have more or less accuracy? Therefore, the distance Figure 3. Example of random points distribution. Second, thematic accuracy was assessed following the same procedure described in Section 3.3.2. Also, the Euclidean distance from each random point to the nearest segment of the OSM feature’s
ISPRS Int. J. Geo-Inf. 2019,8, 116 8 of 18 boundary was computed. The distance values (in ascending order) and the accumulated thematic accuracy were then plotted. The purpose of this process was to ascertain whether the training samples generated for each LULC class near the border of an OSM feature might be inherently less accurate. However, the relationship between the training samples accuracy and their proximity to the corresponding feature boundaries did not consider the influence of each feature area in explaining the degree of accuracy, since the distance to a feature boundary varies across feature areas. Are features with larger or smaller areas likely to have more or less accuracy? Therefore, the distance values from all training samples were standardized using the maximum and minimum distance values of the corresponding OSM feature. Similarly, the accumulated accuracy of all training samples was computed by considering each OSM feature. The relationship between a feature’s boundary proximity and its thematic accuracy is denoted by the Pearson correlation coefficient (R), which was calculated separately for each OSM feature. OSM features with fewer than three training samples were excluded from the analysis (20% of all polygons, for 2015), because when n= 2 the Rcoefficient would always be − 1 or +1, except in the improbable circumstance that both y-values perfectly matched. 4. Results 4.1. OSM Dataset Analysis The identification and selection of OSM data contributed in 2012 and 2015 allowed us to create two different datasets for analysis. However, these datasets presented some logical inconsistencies that had to be resolved. There were 994 and 2,100 non-overlapping features in 2012 and 2015, respectively. After dissolving or removing the overlapping features following the procedure described in Section 3.1, the final datasets were comprised of 1215 and 2863 distinct features in 2012 and 2015, respectively. The descriptive statistics for both years are presented in Table 4. It is notable that the total area covered increased significantly between 2012 and 2015 (+155%). In addition, between 2012 and 2015 the features’ minimum area doubled, while their maximum area increased by 2.5 times. As expected, the standard deviation measure also confirmed high variations between the features’ areas for each annual dataset, as most the natural/anthropological features fit power law (like) distributions. Table 4. Descriptive statistics for 2012 and 2015 OpenStreetMap (OSM) datasets. OSM 2012 2015 Total polygons 1215 2863 Minimum area of features (ha) 0.0003 0.0006 Maximum area of features (ha) 396.45 1000.75 Mean area of features (ha) 3.55 3.84 Standard Deviation area (ha) 17.27 32.97 Total area (ha) 4314.88 10,986.02 In addition, the statistical analysis shows that the 98% of the features in 2012 OSM dataset and the 97% of the features in 2015 OSM dataset have less than 25 ha and the 76.13% of the features in 2012 OSM dataset, as well as the 79.9% of the features in 2015 OSM dataset have less than 1 ha. In this way we are able to verify that the OSM used data are not imported from the reference datasets since the CLC minimum area unit is 25 ha and COS minimum area unit is 1 ha. When we compared the geometry of the 2012 and 2015 OSM features with the CLC and COS datasets, less than 0.1% matching was obtained, proving that no bulk imports were made from these reference data. 4.2. OSM Completeness Accuracy The completeness ratio of the coverage area for the selected study area has been increasing year after year (Table 5). However, in 2012, contributions covered less than 1% of the total district area,
ISPRS Int. J. Geo-Inf. 2019,8, 116 9 of 18 while in 2015 contributions covered about 1%. Following Table 6, Water was the largest LULC class in both annual OSM datasets (45% in 2012 and 46% in 2015). In 2012, agricultural areas represented about 40% of the total dataset, followed by artificial surfaces (16%), while in 2015 both classes were around 20%. In contrast, in 2012 forests were not significantly represented, but in 2015 they accounted for 14% of the total dataset. It is also interesting to note that in 2012 there were only four LULC classes represented (artificial surfaces, agricultural areas, forests, and water), while in 2015 there were five, due to the additional presence of wetlands. However, wetlands accounted for less than 0.001% of total coverage. In addition, the relationship between the total features’ contributions and the corresponding area for each class demonstrated that both agricultural areas and water areas had consistently huge polygons drawn by contributors. Table 5. OSM datasets completeness. Completeness (%) 2011 2012 2013 2014 2015 2016 2017 Total 0.07 0.43 0.81 0.90 1.07 2.8 5.9 Table 6. Descriptive statistics per class of 2012 and 2015 OSM datasets. LULC Class Total Features Area (ha) OSM Class Coverage (%) Completeness (%) 2012 2015 2012 2015 2012 2015 2012 2015 Artificial Surfaces 800 1612 672.42 2187.47 15.58 19.91 0.07 0.21 Agricultural areas 152 392 1705.71 2145.09 39.53 19.53 0.17 0.21 Forest 5 95 5.07 1543.57 0.12 14.05 0 0.15 Wetlands 0 19 0 17.84 0 0.16 0 0 Water 258 745 1931.68 5092.05 44.77 46.35 0.19 0.50 Total 1215 2863 4314.88 10,986.02 100 100 0.43 1.07 4.3. OSM Thematic Accuracy A confusion matrix was computed for each annual dataset (Tables 7and 8) in order to evaluate its thematic accuracy. The results showed that thematic accuracy was 77.3% in 2012 and 91.9% in 2015, indicating a very high agreement (i.e., accuracy) between the 2015 OSM dataset and the reference dataset (COS). However, there was a wide variation among the accuracy values for each class. In particular, agricultural areas were highly accurate, with a 99.8% user accuracy rate in 2012 and a 94.5% user accuracy rate in 2015, suggesting that the areas classified by contributors as agricultural areas closely matched those in the reference datasets (CLC and COS, respectively). While the producer’s accuracy rates were not quite as high (71.9% in 2012 and 79.1% in 2015), they were still high enough to suggest that this class was correctly shown on the OSM. Artificial surfaces also had consistently high accuracy values; in 2012, the user and producer accuracy rates were 70.9% and 75.5%, respectively, and in 2015 they increased to 80.9% and 95.3%, respectively. Other classes showed a strong increase in user accuracy between 2012 and 2015. User accuracy for forests jumped from 0% in 2012 to more than 97% in 2015, and user accuracy for water increased from about 60% in 2012 to over 94% in 2015. In contrast, producer accuracy for water was very high in both years (about 99.7%). Producer accuracy for forests improved from 0% in 2012 to over 88% in 2015. However, wetlands had null values for both user and producer accuracy in both years (0%). Overall, following the standards set by [ 55 ], the Kappa index indicated substantial agreement between the OSM and reference maps in 2012 (0.65) and an almost perfect agreement in 2015 (0.88).
ISPRS Int. J. Geo-Inf. 2019,8, 116 16 of 18 References 1. Fritz, S.; Bartholomé, E.; Belward, A.; Hartley, A.; Stibig, H.-J.; Eva, H.; Mayaux, P.; Bartalev, S.; Latifovic, R.; Kolmert, S.; et al. Harmonisation, Mosaicing and Production of the Global Land Cover 2000 Database (Beta Version); EC-JRC: Brussels, Belgium, 2003. 2. Büttner, G.; Feranec, J. The CORINE Land Cover Update 2000. Techinical Guidelines; EEA Technical Report, 89; EC-JRC: Copenhagen, Denmark, 2002. 3. Goodchild, M.F. Citizens as sensors: The world of volunteered geography. GeoJournal 2007 ,69, 211–221. [CrossRef] 4. Estima, J.; Painho, M. Flickr geotagged and publicly available photos: Preliminary study of its adequacy for helping quality control of Corine Land Cover. Lect. Notes Comput. Sci. 2013,7974, 205–220. [CrossRef] 5. See, L.; Mooney, P.; Foody, G.; Bastin, L.; Comber, A.; Estima, J.; Fritz, S.; Kerle, N.; Jiang, B.; Laakso, M.; et al. Crowdsourcing, Citizen Science or Volunteered Geographic Information? The Current State of Crowdsourced Geographic Information. ISPRS Int. J. Geo-Inf. 2016,5, 55. [CrossRef] 6. Estima, J.; Painho, M. Exploratory analysis of OpenStreetMap for land use classification. In Proceedings of the Second ACM SIGSPATIAL International Workshop on Crowdsourced and Volunteered Geographic Information, Orlando, FL, USA, 5 November 2013; pp. 39–46. [CrossRef] 7. Estima, J.; Painho, M. Investigating the Potential of OpenStreetMap for Land Use/Land Cover Production: A Case Study for Continental Portugal. In OpenStreetMap in GIScience; Arsanjani, J.J., Zipf, A., Mooney, P., Helbich, M., Eds.; Springer: Cham, Switzerland, 2015; pp. 273–293, ISBN 978-3-319-14279-1. 8. Fonte, C.C.; Martinho, N. Assessing the applicability of OpenStreetMap data to assist the validation of land use/land cover maps. Int. J. Geogr. Inf. Sci. 2017,31, 2382–2400. [CrossRef] 9. Jokar Arsanjani, J.; Vaz, E. An assessment of a collaborative mapping approach for exploring land use patterns for several European metropolises. Int. J. Appl. Earth Obs. Geoinf. 2015,35, 329–337. [CrossRef] 10. Arsanjani, J.J.; Mooney, P.; Zipf, A.; Schauss, A. Quality Assessment of the Contributed Land Use Information from OpenStreetMap Versus Authoritative Datasets. In OpenStreetMap in GIScience: Experiences, Research, and Applications; Springer International Publishing: Cham, Switzerland, 2015; Chapter 3; pp. 37–58. 11. Fonte, C.C.; Patriarca, J.A.; Minghini, M.; Antoniou, V.; See, L.; Brovelli, M.A. Using OpenStreetMap to Create Land Use and Land Cover Maps: Development of an Application. In Volunteered Geographic Information and the Future of Geospatial Data; Campelo, C.E.C., Bertolotto, M., Corcoran, P., Eds.; IGI Global: Hershey, PA, USA, 2017; Volume i, pp. 113–137, ISBN 9781522524465. 12. Estima, J.; Painho, M. Photo Based Volunteered Geographic Information Initiatives. Int. J. Agric. Environ. Inf. Syst. 2014,5, 73–89. [CrossRef] 13. See, L.; Estima, J.; P˝odör, A.; Jokar, J.; Laso-Bayas, J.; Vatseva, R. Sources of VGI for Mapping. In Mapping and the Citizen Sensor; Foody, G., See, L., Fritz, S., Mooney, P., Olteanu-Raimond, A.-M., Fonte, C.C., Antoniou, V., Eds.; Ubiquity Press: London, UK, 2016; pp. 13–35, ISBN 9781911529163. 14. Sui, D.; Goodchild, M.; Elwood, S. Volunteered Geographic Information, the Exaflood, and the Growing Digital Divide BT. In Crowdsourcing Geographic Knowledge: Volunteered Geographic Information (VGI) in Theory and Practice; Sui, D., Elwood, S., Goodchild, M., Eds.; Springer: Dordrecht, The Netherlands, 2013; pp. 1–12, ISBN 978-94-007-4587-2. 15. Goodchild, M.F.; Li, L. Assuring the quality of volunteered geographic information. Spat. Stat. 2012 ,1, 110–120. [CrossRef] 16. Haklay, M. How good is volunteered geographical information? A comparative study of OpenStreetMap and ordnance survey datasets. Environ. Plan. B Plan. Des. 2010,37, 682–703. [CrossRef] 17. Mooney, P.; Corcoran, P. Analysis of interaction and co-editing patterns amongst openstreetmap contributors. Trans. GIS 2014,18, 633–659. [CrossRef] 18. Zook, M.; Graham, M.; Shelton, T.; Gorman, S. Volunteered Geographic Information and Crowdsourcing Disaster Relief: A Case Study of the Haitian Earthquake. World Med. Heal. Policy 2010,2, 7–33. [CrossRef] 19. Yang, D.; Fu, C.S.; Smith, A.C.; Yu, Q. Open land-use map: A regional land-use mapping strategy for incorporating OpenStreetMap with earth observations. Geo-Spat. Inf. Sci. 2017,20, 269–281. [CrossRef] 20. Johnson, B.A.; Iizuka, K. Integrating OpenStreetMap crowdsourced data and Landsat time-series imagery for rapid land use/land cover (LULC) mapping: Case study of the Laguna de Bay area of the Philippines. Appl. Geogr. 2016,67, 140–149. [CrossRef]
ISPRS Int. J. Geo-Inf. 2019,8, 116 17 of 18 21. Jokar Arsanjani, J.; Helbich, M.; Bakillah, M.; Hagenauer, J.; Zipf, A. Toward mapping land-use patterns from volunteered geographic information. Int. J. Geogr. Inf. Sci. 2013,27, 2264–2278. [CrossRef] 22. Neis, P.; Zielstra, D.; Zipf, A. The Street Network Evolution of Crowdsourced Maps: OpenStreetMap in Germany 2007–2011. Future Internet 2012,4, 1–21. [CrossRef] 23. Hecht, R.; Kunze, C.; Hahmann, S.; Hecht, R.; Kunze, C.; Hahmann, S. Measuring Completeness of Building Footprints in OpenStreetMap over Space and Time. ISPRS Int. J. Geo-Inf. 2013,2, 1066–1091. [CrossRef] 24. Helbich, M.; Amelunxen, C.; Neis, P.; Zipf, A. Comparative Spatial Analysis of Positional Accuracy of OpenStreetMap and Proprietary Geodata. Proc. GI_Forum 2012, 24–33. [CrossRef] 25. Mooney, P.; Corcoran, P.; Mooney, P.; Corcoran, P. Characteristics of Heavily Edited Objects in OpenStreetMap. Future Internet 2012,4, 285–305. [CrossRef] 26. See, L.; Comber, A.; Salk, C.; Fritz, S.; van der Velde, M.; Perger, C.; Schill, C.; McCallum, I.; Kraxner, F.; Obersteiner, M. Comparing the Quality of Crowdsourced Data Contributed by Expert and Non-Experts. PLoS ONE 2013,8, 1–11. [CrossRef] [PubMed] 27. Fonte, C.C.; Bastin, L.; Foody, G.; Kellenberger, T.; Kerle, N.; Mooney, P.; Olteanu-Raimond, A.-M.; See, L. Vgi Quality Control. In ISPRS Annals of Photogrammetry, Remote Sensing and Spatial Information Sciences; ISPR: La Grande Motte, France, 2015; Volume II-3/W5, pp. 317–324. 28. Comber, A.; Brunsdon, V.; See, L.; Fritz, S. Evaluating Global Land Cover Datasets: Comparing VGI on Cropland with Formal Data. In Proceedings of the GI_Forum 2013. Creat. GISociety, Salzburg, Austria, 2–5 July 2013; pp. 91–95. [CrossRef] 29. Almendros-Jiménez, J.; Becerra-Terón, A. Analyzing the Tagging Quality of the Spanish OpenStreetMap. ISPRS Int. J. Geo-Inf. 2018,7, 323. [CrossRef] 30. Tracewski, L.; Bastin, L.; Fonte, C.C. Repurposing a deep learning network to filter and classify volunteered photographs for land cover and land use characterization. Geo-Spat. Inf. Sci. 2017,20, 252–268. [CrossRef] 31. Antoniou, V.; Skopeliti, A. Measures and indicators of VGI quality: An overview. In ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences; ISPR: La Grande Motte, France, 2015; pp. 345–351. 32. Senaratne, H.; Mobasheri, A.; Ali, A.L.; Capineri, C.; Haklay, M. (Muki) A review of volunteered geographic information quality assessment methods. Int. J. Geogr. Inf. Sci. 2017,31, 139–167. [CrossRef] 33. Dorn, H.; Törnros, T.; Zipf, A.; Dorn, H.; Törnros, T.; Zipf, A. Quality Evaluation of VGI Using Authoritative Data—A Comparison with Land Use Data in Southern Germany. ISPRS Int. J. Geo-Inf. 2015 ,4, 1657–1671. [CrossRef] 34. Arsanjani, J.J.; Fonte, C.C. On the Contribution of Volunteered Geographic Information to Land Monitoring Efforts. In European Handbook of Crowdsourced Geographic Information; Ubiquity Press: London, UK, 2016; pp. 269–284. 35. Jokar Arsanjani, J.; Mooney, P.; Zipf, A.; Schauss, A. Quality Assessment of the Contributed Land Use Information from OpenStreetMap Versus Authoritative Datasets. In OpenStreetMap in GIScience; Springer: Cham, Switzerland, 2015; pp. 37–58. 36. Guptill, S.C.; Morrison, J.L.; International Cartographic Association. Elements of Spatial Data Quality; Elsevier Science: Amsterdam, The Netherlands, 1995; ISBN 9780080424323. 37. Foody, G.M.; See, L.; Fritz, S.; Van der Velde, M.; Perger, C.; Schill, C.; Boyd, D.S. Assessing the Accuracy of Volunteered Geographic Information arising from Multiple Contributors to an Internet Based Collaborative Project. Trans. GIS 2013,17, 847–860. [CrossRef] 38. Yagoub, M.M. Assessment of OpenStreetMap (OSM) Data: The Case of Abu Dhabi City, United Arab Emirates. J. Map Geogr. Libr. 2017,13, 300–319. [CrossRef] 39. Barron, C.; Neis, P.; Zipf, A. A Comprehensive Framework for Intrinsic OpenStreetMap Quality Analysis. Trans. GIS 2014,18, 877–895. [CrossRef] 40. Nasiri, A.; Ali Abbaspour, R.; Chehreghan, A.; Jokar Arsanjani, J. Improving the Quality of Citizen Contributed Geodata through Their Historical Contributions: The Case of the Road Network in OpenStreetMap. ISPRS Int. J. Geo-Inf. 2018,7, 253. [CrossRef] 41. Rehrl, K.; Gröchenig, S.; Rehrl, K.; Gröchenig, S. A Framework for Data-Centric Analysis of Mapping Activity in the Context of Volunteered Geographic Information. ISPRS Int. J. Geo-Inf. 2016,5, 37. [CrossRef] 42. Brovelli, M.; Zamboni, G. A New Method for the Assessment of Spatial Accuracy and Completeness of OpenStreetMap Building Footprints. ISPRS Int. J. Geo-Inf. 2018,7, 289. [CrossRef]
ISPRS Int. J. Geo-Inf. 2019,8, 116 18 of 18 43. Antoniou, V.; Touya, G.; Raimond, A.-M. Quality analysis of the Parisian OSM toponyms evolution. In European Handbook of Crowdsourced Geographic Information; Ubiquity Press: London, UK, 2016; pp. 97–112. 44. Jonietz, D.; Zipf, A. Defining Fitness-for-Use for Crowdsourced Points of Interest (POI). ISPRS Int. J. Geo-Inf. 2016,5, 149. [CrossRef] 45. Sehra, S.; Singh, J.; Rai, H.; Sehra, S.S.; Singh, J.; Rai, H.S. Assessing OpenStreetMap Data Using Intrinsic Quality Indicators: An Extension to the QGIS Processing Toolbox. Future Internet 2017,9, 15. [CrossRef] 46. Keßler, C.; Trame, J.; Kauppinen, T. Tracking Editing Processes in Volunteered Geographic Information: The Case of OpenStreetMap. In Proceedings of the Workshop on Icentifying Objects, Processes and Events in Spatio-Temporally Distributed Data (IOPE 2011), Workshop at COSIT 2011, Belfast, Maine, 12–16 September 2011. 47. Viana, C.M.; Rocha, J. Spatiotemporal analysis and scenario simulation of agricultural land use land cover using GIS and a Markov chain model. In Geospatial Technologies for All: Short Papers, Posters and Poster Abstracts of the 21th AGILE Conference on Geographic Information Science; Mansourian, A., Pilesjö, P., Harrie, L., von Lammeren, R., Eds.; Lund University: Lund, Sweden, 12–15 June 2018. 48. Allen, H.; Simonson, W.; Parham, E.; de Basto E Santos, E.; Hotham, P. Satellite remote sensing of land cover change in a mixed agro-silvo-pastoral landscape in the Alentejo, Portugal. Int. J. Remote Sens. 2018 ,39, 4663–4683. [CrossRef] 49. Viana, C.M.; Oliveira, S.; Oliveira, S.C.; Rocha, J. Land Use/Land Cover Change Detection and Urban Sprawl Analysis. In Spatial Modeling in GIS and R for Earth and Environmental Sciences; Pourghasemi, H.R., Gokceoglu, C., Eds.; Elsevier: Amsterdam, The Netherlands, 2019; pp. 621–651, ISBN 9780128152263. 50. INE. Censos 2011 Resultados Definitivos—Região Alentejo; Instituto Nacional de Estatística: Lisboa, Portugal, 2012; ISBN 978-989-25-0182-6. 51. Mooney, P.; Corcoran, P. Using OSM for LBS–An analysis of changes to attributes of spatial objects. In Advances in Location-Based Services, Lecturenotes in Geoinformation and Cartography; Ortag, G., Gartner, F., Eds.; Springer-Verlag: Berlin/Heidelberg, Germany, 2012; pp. 165–179. 52. Haack, B.; Mahabir, R.; Kerkering, J. Remote sensing-derived national land cover land use maps: A comparison for Malawi. Geocarto Int. 2015,30, 270–292. [CrossRef] 53. Zielstra, D.; Zipf, A. A Comparative Study of Proprietary Geodata and Volunteered Geographic Information for Germany. In Proceedings of the 13th AGILE International Conference on Geographic Information Science, Guimarães, Portugal, 11–14 May 2010; pp. 1–15. 54. Herold, M.; Mayaux, P.; Woodcock, C.E.; Baccini, A.; Schmullius, C. Some challenges in global land cover mapping: An assessment of agreement and accuracy in existing 1 km datasets. Remote Sens. Environ. 2008 , 112, 2538–2556. [CrossRef] 55. Landis, J.R.; Koch, G.G. Measurement of observer agreement for categorical data. Biometrics 1977 ,33, 159–174. [CrossRef] [PubMed] 56. Meneses, B.; Reis, E.; Reis, R.; Vale, M.; Meneses, B.M.; Reis, E.; Reis, R.; Vale, M.J. The Effects of Land Use and Land Cover Geoinformation Raster Generalization in the Analysis of LUCC in Portugal. ISPRS Int. J. Geo-Inf. 2018,7, 390. [CrossRef] 57. Meneses, B.M.; Reis, E.; Vale, M.J.; Reis, R. Modelling the Land Use and Land cover changes in Portugal: A multi-scale and multi-temporal approach. Finisterra 2018,53. [CrossRef] 58. García-Álvarez, D. The Influence of Scale in LULC Modeling. A Comparison Between Two Different LULC Maps (SIOSE and CORINE). In Geomatic Simulations and Scenarios for Modelling LUCC. A Review and Comparison of Modelling Techniques; Camacho Olmedo, M.T., Paegelow, M., Mas, J.F., Escobar, F., Eds.; Springer: Cham, Switzerland, 2018; pp. 187–213. © 2019 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).