Full text
Spatial model of land value for the Brazilian territory with 1km resolution Francisco d’Albertas*, Diogo Rocha, Eduardo Lacerda, Julia Niemeyer, Luiz Gustavo Oliveira, Renata Capelão Bruna Fatiche Pavani Juliana Almeida-Rocha Rafael Loyola 2024-08-05 Authors’ affiliation: International Institute for Sustainability Leading author email: *[email protected] Institutional email address for correspondence: [email protected] 1 Introduction The comprehensive assessment of land value has become increasingly pivotal for understanding land-use dynamics and informing sustainable land management strategies, particularly in the context of prioritization for conservation and restoration sciences (e.g. Strassburg et al. 2020; Strassburg et al. 2019 ) . Despite its importance, accurate information on land value is rarely available at finer scales or in formats easily usable in research, such as raster datasets. Since 2019, the Brazilian Government has annually published land value data (reais ha−1) for a substantial number of municipalities and States across all regions of the country1. The provided information encompasses the following land use/cover categories: cropland with good, regular, and restricted suitability; planted pastureland; silviculture or natural pastureland; and natural vegetation. However, these data are presented in PDF format, with limited spatial coverage, hindering their application in comprehensive, nationwide studies. Given this context, we present a rigorous methodology that integrates advanced geospatial analyses, agricultural production estimation, and machine learning techniques to predict land values across Brazil. Leveraging annual land use/cover data and agricultural suitability information at high resolutions, together with a diverse set of socioeconomic data, the outcomes of this study promise to enrich our understanding of the spatial patterns of land values in Brazil, thereby contributing valuable insights for land use planning, natural resource management, and informed decision-making, with a specific focus on conservation and restoration sciences. 2 Methodological description 2.1 Land value The first step in our analyses was to standardize the data on land value provided by the Government by converting it from PDF into tabular format. We gathered available information for the years 2019, 2021, 2022, and 2023, adjusting monetary values to 2023 by accounting for accumulated inflation based on the Extended National Consumer Price Index (IPCA) (Banco Central do Brasil 2023). Owing to technical constraints, we excluded data from 2020 as it was presented in a non-convertible image format. For instances 1https://www.gov.br/receitafederal/pt-br/centrais-de-conteudo/publicacoes/documentos-tecnicos/vtn 1
of missing information within a specific land use/cover class, we applied the mean municipality value for silviculture, planted pasture, and natural vegetation. Regarding cropland, we substituted missing values at the municipal level, prioritizing the most conservative estimate. For example, if data for good suitability cropland was absent in a municipality, we used information from regular suitability for input. 2.2 Land use and land cover and agricultural suitability We used 2020 land use/cover data with 30 m resolution provided by Mapbiomas (MapBiomas 2023), reclassified into the following categories (Table 1): cropland, pastureland, natural vegetation, silviculture, ignored, other natural lands. we calculated the fraction coverage for each category within a 1 km resolution grid spanning the entirety of Brazil, employing the Mollweide projection. For the spatial information regarding agricultural suitability, we relied on the work of Safanelli et al. (2023), who conducted a comprehensive study modelling soil, climate, and relief suitability for agriculture. Their dataset, characterized by a 30 m resolution, was used to calculate the mean suitability per cell of our 1km grid. The integration of land use/cover data and suitability information enabled us to spatialize the land value data. This spatialized data served as the response variable in our land value model (Figure 1). 2.3 Independent variables We computed a series of co-variables aimed at predicting the distribution of land value across the Brazilian territory (Table 2). Notably, the determination of the production value, or the estimated value derived from agricultural and livestock activities, involved a more intricate calculation. The detailed steps for obtaining this value are detailed in the subsequent section. 2.3.1 Production value We determined the agricultural production value using data sourced from Brazilian Institute of Geography and Statistics (IBGE 2021b). Our methodology involved calculating the mean production in reais per hectare (ha−1) for each Brazilian municipality weighted by the municipality’s area dedicated to both temporary and permanent crops. Subsequently, we assigned a production value to each 1 km grid cell, based on the spatial location of each municipality and multiplied it by the fraction of agriculture covering the respective cell. A similar approach was applied to estimate livestock (IBGE 2021a) and silviculture (IBGE 2021c) values. For livestock, we aggregated the production values of wool, beef, and milk per municipality per hectare, multiplied by the fraction of each grid cell covered with pastureland. Regarding silviculture, we utilized the production value of silviculture, multiplying it by the coverage of silviculture within each cell of the grid. 2.4 Fitting a predictive model of land value We employed a Random Forest approach to fine-tune a regression model, establishing a relationship between the land value data and the independent variables detailed in Table 2. To enhance efficiency, we stratified the data across the five Brazilian territories: North (Norte), Northeast (Nordeste), South (Sul), Southeast (Sudeste), and Central-West (Centro-Oeste), considering their socioeconomic similarities to bolster model accuracy (original Portuguese names in parentheses). Within each region, we conducted a stratified sample, factoring in the municipality’s land value data area, selecting 30% of the dataset (except for Northeast and North, where we sampled 50% due to limited coverage). Subsequently, we partitioned the data into a training set (70%) for model fitting and a testing set (30%) to assess accuracy. To streamline the model, we computed a correlation matrix among potential numeric predictive variables, incorporating only those with a correlation value below 0.7. The Random Forest models were fitted using the ‘randomForestSRC’ package (2008) with 200 trees and nodes set to n=20. Model accuracy was evaluated through R2and RSME. Following model generation, we predicted land use values for all cells within our 1 km grid for each region, which were then consolidated into 2
Table 1: Reclassification of Mapbiomas land use/cover classes into our classes of interest to spatialize land value Land use/cover Mapiomas ID final class forest 1 natural vegetation natural forest 2 natural vegetation forest formation 3 natural vegetation savana 4 natural vegetation mangrove 5 natural vegetation wooded restinga 49 natural vegetation wetlands 11 natural vegetation grassland 12 natural vegetation salt flat 32 other natural lands rocky outcrop 29 ignored herbaceous sandbank vegetation 50 other natural lands other non Forest Formations 13 other natural lands pasture 15 pastureland agriculture 18 agriculture temporary crop 19 agriculture soybean 39 agriculture sugar cane 20 agriculture rice 40 agriculture cotton 62 agriculture other temporary Crops 41 agriculture perennial Crop 36 agriculture coffee 46 agriculture citrus 47 agriculture other perennial crop 48 agriculture forest plantaion 9 silviculture mosaic of uses 21 agriculture beach, dune, sand 23 other natural lands urban 24 ignored mining 30 ignored other non-vegetated areas 25 ignored water 26 ignored river/lake 33 ignored aquaculture 31 ignored non-observed 27 ignored 3
Figure 1: Available land value data provided by the Brazilian Federal Government for the years 2019,2021,2022 and 2023 combined, with values updated to 2023. The gray shade represents areas where there were no data available. 4
Table 2: Independent variables used to fit our land value model variable Description Year Source Urbanization Proportion of muncipality covered by urban areas (0-1) 2017 IBGE Distance to cities Euclidian distance to cities >500 thousand habitants (km) 2010 IBGE Pasture proportion Proportion of each grid cell covered by pastureland (0-1) 2020 Mabiomas collection 7 Agriculture proportion Proportion of each grid cell covered by agriculture (0-1) 2020 Mabiomas collection 7 Natural land proportion Proportion of each grid cell covered by natural cover (0-1) 2020 Mabiomas collection 7 Relief Relief agriculture suitability index (0-100; in which higher values indicate higher suitability) 2023 Safanelli et al. Climate Climate agriculture suitability index (0-100; in which higher values indicate higher suitability) 2023 Safanelli et al. Soil Soil agriculture suitability index (0-100; in which higher values indicate higher suitability) 2023 Safanelli et al. Proportion properties > 100 ha Proportion of municipalities covered by properties > 100ha(0-1) 2017 IBGE Agriculture credit Agricultureal credit conceived to the municipality weighted by total population (reais/habitant) 2020 IBGE; BACEN Proportion agriculture GDP Proportion municipality agriculture GDP / total GDP (0-1) 2020 IBGE GDP per capta Municipal GDP per capta (thousand reais) 2020 IBGE Number of tractors Number of tractor per municipality (thousand units) 2020 IBGE Number of people employed Number of people employed in the rural sector (thousand people) 2017 IBGE Storage capacity Municipality storage capacity (thousand tonnes) 2017 IBGE Proportion power Proportion of properties in each municipality with acess to poewe (0-1) 2017 IBGE Proportion higher education Proportion of landowners in each municipality with higher education(0-1) 2017 IBGE Distance to Federal roads Euclidean distance to Federal roads (calculated for each cell of the grid; km) 2023 Mapbiomas Distance to State roads Euclidean distance to Estate roads (calculated for each cell of the grid; km) 2023 Mapbiomas Distance to ports Euclidean distance to Ports(calculated for each cell of the grid kms). 2023 Mapbiomas Landowner proportion Number of landowners/ total rural workers of each municipality 2017 IBGE Production value Production value of main agriculture and livestock activities (R$/ha) 2021 IBGE; Mapbiomas collection 7 Conflicts Number of Land-tenure conflicts between 2011-2020 2020 Agencia Publica Agressions Number of hospitalizations due to agression bewteen 2011-2020 2020 Agencia Publica HDI Muncipal Human-development Index 2010 IBGE Distance to UCs Euclidean distance to Protected Areas-Ucs (kms) 2023 Mapbiomas Distance to TIs Euclidean distance to Indigenous Lands (kms) 2023 Mapbiomas Distance to mining sites Euclidean distance to small-scale mining sites (kms) 2023 Mapbiomas Proportion of private lands Proportion of the grid area classified as private property (0-1) 2020 Imaflora Note: We built our table based on Instituto Escolhas(2023) and Marques et al., (2023) 5
Table 3: Evaluation of the perfomance of our regional models região R-squared RMSE North 0.94 0.19 Northeast 0.96 0.27 Southeast 0.92 0.23 South 0.95 1.60 Central-West 0.96 0.19 a unified raster for the entire country. All data handling and model fitting was done in R (R Core Team 2023) 3 Results Our exploration of variable correlations within each of the five Brazilian territories revealed collinearity among several variables in all regions (Figure 2). Notably, this examination guided our model refinement process, ensuring that only variables with correlation values below 0.7 were included in the final models. We removed: distance from mining sites, proportion of agriculture GDP and number of tractors (South region); urbanization, proportion of agriculture GDP (Southeast); distance from mining sites, distance to ports, HDI, agriculture subsidies, proportion of agriculture GDP, number of tractors (Northeast); production value, HDI,proportion of agriculture GDP (central-West); proportion of agriculture GDP, distance to ports, HDI(North). The Random Forest models exhibited remarkable accuracy across all regions, with consistently high Rsquared values and low Root Mean Square Error (RMSE) (Table 2). The models achieved R2values exceeding 0.9, highlighting their robust ability to explain the variance in land values, and RMSE below 2, underling the accuracy in predicting land values. The outcomes of my model could suggest potential overfitting. However, this is not a concern in this context, as the model is being used to interpolate values and fill missing data within the same region from which the data originates. The model is not being used to predict values in unsampled regions, but rather to provide complete coverage of the existing dataset. Our choice of tree number was also adequate and allowed fast processing with high accuracy (Figure 3). Variable importance analysis provided additional insights into the factors driving land value predictions. Key variables, such as agricultural production value estimates, land use/cover categories, and socioeconomic indicators, consistently emerged as significant contributors. (Figure 4). Applying our Random Forest models to extrapolate land values across Brazil resulted in detailed spatial predictions for each region. We extrapolated the land values for each region, providing a visual representation of the spatial distribution predicted by our models (Figure 5). These extrapolated values were subsequently mosaiced into a unified raster, offering a comprehensive nationwide overview of land values (Figure 6). The creation of a spatially explicit raster layer of land value for the Brazilian territory at a 1km resolution, based on Brazilian Goverment official data and relevant covariables, provides significant advantages over traditional PDF reports with incomplete coverage. This comprehensive layer allows for a uniform and continuous assessment of land values across the entire country, facilitating better-informed decision-making in spatial planning and resource allocation. The use of covariables enhances the model’s accuracy and extrapolation capabilities, capturing the intricate variations in land value influenced by multiple factors. However, this approach is not without limitations. One notable challenge is the lack of available land value data for certain regions, which can affect the model’s accuracy in those areas. Additionally, the values provided by municipalities or states tend to be underestimated since they are often calculated for tax purposes, which may not reflect the true market value. Nonetheless, this issue is likely uniformly distributed across regions, minimizing any significant spatial bias. Despite these limitations, the raster layer serves as a valuable asset for indicating relative land value across 6
Figure 2: Correlation matrixes among independent variables. Abbreviated variables are as follows: DstC500 = distance to cities > 500k habitants; PropPst = proportion of pastureland ; PropAgr = proportion of agriculture; propurb = proportion of urban areas; Prpr100 = proportion of properties over 100ha; agr2010 = agriculture subsidy; PrpAGDP = proportion of agriculture GDP/total GDP; vIBGE20 = production value; gdpagr = agriculture GDP; gdpprcp = GDP per capta; nmmqnrs = number of tractors; nmcpdsm = number of people employed; cpcddrm = storage capacity; prpcmnr. = proportion of properties with power; prpcmns.= proportion of landowners with higher studies; dstrdvsf = distance to federal roads; dstrdvss = distance to state roads; dstprts = distance to ports; PrpNtVg = proportion of native vegetation; prpprpr = proportion of landowners; DstGrmp = distance to mining sites; PrpPrvd = proportion of private lands; IDH2010 = Human development index. For more details on the variables, refer to Table 2. 7
Figure 3: Error rate estimation for the regional models. 8
Figure 4: Relative importance of the independent variable in our regional models. The importance of all variables summed equals to one. Abreviated variables: Agr.prop. = Agriculture proportion; Prop.propertiesgreater100ha = Proportion properties greater than 100 ha. 9