Appendix 4 to "The Status of Biological Invasions and their Management in South Africa in 2022"—Workflows
Abstract
SANBI and CIB 2023. Appendix 4 to "The Status of Biological Invasions and their Management in South Africa in 2022"—Workflows. South African National Biodiversity Institute, Kirstenbosch and DSI-NRF Centre of Excellence for Invasion Biology, Stellenbosch. http://dx.doi.org/10.5281/zenodo.8217222 For more details see: http://iasreport.sanbi.org.za For the report itself see: http://dx.doi.org/10.5281/zenodo.8217182
Full text
1 Appendix 4—Workflows for the report “The status of biological invasions and their management in South Africa in 2022” This appendix contains workflows used in the report (the report itself is available at http://dx.doi.org/10.5281/zenodo.8217182) Suggested citation: SANBI and CIB 2023. Appendix 4 to "The Status of Biological Invasions and their Management in South Africa in 2022"—Workflows. South African National Biodiversity Institute, Kirstenbosch and DSI-NRF Centre of Excellence for Invasion Biology, Stellenbosch. http://dx.doi.org/10.5281/zenodo.8217222 Contents: Introduction pathway prominence ........................................................................................................ 2 Tracking data sources ......................................................................................................................... 106 Adding alien taxa and enrichment data to the species list ............................................................... 107 Updating the permit database ........................................................................................................... 112 Money spent ....................................................................................................................................... 117 Alien taxa impact assessment ............................................................................................................ 125 Sourcing, capturing, and reporting information for the Prince Edward Islands .............................. 127
2 Introduction pathway prominence Suggested citation: SANBI and CIB 2023. Appendix 4 to "The Status of Biological Invasions and their Management in South Africa in 2022"—Workflows: Introduction pathway prominence. South African National Biodiversity Institute, Kirstenbosch and DSI-NRF Centre of Excellence for Invasion Biology, Stellenbosch. http://dx.doi.org/10.5281/zenodo.8217222 This workflow can be used to prepare, process, and plot the socio-economic data required to populate the indicator introduction pathway prominence. The indicator concerns the pathways that could facilitate the introduction of alien species to a country from another region. The indicator considers the opportunities available for introduction (how active the pathways are socioeconomically) and does not take into account how many introductions occur as a result of those opportunities. For further details see Wilson et al. (2018) and the factsheet for this indicator: https://besjournals.onlinelibrary.wiley.com/action/downloadSupplement?doi=10.1111%2F13652664.13251&file=jpe13251-sup-0004-MaterialS4.docx. Pathways are as per the Convention on Biological Diversity’s pathway classification framework (CBD 2014). See Figure 1 for an overview of the workflow. 1. Requirements A stable internet connection, access to the processed data used in previous assessments of the indicator, and access to published literature is required to obtain the input data for the workflow. R software is required to execute the workflow. The R code required for this workflow is provided in a separate R markdown document. The processed data used in previous assessments should be saved in a folder called ‘Processed_data_previous_report’.
3 Figure 1. Overview of the workflow for the indicator introduction pathway prominence. The steps in the box with the stippled grey line are automated. Automated steps are performed in R, and details on these steps and the required R code are provided in an accompanying R markdown document. 2. Inputs Socio-economic data for the introduction pathways and conversion factor tables for US$ are the required inputs. See below for details on how to obtain these inputs. 2.1. Conversion factor tables Conversion factor tables are required to adjust time series data in current US$ for inflation. Conversion factor tables can be obtained online: https://liberalarts.oregonstate.edu/spp/polisci/research/inflation-conversion-factors. The tables provided at the suggested website present predicted values for 2018 onwards. These can be replaced with actual values using information from: https://www.bls.gov/data/inflation_calculator.htm. 2.2. Socio-economic data for pathways Socio-economic data for introduction pathways can be obtained from the published peer-reviewed and grey literature, and downloaded from online databases. Peer reviewed or grey literature published since the cut-off date for the previous assessment (e.g. 31 December 2016 for the first assessment in South Africa) up until the cut-off date for the current assessment (e.g. 31 December 2019 for the second assessment in South Africa) must be searched for relevant information. This could, for example, be information on the trade in alien species published in research on the pet trade and ornamental plant trade, information on the number of biological control agents under evaluation, and reports or other official documents. Examples from South
4 Africa include the National Ports Plans by the Transnet National Ports Authority, the South African Grain Seeds Market Report by the Department of Agriculture, Land Reform and Rural Development, and the State of the World’s Fisheries and Aquaculture by the Food and Agriculture Organisation of the United Nations. Socio-economic data for many pathways can be downloaded from the websites of international and national organisations. These data are often open-access. All data available are downloaded (i.e. data for all years), as this allows for change over time to be tracked. Data for the period before the data cut off point for the assessment (e.g. December 2019 for the second assessment in South Africa) are required, but data up until this point may not be available in all datasets. This is because some datasets are updated less frequently than the indicator is updated. Create a folder ‘Inputs’ to store the unedited versions of the downloaded data. The name of the saved file must contain information on the pathway, the database the data were downloaded from, and the date the data were downloaded. In instances where data for specific periods of time are downloaded in separate files, the file name must additionally contain information on the period for which the data were downloaded. Below are details on the steps that are followed to obtain online data that have been used in assessments of biological invasions in South Africa. Examples of file names are also provided. 2.2.1 Stowaway: Land vehicles Data are downloaded from the United Nations Comtrade database (https://comtradeplus.un.org/). These are data on the yearly value of vehicle imports to South Africa from the rest of the world. The data are in US$. A csv file containing the data is downloaded. Login using a google account. Select ‘Data’ at the top of the screen, and then ‘Trade data’ from the drop-down menu. The following selections need to be made on the webpage to download the correct data: Type of product: goods Frequency: annual Classification: HS (as reported) HS (as reported commodity codes): 87: Vehicles; other than railway or tramway rolling stock, and parts and accessories thereof 86: 'Railway or tramway locomotives, rolling-stock and parts thereof Reporters: South Africa Trade flows: Import Periods: Select each year for which data are required Partners: World Modes of transport: TOTAL modes of transport 2 nd partner: World Customs code: TOTAL customs procedure codes Breakdown mode: Classic Aggregate by: None File name: Stowaway_land_vehicles_COMTRADE_day_month_year (e.g. Stowaway_land_vehicles_COMTRADE_08_03_2021)
5 2.2.2 Escape: Horticulture; Contaminant: Contaminant nursery material; Contaminant: Contaminant on plants; Contaminant: Parasite on plants Data are downloaded from the United Nations Comtrade database (https://comtradeplus.un.org/). These are data on the yearly value of live plant imports to South Africa from the rest of the world. These data are in US$. A csv file containing the data is downloaded. Login using a google account. Select ‘Data’ at the top of the screen, and then ‘Trade data’ from the drop-down menu. The following selections need to be made on the webpage to download the correct data: Type of product: goods Frequency: annual Classification: HS (as reported) HS (as reported commodity codes): 06: Trees and other plants, live; bulbs, roots and the like; cut flowers and ornamental foliage Reporters: South Africa Trade flows: Import Periods: Select each year for which data are required Partners: World Modes of transport: TOTAL modes of transport 2 nd partner: World Customs code: TOTAL customs procedure codes Breakdown mode: Classic Aggregate by: None File name: Escape_Contaminant_plants_COMTRADE_day_month_year (e.g. Escape_Contaminant_plants_COMTRADE_08_03_2021) 2.2.3 Escape: Farmed animals Data are downloaded from the FAOstat database of the Food and Agricultural Organisation of the United Nations (http://www.fao.org/faostat/en/#data). These are data on the number of animals, of various types (cattle, horses pigs etc), produced yearly in South Africa. A csv file containing the data is downloaded. The following selections need to be made on the webpage to download the correct data: Production: Crops and livestock products This selection will take you to a second webpage where the following selections need to be made to download the correct data: Countries: South Africa Elements: Stocks Items: Live animals Years: Select all File name: Escape_Farmed_animals_FAOSTAT_day_month_year (e.g. Escape_farmed_animals_FAOSTAT_09_03_2021) 2.2.4. Contaminant: Contaminant on animals; Contaminant: Parasites on animals Data are downloaded from the FAOstat database of the Food and Agricultural Organisation of the United Nations (http://www.fao.org/faostat/en/#data). These are data on the number of live
6 animals, of various types, imported yearly into South Africa. A csv file containing the data is downloaded. The following selections need to be made on the webpage to download the correct data: Trade: Crops and livestock products This selection will take you to a second webpage where the following selections need to be made to download the correct data: Countries: South Africa Elements: Import quantity Items: Live animals Years: Select all File name: Contaminant_animals_FAOSTAT_day_month_year (e.g. Contaminant_animals_FAOSTAT_09_03_2021) 2.2.5. Escape: Forestry Data are downloaded from the FAOstat database of the Food and Agricultural Organisation of the United Nations (http://www.fao.org/faostat/en/#data). These are data on the quantity of forestry products produced yearly in South Africa. A csv file containing the data is downloaded. The following selections need to be made on the webpage to download the correct data: Forestry: forestry production and trade This selection will take you to a second webpage where the following selections need to be made to download the correct data: Countries: South Africa Elements: Production quantity Items: Select all Years: Select all File name: Escape_forestry_FAOSTAT_day_month_year (e.g. Escape_forestry_FAOSTAT_09_03_2021) 2.2.6. Contaminant: Timber trade Data are downloaded from the FAOstat database of the Food and Agricultural Organisation of the United Nations (http://www.fao.org/faostat/en/#data). These are data on the value of forestry imports to South Africa each year in 1000 US$. A csv file containing the data is downloaded. The following selections need to be made on the webpage to download the correct data: Forestry: forestry production and trade This selection will take you to a second webpage where the following selections need to be made to download the correct data: Countries: South Africa Elements: Import value Items: Select all Years: Select all File name: Contaminant_timber_FAOSTAT_day_month_year (e.g. Contaminant_timber_FAOSTAT_10_03_2021)
7 2.2.7. Contaminant: Food contaminant Data are downloaded from the FAOstat database of the Food and Agricultural Organisation of the United Nations (http://www.fao.org/faostat/en/#data). These are data on the quantity of various types of food imported to South Africa each year, in 1000 tonnes. A csv file containing the data is downloaded. The following selections need to be made on the webpage to download the correct data: Food balances (2010-) This selection will take you to a second webpage where the following selections need to be made to download the correct data: Countries: South Africa Elements: Import quantity Items: Select all Years: Select all File name: Contaminant_food_FAOSTAT_day_month_year (e.g. Contaminant_food_FAOSTAT_10_03_2021) 2.2.8. Escape: Agriculture Data are downloaded from the FAOstat database of the Food and Agricultural Organisation of the United Nations (http://www.fao.org/faostat/en/#data). These are yearly data on the quantity of crops produced (in tonnes) and the area harvested (in ha) in South Africa. A csv file containing the data is downloaded. The following selections need to be made on the webpage to download the correct data: • Production: Crops and livestock products This selection will take you to a second webpage where the following selections need to be made to download the correct data: Countries: South Africa Elements: Production quantity AND Area harvested Items: Crops Primary Years: Select all File name: Escape_agriculture_FAOSTAT_day_month_year (e.g. Escape_agriculture_FAOSTAT_10_03_2021) 2.2.9. Stowaway: People and their luggage Data are downloaded from the World Tourism and Travel Council (https://wttc.org/Research/Economic-Impact/Data-Visualisation). These are data on the total contribution of tourism and travel to South Africa’s GDP each year. Data are in bn US$ at real prices. A csv file containing the data is downloaded. [Note: The website of the WTTC is being updated, and this information will need to be updated once it is up] File name: Stowaway_people_WTTC_day_month_year (e.g. Stowaway_people_WTTC_10_03_2021). 2.2.10. Stowaway: Container/bulk
8 Data are downloaded from the Transnet National Ports Authority (http://www.transnetnationalportsauthority.net/Commercial%20and%20Marketing/Pages/PortStatistics.aspx). These are data on the number of deepsea containers landed at South African ports. Data are available for each month or year. Yearly data are required, so if the data for each month is downloaded, then work is required to get the yearly values (see below). Pdf(s) containing the data is downloaded. The following selections need to be made on the webpage to download the correct data: TEUs: Select the value for the period of interest (e.g. Calendar year 2020) File name: Stowaway_cargo_Year_TRANSNET_day_month_year (e.g. Stowaway_cargo_2020_TRANSNET_04_03_2021), OR Stowaway_cargo_MonthYear_TRANSNET_day_month_year (e.g. Stowaway_cargo_Jan2020_TRANSNET_04_03_2021). 2.2.11. Stowaway: Ship/boat hull fouling; Stowaway: Ship/boat ballast water; Stowaway: Hitchhikers on ship/boat; Stowaway: Angling/fishing equipment Data are downloaded from the Transnet National Ports Authority (http://www.transnetnationalportsauthority.net/Commercial%20and%20Marketing/Pages/PortStatistics.aspx). These are data on the number of vessel arrivals at South African ports. Data are available for each month or year. Yearly data are required, so if the data for each month is downloaded, then work is required to get the yearly values (see below). Pdf(s) containing the data is downloaded. The following selections need to be made on the webpage to download the correct data: Vessels Arrivals: Select the value for the period of interest (e.g. Calendar year 2020) File name: Stowaway_ship_Year_TRANSNET_day_month_year (e.g. Stowaway_ship_2020_TRANSNET_04_03_2021), OR Stowaway_ship_MonthYear_TRANSNET_day_month_year (e.g. Stowaway_ship_Jan2020_TRANSNET_04_03_2021). 2.2.12. Stowaway: People and their luggage/equipment; Stowaway: Vehicles Data are downloaded from Statistics South Africa (http://www.statssa.gov.za/?page_id=1866&PPN=P0351&SCH=6717&page=1). These are monthly data on tourism and migration for South Africa. A pdf file containing the data for each month is downloaded. The following selections need to be made on the webpage to download the correct data: Archived publications for P0351: Select the data for the month of interest (e.g. P0351 - Tourism and Migration, December 2020) This selection will take you to a second webpage where the following selections need to be made to download the correct data: Main publication: Download the entire publication File name: Stowaway_people_vehicles_MonthYear_STATSSA_day_month_year (e.g. Stowaway_people_vehicles_Jan2019_STATSSA_11_03_2021). 2.2.13. Stowaway: Airplane
9 Data are downloaded from Airports Company South Africa (http://www.airports.co.za/business/statistics/aircraft-and-passenger). These are data on the total consolidated aircraft movements at ACSA run airports each financial year. A pdf file containing the data is downloaded. The following selections need to be made on the webpage to download the correct data: Passenger and aircraft charts: Airport Company South Africa Total Consolidated Aircraft Movements File name: Stowaway_airplane_ACSA_day_month_year (e.g. Stowaway_airplane_ACSA_12_03_2021). 2.2.14. Escape: Research Data are downloaded from CITES (https://trade.cites.org/). These are data on the number of live animals imported each year to South Africa for scientific purposes, breeding in captivity or artificial propagation purposes. A csv file containing the data is downloaded. The following selections need to be made on the webpage to download the correct data: Year range: select years of interest (e.g. from: 2005; to: 2020) Exporting countries: All countries Importing countries: South Africa Source: All sources Purpose: B-Breeding in captivity or artificially propagation AND S-Scientific Trade terms: LIV-Live Search by Taxon: leave blank This selection will take you to a second webpage where the following selections need to be made to download the correct data: Select output types: comma separated Select report type: Gross/Net Trade Tabulations: Gross imports File name: Escape_research_CITES_day_month_year (e.g. Escape_research_CITES_12_03_2021). 2.2.15. Escape: Pet Data are downloaded from CITES (https://trade.cites.org/). These are data on the number of live animals imported to South Africa for commercial or personal purposes each year. A csv file containing the data is downloaded. The following selections need to be made on the webpage to download the correct data: Year range: select years of interest (e.g. from: 2005; to: 2020) Exporting countries: All countries Importing countries: South Africa Source: All sources Purpose: P-Personal AND T-Commercial Trade terms: LIV-Live Search by Taxon: leave blank This selection will take you to a second webpage where the following selections need to be made to download the correct data: Select output types: comma separated Select report type: Gross/Net Trade Tabulations: Gross imports
16 In this automated step (Figure 1), data that are in short-form are converted to long form, unnecessary rows or columns in the datasets are removed, data that are provided at a higher level of detail than required are aggregated, and the most recent data are checked for issues and adjustments made where required. The details and R code required to pre-process the socioeconomic data that have been used in assessments of biological invasions in South Africa are provided in an accompanying R markdown document. 3.4. Processing During this automated step (Figure 1) the data are plotted, compared to the data used in the most recent assessment of the indicator, and calculations are done to determine the percentage change in the data since the previous assessment. This step provides information for estimating the indicator introduction pathway prominence, and tracking change in the indicator. The details and R code required to process the socio-economic data that have been used in assessments of biological invasions in South Africa are provided in an accompanying R markdown document. 3.5. Post-processing The processed data are exported in csv format, information required to track change are saved in a dedicated csv file, and plots are exported as tiffs (Figure 1). All outputs are saved in the folder ‘Processed’ created during automated steps of this workflow. The processed data are exported to the ‘Processed_data’ sub-folder in the ‘Processed’ folder, the plots are exported to the ‘Plots’ subfolder in the ‘Processed’ folder, and the information required to track change are saved in the file ‘Change_tracking.csv’ found in the ‘Processed’ folder. The details and R code required for this automated step to export the outputs used in assessments of biological invasions in South Africa are provided in an accompanying R markdown document. 4. Outputs For each pathway, the outputs of this workflow (Figure 1) and other socio-economic data collected on the pathway through literature searches and from experts are used by an expert to classify the introduction opportunities provided by the pathway into five categories of introduction pathway prominence, and track change. Confidence in the overall assessment as well as that for each pathway is also estimated. For details on calculation procedures and estimating confidence limits see the factsheet for this indicator: https://besjournals.onlinelibrary.wiley.com/action/downloadSupplement?doi=10.1111%2F13652664.13251&file=jpe13251-sup-0004-MaterialS4.docx 5. References CBD (2014) Pathways of introduction of invasive species, their prioritization and management. Montreal Available from: https://www.cbd.int/doc/meetings/sbstta/sbstta-18/official/sbstta18-09-add1-en.pdf (February 7, 2019). Wilson JRU, Faulkner KT, Rahlao SJ, Richardson DM, Zengeya TA, van Wilgen BW (2018) Indicators for monitoring biological invasions at a national level. Journal of Applied Ecology 55: 2612– 2620. https://doi.org/10.1111/1365-2664.13251
17 6. R-markdown code Katelyn Faulkner 05/05/2023 Conversion factor tables and socio-economic data are obtained from various sources to populate the indicator: Introduction Pathway Prominence. The workflow for this indicator includes manually preparing some of the data for processing (e.g. digetising data). This document contains the R code for the automated processes of the workflow. The R code was specifically developed to process the data obtained for assessments of biological invasions in South Africa, but the steps will be appropriate elsewhere, and the code can be adapted. Note that the file names of the csvs used in the code in this document will need to be edited. This is because these file names include the date the data were downloaded. See the pdf on the workflow for further details. Set working directory The folder set as the working directory must contain the folders created in earlier steps of the workflow (e.g. ‘Inputs’ and ‘Prepared_inputs’). See pdf on the workflow for further details. The function setwd() is used to set the working directory. Paste the file path to your working directory in the quotation marks in the code chuck below (i.e. replace the file path in the brackets with that for your working directory). Insert double backslashes where required. setwd("C:\\National_Status_Report_2023\\Introduction_pathway_prominence") Create folders for outputs Create a folder for all processed files or the outputs of the automated part of the workflow. This folder will be in your working directory. dir.create("Processed") Create a sub-folder for plots. This sub-folder will be in the “Processed” folder created using the above chunk of code. dir.create("Processed//Plots") Create a sub-folder for processed data. This sub-folder will be in the “Processed” folder created using the above chunk of code. dir.create("Processed//Processed_data") Create a csv file for outputs on change This is an empty csv file to which outputs that will assist with tracking changes to the indicator will be written. This file will be in the “Processed” folder created using the above chunk of code. names<-data.frame(matrix(ncol=4,nrow=0, dimnames=list(NULL, c("Pathway", "ChangeScenario", "PercentageChange", "BaseChange")))) write.csv(names, "Processed//Change_tracking.csv", row.names = FALSE) Read data
18 Read data that were manually prepared Manually prepared data were saved as csv files in the “Prepared_inputs” folder in the working directory. See the pdf on the workflow for further details. # get the list of files list.files("Prepared_inputs", pattern = ".csv") ## [1] "Contaminant_seeds_DALRRD_prepared.csv" ## [2] "Conversion_factor_2017USD.csv" ## [3] "Stowaway_airplane_ACSA_prepared.csv" ## [4] "Stowaway_cargo_TRANSNET_prepared.csv" ## [5] "Stowaway_fishing_TRANSNET_prepared.csv" ## [6] "Stowaway_people_vehicles_STATSSA_prepared.csv" ## [7] "Stowaway_ship_TRANSNET_prepared.csv" ## [8] "Unaided_IMF_prepared.csv" Read in each of the files, giving each a meaningful name. The assigned name should indicate the pathway for which the data are relevant. Note the assigned names must be unique, and will be used in subsequent steps of the workflow. In the code chunk below, for each file, paste the file name in the quotation marks after “Prepared_inputs\”. See below for examples. # conversion factor table CFDat<-read.csv("Prepared_inputs\\Conversion_factor_2017USD.csv") # airplane data from acsa StowAir<-read.csv("Prepared_inputs\\Stowaway_airplane_ACSA_prepared.csv") # cargo data from Transnet StowCargo<-read.csv("Prepared_inputs\\Stowaway_cargo_TRANSNET_prepared.csv") # fishing vessel data from Transnet StowFishVes<-read.csv("Prepared_inputs\\Stowaway_fishing_TRANSNET_prepared.csv") # migration data from StatsSA StowPeopVeh<-read.csv("Prepared_inputs\\Stowaway_people_vehicles_STATSSA_prepared.csv") # ship arrival data from Transnet StowShip<-read.csv("Prepared_inputs\\Stowaway_ship_TRANSNET_prepared.csv") # Merchandise import data for continental African countries from IMF UnaidedIMF<-read.csv("Prepared_inputs\\Unaided_IMF_prepared.csv") # Seed import data from DALRRD ContSeed<-read.csv("Prepared_inputs\\Contaminant_seeds_DALRRD_prepared.csv") Read data that did not require manual preparation These data were saved as csv files in the “Inputs” folder in the working directory. See the pdf on the workflow for further details.
19 # get the list of files list.files("Inputs", pattern = ".csv") ## [1] "Contaminant_animals_FAOSTAT_21_04_2023.csv" ## [2] "Contaminant_food_FAOSTAT_21_04_2023.csv" ## [3] "Contaminant_timber_FAOSTAT_21_04_2023.csv" ## [4] "Escape_agriculture_FAOSTAT_21_04_2023.csv" ## [5] "Escape_aquaculture_FISHSTATJ_21_04_2023.csv" ## [6] "Escape_botgarden_zoo_CITES_21_04_2023.csv" ## [7] "Escape_Contaminant_plants_COMTRADE_21_04_2023.csv" ## [8] "Escape_farmed_animals_FAOSTAT_21_04_2023.csv" ## [9] "Escape_forestry_FAOSTAT_21_04_2023.csv" ## [10] "Escape_pet_CITES_21_04_2023.csv" ## [11] "Escape_research_CITES_21_04_2023.csv" ## [12] "Stowaway_fishing_FISHSTATJ_21_04_2023.csv" ## [13] "Stowaway_land_vehicles_COMTRADE_21_04_2023.csv" ## [14] "Stowaway_people_WTTC_23_08_2019.csv" ## [15] "Unaided_WTO_21_04_2023.csv" Read in each of the files, giving each a meaningful name. The assigned name should indicate the pathway for which the data are relevant. Note the assigned names must be unique, and will be used in subsequent steps of the workflow. In the code chunk below, for each file, paste the file name in the quotation marks after “Inputs\”. See below for examples. # animal import data from the FAO ContAni<-read.csv("Inputs\\Contaminant_animals_FAOSTAT_21_04_2023.csv") # food import data from the FAO ContFood<-read.csv("Inputs\\Contaminant_food_FAOSTAT_21_04_2023.csv") # timber import data from the FAO ContTimb<-read.csv("Inputs\\Contaminant_timber_FAOSTAT_21_04_2023.csv") # animal import data from the FAO ContAni<-read.csv("Inputs\\Contaminant_animals_FAOSTAT_21_04_2023.csv") # agriculture production data from the FAO EscAgri<-read.csv("Inputs\\Escape_agriculture_FAOSTAT_21_04_2023.csv") # aquaculture production data from FishStatJ EscAqua<-read.csv("Inputs\\Escape_aquaculture_FISHSTATJ_21_04_2023.csv") # animal imports for botanical gardens and zoos from CITES EscBotZoo<-read.csv("Inputs\\Escape_botgarden_zoo_CITES_21_04_2023.csv") # plant import data from the Comtrade EscContPlant<-read.csv("Inputs\\Escape_Contaminant_plants_COMTRADE_21_04_2023.csv") # farm animal production data from the FAO EscFarm<-read.csv("Inputs\\Escape_farmed_animals_FAOSTAT_21_04_2023.csv")
20 # forestry production data from the FAO EscForest<-read.csv("Inputs\\Escape_forestry_FAOSTAT_21_04_2023.csv") # data on organisms imported for personal or commercial reasons from CITES EscPet<-read.csv("Inputs\\Escape_pet_CITES_21_04_2023.csv") # data on organisms imported for research from CITES EscRes<-read.csv("Inputs\\Escape_research_CITES_21_04_2023.csv") # fishing catch data from FishStatJ StowFishCatch<-read.csv("Inputs\\Stowaway_fishing_FISHSTATJ_21_04_2023.csv") # vehicle import data from Comtrade StowLand<-read.csv("Inputs\\Stowaway_land_vehicles_COMTRADE_21_04_2023.csv") # imports data for Africa from WTO UnaidedWTO<-read.csv("Inputs\\Unaided_WTO_21_04_2023.csv") # tourism data from WTTC StowPeopWTTC<-read.csv("Inputs\\Stowaway_people_WTTC_23_08_2019.csv") All the required files were read into R during the previous step. In the subsequent steps each file is processed separately, following the remaining steps of the automated part of the workflow: standardisation, pre-processing, processing, post-processing. See below for the detailed steps for each of the datasets obtained for assessments of biological invasions in South Africa. Stowaway: Land vehicles (Comtrade) Data on the value of vehicle imports to South Africa were downloaded from the UN Comtrade database. The value of various vehicle imports per year were downloaded. Therefore, there will be more than one value per year. The data are in current US$. View data structure str(StowLand) # data structure ## 'data.frame': 46 obs. of 47 variables: ## $ TypeCode : chr "C" "C" "C" "C" ... ## $ FreqCode : chr "A" "A" "A" "A" ... ## $ RefPeriodId : int 20220101 20220101 20210101 20210101 20200101 20200101 20190101 20190101 20180101 20180101 ... ## $ RefYear : int 2022 2022 2021 2021 2020 2020 2019 2019 2018 2018 ... ## $ RefMonth : int 52 52 52 52 52 52 52 52 52 52 ... ## $ Period : int 2022 2022 2021 2021 2020 2020 2019 2019 2018 2018 ... ## $ ReporterCode : int 710 710 710 710 710 710 710 710 710 710 ... ## $ ReporterISO : chr "ZAF" "ZAF" "ZAF" "ZAF" ... ## $ ReporterDesc : chr "South Africa" "South Africa" "South Africa" "South Africa" ... ## $ FlowCode : chr "M" "M" "M" "M" ... ## $ FlowDesc : chr "Import" "Import" "Import" "Import" ... ## $ PartnerCode : int 0 0 0 0 0 0 0 0 0 0 ...
21 ## $ PartnerISO : chr "W00" "W00" "W00" "W00" ... ## $ PartnerDesc : chr "World" "World" "World" "World" ... ## $ Partner2Code : int 0 0 0 0 0 0 0 0 0 0 ... ## $ Partner2ISO : chr "W00" "W00" "W00" "W00" ... ## $ Partner2Desc : chr "World" "World" "World" "World" ... ## $ ClassificationCode : chr "H6" "H6" "H5" "H5" ... ## $ ClassificationSearchCode: chr "HS" "HS" "HS" "HS" ... ## $ IsOriginalClassification: logi TRUE TRUE TRUE TRUE TRUE TRUE ... ## $ CmdCode : int 87 86 86 87 86 87 86 87 86 87 ... ## $ CmdDesc : chr "Vehicles; other than railway or tramway rolling stock, and parts and accessories thereof" "Railway, tramway locomotives, rolling-stock and parts thereof; railway or tramway track fixtures and fittings a"| __truncated__ "Railway, tramway locomotives, rolling-stock and parts thereof; railway or tramway track fixtures and fittings a"| __truncated__ "Vehicles; other than railway or tramway rolling stock, and parts and accessories thereof" ... ## $ AggrLevel : int 2 2 2 2 2 2 2 2 2 2 ... ## $ IsLeaf : logi FALSE FALSE FALSE FALSE FALSE FALSE ... ## $ CustomsCode : chr "C00" "C00" "C00" "C00" ... ## $ CustomsDesc : chr "TOTAL CPC" "TOTAL CPC" "TOTAL CPC" "TOTAL CPC" ... ## $ MosCode : int 0 0 0 0 0 0 0 0 0 0 ... ## $ MotCode : int 0 0 0 0 0 0 0 0 0 0 ... ## $ MotDesc : chr "TOTAL MOT" "TOTAL MOT" "TOTAL MOT" "TOTAL MOT" ... ## $ QtyUnitCode : int -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 ... ## $ QtyUnitAbbr : chr "N/A" "N/A" "N/A" "N/A" ... ## $ Qty : int 0 0 0 0 0 0 0 0 0 0 ... ## $ IsQtyEstimated : logi FALSE FALSE FALSE FALSE FALSE FALSE ... ## $ AltQtyUnitCode : int -1 -1 -1 -1 -1 -1 -1 -1 -1 -1 ... ## $ AltQtyUnitAbbr : chr "N/A" "N/A" "N/A" "N/A" ... ## $ AtlQty : int 0 0 0 0 0 0 0 0 0 0 ... ## $ IsAltQtyEstimated : logi FALSE FALSE FALSE FALSE FALSE FALSE ... ## $ NetWgt : int 0 0 0 0 0 0 0 0 0 0 ... ## $ IsNetWgtEstimated : logi TRUE FALSE FALSE TRUE TRUE TRUE ... ## $ GrossWgt : int 0 0 0 0 0 0 0 0 0 0 ... ## $ IsGrossWgtEstimated : logi FALSE FALSE FALSE FALSE FALSE FALSE ... ## $ Cifvalue : num 0 0 0 0 0 0 0 0 0 0 ... ## $ Fobvalue : num 8.33e+09 1.17e+08 1.26e+08 6.30e+09 9.45e+07 ... ## $ PrimaryValue : num 8.33e+09 1.17e+08 1.26e+08 6.30e+09 9.45e+07 ... ## $ LegacyEstimationFlag : int 4 0 0 4 4 4 4 4 4 4 ... ## $ IsReported : logi FALSE FALSE FALSE FALSE FALSE FALSE ... ## $ IsAggregate : logi TRUE TRUE TRUE TRUE TRUE TRUE ... Standardisation Convert data into 2017 USD using the conversion factor table. # Merge the dataset with the dataset containing conversion factors StowLandS <- merge(StowLand, CFDat,by.x = c("RefYear"), by.y = c("Year")) # merge datasets by year head (StowLandS) # header of data ## RefYear TypeCode FreqCode RefPeriodId RefMonth Period ReporterCode ## 1 2000 C A 20000101 52 2000 710 ## 2 2000 C A 20000101 52 2000 710 ## 3 2001 C A 20010101 52 2001 710
22 ## 4 2001 C A 20010101 52 2001 710 ## 5 2002 C A 20020101 52 2002 710 ## 6 2002 C A 20020101 52 2002 710 ## ReporterISO ReporterDesc FlowCode FlowDesc PartnerCode PartnerISO PartnerDesc ## 1 ZAF South Africa M Import 0 W00 World ## 2 ZAF South Africa M Import 0 W00 World ## 3 ZAF South Africa M Import 0 W00 World ## 4 ZAF South Africa M Import 0 W00 World ## 5 ZAF South Africa M Import 0 W00 World ## 6 ZAF South Africa M Import 0 W00 World ## Partner2Code Partner2ISO Partner2Desc ClassificationCode ## 1 0 W00 World H1 ## 2 0 W00 World H1 ## 3 0 W00 World H1 ## 4 0 W00 World H1 ## 5 0 W00 World H2 ## 6 0 W00 World H2 ## ClassificationSearchCode IsOriginalClassification CmdCode ## 1 HS TRUE 86 ## 2 HS TRUE 87 ## 3 HS TRUE 86 ## 4 HS TRUE 87 ## 5 HS TRUE 86 ## 6 HS TRUE 87 ## CmdDesc ## 1 Railway, tramway locomotives, rolling-stock and parts thereof; railway or tramway track fixtures and fittings and parts thereof; mechanical (including electro-mechanical) traffic signalling equipment of all kinds ## 2 Vehicles; other than railway or tramway rolling stock, and parts and accessories thereof ## 3 Railway, tramway locomotives, rolling-stock and parts thereof; railway or tramway track fixtures and fittings and parts thereof; mechanical (including electro-mechanical) traffic signalling equipment of all kinds ## 4 Vehicles; other than railway or tramway rolling stock, and parts and accessories thereof ## 5 Railway, tramway locomotives, rolling-stock and parts thereof; railway or tramway track fixtures and fittings and parts thereof; mechanical (including electro-mechanical) traffic signalling equipment of all kinds ## 6 Vehicles; other than railway or tramway rolling stock, and parts and accessories thereof ## AggrLevel IsLeaf CustomsCode CustomsDesc MosCode MotCode MotDesc ## 1 2 FALSE C00 TOTAL CPC 0 0 TOTAL MOT ## 2 2 FALSE C00 TOTAL CPC 0 0 TOTAL MOT ## 3 2 FALSE C00 TOTAL CPC 0 0 TOTAL MOT ## 4 2 FALSE C00 TOTAL CPC 0 0 TOTAL MOT ## 5 2 FALSE C00 TOTAL CPC 0 0 TOTAL MOT ## 6 2 FALSE C00 TOTAL CPC 0 0 TOTAL MOT ## QtyUnitCode QtyUnitAbbr Qty IsQtyEstimated AltQtyUnitCode AltQtyUnitAbbr ## 1 -1 N/A NA FALSE -1 N/A ## 2 -1 N/A NA FALSE -1 N/A
23 ## 3 -1 N/A NA FALSE -1 N/A ## 4 -1 N/A NA FALSE -1 N/A ## 5 -1 N/A NA FALSE -1 N/A ## 6 -1 N/A NA FALSE -1 N/A ## AtlQty IsAltQtyEstimated NetWgt IsNetWgtEstimated GrossWgt ## 1 NA FALSE NA FALSE NA ## 2 NA FALSE NA FALSE NA ## 3 NA FALSE NA FALSE NA ## 4 NA FALSE NA FALSE NA ## 5 NA FALSE NA FALSE NA ## 6 NA FALSE NA FALSE NA ## IsGrossWgtEstimated Cifvalue Fobvalue PrimaryValue LegacyEstimationFlag ## 1 FALSE 14925347 NA 14925347 0 ## 2 FALSE 1515226062 NA 1515226062 0 ## 3 FALSE 20026626 NA 20026626 0 ## 4 FALSE 1610085158 NA 1610085158 0 ## 5 FALSE 22063583 NA 22063583 0 ## 6 FALSE 1792848346 NA 1792848346 0 ## IsReported IsAggregate Conversion_factor ## 1 FALSE FALSE 0.703 ## 2 FALSE FALSE 0.703 ## 3 FALSE FALSE 0.723 ## 4 FALSE FALSE 0.723 ## 5 TRUE FALSE 0.734 ## 6 TRUE FALSE 0.734 # Convert the data by dividing the value for each year by the conversion factor for that year StowLandS$Value_2017USD<-StowLandS$PrimaryValue/StowLandS$Conversion_factor # convert values head (StowLandS) # header of data ## RefYear TypeCode FreqCode RefPeriodId RefMonth Period ReporterCode ## 1 2000 C A 20000101 52 2000 710 ## 2 2000 C A 20000101 52 2000 710 ## 3 2001 C A 20010101 52 2001 710 ## 4 2001 C A 20010101 52 2001 710 ## 5 2002 C A 20020101 52 2002 710 ## 6 2002 C A 20020101 52 2002 710 ## ReporterISO ReporterDesc FlowCode FlowDesc PartnerCode PartnerISO PartnerDesc ## 1 ZAF South Africa M Import 0 W00 World ## 2 ZAF South Africa M Import 0 W00 World ## 3 ZAF South Africa M Import 0 W00 World ## 4 ZAF South Africa M Import 0 W00 World ## 5 ZAF South Africa M Import 0 W00 World ## 6 ZAF South Africa M Import 0 W00 World ## Partner2Code Partner2ISO Partner2Desc ClassificationCode ## 1 0 W00 World H1 ## 2 0 W00 World H1 ## 3 0 W00 World H1 ## 4 0 W00 World H1 ## 5 0 W00 World H2 ## 6 0 W00 World H2
24 ## ClassificationSearchCode IsOriginalClassification CmdCode ## 1 HS TRUE 86 ## 2 HS TRUE 87 ## 3 HS TRUE 86 ## 4 HS TRUE 87 ## 5 HS TRUE 86 ## 6 HS TRUE 87 ## CmdDesc ## 1 Railway, tramway locomotives, rolling-stock and parts thereof; railway or tramway track fixtures and fittings and parts thereof; mechanical (including electro-mechanical) traffic signalling equipment of all kinds ## 2 Vehicles; other than railway or tramway rolling stock, and parts and accessories thereof ## 3 Railway, tramway locomotives, rolling-stock and parts thereof; railway or tramway track fixtures and fittings and parts thereof; mechanical (including electro-mechanical) traffic signalling equipment of all kinds ## 4 Vehicles; other than railway or tramway rolling stock, and parts and accessories thereof ## 5 Railway, tramway locomotives, rolling-stock and parts thereof; railway or tramway track fixtures and fittings and parts thereof; mechanical (including electro-mechanical) traffic signalling equipment of all kinds ## 6 Vehicles; other than railway or tramway rolling stock, and parts and accessories thereof ## AggrLevel IsLeaf CustomsCode CustomsDesc MosCode MotCode MotDesc ## 1 2 FALSE C00 TOTAL CPC 0 0 TOTAL MOT ## 2 2 FALSE C00 TOTAL CPC 0 0 TOTAL MOT ## 3 2 FALSE C00 TOTAL CPC 0 0 TOTAL MOT ## 4 2 FALSE C00 TOTAL CPC 0 0 TOTAL MOT ## 5 2 FALSE C00 TOTAL CPC 0 0 TOTAL MOT ## 6 2 FALSE C00 TOTAL CPC 0 0 TOTAL MOT ## QtyUnitCode QtyUnitAbbr Qty IsQtyEstimated AltQtyUnitCode AltQtyUnitAbbr ## 1 -1 N/A NA FALSE -1 N/A ## 2 -1 N/A NA FALSE -1 N/A ## 3 -1 N/A NA FALSE -1 N/A ## 4 -1 N/A NA FALSE -1 N/A ## 5 -1 N/A NA FALSE -1 N/A ## 6 -1 N/A NA FALSE -1 N/A ## AtlQty IsAltQtyEstimated NetWgt IsNetWgtEstimated GrossWgt ## 1 NA FALSE NA FALSE NA ## 2 NA FALSE NA FALSE NA ## 3 NA FALSE NA FALSE NA ## 4 NA FALSE NA FALSE NA ## 5 NA FALSE NA FALSE NA ## 6 NA FALSE NA FALSE NA ## IsGrossWgtEstimated Cifvalue Fobvalue PrimaryValue LegacyEstimationFlag ## 1 FALSE 14925347 NA 14925347 0 ## 2 FALSE 1515226062 NA 1515226062 0 ## 3 FALSE 20026626 NA 20026626 0 ## 4 FALSE 1610085158 NA 1610085158 0 ## 5 FALSE 22063583 NA 22063583 0
25 ## 6 FALSE 1792848346 NA 1792848346 0 ## IsReported IsAggregate Conversion_factor Value_2017USD ## 1 FALSE FALSE 0.703 21230935 ## 2 FALSE FALSE 0.703 2155371354 ## 3 FALSE FALSE 0.723 27699344 ## 4 FALSE FALSE 0.723 2226950426 ## 5 TRUE FALSE 0.734 30059377 ## 6 TRUE FALSE 0.734 2442572678 Pre-processing Aggregate data. Data for various vehicle imports per year have been downloaded. Calculate the total value for vehicle imports per year. StowLandPP <- aggregate(StowLandS$Value_2017USD, by = list(StowLandS$RefYear), FUN = sum) colnames(StowLandPP)<-c("Year", "Value_2017USD") # set column names head(StowLandPP) # header of data ## Year Value_2017USD ## 1 2000 2176602289 ## 2 2001 2254649770 ## 3 2002 2472632056 ## 4 2003 3443196812 ## 5 2004 5226849958 ## 6 2005 7156593620 Check for issues with recent data and make nessesary adjustments. For example, data for incomplete years included, or data values indicate that the data are not accurate (e.g. not all data has been submitted) tail(StowLandPP) # tail of data ## Year Value_2017USD ## 18 2017 7284778703 ## 19 2018 7056557704 ## 20 2019 6749647833 ## 21 2020 4228938463 ## 22 2021 5947713413 ## 23 2022 7349015682 No adjustments required Processing Plot data Data are presented as millions of US dollars windows(6,4) par(mar = c(4,5,1,1)) plot(StowLandPP$Year, StowLandPP$Value_2017USD/1000000, pch = 18, ylab = "Value of imports\n(millions of US dollars)", xlab = "Year", ylim = c(0, 11000))
32 Plot data Data are presented as millions of US dollars. windows(6,4) par(mar = c(4,5,1,1)) plot(EscContPlantPP$Year, EscContPlantPP$Value_2017USD/1000000, pch = 18, ylab = "Live plant imports\n(millions of US dollars)", xlab = "Year", ylim = c(0, 25)) Perform calculations Read in processed data used in previous report. These data have been saved in the ‘Processed_data_previous_report’ folder. See the workflow pdf for further details. Paste the file path to the required file in the quotation marks in the code below (i.e. replace the file path in the brackets with that for the required file). Insert double backslashes where required PrevDat<-read.csv( "Processed_data_previous_report\\Escape_Contaminant_plants_COMTRADE_processed.csv") See if new data have become available, if the baseline from the previous report has changed, and if new data are available, calculate the percentage change. Endp<-max(PrevDat$Year) Endn<-max(EscContPlantPP$Year) x<-Endn-Endp if(x>0){ ChangeScenario<-"NewDataForRecentYears" y<-which(EscContPlantPP$Year == Endp) z<-which(EscContPlantPP$Year == Endn) PercentageChange<-((EscContPlantPP$Value_2017USD[z]- EscContPlantPP$Value_2017USD[y])/EscContPlantPP$Value_2017USD[y])*100 b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Value_2017USD[b] != EscContPlantPP$Value_2017USD[y] } if(x==0){ ChangeScenario<-"NoNewDataForRecentYears" PercentageChange<-NA y<-which(EscContPlantPP$Year == Endp) b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Value_2017USD[b] != EscContPlantPP$Value_2017USD[y] } Pathway<-"escape/contaminant:plantsComtrade" results<-cbind(Pathway, ChangeScenario, PercentageChange, BaseChange) Post-processing The processed data are saved in csv format in the folder ‘Processed_data’ which was created in previous steps
33 write.csv(EscContPlantPP, "Processed\\Processed_data\\\\Escape_Contaminant_plants_COMTRADE_processed.csv", row.names = FALSE) The outputs on changes to the data since the last report are written to the ‘Change_tracking’ csv in the ‘Processed’ folder. write.table(results,"Processed\\Change_tracking.csv", row.names=F, append=T, quote= FALSE, sep=",", col.names = F) The plot is saved as a tiff in the plots folder created in previous steps. tiff(filename = "Processed\\Plots\\Escape_Contaminant_plants_COMTRADE.tiff", width = 6, height = 4, units = "in", compression = "lzw", bg = "white", res = 1000) par(mar = c(4,5,1,1)) plot(EscContPlantPP$Year, EscContPlantPP$Value_2017USD/1000000, pch = 18, ylab = "Live plant imports\n(millions of US dollars)", xlab = "Year", ylim = c(0, 25)) graphics.off() Escape: Farmed animals (FAO) Data on the number of animals produced in South Africa were downloaded from the FAOSTAT database. These data are the number of various types of animals produced per year and, therefore, there will be more than one value per year. The units the data are recorded in also differ (number, head or 1000 head). View data structure str(EscFarm) # data structure ## 'data.frame': 732 obs. of 14 variables: ## $ ï..Domain.Code : chr "QCL" "QCL" "QCL" "QCL" ... ## $ Domain : chr "Crops and livestock products" "Crops and livestock products" "Crops and livestock products" "Crops and livestock products" ... ## $ Area.Code..M49. : int 710 710 710 710 710 710 710 710 710 710 ... ## $ Area : chr "South Africa" "South Africa" "South Africa" "South Africa" ... ## $ Element.Code : int 5111 5111 5111 5111 5111 5111 5111 5111 5111 5111 ... ## $ Element : chr "Stocks" "Stocks" "Stocks" "Stocks" ... ## $ Item.Code..CPC. : int 2132 2132 2132 2132 2132 2132 2132 2132 2132 2132 ... ## $ Item : chr "Asses" "Asses" "Asses" "Asses" ... ## $ Year.Code : int 1961 1962 1963 1964 1965 1966 1967 1968 1969 1970 ... ## $ Year : int 1961 1962 1963 1964 1965 1966 1967 1968 1969 1970 ... ## $ Unit : chr "Head" "Head" "Head" "Head" ... ## $ Value : int 345000 343000 340000 335000 330000 310000 280000 250000 230000 215000 ... ## $ Flag : chr "A" "E" "E" "E" ... ## $ Flag.Description: chr "Official figure" "Estimated value" "Estimated value" "Estimated value" ... Standardisation Convert the data where required so that they are all in the same units
34 unique(EscFarm$Unit) # units data are recorded in ## [1] "Head" "No" "1000 Head" EscFarmS<-EscFarm tmp<-which(EscFarmS$Unit == "1000 Head") # rows with data in 1000 Head EscFarmS$Value[tmp]<-EscFarmS$Value[tmp]*1000 # convert the values to head EscFarmS$Unit[tmp]<-"Head" # change unit unique(EscFarmS$Unit) # units of data ## [1] "Head" "No" Pre-processing Remove columns in dataset that are not required. Only retain columns with information on the year, the unit the data were recorded in, and the value. EscFarmPP<-EscFarmS[,c("Year", "Unit", "Value")] head(EscFarmPP) # header of data ## Year Unit Value ## 1 1961 Head 345000 ## 2 1962 Head 343000 ## 3 1963 Head 340000 ## 4 1964 Head 335000 ## 5 1965 Head 330000 ## 6 1966 Head 310000 Aggregate the data Data for various farmed animals per year have been downloaded. Calculate the total number of animals produced per year EscFarmPP<- aggregate(EscFarmPP$Value, by = list(EscFarmPP$Year), FUN = sum) colnames(EscFarmPP)<-c("Year", "Number_animals") # set column names head(EscFarmPP) # header of data ## Year Number_animals ## 1 1961 76910008 ## 2 1962 77616000 ## 3 1963 78110000 ## 4 1964 77299008 ## 5 1965 78006000 ## 6 1966 79790748 Check for issues with recent data and make necessary adjustments. For example, data for incomplete years included, or data values indicate that the data are not accurate (e.g. not all data has been submitted) tail(EscFarmPP) # tail of data ## Year Number_animals ## 56 2016 209324256 ## 57 2017 211274152 ## 58 2018 217160820
35 ## 59 2019 213232190 ## 60 2020 212577753 ## 61 2021 212580950 No adjustments required Processing Plot data Data are presented as millions of animals. windows(6,4) par(mar = c(4,5,1,1)) plot(EscFarmPP$Year, EscFarmPP$Number_animals/1000000, pch = 18, ylab = "Number of animals\n(millions)", xlab = "Year", ylim = c(0, 250)) Perform calculations Read in processed data used in previous report. These data have been saved in the ‘Processed_data_previous_report’ folder. Paste the file path to your file in the quotation marks in the code below (i.e. replace the file path in the brackets with that for your file). Insert double backslashes where required PrevDat<- read.csv("Processed_data_previous_report\\Escape_farmed_animals_FAOSTAT_processed.csv") See if new data have become available, if the baseline from the previous report has changed, and if new data are available, calculate the percentage change. Endp<-max(PrevDat$Year) Endn<-max(EscFarmPP$Year) x<-Endn-Endp if(x>0){ ChangeScenario<-"NewDataForRecentYears" y<-which(EscFarmPP$Year == Endp) z<-which(EscFarmPP$Year == Endn) PercentageChange<-((EscFarmPP$Number_animals[z]- EscFarmPP$Number_animals[y])/EscFarmPP$Number_animals[y])*100 b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Number_animals[b] != EscFarmPP$Number_animals[y] } if(x==0){ ChangeScenario<-"NoNewDataForRecentYears" PercentageChange<-NA y<-which(EscFarmPP$Year == Endp) b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Number_animals[b] != EscFarmPP$Number_animals[y] } Pathway<-"escape:farmFAO" results<-cbind(Pathway, ChangeScenario, PercentageChange, BaseChange)
36 Post-processing The processed data are saved in csv format in the folder ‘Processed_data’ which was created in previous steps write.csv(EscFarmPP, "Processed\\Processed_data\\Escape_farmed_animals_FAOSTAT_processed.csv", row.names = FALSE) The outputs on changes to the data since the last report are written to the ‘Change_tracking’ csv in the ‘Processed’ folder. write.table(results,"Processed\\Change_tracking.csv", row.names=F, append=T, quote= FALSE, sep=",", col.names = F) The plot is saved as a tiff in the plots folder created in previous steps. tiff(filename = "Processed\\Plots\\Escape_farmed_animals_FAOSTAT.tiff", width = 6, height = 4, units = "in", compression = "lzw", bg = "white", res = 1000) par(mar = c(4,5,1,1)) plot(EscFarmPP$Year, EscFarmPP$Number_animals/1000000, pch = 18, ylab = "Number of animals\n(millions)", xlab = "Year", ylim = c(0, 250)) graphics.off() Contaminant: Contaminant on animals; Contaminant: Parasites on animals (FAO) Data on the number of animals imported to South Africa were downloaded from the FAOSTAT database. These data are the number of various types of animals imported per year. Therefore, there will be more than one value per year. The units the data are recorded in also differ (number, head or 1000 head). The dataset also contains missing data. View data structure str(ContAni) # data structure ## 'data.frame': 501 obs. of 14 variables: ## $ ï..Domain.Code : chr "TCL" "TCL" "TCL" "TCL" ... ## $ Domain : chr "Crops and livestock products" "Crops and livestock products" "Crops and livestock products" "Crops and livestock products" ... ## $ Area.Code..M49. : int 710 710 710 710 710 710 710 710 710 710 ... ## $ Area : chr "South Africa" "South Africa" "South Africa" "South Africa" ... ## $ Element.Code : int 5608 5608 5608 5608 5608 5608 5608 5608 5608 5608 ... ## $ Element : chr "Import Quantity" "Import Quantity" "Import Quantity" "Import Quantity" ... ## $ Item.Code..CPC. : num 2132 2132 2132 2132 2132 ... ## $ Item : chr "Asses" "Asses" "Asses" "Asses" ... ## $ Year.Code : int 1996 1997 1998 1999 2000 2001 2002 2003 2004 2005 ... ## $ Year : int 1996 1997 1998 1999 2000 2001 2002 2003 2004 2005 ... ## $ Unit : chr "Head" "Head" "Head" "Head" ... ## $ Value : int 2 0 4 0 0 0 0 0 0 0 ... ## $ Flag : chr "A" "T" "A" "T" ... ## $ Flag.Description: chr "Official figure" "Unofficial figure" "Official figure" "Unofficial figure" ...
37 Standardisation Convert units. Convert the data, where required, so that they are all in the same units. unique(ContAni$Unit) # units data are recorded in ## [1] "Head" "1000 Head" ContAniS<-ContAni tmp<-which(ContAniS$Unit == "1000 Head") # rows with data in 1000 Head ContAniS$Value[tmp]<-ContAniS$Value[tmp]*1000 # convert the values ContAniS$Unit[tmp]<-"Head" # change unit unique(ContAniS$Unit) # units of data ## [1] "Head" Pre-processing Aggregate the data and remove columns that are not required. Data for various imported animals per year have been downloaded. Calculate the total number of animals imported per year. During this process columns in dataset that are not required are removed. ContAniPP<- aggregate(ContAniS$Value, by = list(ContAniS$Year), FUN = sum) colnames(ContAniPP)<-c("Year", "Number_animals") # set column names head(ContAniPP) # header of data ## Year Number_animals ## 1 1961 597605 ## 2 1962 490834 ## 3 1963 614001 ## 4 1964 606405 ## 5 1965 417766 ## 6 1966 473615 Check for issues with recent data and make necessary adjustments. For example, data for incomplete years included, or data values indicate that the data are not accurate (e.g. not all data has been submitted) tail(ContAniPP) # tail of data ## Year Number_animals ## 56 2016 1174950 ## 57 2017 1288017 ## 58 2018 1406694 ## 59 2019 1255753 ## 60 2020 1033616 ## 61 2021 1228880 No adjustments required Processing Plot data
38 Data are presented as millions of animals. windows(6,4) par(mar = c(4,5,1,1)) plot(ContAniPP$Year, ContAniPP$Number_animals/1000000, pch = 18, ylab = "Number of animals\n(millions)", xlab = "Year") Perform calculations Read in processed data used in previous report. These data have been saved in the ‘Processed_data_previous_report’ folder. Paste the file path to the required file in the quotation marks in the code below (i.e. replace the file path in the brackets with that for the required file). Insert double backslashes where required PrevDat<- read.csv("Processed_data_previous_report\\Contaminant_animals_FAOSTAT_processed.csv") See if new data have become available, if the baseline from the previous report has changed, and if new data are available, calculate the percentage change. Endp<-max(PrevDat$Year) Endn<-max(ContAniPP$Year) x<-Endn-Endp if(x>0){ ChangeScenario<-"NewDataForRecentYears" y<-which(ContAniPP$Year == Endp) z<-which(ContAniPP$Year == Endn) PercentageChange<-((ContAniPP$Number_animals[z]- ContAniPP$Number_animals[y])/ContAniPP$Number_animals[y])*100 b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Number_animals[b] != ContAniPP$Number_animals[y] } if(x==0){ ChangeScenario<-"NoNewDataForRecentYears" PercentageChange<-NA y<-which(ContAniPP$Year == Endp) b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Number_animals[b] != ContAniPP$Number_animals[y] } Pathway<-"contaminant:animalsFAO" results<-cbind(Pathway, ChangeScenario, PercentageChange, BaseChange) Post-processing The processed data are saved in csv format in the folder ‘Processed_data’ which was created in previous steps write.csv(ContAniPP, "Processed\\Processed_data\\Contaminant_animals_FAOSTAT_processed.csv", row.names = FALSE)
39 The outputs on changes to the data since the last report are written to the ‘Change_tracking’ csv in the ‘Processed’ folder. write.table(results,"Processed\\Change_tracking.csv", row.names=F, append=T, quote= FALSE, sep=",", col.names = F) The plot is saved as a tiff in the plots folder created in previous steps. tiff(filename = "Processed\\Plots\\Contaminant_animals_FAOSTAT.tiff", width = 6, height = 4, units = "in", compression = "lzw", bg = "white", res = 1000) par(mar = c(4,5,1,1)) plot(ContAniPP$Year, ContAniPP$Number_animals/1000000, pch = 18, ylab = "Number of animals\n(millions)", xlab = "Year") graphics.off() Escape: Forestry (FAO) Data on the amount of forestry products produced in South Africa were downloaded from the FAOSTAT database. These data are the amount of various types of forestry products produced per year. Therefore, there will be more than one value per year. The units the data are recorded in also differ (݉ ଷ and tonnes). View data structure str(EscForest) # data structure ## 'data.frame': 2049 obs. of 14 variables: ## $ ï..Domain.Code : chr "FO" "FO" "FO" "FO" ... ## $ Domain : chr "Forestry Production and Trade" "Forestry Production and Trade" "Forestry Production and Trade" "Forestry Production and Trade" ... ## $ Area.Code..M49. : int 710 710 710 710 710 710 710 710 710 710 ... ## $ Area : chr "South Africa" "South Africa" "South Africa" "South Africa" ... ## $ Element.Code : int 5516 5516 5516 5516 5516 5516 5516 5516 5516 5516 ... ## $ Element : chr "Production" "Production" "Production" "Production" ... ## $ Item.Code : int 1627 1627 1627 1627 1627 1627 1627 1627 1627 1627 ... ## $ Item : chr "Wood fuel, coniferous" "Wood fuel, coniferous" "Wood fuel, coniferous" "Wood fuel, coniferous" ... ## $ Year.Code : int 1961 1962 1963 1964 1965 1966 1967 1968 1969 1970 ... ## $ Year : int 1961 1962 1963 1964 1965 1966 1967 1968 1969 1970 ... ## $ Unit : chr "m3" "m3" "m3" "m3" ... ## $ Value : int 62000 41000 43000 52000 70000 676000 678000 680000 682000 684000 ... ## $ Flag : chr "A" "A" "A" "A" ... ## $ Flag.Description: chr "Official figure" "Official figure" "Official figure" "Official figure" ... Standardisation Remove rows with data recorded in tonnes. These data include products such as wood palp and paper etc. Therefore, only production measured in ݉ ଷ will be retained, this includes products like sawnwood, plywood, wood boards etc. unique(EscForest$Unit) # units of data
40 ## [1] "m3" "tonnes" tmp<-which(EscForest$Unit == "tonnes") # rows with data in tonnes EscForestS<-EscForest[-tmp,] # remove rows with data in tonnes Pre-processing Aggregate data and remove columns that are not required. Data for various forestry products per year have been downloaded. Calculate the total forestry production per year. During this process only required columns are retained. EscForestPP<- aggregate(EscForestS$Value, by = list(EscForestS$Year), FUN = sum) colnames(EscForestPP)<-c("Year", "Production_m3") # set column names head(EscForestPP) # header of data ## Year Production_m3 ## 1 1961 6228900 ## 2 1962 6809600 ## 3 1963 7396200 ## 4 1964 7668100 ## 5 1965 9488000 ## 6 1966 15783816 Check for issues with recent data and make necessary adjustments. For example, data for incomplete years included, or data values indicate that the data are not accurate (e.g. not all data has been submitted) tail(EscForestPP) # tail of data ## Year Production_m3 ## 56 2016 32410479 ## 57 2017 31122288 ## 58 2018 34039529 ## 59 2019 34525442 ## 60 2020 34633178 ## 61 2021 34983178 No adjustments required Processing Plot data Data are presented as production in millions ݉ ଷ . windows(6,4) par(mar = c(4,5,1,1)) plot(EscForestPP$Year, EscForestPP$Production_m3/1000000, pch = 18, ylab = expression(Forestry ~ production ~ (million ~ m^{3})), xlab = "Year", ylim = c(0,55)) Perform calculations Read in processed data used in previous report. These data have been saved in the ‘Processed_data_previous_report’ folder.
41 Paste the file path to the required file in the quotation marks in the code below (i.e. replace the file path in the brackets with that for the required file). Insert double backslashes where required PrevDat<-read.csv("Processed_data_previous_report\\Escape_forestry_FAOSTAT_processed.csv") See if new data have become available, if the baseline from the previous report has changed, and if new data are available, calculate the percentage change. Endp<-max(PrevDat$Year) Endn<-max(EscForestPP$Year) x<-Endn-Endp if(x>0){ ChangeScenario<-"NewDataForRecentYears" y<-which(EscForestPP$Year == Endp) z<-which(EscForestPP$Year == Endn) PercentageChange<-((EscForestPP$Production_m3[z]- EscForestPP$Production_m3[y])/EscForestPP$Production_m3[y])*100 b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Production_m3[b] != EscForestPP$Production_m3[y] } if(x==0){ ChangeScenario<-"NoNewDataForRecentYears" PercentageChange<-NA y<-which(EscForestPP$Year == Endp) b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Production_m3[b] != EscForestPP$Production_m3[y] } Pathway<-"escape:forestryFAO" results<-cbind(Pathway, ChangeScenario, PercentageChange, BaseChange) Post-processing The processed data are saved in csv format in the folder ‘Processed_data’ which was created in previous steps write.csv(EscForestPP, "Processed\\Processed_data\\Escape_forestry_FAOSTAT_processed.csv", row.names = FALSE) The outputs on changes to the data since the last report are written to the ‘Change_tracking’ csv in the ‘Processed’ folder. write.table(results,"Processed\\Change_tracking.csv", row.names=F, append=T, quote= FALSE, sep=",", col.names = F) The plot is saved as a tiff in the plots folder created in previous steps. tiff(filename = "Processed\\Plots\\Escape_forestry_FAOSTAT.tiff", width = 6, height = 4, units = "in", compression = "lzw", bg = "white", res = 1000) par(mar = c(4,5,1,1)) plot(EscForestPP$Year, EscForestPP$Production_m3/1000000, pch = 18, ylab = expression(Forestry ~ production ~ (million ~ m^{3})),
48 x<-Endn-Endp if(x>0){ ChangeScenario<-"NewDataForRecentYears" y<-which(ContFoodPP$Year == Endp) z<-which(ContFoodPP$Year == Endn) PercentageChange<-((ContFoodPP$Tonnes[z]- ContFoodPP$Tonnes[y])/ContFoodPP$Tonnes[y])*100 b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Tonnes[b] != ContFoodPP$Tonnes[y] } if(x==0){ ChangeScenario<-"NoNewDataForRecentYears" PercentageChange<-NA y<-which(ContFoodPP$Year == Endp) b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Tonnes[b] != ContFoodPP$Tonnes[y] } Pathway<-"contaminant:foodFAO" results<-cbind(Pathway, ChangeScenario, PercentageChange, BaseChange) Post-processing The processed data are saved in csv format in the folder ‘Processed_data’ which was created in previous steps write.csv(ContFoodPP, "Processed\\Processed_data\\Contaminant_food_FAOSTAT_processed.csv", row.names = FALSE) The outputs on changes to the data since the last report are written to the ‘Change_tracking’ csv in the ‘Processed’ folder. write.table(results,"Processed\\Change_tracking.csv", row.names=F, append=T, quote= FALSE, sep=",", col.names = F) The plot is saved as a tiff in the plots folder created in previous steps. tiff(filename = "Processed\\Plots\\Contaminant_food_FAOSTAT.tiff", width = 6, height = 4, units = "in", compression = "lzw", bg = "white", res = 1000) par(mar = c(4,5,1,1)) plot(ContFoodPP$Year, ContFoodPP$Tonnes/1000000, pch = 18, ylab = "Quantity of imported food\n(million tonnes)", xlab = "Year", ylim = c(0, 11)) graphics.off() Escape: Agriculture (FAO) Data on the quantity of crops produced and area harvested in South Africa were downloaded from the FAOSTAT database. These data are the amount produced and area harvested for various types of crops per year. Therefore, there will be more than one value per year, for both production and area
49 harvested. Data for area harvested are in ha, and data for production are in tonnes. The dataset contains missing data. View data structure str(EscAgri) # data structure ## 'data.frame': 8481 obs. of 14 variables: ## $ ï..Domain.Code : chr "QCL" "QCL" "QCL" "QCL" ... ## $ Domain : chr "Crops and livestock products" "Crops and livestock products" "Crops and livestock products" "Crops and livestock products" ... ## $ Area.Code..M49. : int 710 710 710 710 710 710 710 710 710 710 ... ## $ Area : chr "South Africa" "South Africa" "South Africa" "South Africa" ... ## $ Element.Code : int 5312 5312 5312 5312 5312 5312 5312 5312 5312 5312 ... ## $ Element : chr "Area harvested" "Area harvested" "Area harvested" "Area harvested" ... ## $ Item.Code..CPC. : num 1341 1341 1341 1341 1341 ... ## $ Item : chr "Apples" "Apples" "Apples" "Apples" ... ## $ Year.Code : int 1961 1962 1963 1964 1965 1966 1967 1968 1969 1970 ... ## $ Year : int 1961 1962 1963 1964 1965 1966 1967 1968 1969 1970 ... ## $ Unit : chr "ha" "ha" "ha" "ha" ... ## $ Value : num 7000 8000 8000 10000 10000 12000 12000 15000 15000 15000 ... ## $ Flag : chr "E" "E" "E" "E" ... ## $ Flag.Description: chr "Estimated value" "Estimated value" "Estimated value" "Estimated value" ... Standardisation Data do not need to be standardised Pre-processing Remove columns in dataset that are not required. Only retain columns with information on the year, the type of data (production or area harvested), the units the data were recorded in, and the value. EscAgriPP<-EscAgri[,c("Element","Year", "Unit", "Value")] head(EscAgriPP) # header of data ## Element Year Unit Value ## 1 Area harvested 1961 ha 7000 ## 2 Area harvested 1962 ha 8000 ## 3 Area harvested 1963 ha 8000 ## 4 Area harvested 1964 ha 10000 ## 5 Area harvested 1965 ha 10000 ## 6 Area harvested 1966 ha 12000 Separate into two datasets, one for area harvested, and one for production unique(EscAgriPP$Element) ## [1] "Area harvested" "Production" EscAgriProd<-EscAgriPP[EscAgriPP$Element == "Production",] # production dataset EscAgriArea<-EscAgriPP[EscAgriPP$Element == "Area harvested",] # area harvested dataset Aggregate data.
50 Data for various crops per year have been downloaded. Calculate the total crops produced or area harvested per year EscAgriProd<- aggregate(EscAgriProd$Value, by = list(EscAgriProd$Year), FUN = sum) colnames(EscAgriProd)<-c("Year", "Tonnes") # set column names head(EscAgriProd) # header of data for production ## Year Tonnes ## 1 1961 18633862 ## 2 1962 20199198 ## 3 1963 21115101 ## 4 1964 20291592 ## 5 1965 18199915 ## 6 1966 24290547 EscAgriArea<- aggregate(EscAgriArea$Value, by = list(EscAgriArea$Year), FUN = sum) colnames(EscAgriArea)<-c("Year", "ha") # set column names head(EscAgriArea) # header of data for area harvested ## Year ha ## 1 1961 7257879 ## 2 1962 7406680 ## 3 1963 7973507 ## 4 1964 7671384 ## 5 1965 7710539 ## 6 1966 7461358 Create one dataset EscAgriPP <- merge(EscAgriArea,EscAgriProd,by="Year") # merge datasets by year head (EscAgriPP) # header of data ## Year ha Tonnes ## 1 1961 7257879 18633862 ## 2 1962 7406680 20199198 ## 3 1963 7973507 21115101 ## 4 1964 7671384 20291592 ## 5 1965 7710539 18199915 ## 6 1966 7461358 24290547 Check for issues with recent data and make necessary adjustments. For example, data for incomplete years included, or data values indicate that the data are not accurate (e.g. not all data has been submitted) tail(EscAgriPP) # tail of data ## Year ha Tonnes ## 56 2016 5124186 38863809 ## 57 2017 5951579 51745855 ## 58 2018 5854130 50201568 ## 59 2019 5746296 48173091 ## 60 2020 5997125 52471104 ## 61 2021 6324036 54211340
51 No adjustments required Processing Plot data Data are presented as millions tonnes, and millions of ha. windows(6,4) par(mar = c(4,5,1,5)+ 0.3) plot(EscAgriPP$Year, EscAgriPP$ha/1000000, pch = 18, ylab = "Area harvested (million ha)", xlab = "Year", ylim = c(0, 10)) # Create first plot. par(new = TRUE) plot(EscAgriPP$Year, EscAgriPP$Tonnes/1000000, pch = 15, col = "gray", axes = FALSE, xlab = "", ylab = "", ylim = c(0, 55)) # Create second plot without axes axis(side = 4, at = pretty(range(EscAgriPP$Tonnes/1000000))) # Add second axis axis(side = 4, at = seq(0, 15, by = 5), labels = c("", "5", "10", "15"))# Add labels mtext("Production (million tonnes)", side = 4, line = 3) # Add second axes label Perform calculations Read in processed data used in previous report. These data have been saved in the ‘Processed_data_previous_report’ folder. Paste the file path to the required file in the quotation marks in the code below (i.e. replace the file path in the brackets with that for the required file). Insert double backslashes where required PrevDat<-read.csv("Processed_data_previous_report\\Escape_agriculture_FAOSTAT_processed.csv") See if new data have become available, if the baseline from the previous report has changed, and if new data are available, calculate the percentage change. Endp<-max(PrevDat$Year) Endn<-max(EscAgriPP$Year) x<-Endn-Endp if(x>0){ ChangeScenario<-"NewDataForRecentYears" y<-which(EscAgriPP$Year == Endp) z<-which(EscAgriPP$Year == Endn) PercentageChange<-((EscAgriPP$Tonnes[z]-EscAgriPP$Tonnes[y])/EscAgriPP$Tonnes[y])*100 b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Tonnes[b] != EscAgriPP$Tonnes[y] } if(x==0){ ChangeScenario<-"NoNewDataForRecentYears" PercentageChange<-NA y<-which(EscAgriPP$Year == Endp) b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Tonnes[b] != EscAgriPP$Tonnes[y] }
52 Pathway<-"escape:agricultureProductionFAO" results1<-cbind(Pathway, ChangeScenario, PercentageChange, BaseChange) Endp<-max(PrevDat$Year) Endn<-max(EscAgriPP$Year) x<-Endn-Endp if(x>0){ ChangeScenario<-"NewDataForRecentYears" y<-which(EscAgriPP$Year == Endp) z<-which(EscAgriPP$Year == Endn) PercentageChange<-((EscAgriPP$ha[z]-EscAgriPP$ha[y])/EscAgriPP$ha[y])*100 b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$ha[b] != EscAgriPP$ha[y] } if(x==0){ ChangeScenario<-"NoNewDataForRecentYears" PercentageChange<-NA y<-which(EscAgriPP$Year == Endp) b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$ha[b] != EscAgriPP$ha[y] } Pathway<-"escape:agriultureAreaFAO" results2<-cbind(Pathway, ChangeScenario, PercentageChange, BaseChange) Post-processing The processed data are saved in csv format in the folder ‘Processed_data’ which was created in previous steps write.csv(EscAgriPP, "Processed\\Processed_data\\Escape_agriculture_FAOSTAT_processed.csv", row.names = FALSE) The outputs on changes to the data since the last report are written to the ‘Change_tracking’ csv in the ‘Processed’ folder. write.table(results1,"Processed\\Change_tracking.csv", row.names=F, append=T, quote= FALSE, sep=",", col.names = F) write.table(results2,"Processed\\Change_tracking.csv", row.names=F, append=T, quote= FALSE, sep=",", col.names = F) The plot is saved as a tiff in the plots folder created in previous steps. tiff(filename = "Processed\\Plots\\Escape_agriculture_FAOSTAT.tiff", width = 6, height = 4, units = "in", compression = "lzw", bg = "white", res = 1000) par(mar = c(4,5,1,5)+ 0.3) plot(EscAgriPP$Year, EscAgriPP$ha/1000000, pch = 18, ylab = "Area harvested (million ha)", xlab = "Year", ylim = c(0, 10)) # Create first plot. par(new = TRUE) plot(EscAgriPP$Year, EscAgriPP$Tonnes/1000000, pch = 15,
53 col = "gray", axes = FALSE, xlab = "", ylab = "", ylim = c(0, 55)) # Create second plot without axes axis(side = 4, at = pretty(range(EscAgriPP$Tonnes/1000000))) # Add second axis axis(side = 4, at = seq(0, 15, by = 5), labels = c("", "5", "10", "15"))# Add labels mtext("Production (million tonnes)", side = 4, line = 3) # Add second axes label graphics.off() Stowaway: Container/bulk (Transnet) Data on the number of deepsea containers landed in South Africa yearly were obtained from Transnet National Ports Authority. The original data were downloaded as a pdf, but were manually extracted and inputted into a csv in previous steps of this workflow. View data structure str(StowCargo) # data structure ## 'data.frame': 13 obs. of 2 variables: ## $ Year : int 2009 2010 2011 2012 2014 2015 2016 2017 2018 2019 ... ## $ Total_teus: int 1618378 1510837 1619738 1686291 1728135 1769498 1670693 NA 1872412 2319445 ... Standardisation These data do not need to be standardised Pre-processing Check for issues with recent data and make necessary adjustments. For example, data for incomplete years included, or data values indicate that the data are not accurate (e.g. not all data has been submitted) tail(StowCargo) # tail of data ## Year Total_teus ## 8 2017 NA ## 9 2018 1872412 ## 10 2019 2319445 ## 11 2020 1605458 ## 12 2021 1790298 ## 13 2022 1774171 No adjustments required Processing Plot data Data are presented as millions of teus. windows(6,4) par(mar = c(4,5,1,1)) plot(StowCargo$Year, StowCargo$Total_teus/1000000, pch = 18, ylab = "Number of deepsea containers landed\n(million teus)", xlab = "Year", ylim = c(0,2.5))
54 Perform calculations Read in processed data used in previous report. These data have been saved in the ‘Processed_data_previous_report’ folder. Paste the file path to your file in the quotation marks in the code below (i.e. replace the file path in the brackets with that for your file). Insert double backslashes where required PrevDat<-read.csv("Processed_data_previous_report\\Stowaway_cargo_TRANSNET_processed.csv") See if new data have become available, if the baseline from the previous report has changed, and if new data are available, calculate the percentage change. Endp<-max(PrevDat$Year) Endn<-max(StowCargo$Year) x<-Endn-Endp if(x>0){ ChangeScenario<-"NewDataForRecentYears" y<-which(StowCargo$Year == Endp) z<-which(StowCargo$Year == Endn) PercentageChange<-((StowCargo$Total_teus[z]- StowCargo$Total_teus[y])/StowCargo$Total_teus[y])*100 b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Total_teus[b] != StowCargo$Total_teus[y] } if(x==0){ ChangeScenario<-"NoNewDataForRecentYears" PercentageChange<-NA y<-which(StowCargo$Year == Endp) b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Total_teus[b] != StowCargo$Total_teus[y] } Pathway<-"stowaway:cargoTransnet" results<-cbind(Pathway, ChangeScenario, PercentageChange, BaseChange) Post-processing The processed data are saved in csv format in the folder ‘Processed_data’ which was created in previous steps write.csv(StowCargo, "Processed\\Processed_data\\Stowaway_cargo_TRANSNET_processed.csv", row.names = FALSE) The outputs on changes to the data since the last report are written to the ‘Change_tracking’ csv in the ‘Processed’ folder. write.table(results,"Processed\\Change_tracking.csv", row.names=F, append=T, quote= FALSE, sep=",", col.names = F) The plot is saved as a tiff in the plots folder created in previous steps. tiff(filename = "Processed\\Plots\\Stowaway_cargo_TRANSNET.tiff", width = 6, height = 4, units = "in",
55 compression = "lzw", bg = "white", res = 1000) par(mar = c(4,5,1,1)) plot(StowCargo$Year, StowCargo$Total_teus/1000000, pch = 18, ylab = "Number of deepsea containers landed\n(million teus)", xlab = "Year", ylim = c(0,2.5)) graphics.off() Stowaway: Ship/boat hull fouling; Stowaway: Ship/boat ballast water; Stowaway: Hitchhikers on ship/boat (Transnet) Data on the total number of ocean going vessels arriving in South Africa yearly were downloaded from Transnet National Ports Authority. The original data were downloaded as a pdf, but were extracted and inputted into a csv during manual processeing in previous steps of this workflow. View data structure str(StowShip) # data structure ## 'data.frame': 19 obs. of 2 variables: ## $ Year : int 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 ... ## $ Total_arrivals: int 9405 8833 8929 9200 9099 9331 9251 NA NA NA ... Standardisation Data does not need to be standardised Pre-processing Check for issues with recent data and make necessary adjustments. For example, data for incomplete years included, or data values indicate that the data are not accurate (e.g. not all data has been submitted) tail(StowShip) # tail of data ## Year Total_arrivals ## 14 2017 NA ## 15 2018 7622 ## 16 2019 8041 ## 17 2020 7654 ## 18 2021 7094 ## 19 2022 7096 No adjustments required Processing Plot data windows(6,4) par(mar = c(4,5,1,1)) plot(StowShip$Year, StowShip$Total_arrivals, pch = 18, ylab = "Number of ocean going vessel arrivals", xlab = "Year", ylim = c(0, 10000)) Perform calculations
56 Read in processed data used in previous report. These data have been saved in the ‘Processed_data_previous_report’ folder. Paste the file path to your file in the quotation marks in the code below (i.e. replace the file path in the brackets with that for your file). Insert double backslashes where required PrevDat<-read.csv("Processed_data_previous_report\\Stowaway_ship_TRANSNET_processed.csv") See if new data have become available, if the baseline from the previous report has changed, and if new data are available, calculate the percentage change. Endp<-max(PrevDat$Year) Endn<-max(StowShip$Year) x<-Endn-Endp if(x>0){ ChangeScenario<-"NewDataForRecentYears" y<-which(StowShip$Year == Endp) z<-which(StowShip$Year == Endn) PercentageChange<-((StowShip$Total_arrivals[z]- StowShip$Total_arrivals[y])/StowShip$Total_arrivals[y])*100 b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Total_arrivals[b] != StowShip$Total_arrivals[y] } if(x==0){ ChangeScenario<-"NoNewDataForRecentYears" PercentageChange<-NA y<-which(StowShip$Year == Endp) b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Total_arrivals[b] != StowShip$Total_arrivals[y] } Pathway<-"stowaway:shipTransnet" results<-cbind(Pathway, ChangeScenario, PercentageChange, BaseChange) Post-processing The processed data are saved in csv format in the folder ‘Processed_data’ which was created in previous steps write.csv(StowShip, "Processed\\Processed_data\\Stowaway_ship_TRANSNET_processed.csv", row.names = FALSE) The outputs on changes to the data since the last report are written to the ‘Change_tracking’ csv in the ‘Processed’ folder. write.table(results,"Processed\\Change_tracking.csv", row.names=F, append=T, quote= FALSE, sep=",", col.names = F) The plot is saved as a tiff in the plots folder created in previous steps. tiff(filename = "Processed\\Plots\\Stowaway_ship_TRANSNET.tiff", width = 6, height = 4, units = "in", compression = "lzw", bg = "white", res = 1000)
57 par(mar = c(4,5,1,1)) plot(StowShip$Year, StowShip$Total_arrivals, pch = 18, ylab = "Number of ocean going vessel arrivals", xlab = "Year", ylim = c(0, 10000)) graphics.off() Stowaway: Angling/fishing equipment (Transnet) Data on the total number of foreign and South African trawlers in South Africa yearly were obtained from Transnet National Ports Authority. The original data were downloaded as a pdf, but were manually processed in previous steps of this workflow. View data structure str(StowFishVes) # data structure ## 'data.frame': 18 obs. of 3 variables: ## $ Year : int 2014 2015 2016 2017 2018 2019 2020 2021 2022 2014 ... ## $ Trawler: chr "South_African" "South_African" "South_African" "South_African" ... ## $ Number : int 1186 1033 879 NA 468 521 385 542 NA 533 ... Standardisation Data does not need to be standardised Pre-processing Aggregate data. Calculate the total number of fishing vessels per year. StowFishVesPP<- aggregate(StowFishVes$Number, by = list(StowFishVes$Year), FUN = sum) colnames(StowFishVesPP)<-c("Year", "Number") # set column names head(StowFishVesPP) # header of data for area harvested ## Year Number ## 1 2014 1719 ## 2 2015 1518 ## 3 2016 1245 ## 4 2017 NA ## 5 2018 757 ## 6 2019 891 Check for issues with recent data and make necessary adjustments. For example, data for incomplete years included, or data values indicate that the data are not accurate (e.g. not all data has been submitted) tail(StowFishVesPP) # tail of data ## Year Number ## 4 2017 NA ## 5 2018 757 ## 6 2019 891 ## 7 2020 682
64 tmp<-which(EscResPP$Year > 2020) EscResPP<-EscResPP[-tmp,] Processing Plot data windows(6,4) par(mar = c(4,5,1,1)) plot(EscResPP$Year, EscResPP$Number, pch = 18, ylab = "Number of organisms", xlab = "Year") Perform calculations Read in processed data used in previous report. These data have been saved in the ‘Processed_data_previous_report’ folder. Paste the file path to the required file in the quotation marks in the code below (i.e. replace the file path in the brackets with that for the required file). Insert double backslashes where required PrevDat<-read.csv("Processed_data_previous_report\\Escape_research_CITES_processed.csv") See if new data have become available, if the baseline from the previous report have changed, and if new data are available calculate the percentage change. Endp<-max(PrevDat$Year) Endn<-max(EscResPP$Year) x<-Endn-Endp if(x>0){ ChangeScenario<-"NewDataForRecentYears" y<-which(EscResPP$Year == Endp) z<-which(EscResPP$Year == Endn) PercentageChange<-((EscResPP$Number[z]-EscResPP$Number[y])/EscResPP$Number[y])*100 b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Number[b] != EscResPP$Number[y] } if(x==0){ ChangeScenario<-"NoNewDataForRecentYears" PercentageChange<-NA y<-which(EscResPP$Year == Endp) b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Number[b] != EscResPP$Number[y] } Pathway<-"escape:researchCITES" results<-cbind(Pathway, ChangeScenario, PercentageChange, BaseChange) Post-processing The processed data are saved in csv format in the folder ‘Processed_data’ which was created in previous steps write.csv(EscResPP, "Processed\\Processed_data\\Escape_research_CITES_processed.csv", row.names = FALSE)
65 The outputs on changes to the data since the last report are written to the ‘Change_tracking’ csv in the ‘Processed’ folder. write.table(results,"Processed\\Change_tracking.csv", row.names=F, append=T, quote= FALSE, sep=",", col.names = F) The plot is saved as a tiff in the plots folder created in previous steps. tiff(filename = "Processed\\Plots\\Escape_research_CITES.tiff", width = 6, height = 4, units = "in", compression = "lzw", bg = "white", res = 1000) par(mar = c(4,5,1,1)) plot(EscResPP$Year, EscResPP$Number, pch = 18, ylab = "Number of organisms", xlab = "Year") graphics.off() Escape: Pet Data on animals imported for personal or commercial purposes per year were downloaded from CITES. Therefore, there will be more than one value per year. Data are in short form. In a few instances units are grams/kilograms or other units. View data structure str(EscPet) # data structure ## 'data.frame': 5155 obs. of 53 variables: ## $ App. : chr "I" "I" "I" "I" ... ## $ Taxon : chr "Acinonyx jubatus" "Addax nasomaculatus" "Aerangis ellisii" "Aloe bellatula" ... ## $ Term : chr "live" "live" "live" "live" ... ## $ Unit : chr "" "" "" "" ... ## $ Country: chr "ZA" "ZA" "ZA" "ZA" ... ## $ X1975 : logi NA NA NA NA NA NA ... ## $ X1976 : logi NA NA NA NA NA NA ... ## $ X1977 : logi NA NA NA NA NA NA ... ## $ X1978 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1979 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1980 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1981 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1982 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1983 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1984 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1985 : int 11 NA NA NA NA NA NA NA NA NA ... ## $ X1986 : int 15 NA NA NA NA NA NA NA NA NA ... ## $ X1987 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1988 : int 23 NA NA NA NA NA NA NA NA NA ... ## $ X1989 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1990 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1991 : int 2 NA NA NA NA NA NA NA NA NA ... ## $ X1992 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1993 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1994 : int 4 NA NA NA NA NA NA NA NA NA ... ## $ X1995 : num 26 NA NA 2 NA 2 4 2 2 NA ...
66 ## $ X1996 : int 1 NA NA NA NA NA NA NA NA NA ... ## $ X1997 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1998 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1999 : int 4 25 NA NA NA NA NA NA NA NA ... ## $ X2000 : int 1 NA NA NA NA NA NA NA NA NA ... ## $ X2001 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2002 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2003 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2004 : int 2 NA NA NA NA NA NA NA NA NA ... ## $ X2005 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2006 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2007 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2008 : num 2 NA NA NA NA NA NA NA NA NA ... ## $ X2009 : int NA NA NA NA NA 4 5 4 NA NA ... ## $ X2010 : int NA NA NA NA 8 NA NA NA NA NA ... ## $ X2011 : int NA 5 NA NA 6 NA NA NA NA 5 ... ## $ X2012 : num NA NA NA NA NA NA NA NA NA NA ... ## $ X2013 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2014 : int 1 NA NA NA NA NA NA NA NA NA ... ## $ X2015 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2016 : int NA NA 2 NA NA NA NA NA NA NA ... ## $ X2017 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2018 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2019 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2020 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2021 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2022 : logi NA NA NA NA NA NA ... Standardisation Remove data collected in grams/kg or other units unique(EscPet$Unit) # units data were recorded in ## [1] "" "Number of specimens" "kg" ## [4] "flasks" "g" "sets" tmp<-which(EscPet$Unit == "g") # which rows contain data in g tmp2<-which(EscPet$Unit == "kg") # which rows contain data in kg tmp3<-which(EscPet$Unit == "sets") # which rows contain data in sets tmp4<-which(EscPet$Unit == "flasks") # which rows contain data in flasks EscPetPP<-EscPet[-c(tmp, tmp2, tmp3, tmp4),] # remove rows Pre-processing Aggregate data. Calculate the total number of organisms for each year names(EscPetPP) # column names ## [1] "App." "Taxon" "Term" "Unit" "Country" "X1975" "X1976" ## [8] "X1977" "X1978" "X1979" "X1980" "X1981" "X1982" "X1983" ## [15] "X1984" "X1985" "X1986" "X1987" "X1988" "X1989" "X1990" ## [22] "X1991" "X1992" "X1993" "X1994" "X1995" "X1996" "X1997"
67 ## [29] "X1998" "X1999" "X2000" "X2001" "X2002" "X2003" "X2004" ## [36] "X2005" "X2006" "X2007" "X2008" "X2009" "X2010" "X2011" ## [43] "X2012" "X2013" "X2014" "X2015" "X2016" "X2017" "X2018" ## [50] "X2019" "X2020" "X2021" "X2022" EscPetPP<-EscPetPP[,c(6:ncol(EscPetPP))] # restrict data to columns with counts EscPetPP<-colSums(EscPetPP, na.rm = TRUE) # calculate the total per year Create a long form dataset Year<-strsplit(names(EscPetPP), "X") # remove 'X' from the years Year<-unlist(Year) # create vector with the years tmp3<-which(Year == "") # identify missing values Year<-Year[-tmp3] # remove missing values EscPetPP<-as.data.frame(EscPetPP) EscPetPP<-cbind(Year, EscPetPP) # bind the counts to the years row.names(EscPetPP)<-NULL colnames(EscPetPP)[2]<-"Number" # column name head(EscPetPP) # header of data ## Year Number ## 1 1975 0 ## 2 1976 0 ## 3 1977 0 ## 4 1978 2 ## 5 1979 332 ## 6 1980 2105 EscPetPP$Year<-as.numeric(EscPetPP$Year) Check for issues with recent data and make necessary adjustments. For example, data for incomplete years included, or data values indicate that the data are not accurate (e.g. not all data has been submitted) tail(EscPetPP) # tail of data ## Year Number ## 43 2017 141809 ## 44 2018 68540 ## 45 2019 95411 ## 46 2020 190310 ## 47 2021 105689 ## 48 2022 0 Remove the two most recent years as data does not appear to be complete tmp<-which(EscPetPP$Year > 2020) EscPetPP<-EscPetPP[-tmp,] Processing Plot data Data are presented as thousands of animals
68 windows(6,4) par(mar = c(4,5,1,1)) plot(EscPetPP$Year, EscPetPP$Number/1000, pch = 18, ylab = "Number of organisms\n(thousands)", xlab = "Year") Perform calculations Read in processed data used in previous report. These data have been saved in the ‘Processed_data_previous_report’ folder. Paste the file path to the required file in the quotation marks in the code below (i.e. replace the file path in the brackets with that for your file). Insert double backslashes where required PrevDat<-read.csv("Processed_data_previous_report\\Escape_pet_CITES_processed.csv") See if new data have become available, if the baseline from the previous report has changed, and if new data are available, calculate the percentage change. Endp<-max(PrevDat$Year) Endn<-max(EscPetPP$Year) x<-Endn-Endp if(x>0){ ChangeScenario<-"NewDataForRecentYears" y<-which(EscPetPP$Year == Endp) z<-which(EscPetPP$Year == Endn) PercentageChange<-((EscPetPP$Number[z]-EscPetPP$Number[y])/EscPetPP$Number[y])*100 b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Number[b] != EscPetPP$Number[y] } if(x==0){ ChangeScenario<-"NoNewDataForRecentYears" PercentageChange<-NA y<-which(EscPetPP$Year == Endp) b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Number[b] != EscPetPP$Number[y] } Pathway<-"escape:petCITES" results<-cbind(Pathway, ChangeScenario, PercentageChange, BaseChange) Post-processing The processed data are saved in csv format in the folder ‘Processed_data’ which was created in previous steps write.csv(EscPetPP, "Processed\\Processed_data\\Escape_pet_CITES_processed.csv", row.names = FALSE) The outputs on changes to the data since the last report are written to the ‘Change_tracking’ csv in the ‘Processed’ folder.
69 write.table(results,"Processed\\Change_tracking.csv", row.names=F, append=T, quote= FALSE, sep=",", col.names = F) The plot is saved as a tiff in the plots folder created in previous steps. tiff(filename = "Processed\\Plots\\Escape_pet_CITES.tiff", width = 6, height = 4, units = "in", compression = "lzw", bg = "white", res = 1000) par(mar = c(4,5,1,1)) plot(EscPetPP$Year, EscPetPP$Number/1000, pch = 18, ylab = "Number of organisms\n(thousands)", xlab = "Year") graphics.off() Escape: Botanical garden and zoo (CITES) Data on the amount of live animals imported per year for botanical garden or zoo purposes were downloaded from CITES. Therefore, there will be more than one value per year. Data are in short form. View data structure str(EscBotZoo) # data structure ## 'data.frame': 428 obs. of 53 variables: ## $ App. : chr "I" "I" "I" "I" ... ## $ Taxon : chr "Acinonyx jubatus" "Acrantophis dumerili" "Addax nasomaculatus" "Ailurus fulgens" ... ## $ Term : chr "live" "live" "live" "live" ... ## $ Unit : chr "" "" "" "" ... ## $ Country: chr "ZA" "ZA" "ZA" "ZA" ... ## $ X1975 : logi NA NA NA NA NA NA ... ## $ X1976 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1977 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1978 : logi NA NA NA NA NA NA ... ## $ X1979 : logi NA NA NA NA NA NA ... ## $ X1980 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1981 : int NA 2 NA NA NA NA NA NA NA NA ... ## $ X1982 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1983 : int 2 NA NA NA NA NA NA NA NA NA ... ## $ X1984 : int 11 NA NA NA NA NA NA NA NA NA ... ## $ X1985 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1986 : int 3 NA NA NA NA NA NA NA NA NA ... ## $ X1987 : int 11 NA NA NA NA NA NA NA NA NA ... ## $ X1988 : int 10 1 NA NA NA NA NA NA NA NA ... ## $ X1989 : int 27 NA NA NA 4 NA NA NA NA NA ... ## $ X1990 : int 9 NA NA NA NA 6 NA NA NA NA ... ## $ X1991 : int 11 NA NA NA NA NA 2 NA NA 2 ... ## $ X1992 : int 7 NA NA NA NA 9 NA 4 1 NA ... ## $ X1993 : int 4 NA NA NA NA NA NA NA NA NA ... ## $ X1994 : int 8 NA NA NA NA NA NA 2 NA NA ... ## $ X1995 : int 43 NA NA NA NA NA NA NA NA NA ...
70 ## $ X1996 : int 1 NA NA 2 NA NA NA NA NA NA ... ## $ X1997 : int 6 NA NA NA NA NA NA NA NA NA ... ## $ X1998 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X1999 : int 2 NA NA NA NA NA NA NA NA NA ... ## $ X2000 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2001 : int 1 NA NA NA NA NA NA NA NA NA ... ## $ X2002 : int NA NA NA NA NA NA 1 NA NA NA ... ## $ X2003 : int NA NA NA NA NA NA 1 NA NA NA ... ## $ X2004 : int NA NA NA NA NA NA NA NA 46 NA ... ## $ X2005 : int NA NA NA NA NA 6 NA NA NA NA ... ## $ X2006 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2007 : int NA NA 1 NA NA NA NA NA NA NA ... ## $ X2008 : int 2 NA 1 NA NA NA NA NA NA NA ... ## $ X2009 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2010 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2011 : int NA NA NA NA NA NA 15 NA NA NA ... ## $ X2012 : int NA NA NA NA NA NA 5 NA NA NA ... ## $ X2013 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2014 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2015 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2016 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2017 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2018 : int NA NA NA NA NA 4 NA NA NA NA ... ## $ X2019 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2020 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2021 : int NA NA NA NA NA NA NA NA NA NA ... ## $ X2022 : logi NA NA NA NA NA NA ... Standardisation The data do not need to be standardised Pre-processing Aggregate data. Calculate the total number of organisms for each year names(EscBotZoo) # column names ## [1] "App." "Taxon" "Term" "Unit" "Country" "X1975" "X1976" ## [8] "X1977" "X1978" "X1979" "X1980" "X1981" "X1982" "X1983" ## [15] "X1984" "X1985" "X1986" "X1987" "X1988" "X1989" "X1990" ## [22] "X1991" "X1992" "X1993" "X1994" "X1995" "X1996" "X1997" ## [29] "X1998" "X1999" "X2000" "X2001" "X2002" "X2003" "X2004" ## [36] "X2005" "X2006" "X2007" "X2008" "X2009" "X2010" "X2011" ## [43] "X2012" "X2013" "X2014" "X2015" "X2016" "X2017" "X2018" ## [50] "X2019" "X2020" "X2021" "X2022" EscBotZooPP<-EscBotZoo[,c(6:ncol(EscBotZoo))] # restrict data to columns with counts EscBotZooPP<-colSums(EscBotZooPP, na.rm = TRUE) # calculate the total per year Create a long form dataset Year<-strsplit(names(EscBotZooPP), "X") # remove 'X' from the years Year<-unlist(Year) # create vector with the years
71 tmp3<-which(Year == "") # identify missing values Year<-Year[-tmp3] # remove missing values EscBotZooPP<-as.data.frame(EscBotZooPP) EscBotZooPP<-cbind(Year, EscBotZooPP) # bind the counts to the years row.names(EscBotZooPP)<-NULL colnames(EscBotZooPP)[2]<-"Number" # column name head(EscBotZooPP) # header of data ## Year Number ## 1 1975 0 ## 2 1976 8 ## 3 1977 1 ## 4 1978 0 ## 5 1979 0 ## 6 1980 7 EscBotZooPP$Year<-as.numeric(EscBotZooPP$Year) Check for issues with recent data and make necessary adjustments. For example, data for incomplete years included, or data values indicate that the data are not accurate (e.g. not all data has been submitted) tail(EscBotZooPP) # tail of data ## Year Number ## 43 2017 71 ## 44 2018 31 ## 45 2019 79 ## 46 2020 30 ## 47 2021 151 ## 48 2022 0 Remove the two most recent years as data are incomplete for these years tmp<-which(EscBotZooPP$Year > 2020) EscBotZooPP<-EscBotZooPP[-tmp,] Processing Plot data windows(6,4) par(mar = c(4,5,1,1)) plot(EscBotZooPP$Year, EscBotZooPP$Number, pch = 18, ylab = "Number of organisms", xlab = "Year") Perform calculations Read in processed data used in previous report. These data have been saved in the ‘Processed_data_previous_report’ folder. Paste the file path to the required file in the quotation marks in the code below (i.e. replace the file path in the brackets with that for the required file). Insert double backslashes where required
72 PrevDat<- read.csv("Processed_data_previous_report\\Escape_botgarden_zoo_CITES_processed.csv") See if new data have become available, if the baseline from the previous report has changed, and if new data are available, calculate the percentage change. Endp<-max(PrevDat$Year) Endn<-max(EscBotZooPP$Year) x<-Endn-Endp if(x>0){ ChangeScenario<-"NewDataForRecentYears" y<-which(EscBotZooPP$Year == Endp) z<-which(EscBotZooPP$Year == Endn) PercentageChange<-((EscBotZooPP$Number[z]- EscBotZooPP$Number[y])/EscBotZooPP$Number[y])*100 b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Number[b] != EscBotZooPP$Number[y] } if(x==0){ ChangeScenario<-"NoNewDataForRecentYears" PercentageChange<-NA y<-which(EscBotZooPP$Year == Endp) b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Number[b] != EscBotZooPP$Number[y] } Pathway<-"escape:botzooCITES" results<-cbind(Pathway, ChangeScenario, PercentageChange, BaseChange) Post-processing The processed data are saved in csv format in the folder ‘Processed_data’ which was created in previous steps write.csv(EscBotZooPP, "Processed\\Processed_data\\Escape_botgarden_zoo_CITES_processed.csv", row.names = FALSE) The outputs on changes to the data since the last report are written to the ‘Change_tracking’ csv in the ‘Processed’ folder. write.table(results,"Processed\\Change_tracking.csv", row.names=F, append=T, quote= FALSE, sep=",", col.names = F) The plot is saved as a tiff in the plots folder created in previous steps. tiff(filename = "Processed\\Plots\\Escape_botgarden_zoo_CITES.tiff", width = 6, height = 4, units = "in", compression = "lzw", bg = "white", res = 1000) par(mar = c(4,5,1,1)) plot(EscBotZooPP$Year, EscBotZooPP$Number, pch = 18, ylab = "Number of organisms", xlab = "Year") graphics.off()
73 Unaided (WTO) Data on the yearly value of merchandise imports for each continental African country, except South Africa, were downloaded from the World Trade Organisation. Therefore, there will be more than one value per year. Data are million US$. No indication in the metadata that these data have been adjusted for inflation. View data structure str(UnaidedWTO) # data structure ## 'data.frame': 3273 obs. of 24 variables: ## $ Indicator.Category : chr "Merchandise trade values" "Merchandise trade values" "Merchandise trade values" "Merchandise trade values" ... ## $ Indicator.Code : chr "ITS_MTV_AM" "ITS_MTV_AM" "ITS_MTV_AM" "ITS_MTV_AM" ... ## $ Indicator : chr "Merchandise imports by product group – annual" "Merchandise imports by product group – annual" "Merchandise imports by product group – annual" "Merchandise imports by product group – annual" ... ## $ Reporting.Economy.Code : int 12 12 12 12 12 12 12 12 12 12 ... ## $ Reporting.Economy.ISO3A.Code : chr "DZA" "DZA" "DZA" "DZA" ... ## $ Reporting.Economy : chr "Algeria" "Algeria" "Algeria" "Algeria" ... ## $ Partner.Economy.Code : int 0 0 0 0 0 0 0 0 0 0 ... ## $ Partner.Economy.ISO3A.Code : logi NA NA NA NA NA NA ... ## $ Partner.Economy : chr "World" "World" "World" "World" ... ## $ Product.Sector.Classification.Code: chr "SITC3" "SITC3" "SITC3" "SITC3" ... ## $ Product.Sector.Classification : chr "Merchandise - SITC Revision 3 (aggregates)" "Merchandise - SITC Revision 3 (aggregates)" "Merchandise - SITC Revision 3 (aggregates)" "Merchandise - SITC Revision 3 (aggregates)" ... ## $ Product.Sector.Code : chr "TO" "TO" "TO" "TO" ... ## $ Product.Sector : chr "Total merchandise" "Total merchandise" "Total merchandise" "Total merchandise" ... ## $ Period.Code : chr "A" "A" "A" "A" ... ## $ Period : chr "Annual" "Annual" "Annual" "Annual" ... ## $ Frequency.Code : chr "A" "A" "A" "A" ... ## $ Frequency : chr "Annual" "Annual" "Annual" "Annual" ... ## $ Unit.Code : chr "USM" "USM" "USM" "USM" ... ## $ Unit : chr "Million US dollar" "Million US dollar" "Million US dollar" "Million US dollar" ... ## $ Year : int 2001 1994 2002 2000 1993 1992 1990 1995 2006 1979 ... ## $ Value.Flag.Code : chr "" "" "" "" ... ## $ Value.Flag : chr "" "" "" "" ... ## $ Text.Value : logi NA NA NA NA NA NA ... ## $ Value : num 9940 9154 11969 9171 8785 ... Standardisation Convert the data, which is in 1 million USD, into USD unique(UnaidedWTO$Unit) # units of data ## [1] "Million US dollar"
80 These data do not require standardisation Pre-processing Replace odd values with NA, and remove commas names(UnaidedIMF) ## [1] "Country" "Subject.Descriptor" ## [3] "Units" "Scale" ## [5] "Country.Series.specific.Notes" "X1980" ## [7] "X1981" "X1982" ## [9] "X1983" "X1984" ## [11] "X1985" "X1986" ## [13] "X1987" "X1988" ## [15] "X1989" "X1990" ## [17] "X1991" "X1992" ## [19] "X1993" "X1994" ## [21] "X1995" "X1996" ## [23] "X1997" "X1998" ## [25] "X1999" "X2000" ## [27] "X2001" "X2002" ## [29] "X2003" "X2004" ## [31] "X2005" "X2006" ## [33] "X2007" "X2008" ## [35] "X2009" "X2010" ## [37] "X2011" "X2012" ## [39] "X2013" "X2014" ## [41] "X2015" "X2016" ## [43] "X2017" "X2018" ## [45] "X2019" "X2020" ## [47] "X2021" "X2022" ## [49] "X2023" "X2024" ## [51] "X2025" "X2026" ## [53] "X2027" "X2028" ## [55] "Estimates.Start.After" tmp<-length(names(UnaidedIMF))-1 UnaidedIMFPP<-UnaidedIMF[,c(6:tmp)] # restrict to columns with values UnaidedIMFPP[UnaidedIMFPP == "" ] <- NA # replace missing value UnaidedIMFPP[UnaidedIMFPP == "n/a" ] <- NA # replace odd value UnaidedIMFPP[UnaidedIMFPP == "--" ] <- NA # replace odd value UnaidedIMFPP<-as.data.frame(lapply(UnaidedIMFPP, function(x) gsub(",", "", x))) # remove comma in one record tmp<-sapply(UnaidedIMFPP, class) # class of data tmp2<-which(tmp == "character") UnaidedIMFPP[tmp2] <- sapply(UnaidedIMFPP[tmp2],as.numeric, na.rm = TRUE) # convert to numeric Aggregate data. For each year calculate the average percentage change across the countries
81 UnaidedIMFPP<-colMeans(UnaidedIMFPP, na.rm = T) Create a long form dataset Year<-names(UnaidedIMFPP) # vector of years Year<-unlist(strsplit(Year, "X")) # remove x's from vector tmp3<-which(Year == "") # identify missing values Year<-Year[-tmp3] # remove missing values UnaidedIMFPP<-as.data.frame(UnaidedIMFPP) UnaidedIMFPP<-cbind(Year, UnaidedIMFPP) # bind values to years row.names(UnaidedIMFPP)<-NULL colnames(UnaidedIMFPP)[2]<-"Percent_change" # column name UnaidedIMFPP$Year<-as.numeric(UnaidedIMFPP$Year) head(UnaidedIMFPP) # header of data ## Year Percent_change ## 1 1980 2.464194 ## 2 1981 5.917222 ## 3 1982 1.040444 ## 4 1983 -4.093417 ## 5 1984 2.708750 ## 6 1985 5.737000 Check for issues with recent data and make necessary adjustments. For example, data for incomplete years included, or data values indicate that the data are not accurate (e.g. not all data has been submitted) tail(UnaidedIMFPP) # tail of data ## Year Percent_change ## 44 2023 7.123733 ## 45 2024 5.311667 ## 46 2025 4.301844 ## 47 2026 3.825378 ## 48 2027 3.615644 ## 49 2028 4.280800 Most estimates start after 2020, so remove data for years after this tmp<-which(UnaidedIMFPP$Year > 2020) UnaidedIMFPP<-UnaidedIMFPP[-tmp,] Processing Plot data windows(6,4) par(mar = c(4,5,1,1)) plot(UnaidedIMFPP$Year[UnaidedIMFPP$Year >=2010], UnaidedIMFPP$Percent_change[UnaidedIMFPP$Year >=2010], pch = 18, ylab = "Percent change in the volume\nof imported goods", xlab = "Year") Perform calculations
82 Read in processed data used in previous report. These data have been saved in the ‘Processed_data_previous_report’ folder. Paste the file path to the required file in the quotation marks in the code below (i.e. replace the file path in the brackets with that for the required file). Insert double backslashes where required PrevDat<-read.csv("Processed_data_previous_report\\Unaided_IMF_processed.csv") See if new data have become available, if the baseline from the previous report has changed, and if new data are available, calculate the percentage change. Endp<-max(PrevDat$Year) Endn<-max(UnaidedIMFPP$Year) x<-Endn-Endp if(x>0){ ChangeScenario<-"NewDataForRecentYears" y<-which(UnaidedIMFPP$Year == Endp) z<-which(UnaidedIMFPP$Year == Endn) PercentageChange<-((UnaidedIMFPP$Percent_change[z]- UnaidedIMFPP$Percent_change[y])/UnaidedIMFPP$Percent_change[y])*100 b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Percent_change[b] != UnaidedIMFPP$Percent_change[y] } if(x==0){ ChangeScenario<-"NoNewDataForRecentYears" PercentageChange<-NA y<-which(UnaidedIMFPP$Year == Endp) b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Percent_change[b] != UnaidedIMFPP$Percent_change[y] } Pathway<-"unaided:IMF" results<-cbind(Pathway, ChangeScenario, PercentageChange, BaseChange) Post-processing The processed data are saved in csv format in the folder ‘Processed_data’ which was created in previous steps write.csv(UnaidedIMFPP, "Processed\\Processed_data\\Unaided_IMF_processed.csv", row.names = FALSE) The outputs on changes to the data since the last report are written to the ‘Change_tracking’ csv in the ‘Processed’ folder. write.table(results,"Processed\\Change_tracking.csv", row.names=F, append=T, quote= FALSE, sep=",", col.names = F) The plot is saved as a tiff in the plots folder created in previous steps. tiff(filename = "Processed\\Plots\\Unaided_IMF.tiff", width = 6, height = 4, units = "in", compression = "lzw", bg = "white", res = 1000)
83 par(mar = c(4,5,1,1)) plot(UnaidedIMFPP$Year[UnaidedIMFPP$Year >=2010], UnaidedIMFPP$Percent_change[UnaidedIMFPP$Year >=2010], pch = 18, ylab = "Percent change in the volume\nof imported goods", xlab = "Year") graphics.off() Plot data for unaided from both the IMF and WTO. Data on value of imports are presented as billions of US dollars windows(12,16) par(mfrow = c(2,1)) par(mar = c(4,5,1,1)) plot(UnaidedWTOPP$Year, UnaidedWTOPP$Value_2017USD/1000000000, pch = 18, ylab = "Total merchandise imports\n(billions of US dollars)", xlab = "Year") mtext("A", side = 3, line = 0.2, adj = -0.2) plot(UnaidedIMFPP$Year[UnaidedIMFPP$Year >=2010], UnaidedIMFPP$Percent_change[UnaidedIMFPP$Year >=2010], pch = 18, ylab = "Percent change in the volume\nof imported goods", xlab = "Year") mtext("B", side = 3, line = 0.2, adj = -0.2) Saved as a tiff in the plots folder created in previous steps. tiff(filename = "Processed\\Plots\\Unaided_IMF_WTO.tiff", width = 12, height = 16, units = "in", compression = "lzw", bg = "white", res = 1000) par(mfrow = c(2,1)) par(mar = c(4,5,1,1)) plot(UnaidedWTOPP$Year, UnaidedWTOPP$Value_2017USD/1000000000, pch = 18, ylab = "Total merchandise imports\n(billions of US dollars)", xlab = "Year") mtext("A", side = 3, line = 0.2, adj = -0.2) plot(UnaidedIMFPP$Year[UnaidedIMFPP$Year >=2010], UnaidedIMFPP$Percent_change[UnaidedIMFPP$Year >=2010], pch = 18, ylab = "Percent change in the volume\nof imported goods", xlab = "Year") mtext("B", side = 3, line = 0.2, adj = -0.2) graphics.off() Escape: Aquaculture (FishStatJ) Data on the quantity of aquaculture production in South Africa were downloaded from the FishStatJ database of the UN’s Food and Agriculture Organisation. These are yearly data in tonnes. Data for various types of organisms were downloaded and so there is more than one value per year, but one row of the dataset includes the totals for each year. Data are in short form. View data structure str(EscAqua) # data structure
84 ## 'data.frame': 3 obs. of 150 variables: ## $ Country..Name. : chr "South Africa" "Totals - Tonnes - live weight" "FAO. 2023. Fishery and Aquaculture Statistics. Global aquaculture production 1950-2021 (FishStatJ). In: FAO Fis"| __truncated__ ## $ ASFIS.species..Name. : chr "All" "" "" ## $ FAO.major.fishing.area..Name.: chr "All" "" "" ## $ Environment..Name. : chr "All" "" "" ## $ Unit..Name. : chr "Tonnes - live weight" "" "" ## $ Unit : chr "TLW" "" "" ## $ X.1950. : int 0 0 NA ## $ S : chr "..." "" "" ## $ X.1951. : int 0 0 NA ## $ S.1 : chr "..." "" "" ## $ X.1952. : int 0 0 NA ## $ S.2 : chr "..." "" "" ## $ X.1953. : int 0 0 NA ## $ S.3 : chr "..." "" "" ## $ X.1954. : int 0 0 NA ## $ S.4 : chr "..." "" "" ## $ X.1955. : int 0 0 NA ## $ S.5 : chr "..." "" "" ## $ X.1956. : int 0 0 NA ## $ S.6 : chr "..." "" "" ## $ X.1957. : int 0 0 NA ## $ S.7 : chr "..." "" "" ## $ X.1958. : int 0 0 NA ## $ S.8 : chr "..." "" "" ## $ X.1959. : int 0 0 NA ## $ S.9 : chr "..." "" "" ## $ X.1960. : int 0 0 NA ## $ S.10 : chr "..." "" "" ## $ X.1961. : int 0 0 NA ## $ S.11 : chr "..." "" "" ## $ X.1962. : int 0 0 NA ## $ S.12 : chr "..." "" "" ## $ X.1963. : int 0 0 NA ## $ S.13 : chr "..." "" "" ## $ X.1964. : int 0 0 NA ## $ S.14 : chr "..." "" "" ## $ X.1965. : int 0 0 NA ## $ S.15 : chr "..." "" "" ## $ X.1966. : int 0 0 NA ## $ S.16 : chr "..." "" "" ## $ X.1967. : int 0 0 NA ## $ S.17 : chr "..." "" "" ## $ X.1968. : int 0 0 NA ## $ S.18 : chr "..." "" "" ## $ X.1969. : int 0 0 NA ## $ S.19 : chr "..." "" "" ## $ X.1970. : int 0 0 NA ## $ S.20 : chr "N" "" ""
85 ## $ X.1971. : int 0 0 NA ## $ S.21 : chr "N" "" "" ## $ X.1972. : int 0 0 NA ## $ S.22 : chr "N" "" "" ## $ X.1973. : int 0 0 NA ## $ S.23 : chr "N" "" "" ## $ X.1974. : int 9 9 NA ## $ S.24 : logi NA NA NA ## $ X.1975. : int 9 9 NA ## $ S.25 : logi NA NA NA ## $ X.1976. : int 9 9 NA ## $ S.26 : logi NA NA NA ## $ X.1977. : int 12 12 NA ## $ S.27 : logi NA NA NA ## $ X.1978. : int 14 14 NA ## $ S.28 : logi NA NA NA ## $ X.1979. : int 14 14 NA ## $ S.29 : logi NA NA NA ## $ X.1980. : int 13 13 NA ## $ S.30 : logi NA NA NA ## $ X.1981. : int 11 11 NA ## $ S.31 : logi NA NA NA ## $ X.1982. : int 10 10 NA ## $ S.32 : logi NA NA NA ## $ X.1983. : int 6 6 NA ## $ S.33 : logi NA NA NA ## $ X.1984. : int 366 366 NA ## $ S.34 : chr "E" "" "" ## $ X.1985. : int 615 615 NA ## $ S.35 : chr "E" "" "" ## $ X.1986. : int 789 789 NA ## $ S.36 : chr "E" "" "" ## $ X.1987. : int 1096 1096 NA ## $ S.37 : chr "E" "" "" ## $ X.1988. : int 1376 1376 NA ## $ S.38 : logi NA NA NA ## $ X.1989. : int 2037 2037 NA ## $ S.39 : logi NA NA NA ## $ X.1990. : int 3613 3613 NA ## $ S.40 : logi NA NA NA ## $ X.1991. : int 4882 4882 NA ## $ S.41 : logi NA NA NA ## $ X.1992. : int 4288 4288 NA ## $ S.42 : logi NA NA NA ## $ X.1993. : int 4021 4021 NA ## $ S.43 : logi NA NA NA ## $ X.1994. : int 4429 4429 NA ## $ S.44 : logi NA NA NA ## $ X.1995. : int 3530 3530 NA ## $ S.45 : logi NA NA NA
86 ## $ X.1996. : int 3013 3013 NA ## [list output truncated] Standardisation These data do not need to be standardised Pre-processing Restrict dataset to required rows and columns. Some columns do not contain data, and only the row of data with the total quantity produced per year is required. names(EscAqua) # column names ## [1] "Country..Name." "ASFIS.species..Name." ## [3] "FAO.major.fishing.area..Name." "Environment..Name." ## [5] "Unit..Name." "Unit" ## [7] "X.1950." "S" ## [9] "X.1951." "S.1" ## [11] "X.1952." "S.2" ## [13] "X.1953." "S.3" ## [15] "X.1954." "S.4" ## [17] "X.1955." "S.5" ## [19] "X.1956." "S.6" ## [21] "X.1957." "S.7" ## [23] "X.1958." "S.8" ## [25] "X.1959." "S.9" ## [27] "X.1960." "S.10" ## [29] "X.1961." "S.11" ## [31] "X.1962." "S.12" ## [33] "X.1963." "S.13" ## [35] "X.1964." "S.14" ## [37] "X.1965." "S.15" ## [39] "X.1966." "S.16" ## [41] "X.1967." "S.17" ## [43] "X.1968." "S.18" ## [45] "X.1969." "S.19" ## [47] "X.1970." "S.20" ## [49] "X.1971." "S.21" ## [51] "X.1972." "S.22" ## [53] "X.1973." "S.23" ## [55] "X.1974." "S.24" ## [57] "X.1975." "S.25" ## [59] "X.1976." "S.26" ## [61] "X.1977." "S.27" ## [63] "X.1978." "S.28" ## [65] "X.1979." "S.29" ## [67] "X.1980." "S.30" ## [69] "X.1981." "S.31" ## [71] "X.1982." "S.32" ## [73] "X.1983." "S.33"
87 ## [75] "X.1984." "S.34" ## [77] "X.1985." "S.35" ## [79] "X.1986." "S.36" ## [81] "X.1987." "S.37" ## [83] "X.1988." "S.38" ## [85] "X.1989." "S.39" ## [87] "X.1990." "S.40" ## [89] "X.1991." "S.41" ## [91] "X.1992." "S.42" ## [93] "X.1993." "S.43" ## [95] "X.1994." "S.44" ## [97] "X.1995." "S.45" ## [99] "X.1996." "S.46" ## [101] "X.1997." "S.47" ## [103] "X.1998." "S.48" ## [105] "X.1999." "S.49" ## [107] "X.2000." "S.50" ## [109] "X.2001." "S.51" ## [111] "X.2002." "S.52" ## [113] "X.2003." "S.53" ## [115] "X.2004." "S.54" ## [117] "X.2005." "S.55" ## [119] "X.2006." "S.56" ## [121] "X.2007." "S.57" ## [123] "X.2008." "S.58" ## [125] "X.2009." "S.59" ## [127] "X.2010." "S.60" ## [129] "X.2011." "S.61" ## [131] "X.2012." "S.62" ## [133] "X.2013." "S.63" ## [135] "X.2014." "S.64" ## [137] "X.2015." "S.65" ## [139] "X.2016." "S.66" ## [141] "X.2017." "S.67" ## [143] "X.2018." "S.68" ## [145] "X.2019." "S.69" ## [147] "X.2020." "S.70" ## [149] "X.2021." "S.71" ColNum<-seq(7, length(names(EscAqua)), 2) # column numbers where values are stored EscAquaPP<-EscAqua[2,ColNum] # restrict dataset Create a long form dataset Year<-names(EscAquaPP) # create a vector for year Year<-unlist(strsplit(Year, "X.")) # remove X.s Year<-unlist(strsplit(Year, split=".", fixed = TRUE)) # remove points Year # final year vector ## [1] "1950" "1951" "1952" "1953" "1954" "1955" "1956" "1957" "1958" "1959" ## [11] "1960" "1961" "1962" "1963" "1964" "1965" "1966" "1967" "1968" "1969" ## [21] "1970" "1971" "1972" "1973" "1974" "1975" "1976" "1977" "1978" "1979"
88 ## [31] "1980" "1981" "1982" "1983" "1984" "1985" "1986" "1987" "1988" "1989" ## [41] "1990" "1991" "1992" "1993" "1994" "1995" "1996" "1997" "1998" "1999" ## [51] "2000" "2001" "2002" "2003" "2004" "2005" "2006" "2007" "2008" "2009" ## [61] "2010" "2011" "2012" "2013" "2014" "2015" "2016" "2017" "2018" "2019" ## [71] "2020" "2021" Quantity<-as.vector(EscAquaPP,mode='numeric')# vector of the values EscAquaPP<-cbind(Year, Quantity) # bind year and quantity vectors EscAquaPP<-as.data.frame(EscAquaPP) # data frame EscAquaPP$Year<-as.numeric(EscAquaPP$Year) # convert the data to numeric EscAquaPP$Quantity<-as.numeric(EscAquaPP$Quantity) # convert to numeric head(EscAquaPP) # header ## Year Quantity ## 1 1950 0 ## 2 1951 0 ## 3 1952 0 ## 4 1953 0 ## 5 1954 0 ## 6 1955 0 Check for issues with recent data and make necessary adjustments. For example, data for incomplete years included, or data values indicate that the data are not accurate (e.g. not all data has been submitted) tail(EscAquaPP) # tail of data ## Year Quantity ## 67 2016 8094.27 ## 68 2017 6338.32 ## 69 2018 7991.55 ## 70 2019 9224.32 ## 71 2020 9753.01 ## 72 2021 10525.32 No adjustments required Processing Plot data windows(6,4) par(mar = c(4,5,1,1)) plot(EscAquaPP$Year, EscAquaPP$Quantity, pch = 18, ylab = "Aquaculture production\n(tonnes)", xlab = "Year") Perform calculations Read in processed data used in previous report. These data have been saved in the ‘Processed_data_previous_report’ folder. Paste the file path to the required file in the quotation marks in the code below (i.e. replace the file path in the brackets with that for the required file). Insert double backslashes where required
89 PrevDat<- read.csv("Processed_data_previous_report\\Escape_aquaculture_FISHSTATJ_processed.csv") See if new data have become available, if the baseline from the previous report has changed, and if new data are available, calculate the percentage change. Endp<-max(PrevDat$Year) Endn<-max(EscAquaPP$Year) x<-Endn-Endp if(x>0){ ChangeScenario<-"NewDataForRecentYears" y<-which(EscAquaPP$Year == Endp) z<-which(EscAquaPP$Year == Endn) PercentageChange<-((EscAquaPP$Quantity[z]- EscAquaPP$Quantity[y])/EscAquaPP$Quantity[y])*100 b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Quantity[b] != EscAquaPP$Quantity[y] } if(x==0){ ChangeScenario<-"NoNewDataForRecentYears" PercentageChange<-NA y<-which(EscAquaPP$Year == Endp) b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Quantity[b] != EscAquaPP$Quantity[y] } Pathway<-"escape:aquacultureFishStatJ" results<-cbind(Pathway, ChangeScenario, PercentageChange, BaseChange) Post-processing The processed data are saved in csv format in the folder ‘Processed_data’ which was created in previous steps write.csv(EscAquaPP, "Processed\\Processed_data\\Escape_aquaculture_FISHSTATJ_processed.csv", row.names = FALSE) The outputs on changes to the data since the last report are written to the ‘Change_tracking’ csv in the ‘Processed’ folder. write.table(results,"Processed\\Change_tracking.csv", row.names=F, append=T, quote= FALSE, sep=",", col.names = F) The plot is saved as a tiff in the plots folder created in previous steps. tiff(filename = "Processed\\Plots\\Escape_aquaculture_FISHSTATJ.tiff", width = 6, height = 4, units = "in", compression = "lzw", bg = "white", res = 1000) par(mar = c(4,5,1,1)) plot(EscAquaPP$Year, EscAquaPP$Quantity, pch = 18, ylab = "Aquaculture production\n(tonnes)", xlab = "Year") graphics.off()
96 z<-which(StowFishCatchPP$Year == Endn) PercentageChange<-((StowFishCatchPP$Quantity[z]- StowFishCatchPP$Quantity[y])/StowFishCatchPP$Quantity[y])*100 b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Quantity[b] != StowFishCatchPP$Quantity[y] } if(x==0){ ChangeScenario<-"NoNewDataForRecentYears" PercentageChange<-NA y<-which(StowFishCatchPP$Year == Endp) b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Quantity[b] != StowFishCatchPP$Quantity[y] } Pathway<-"stowaway:fishingFishStatJ" results<-cbind(Pathway, ChangeScenario, PercentageChange, BaseChange) Post-processing The processed data are saved in csv format in the folder ‘Processed_data’ which was created in previous steps write.csv(StowFishCatchPP, "Processed\\Processed_data\\Stowaway_fishing_FISHSTATJ_processed.csv", row.names = FALSE) The outputs on changes to the data since the last report are written to the ‘Change_tracking’ csv in the ‘Processed’ folder. write.table(results,"Processed\\Change_tracking.csv", row.names=F, append=T, quote= FALSE, sep=",", col.names = F) The plot is saved as a tiff in the plots folder created in previous steps. tiff(filename = "Processed\\Plots\\Stowaway_fishing_FISHSTATJ.tiff", width = 6, height = 4, units = "in", compression = "lzw", bg = "white", res = 1000) par(mar = c(4,5,1,1)) plot(StowFishCatchPP$Year, StowFishCatchPP$Quantity/1000, pch = 18, ylab = "Quantity of fish caught\n(thousand tonnes)", xlab = "Year", ylim = c(0,2200)) graphics.off() Stowaway: People and their luggage/equipment; Stowaway: Vehicles (StatsSA) Data on the total number of foreign travelers and South African residents arriving in South Africa monthly were obtained from Statistics South Africa. The data includes the total number arriving, and the number arriving at different ports of entry. Therefore, there is more than one value per year. The original data were downloaded as a pdf, but manually processed in previous steps of this workflow. View data structure str(StowPeopVeh) # data structure
97 ## 'data.frame': 21 obs. of 3 variables: ## $ Year : int 2016 2017 2018 2019 2020 2021 2022 2016 2017 2018 ... ## $ Number : int 5379569 5700909 5866531 5764311 1529488 1211859 3694507 16190439 15921082 15860569 ... ## $ Transport: chr "air" "air" "air" "air" ... Standardisation These data do not need to be standardised. Pre-processing Aggregate data StowPeopVehPP <- aggregate(StowPeopVeh$Number, by = list(StowPeopVeh$Year), FUN = sum) colnames(StowPeopVehPP)<-c("Year", "Number") # set column names Check for issues with recent data and make necessary adjustments. For example, data for incomplete years included, or data values indicate that the data are not accurate (e.g. not all data has been submitted) tail(StowPeopVehPP) # tail of data ## Year Number ## 2 2017 21703731 ## 3 2018 21876833 ## 4 2019 21828864 ## 5 2020 6414436 ## 6 2021 4606742 ## 7 2022 11540047 No adjustments required Processing Plot data Air transport data are presented as millions of arrivals; road transport data are presented as millions of arrivals; and sea transport data are presented as thousands of arrivals. windows(6,12) par(mfrow = c(3,1)) par(mar = c(4,5,1,1)) barplot(StowPeopVeh$Number[StowPeopVeh$Transport == "air"]/1000000, ylab = "Number of arrivals\n(millions)", xlab = "Year", names.arg = unique(StowPeopVeh$Year), col = "black") mtext("A", side = 3, line = -0.5, adj = -0.2) barplot(StowPeopVeh$Number[StowPeopVeh$Transport == "road"]/1000000, ylab = "Number of arrivals\n(millions)", xlab = "Year", names.arg = unique(StowPeopVeh$Year), col = "black") mtext("B", side = 3, line = -0.5, adj = -0.2) barplot(StowPeopVeh$Number[StowPeopVeh$Transport == "sea"]/1000, ylab = "Number of arrivals\n(thousands)", xlab = "Year", names.arg = unique(StowPeopVeh$Year), col = "black") mtext("C", side = 3, line = -0.5, adj = -0.2)
98 Perform calculations Read in processed data used in previous report. These data have been saved in the ‘Processed_data_previous_report’ folder. Paste the file path to the required file in the quotation marks in the code below (i.e. replace the file path in the brackets with that for the required file). Insert double backslashes where required PrevDat<- read.csv("Processed_data_previous_report\\Stowaway_people_vehicles_STATSSA_processed.csv") See if new data have become available, if the baseline from the previous report has changed, and if new data are available, calculate the percentage change. Endp<-max(PrevDat$Year) Endn<-max(StowPeopVehPP$Year) x<-Endn-Endp if(x>0){ ChangeScenario<-"NewDataForRecentYears" y<-which(StowPeopVehPP$Year == Endp) z<-which(StowPeopVehPP$Year == Endn) PercentageChange<-((StowPeopVehPP$Number[z]- StowPeopVehPP$Number[y])/StowPeopVehPP$Number[y])*100 b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Number[b] != StowPeopVehPP$Number[y] } if(x==0){ ChangeScenario<-"NoNewDataForRecentYears" PercentageChange<-NA y<-which(StowPeopVehPP$Year == Endp) b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Number[b] != StowPeopVehPP$Number[y] } Pathway<-"stowaway:people_vehiclesSTATSSA" results<-cbind(Pathway, ChangeScenario, PercentageChange, BaseChange) Post-processing The processed data are saved in csv format in the folder ‘Processed_data’ which was created in previous steps write.csv(StowPeopVehPP, "Processed\\Processed_data\\Stowaway_people_vehicles_STATSSA_processed.csv", row.names = FALSE) The outputs on changes to the data since the last report are written to the ‘Change_tracking’ csv in the ‘Processed’ folder. write.table(results,"Processed\\Change_tracking.csv", row.names=F, append=T, quote= FALSE, sep=",", col.names = F) The plot is saved as a tiff in the plots folder created in previous steps.
99 tiff(filename = "Processed\\Plots\\Stowaway_people_vehicles_STATSSA.tiff", width = 6, height = 12, units = "in", compression = "lzw", bg = "white", res = 1000) par(mfrow = c(3,1)) par(mar = c(4,8,1,1)) barplot(StowPeopVeh$Number[StowPeopVeh$Transport == "air"]/1000000, ylab = "Number of arrivals\n(millions)", xlab = "Year", names.arg = unique(StowPeopVeh$Year), col = "black", cex.axis = 2, cex.names = 2, cex.lab = 2.5) mtext("A", side = 3, line = -1.4, adj = -0.2, cex = 2.5) barplot(StowPeopVeh$Number[StowPeopVeh$Transport == "road"]/1000000, ylab = "Number of arrivals\n(millions)", xlab = "Year", names.arg = unique(StowPeopVeh$Year), col = "black", cex.axis = 2, cex.names = 2, cex.lab = 2.5) mtext("B", side = 3, line = -1.4, adj = -0.2, cex = 2.5) barplot(StowPeopVeh$Number[StowPeopVeh$Transport == "sea"]/1000, ylab = "Number of arrivals\n(thousands)", xlab = "Year", names.arg = unique(StowPeopVeh$Year), col = "black", cex.axis = 2, cex.names = 2, cex.lab = 2.5) mtext("C", side = 3, line = -1.4, adj = -0.2, cex = 2.5) graphics.off() Stowaway: People and their luggage (WTTC) Data on the total, yearly contribution of tourism and travel to South Africa’s GDP were downloaded from the World Tourism and Travel Council. Data are in bn US$ at real prices, and are in short form. View data structure str(StowPeopWTTC) # data structure ## 'data.frame': 3 obs. of 36 variables: ## $ Total.contribution.to.GDP: chr "South Africa" "US$ in bn (Real prices)" "Source: WTTC" ## $ X1995 : num NA 10.9 NA ## $ X1996 : num NA 13.5 NA ## $ X1997 : num NA 15 NA ## $ X1998 : num NA 15.3 NA ## $ X1999 : num NA 15.8 NA ## $ X2000 : num NA 16.5 NA ## $ X2001 : num NA 18.4 NA ## $ X2002 : num NA 21.3 NA ## $ X2003 : num NA 22.1 NA ## $ X2004 : num NA 22.8 NA ## $ X2005 : num NA 25.7 NA ## $ X2006 : num NA 29.9 NA ## $ X2007 : num NA 30.7 NA ## $ X2008 : num NA 30.4 NA ## $ X2009 : num NA 29.5 NA ## $ X2010 : num NA 29.5 NA ## $ X2011 : num NA 28.9 NA ## $ X2012 : num NA 31.3 NA ## $ X2013 : num NA 32.3 NA
100 ## $ X2014 : num NA 33.4 NA ## $ X2015 : num NA 33.2 NA ## $ X2016 : num NA 33.3 NA ## $ X2017 : num NA 32.7 NA ## $ X2018 : num NA 32.1 NA ## $ X2019 : num NA 33.3 NA ## $ X2020 : num NA 34.1 NA ## $ X2021 : num NA 35.1 NA ## $ X2022 : num NA 36.1 NA ## $ X2023 : num NA 37.4 NA ## $ X2024 : num NA 38.6 NA ## $ X2025 : num NA 39.9 NA ## $ X2026 : num NA 41.3 NA ## $ X2027 : num NA 42.9 NA ## $ X2028 : num NA 44.4 NA ## $ X2029 : num NA 46.2 NA Standardisation Convert the data, which is in billion USD, into USD names(StowPeopWTTC) # column names ## [1] "Total.contribution.to.GDP" "X1995" ## [3] "X1996" "X1997" ## [5] "X1998" "X1999" ## [7] "X2000" "X2001" ## [9] "X2002" "X2003" ## [11] "X2004" "X2005" ## [13] "X2006" "X2007" ## [15] "X2008" "X2009" ## [17] "X2010" "X2011" ## [19] "X2012" "X2013" ## [21] "X2014" "X2015" ## [23] "X2016" "X2017" ## [25] "X2018" "X2019" ## [27] "X2020" "X2021" ## [29] "X2022" "X2023" ## [31] "X2024" "X2025" ## [33] "X2026" "X2027" ## [35] "X2028" "X2029" StowPeopWTTCS<-StowPeopWTTC ColNum<-2:length(names(StowPeopWTTCS)) StowPeopWTTCS<-StowPeopWTTCS[2,ColNum] StowPeopWTTCS[1,]<-StowPeopWTTCS[1,]*1000000000 # multiply values by 1 billion Pre-processing Reorganise data so that it is in long form Year<-colnames(StowPeopWTTCS) # convert column names (Years) into a vector Year<-unlist(strsplit(Year, "X")) # remove Xs tmp<-which(Year == "") # which values are missing
101 Year<-Year[-tmp] # remove missing values Year # look at vector ## [1] "1995" "1996" "1997" "1998" "1999" "2000" "2001" "2002" "2003" "2004" ## [11] "2005" "2006" "2007" "2008" "2009" "2010" "2011" "2012" "2013" "2014" ## [21] "2015" "2016" "2017" "2018" "2019" "2020" "2021" "2022" "2023" "2024" ## [31] "2025" "2026" "2027" "2028" "2029" Value<-as.vector(StowPeopWTTCS[1,],mode='numeric') # create a vector of the values StowPeopWTTCPP<-as.data.frame(cbind(as.numeric(Year), as.numeric(Value))) # create longform dataset colnames(StowPeopWTTCPP)<-c("Year", "GDPContribution_USD") # set column names head(StowPeopWTTCPP) # header of data ## Year GDPContribution_USD ## 1 1995 10925800000 ## 2 1996 13485100000 ## 3 1997 14973800000 ## 4 1998 15254700000 ## 5 1999 15770700000 ## 6 2000 16533900000 Check for issues with recent data and make necessary adjustments. For example, data for incomplete years included, or data values indicate that the data are not accurate (e.g. not all data has been submitted) tail(StowPeopWTTCPP) # tail of data ## Year GDPContribution_USD ## 30 2024 38649000000 ## 31 2025 39947300000 ## 32 2026 41328600000 ## 33 2027 42877200000 ## 34 2028 44448900000 ## 35 2029 46242900000 Data from 2018 are estimates, so remove these data tmp<-which(StowPeopWTTCPP$Year>2018) StowPeopWTTCPP<-StowPeopWTTCPP[-tmp,] Processing Plot data Data are presented as billions of US dollars. windows(6,4) par(mar = c(4,5,1,1)) plot(StowPeopWTTCPP$Year, StowPeopWTTCPP$GDPContribution_USD/1000000000, pch = 18, ylab = "Contribution to GDP\n(billion US dollars at real prices)", xlab = "Year", ylim = c(0,50)) Perform calculations
102 Read in processed data used in previous report. These data have been saved in the ‘Processed_data_previous_report’ folder. Paste the file path to the required file in the quotation marks in the code below (i.e. replace the file path in the brackets with that for the required file). Insert double backslashes where required PrevDat<-read.csv("Processed_data_previous_report\\Stowaway_people_WTTC_processed.csv") See if new data have become available, if the baseline from the previous report has changed, and if new data are available, calculate the percentage change. Endp<-max(PrevDat$Year) Endn<-max(StowPeopWTTCPP$Year) x<-Endn-Endp if(x>0){ ChangeScenario<-"NewDataForRecentYears" y<-which(StowPeopWTTCPP$Year == Endp) z<-which(StowPeopWTTCPP$Year == Endn) PercentageChange<-((SStowPeopWTTCPP$GDPContribution_USD[z]- StowPeopWTTCPP$GDPContribution_USD[y])/ StowPeopWTTCPP$GDPContribution_USD[y])*100 b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$GDPContribution_USD[b] != StowPeopWTTCPP$GDPContribution_USD[y] } if(x==0){ ChangeScenario<-"NoNewDataForRecentYears" PercentageChange<-NA y<-which(StowPeopWTTCPP$Year == Endp) b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$GDPContribution_USD[b] != StowPeopWTTCPP$GDPContribution_USD[y] } Pathway<-"stowaway:peopleWTTC" results<-cbind(Pathway, ChangeScenario, PercentageChange, BaseChange) Post-processing The processed data are saved in csv format in the folder ‘Processed_data’ which was created in previous steps write.csv(StowPeopWTTCPP, "Processed\\Processed_data\\Stowaway_people_WTTC_processed.csv", row.names = FALSE) The outputs on changes to the data since the last report are written to the ‘Change_tracking’ csv in the ‘Processed’ folder. write.table(results,"Processed\\Change_tracking.csv", row.names=F, append=T, quote= FALSE, sep=",", col.names = F) The plot is saved as a tiff in the plots folder created in previous steps. tiff(filename = "Processed\\Plots\\Stowaway_people_WTTC.tiff", width = 6, height = 4, units = "in",
103 compression = "lzw", bg = "white", res = 1000) par(mar = c(4,5,1,1)) plot(StowPeopWTTCPP$Year, StowPeopWTTCPP$GDPContribution_USD/1000000000, pch = 18, ylab = "Contribution to GDP\n(billion US dollars at real prices)", xlab = "Year", ylim = c(0,50)) graphics.off() Contaminant: Seeds Data on the quantity of agronomic seed imported every year were obtained from the Department of Agriculture, Land reform and Rural Development. The units is tonnes. The original data were downloaded as a pdf, and were processed manually in previous steps of this workflow. View data structure str(ContSeed) # data structure ## 'data.frame': 11 obs. of 2 variables: ## $ Year : int 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 ... ## $ Quantity: num 5818 91696 38207 39089 274131 ... Standardisation Data do not need to be standardised Pre-processing Check for issues with recent data and make necessary adjustments. For example, data for incomplete years included, or data values indicate that the data are not accurate (e.g. not all data has been submitted) tail(ContSeed) # tail of data ## Year Quantity ## 6 2013 83773.0 ## 7 2014 155535.0 ## 8 2015 80190.0 ## 9 2016 388352.0 ## 10 2017 142856.7 ## 11 2018 39026.0 No adjustments required Processing Plot data Data will be presented in thousand tonnes windows(6,4) par(mar = c(4,5,1,1)) plot(ContSeed$Year, ContSeed$Quantity/1000, pch = 18, ylab = "Quantity of imported seed\n(thousand tonnes)", xlab = "Year", ylim = c(0, 500))
104 Perform calculations Read in processed data used in previous report. These data have been saved in the ‘Processed_data_previous_report’ folder. Paste the file path to the required file in the quotation marks in the code below (i.e. replace the file path in the brackets with that for the required file). Insert double backslashes where required PrevDat<-read.csv("Processed_data_previous_report\\Contaminant_seeds_DALRRD_processed.csv") See if new data have become available, if the baseline from the previous report have changed, and if new data are available calculate the percentage change. Endp<-max(PrevDat$Year) Endn<-max(ContSeed$Year) x<-Endn==Endp if(x == FALSE){ ChangeScenario<-"NewDataForRecentYears" y<-which(ContSeed$Year == Endp) z<-which(ContSeed$Year == Endn) PercentageChange<-((ContSeed$Quantity[z]- ContSeed$Quantity[y])/ContSeed$Quantity[y])*100 b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Quantity[b] != ContSeed$Quantity[y] } if(x==TRUE){ ChangeScenario<-"NoNewDataForRecentYears" PercentageChange<-NA y<-which(ContSeed$Year == Endp) b<-which(PrevDat$Year == Endp) BaseChange<-PrevDat$Quantity[b] != ContSeed$Quantity[y] } Pathway<-"contaminant:seedsDALRRD" results<-cbind(Pathway, ChangeScenario, PercentageChange, BaseChange) Post-processing The processed data are saved in csv format in the folder ‘Processed_data’ which was created in previous steps write.csv(ContSeed, "Processed\\Processed_data\\Contaminant_seeds_DALRRD_processed.csv", row.names = FALSE) The outputs on changes to the data since the last report are written to the ‘Change_tracking’ csv in the ‘Processed’ folder. write.table(results,"Processed\\Change_tracking.csv", row.names=F, append=T, quote= FALSE, sep=",", col.names = F) The plot is saved as a tiff in the plots folder created in previous steps. tiff(filename = "Processed\\Plots\\Contaminant_seeds_DALRRD.tiff", width = 6, height = 4, units = "in",
105 compression = "lzw", bg = "white", res = 1000) par(mar = c(4,5,1,1)) plot(ContSeed$Year, ContSeed$Quantity/1000, pch = 18, ylab = "Quantity of imported seed\n(thousand tonnes)", xlab = "Year", ylim = c(0, 500)) graphics.off()
112 Updating the permit database Suggested citation: SANBI and CIB 2023. Appendix 4 to "The Status of Biological Invasions and their Management in South Africa in 2022"—Workflows: Updating the permit database. South African National Biodiversity Institute, Kirstenbosch and DSI-NRF Centre of Excellence for Invasion Biology, Stellenbosch. http://dx.doi.org/10.5281/zenodo.8217222 1. In January request an updated permit database from the Department of Forestry, Fisheries and the Environment: Environmental Programmes, Biosecurity—Permitting sub directorate 2. Focus only on permits issued. Applications and refusals of applications can be for a number of reasons related to administrative issues and not to do with the risks involved. In future a separate process might be needed to look at those. Manually check the last few permits issued from the previous year with those issued this year so that it is clear which permits to add. 3. Create a worksheet that explains all the data cleaning done in converting the DFFE's database to the permit database (e.g., NEMBA AandIS permits 2014–2022 v20230509 edited files.xlsx). The broad steps involved are described below, but see the file itself for details. A separate file noting queries is captured and sent back to DFFE. edited: Copied without formatting, removed empty lines between months / financial years. In cases where multiple permits are issued to one person make sure the details are in multiple row s(i.e., there are no rows with merged cells); check that there are no empty cells, and if there are either fill with neighbouring rows if the result of merged cells or add NA. Add a notes column to capture any floating notes on the original. edited 2: removed columns "Permit application attended by:". Date Received, Permit issue date, Permit expiry date, in format custom '[$-en-ZA]d mmm yyyy]', visual check of data and correct typos. Check overlap between data to be added and new data, and delete rows already captured. Sort all the dates to see unusually large and small values, also check that received is earlier or the same as issued and the expire is later than issued (e.g., add a new cell and put in a formula subtracting one from the other). Add in missing dates, even if missing from original database as can usually calculate from issued date, expiry date, and length of permit. edited 3: (in future do step 4 first) separated species into unique rows (note the same permit can have several taxa on it). edited 4: (in future do before step 3) checked through permit numbers copied just filled cells across (not whole spreadsheet). Create a column 'permitNumber' and add a P at the front so all stored as text. Check the length of the permit numbers using the function len(). Including the P there should be 18 characters, but in a few cases the last 0 is missing (this was added back in manually) and in a few other cases the permit number was 19 characters (checked with DFFE). Note if the digits go above 15 then they have to be stored as text as otherwise Excel cuts off the number. One permit was issued to Garden Route National Park for possession on 29 Sep 2022, but there are four species codes, PS.2551|BS.2551|MIS.2551|FS.2551, duplication removed Column added 'duration' and validity of permits column converted to months and deleted (in one case it was missing so it was added based on issue and expiry date); added permit issue dates where lost during demerging of cells =IF(OR(S483="years",S483="year"),R483*12,IF(S483="months",R483,NA))
113 edited 5: Split "Type of Permit" into the categories noted below and added columns for each activity and either "TRUE" or "F" (rather than FALSE to ensure it is easy to see) Need to edit the reasons a permit was issued so they align with the regulations as various text is used in the database. From Notice 1 of the NEM:BA A&IS lists 2020 where a–e as defined in the Act; f–l as in Chapter 6 of the Regulations import = a. Importing into the Republic, including introducing from the sea, any specimen of a listed invasive species. possess = b. Having in possession or exercising physical control over any specimen of a listed invasive species. breed = c. Growing, breeding or in any other way propagating any specimen of a listed invasive species, or causing it to multiply. convey = d. Conveying, moving or otherwise translocating any specimen of a listed invasive species. trade = e. Selling or otherwise trading in, buying, receiving, giving, donating or accepting as a gift, or in any way acquiring or disposing of any specimen of a listed invasive species. spread = f. Spreading or allowing the spread of any specimen of a listed invasive species. releasing = g. Releasing any specimen of a listed invasive species. move.fw = h. The transfer or release of a specimen of a listed invasive fresh-water species from one discrete catchment system in which it occurs, to another discrete catchment system in which it does not occur; or, from within a part of a discrete catchment system where it does occur to another part where it does not occur as a result of a natural or artificial barrier. discharge = i. Discharging of or disposing into any waterway or the ocean, water from an aquarium, tank or other receptacle that has been used to keep a specimen of an alien or a listed invasive species. catch.and.release = j. Catch and release of a specimen of a listed invasive fresh-water fish or listed invasive fresh-water invertebrate species. take.to.island = k. The introduction of a specimen of an alien or a listed invasive species to off-shore islands. release.fw = l. The release of a specimen of a listed invasive fresh-water fish species, or of a listed invasive fresh-water invertebrate species, into a discrete catchment system in which it already occurs. research = (1) Despite anything to the contrary in these regulations, a permit may be issued subject to permit conditions, to a scientific institution to carry out a restricted activity involving a specimen of an alien or listed invasive species, and must be issued under the conditions that the specimen must— (a) be kept for identification or research purposes only; biocontrol = (1) Despite anything to the contrary in these regulations, a permit may be issued subject to permit conditions, to a scientific institution to carry out a restricted activity involving a
114 specimen of an alien or listed invasive species, and must be issued under the conditions that the specimen must— (b) form part of a preliminary study into biological control methods; or (c) form part of an effective biological control programme. display = (3) Despite anything to the contrary in these regulations, a permit may be issued, subject to permit conditions, to a zoological or botanical institution to carry out a restricted activity involving a specimen of an alien or listed invasive species, including for display purposes. inter.basin = (5) Despite anything to the contrary in these regulations, a permit may be issued, subject to permit conditions, for the transfer of a specimen of an alien or listed invasive species from one fresh-water system in which it occurs to another freshwater system in which it does not occur through a state inter-basin transfer scheme. Note there is no difference between buy and sell in the regulations, although both terms are used in the permit database If a permit was issued for possession for research and an accompanying permit was issued for conveyance, then the conveyance permit is assumed to also be for research. Research was if flagged except National Zoological Gardens of South Africa for display purposes; and Rhodes University a named person part of the Centre for Biological Control and Agricultural Research Council a named person part of the Plant Health Protection group for biocontrol purposes edited 6: column headings changed to align with the final column headings (and so compatible with R) with the exception the first column is permitName (align with scientificName is the next step). This involves the removal of the columns: Species Code, Applicant, Contact number, Location where the restricted activity will take place, Province, Restricted activity involved, Type of Permit. This is the final stage at which the original order is maintained edited 7: the database is sorted by permitName and aligned to scientificName as per the 2020 regulatory lists where possible, if there is a deviation between the regulatory name and the name on the permit this is flagged in the notes. Check for changes in nomenclature in past names. The database does not include any information on the listing category. Look at <NEMBA_AandIS_Regulatory_Lists_20230503.xlsx> or similar, noting that it is important to recognise that the current listing category might differ from the listing category when the permit was issued. ...As a final check ensure that all the cells with TRUE or F are formatted the same way (including alignment) and do a search and replace for each term (match full cell contents). Otherwise sorting the cells A to Z might not work. 5. Save the changes made, and create an updated file with all the permits of the form <NEMBA_A&IS_permits_issued_2014–20XX.xlsx>. Backup the intermediary edited file, note that this file cannot be made publicly accessible as it contains sensitive information. 6. Log on to Zenodo using the status report's Zenodo account and go to https://dx.doi.org/10.5281/zenodo.3947809. The goal is to update the version on-line, the first step is to reserve a new doi. This doi is then used within the file created and in the metadata on the Zenodo web-site. Once all documents are updated, publish the file on-line.
115 7. Share with DFFE and provide queries / feed-back, edit the file and republish to Zenodo as necessary. 8. When it comes to analysing the database in R, save the file as a text tab-delimited file then replace all apostrophes in authority names (i.e. ') with another character, e.g., an underscore. 9. The following code is to make graphs of permits issued per month over time and to get numbers of permits per species #make sure the text file has removed ? (i.e. accents etc.) setwd("C:\\Users\\jrwilson\\Desktop") # setwd("C:\\docs\\projects\\Listing and lists\\DFFE permit database") permit <- read.table("NEMBA_AandIS_permits_issued_2014–2022.txt",sep="\t",header=T) permit$issued<-as.Date(permit$issued,"%d %b %Y") permit$received<-as.Date(permit$received,"%d %b %Y") permit$expires<-as.Date(permit$expires,"%d %b %Y") #remove permits with the same permit number, note this will remove different species so don't use this for speceis anlayses permit.unique <- permit[!duplicated(permit$permitNumber), ] permit.unique.cat2 <- permit.unique[!(permit.unique$research | permit.unique$biocontrol | permit.unique$display |permit.unique$inter.basin),] #convert to number of permits per calendar month library(lubridate) x2 <- floor_date(permit.unique$issued,unit="month") month(x2) sort(permit.unique$issued) table(x2) write.table(table(x2),"a.txt") library(lubridate); x2 <- floor_date(permit.unique.cat2$issued,unit="month"); month(x2); sort(permit.unique$issued); table(x2); write.table(table(x2),"aCat2.txt") ##################### ###add in missing months in Excel into a file called dates ################################ povert <- read.table("datesALL.txt",sep="\t",header=T) povertCat2 <- read.table("datesCat2.txt",sep="\t",header=T) povertCat2$date<-as.Date(povertCat2$date,format="%Y/%m/%d") vals <- barplot(povert$numpermits,ann=F,ylim=c(0,100),plot=F) windows(6,3) par(mar=c(4,6,0.5,0),las=1) barplot(povert$numpermits,ann=F,col="grey20",ylim=c(0,100)) barplot(povertCat2$numpermits,ann=F,add=T,col="grey80") legend("topright",bty="n",fill=c("grey20","grey80"),c("research, biocontrol, etc..","category 2")) axis(1,at=vals[seq(1,100,12)],labels=c("Jan 2015",2016:2023)) mtext(side=1,line=3,"Month")
116 mtext(side=2,line=4,adj=0.5,"Number\nof\npermits\nissued") ####Number of permits per species (assumes each entry represents a different permit) #number of species with permits issued on them length(unique(permit$scientficName)) xx <- permit$scientficName #xx <- permit$scientficName[permit$research==F & permit$biocontrol==F & permit$display==F & permit$inter.basin==F] xx<-as.factor(xx) out<-data.frame(names(table(xx)),as.numeric(table(xx))) names(out)<-c("species","numpermits") barplot(out$numpermits) x2<-table(out$numpermits) windows(6,3) par(mar=c(4,6,0.5,0),las=1) plot(as.numeric(names(x2)),as.numeric(x2),log="x",type="h",ann=F,bty="n",lwd=3) axis(1,at=c(1,1000),labels=c("","")) axis(2,at=seq(0,100,10)) mtext(side=1,line=3,"Number of permits issued") mtext(side=2,line=4,adj=0.5,"Number\nof\ntaxa") #permits per taxon cumsum(sort(out$numpermits))/sum(out$numpermits) cumsum(x2)/sum(x2) write.table(out,"a.txt") #repeat only for period xx <- permit$scientficName[permit$research==F & permit$biocontrol==F & permit$display==F & permit$inter.basin==F & permit$issued>as.Date("31 Dec 2019","%d %b %Y")] xx<-as.factor(xx) out<-data.frame(names(table(xx)),as.numeric(table(xx))) write.table(out,"a.txt")
117 Money spent Suggested citation: SANBI and CIB 2023. Appendix 4 to "The Status of Biological Invasions and their Management in South Africa in 2022"—Workflows: Money spent. South African National Biodiversity Institute, Kirstenbosch and DSI-NRF Centre of Excellence for Invasion Biology, Stellenbosch. http://dx.doi.org/10.5281/zenodo.8217222 1. Ra6onale Using a workflow to integrate and summarise the cost of managing biological invasions enables a transparent and repeatable means to standardise the repor`ng of money spent at various scales (local, na`onal, or global). With the development of the InvaCost database and many independent studies on costs all using differing approaches, the adop`on of a workflow becomes crucial for enhancing data integra`on and facilita`ng meaningful comparisons (ensuring FAIR data principles are met). Here, we present a workflow that deconstructs the approach used in the third Na`onal Status Report on Biological Invasions in South Africa to assess the money spent on biological invasions in the country. As of yet, code has not been developed to automate this approach; however various func`ons in the InvaCost package (Leroy et al. 2023) can be used to summarise costs. 2. Aim The workflow aims to integrate and evaluate cost data related to managing biological invasions, with a specific emphasis on summarising expenditures based on stakeholders and for three primary aspects of biological invasions (pathways, species, and sites). Here, we present a workflow that deconstructs the approach used in the third Na`onal Status Report on Biological Invasions in South Africa to calculate the cost of managing biological invasions in the country. It aims to ensure consistency and reproducibility throughout the data manipula`on and analysis stages, enabling informed decision-making and facilita`ng meaningful comparisons across various studies or datasets. 3. Structure The workflow consists of five primary steps 1) background research; 2) data collec`on and entry; 3) data processing; 4) data analysis; and 5) summary and communica`on of results. Time will need to be allocated to steps 1 and 2 as collec`ng and colla`ng data from official en``es is omen a laborious process. Furthermore, there are mul`ple instances where costs will not be made available from stakeholders, or where the costs received are an incomplete representa`on of the money spent. Step 3 requires some manual data processing to expand the data to prepare it for analysis, but the remaining steps are rela`vely straight forward and can be done using excel or exis`ng script in R. Each primary step in the workflow is broken down into smaller steps to streamline the process and increase the usability of the workflow. Where needed examples have been provided in boxes to clarify aspects that may otherwise prove difficult to interpret. Sugges`ons for file names have also been provided. 4. Steps Step 1: Background research 1.1. Define the scope a. Clearly define the geographic scope of the study– for which area you will assess the total money spent on managing biological invasions. This maybe local, na`onal, con`nental, global etc.
118 b. Determine the `me frame for the assessment - whether it considers a single year or mul`ple years. 1.2. Iden`fy sectors and stakeholders Iden`fy the relevant stakeholders (and any sub-en``es) that could be involved in managing biological invasions in the study area, such as government agencies and their related departments and programmes, research ins`tu`ons, NGOs, industry associa`ons, private companies, and interna`onal organisa`ons. This process will highlight some of the primary stakeholders from/ for which data will need to be obtained. In addi`on, this process can help approximate the completeness of data received in terms of stakeholder representa`on. Step 2: Data collec`on and entry 2.1. Gather publicly available data Create a primary working folder in which all subfolders and their relevant files will be saved throughout the workflow, naming the folder [cost_data]. Iden`fy and gather publicly available data sources for each stakeholder involved, consult government reports, research studies, project reports, and other sources of informa`on (published or grey literature). Within the primary working folder create a sub-folder named [open_access_cost_data] in which you will save all individual data sources (files) in their original format (excel, pdf, etc.) naming the file [*stakeholder name*_*year*_open_access_original]. Note: it is useful to save all files and folders in lower case leaers as this can improve efficiency when working in R. 2.2. Request data from stakeholders Engage with stakeholders and experts to request addi`onal data and/or a detailed breakdown of the exis`ng data. Data requests may be made telephonically or through official email pathways using a clear and concise data request leaer. Data requests can be tailored to ensure data are received in a preferred format and regarding specific points of interest. In this instance, it would be useful to request the data in excel or csv format, and where possible, to breakdown expenditure into management type (provide examples), and/or as per the aspect of biological invasions covered (pathways, species, and sites). Requested data may include budgetary alloca`ons, expenditures, grants, financial records, budget alloca`ons, and/or funding alloca`ons from various sources. In the [cost_data] folder create a sub-folder named [requested_cost_data] and save all received data (in their original format) as individual files named [*stakeholder name*_*year*_requested_original]. Note: it is likely that data will be sent sporadically, as such it is important to label all files carefully and consistently and to track for which stakeholders you do and do not have cost data. Such informa`on can be later used to highlight any gaps in reported expenditure. 2.3. Create a data template In the [cost_data] folder create a sub-folder named [data_entry]. Create an empty excel spreadsheet for data entry (see Table 1 for the name and number of columns to be used) and save the file as [data_entry_template] in the [data_entry] sub-folder. 2.4 Enter the raw data For each of the original datasets obtained per stakeholder in the folders [requested_cost_data_raw] and [open_access_cost_data_raw] populate the data_entry_template (see Table 1 for metadata) crea`ng one excel file per stakeholder. Save each individual data file with the corresponding name [*stakeholder name*_raw] in the [data_entry] sub-folder.
119 Note: it is important that the data received per stakeholder is entered into and saved as an individual data file as the data will later be expanded. Performing this ac`on in a consolidated data file (i.e., all received data entered into a single file) can prove to be `me consuming and difficult to track.
120 Table 1. Data entry template indica`ng the columns and their associated variables, descriptors, and metadata. If the data received is not specific to the descriptors outlined, the use of ‘unspecified’ as a descriptor for that variable is suggested (i.e., if the data provided does not specify management type one would use the descriptor ‘unspecified’ in that column). Where mul`ple descriptors apply for a single variable pipe delimiters should be used without spacing (e.g., managementType could be entered for a single cost entry as mechanical|chemical|research). Note the descriptor ‘various’ only applies to species as such lists may prove too lengthy to populate a single cell). Purpose Variable Name Column number Variable type Descriptor metadata - Descrip6on Cost iden`fier stakeholderName 1 Descrip`ve The name of the stakeholder for the cost entry sectorType 2 Categorical government - a public en`ty that is owned and controlled by the state or its representa`ves private – private landowners, businesses, and organisa`ons with full autonomy in ownership and decision-making sourceName 3 Descrip`ve The name of the source: can either be the same as the stakeholder name or a reference e.g., Noel et al. (2020) Cost es`mate rawCostYear 4 Numeric The applicable year or `mespan (e.g., 2020 – 2022) of the reported cost rawCost 5 Numeric Cost es`mate directly received from the stakeholder or source costStandardised 6 Numeric The rawCost adjusted for infla`on to reflect the baseline year chosen for repor`ng Cost type managementType 7 Categorical mechanical - using physical methods to remove or control invasive species chemical - using chemical treatments to control invasive species biological - introducing natural enemies or biological control agents to control or regulate popula`ons biosecurity - preventa`ve measures, this includes border control and quaran`ne protocols, etc. integrated – one or more of the approached combined research – all research pertaining to biological invasions outreach – public educa`on and awareness raising overheads – salaries etc. indicatorType 8 Categorical lumpSum – the cost received is at a course resolu`on with no breakdown on pathways/ species/ sites speciesBased – the cost received are or can be broken down into species specific costs sitesBased – are or can be broken down into site (area) specific costs pathwaysBased – are or can be broken down into pathway specific costs Pathways breakdown pathwayCategory 9 Categorical release – control efforts focus on mi`ga`ng invasions introduced via pathways associated with purposeful release into nature – legisla`on??? escape – efforts focus on invasions that have escaped from confinement e.g., legisla`on of traded/ bred species contaminant – efforts focus on preven`ng/ reducing introduc`ons of species as contaminants e.g., border control stowaway - efforts focus on preven`ng/ reducing introduc`ons of species as stowaways e.g., biosecurity Species breakdown speciesName 10 Descrip`ve The full species name: Genus species, or various (where a lump sum is provided for a list of species) kingdom 11 Descrip`ve Plantae, Animalia, Fungi, Pro`sta, Monera family 12 Descrip`ve Taxonomic family of the associated cost entry genus 13 Descrip`ve Taxonomic genus of the associated cost entry Sites breakdown province 14 Descrip`ve Descrip`ve - There are 9 provinces in South Africa, include the province(s) to which the cost pertains. Equivalent of ‘state’ in other countries municipality 15 Descrip`ve Descrip`ve – There are 205 local municipali`es in South Africa list the municipality(ies) associated with the cost landType 16 Categorical protectedArea - areas that are managed and legally protected to conserve and preserve their natural, ecological, and cultural values privateLand - and that is owned by individuals, organiza`ons, or en``es, rather than being owned by the government or held in public trust stateOwned – land held in the public domain that is owned and managed by the government at the na`onal, regional, or local level. communalLand - land that is owned and managed collec`vely by a community or a group of people who have customary or tradi`onal rights other – the land-type associated with the cost entry does not fit into the above listed categories.
121 Step 3: Data processing 3.1. Expand the data (to reflect a single applicable year per cost entry) For each individual raw cost dataset examine if the data need expansion by considering the column rawCostYear (column 4; see Table 1). For data entries where the cost data reflects a `mespan (e.g., 2020-2022) it is necessary to expand the data to ensure one row of data expresses a singular cost for a single year. Note: not all data will require expansion - if the data provided already reflects a single year then no expansion is required and the row of data should be retained as is. Expand per year process: Within each raw dataset iden`fy any cost entry (i.e., row of data) that indicates a singular cost for a `me range (e.g., 2020 – 2021; spanning 2 years). To analyse the data, this cost needs to be parsed out to reflect a single cost per single year. A new blank row needs to be added per year for the `me frame (e.g., 2 new blank rows for a 2-year `mespan - one for 2020 and one for 2021). The total cost for the `mespan should now be divided by the number of years indicated in the raw data (e.g., by two) and entered into the new rows per year (e.g., 1 million spent from 2020–2021 would be re-entered as 500 000 for 2020, and 500 000 for 2021) – see Box 1 for detailed examples. All other suppor`ng informa`on (i.e., species name, land use type, management type, etc.) should be copied into the relevant cells per expanded entry and should not change. It is very important that newly expanded files be saved separately to the raw data files, and the file appropriately named [*stakeholder name*_expanded]. Note: Although this assumes equal spending per year it allows for the assessment of temporal trends. Box 1. An example (focussing on selected columns) of the data expansion process to achieve one row of data per cost per year. Please note the use of pipe-delimiters for categorical variables with more than one descriptor. Raw data base example stakeholderName rawCost Year Management type species landType Department of nature 4 000 000 2020 - 2021 Biological|Chemical Species A|Species protected Invasion busters co. 1 000 000 2021 Mechanical Species A|Species B private|other Expanded database example stakeholderName rawCost Year Management type species landType Department of nature 2 000 000 2020 Biological|Chemical Species A|Species B protected Department of nature 2 000 000 2021 Biological|Chemical Species A|Species B protected Invasion busters co. 1 000 000 2021 Mechanical Species A|Species B private|other 3.2. Merge data into a single dataset Create a new blank data template as per ‘Step 2.3’. Enter into this template all the data from each expanded data file. This process should be rela`vely simple as the template design and metadata should be exactly matched across all individual expanded datasets that have been created per stakeholder. Note: if copying and pas`ng the data from one file to another, ensure to paste as values so that no informa`on is lost. Save the new file in the [cost_data] folder naming the file [all_costs_merged]. 3.3. Standardise the raw cost to a baseline year Standardising costs to a baseline year improves the comparability of results over `me and brings all cost data to a baseline to improve the accuracy of the analysis. To standardise all received costs to a single baseline year, adjust the data for infla`on. For example, if the chosen baseline year is 2021 and
- Abundance: Plants: le Roux et al. (2013), le Roux and Greve unpublished data (collected between 2018-2020). Invertebrates (for Figures 5.5 and 5.6): no total abundance estimates, but number of individuals or per m2 or km2 were obtained from: Gabriel et al. (2001): abundance data on five springtail (Collembola) species currently present on Marion Island (Ceratophysella denticulata, Isotomurus maculatus, Megalothorax minimus, Parisotoma notabilis and Pogonognathellus flavescens), out of six species (one species currently absent, therefore no data) Barendse et al. (2002): abundance data on three taxa of mites (Cillibidae, Dendrolaelaps and Pygmephoridae = Bakerdania) out of seven taxa present on Marion Island. Hugo et al. (2006) for PEI microarthropods: two cryptogenic mite taxa (Cillibidae and Dendrolaelaps), and five springtail (Collembola) species Mice: McLelland et al. (2018) presents data on mouse density through live trapping and details on sex, breeding status, age (juvenile or subadult/adult). These data were the most up to date estimates as of May 2023 ( according to expert Anton Wolfaardt). Species list The species list for the PEIs follows instructions in the PEIs metadata to determine possible values in each column. The PEIs metadata is different from the mainland one, as for example different realms apply. The first draft of the PEIs’ species list was developed following Greve et. al. (2017). Additional taxa were added from a number of different data sources by scouring the literature. Greve et al. (2017) did not have a comprehensive list of species that were recorded only once; therefore, most species added were species that were transient. E.g. 14 taxa were added from Watkins and Cooper (1986) and the list was compared with Pagad (2020) (PEIs species list from GBIF), but all records were present already in the report’s list. Finally, the list was compared to that from Leihy et. al. (2023), and three more taxa were added (Astemolaelaps, Blattella germanica, and Passer domesticus). This last source also has information on occurrence, eradication, introduction status, first date observed and estimated date of introduction. ● For species for which occurrencestatus=doubtful or cryptogenic, degreeOfEstablishment and introductionStatus will be NA, together with confidence and source. The information regarding what degree of establishment a taxon reached while present will be reflected in the historyOfInvasion column (last column), which is only present in the PEIs species list (not mainland). ● If a taxon appears in the bibliography as ‘transient’ or ‘not established’, then it is considered as ‘absent’ in the ‘occurrenceStatus’ column. Pathways For each species, the literature was exhaustively searched for information on the mode of introduction. Greve et al. (2017) had some information on pathways, but for several species, pathway information was fairly unspecific. E.g. many species are thought to have been introduced accidentally as
contaminant, but the literature was unclear about what the suspected contaminant was. Therefore, for many species, the pathway could not be determined, and has been categorised as ‘unknown’. Effectiveness in managing invasions A ‘First party assessment’ is performed by the ECOs (they assess their own work), and they produce one report every year where they report on their management actions and findings. Then, a ‘Second party assessment’ occurs every three years to perform quality assurance: a DFFE member goes to the island and checks that everything is properly used/done. This assessment involves field assessments and checking ECOs' reports. Money spent A budget was provided by Debbie Muir with expenses on herbicide, pesticide and personal protective equipment (PPE). However, the time invested by the ECOs on managing invasive species has not been included in the calculations of money spent, as there was no information provided on either the salary of the ECOs or how many hours per month or year they invest in controlling invasive taxa.