Four Decades of Fire Science in the Amazon: Bibliometric Analysis Dataset and R Scripts (1980–2024)
Dutra, Débora Joana; Sánchez, Alber; Mataveli, Guilherme; Ferreira, Igor José Malfetoni; dos Santos Junior, Marcelo Augusto; de Freitas, Ana Larissa Ribeiro; Leão, Henrique; de Medeiros, Thaís Pereira; Ferro, Poliana Domingos; Magalhães, Deila da Silva;
- Publisher
- Zenodo
- Language
- en
Abstract
Title:Four Decades of Fire Science in the Amazon: Bibliometric Analysis Dataset and R Scripts (1980–2024) Description:This dataset and accompanying R scripts provide a comprehensive bibliometric analysis of scientific research on fire dynamics in the Amazon over four decades (1980–2024). The compilation integrates data from Scopus and Web of Science, covering publications, authors, affiliations, countries, journals, keywords, and methodologies. The repository enables the reproduction of key analyses, including: Temporal evolution of publications, highlighting key environmental events and policy milestones. Co-authorship networks at the author and country levels. Thematic mapping of research trends and emerging topics. International collaborations and network visualization. Methodological evolution in Amazon fire science research. Top journals and publication counts per decade. All scripts are written in R, leveraging packages such as bibliometrix, tidyverse, igraph, ggraph, sf, and others, with detailed documentation to facilitate replication or adaptation for similar studies. This work provides a valuable resource for researchers, policymakers, and stakeholders interested in understanding the scientific landscape of Amazon fire research, identifying knowledge gaps, and supporting evidence-based management of fire, biodiversity, and carbon dynamics in tropical forests. Keywords:Amazon, fire dynamics, bibliometric analysis, R, bibliometrix, co-authorship network, thematic mapping, scientific trends, research evolution, environmental science
Full text
Bibliometric Analysis Synthesizing Four Decades of Fire Science in the Amazon Débora Joana Dutra1, Alber Sánchez1, Guilherme Mataveli1, Igor José Malfetoni Ferreira1, Marcelo Augusto dos Santos Junior1, Ana Larissa Ribeiro de Freitas1, Henrique Leão1, Thaís Pereira de Medeiros1, Poliana Domingos Ferro1, Deila da Silva Magalhães1, Daniel Braga1, and Liana Oighenstein Anderson1,2 1Remote Sensing Postgraduate Program (PGSER), Coordination for Education, Research and Outreach (COEPE), Brazil’s National Institute for Space Research (INPE) 2Earth Observation and Geoinformatics Division (DIOTG), Earth Sciences General Coordination (CGCT), Brazil’s National Institute for Space Research (INPE) This repository contains R scripts and processed datasets for a comprehensive bibliometric analysis of Amazon fire science publications (1980–2024). The analysis includes temporal evolution, co-authorship networks, thematic mapping, methodology trends, and international collaborations. 1. Repository Structure /bibliometria ├─ scopus.csv # Raw Scopus export ├─ wos.txt # Raw Web of Science export ├─ dados_total.csv # Combined and deduplicated dataset ├─ Table_Top20_Journals.csv # Top journals by publications ├─ Tabela1_Publicacoes_por_Decada.csv # Publications per decade ├─ Publications_by_Year_PPCDAM_clean_no2025.png ├─ Top_Coauthorship_Network.png ├─ Coauthorship_All_Decades.png ├─ Thematic_Map_Bibliometry_Global.png ├─ Thematic_Map_1980s.png ├─ ... # Thematic maps for other decades ├─ Mapa_Collaborations_World.png ├─ method_evolution_dotplot.png ├─ Figure4_Top_Journals.png └─ scripts.R # Main R script 2. Installation and Library Functions
Below are the R packages used, with descriptions and why they are necessary: Core Bibliometric and Data Manipulation Package Purpose bibliometrix Main package for bibliometric analysis, converting Scopus/WoS files, generating co-authorship, thematic maps, and basic statistics. Functions: convert2df(), mergeDbSources(), biblioAnalysis(), summary(). tidyverse Suite of packages for data wrangling: dplyr (filter, summarize), tidyr (reshape data), readr (CSV/TSV handling). ggplot2 Visualization: create barplots, dot plots, scatterplots with extensive customization. dplyr Part of tidyverse, simplifies filtering, grouping, and summarizing large datasets efficiently. Network Analysis and Graph Visualization Package Purpose igraph Core graph/network analysis: nodes, edges, centrality measures. Used for co-authorship and country networks. ggraph Extension of ggplot2 for graph visualization: plotting nodes and edges with custom layouts. RColorBrew er Provides qualitative and sequential color palettes for network and heatmap visualization. patchwork Combines multiple ggplot2 plots into a single figure. Useful for decade-by-decade comparisons. Geospatial Analysis and Mapping Package Purpose sf Modern spatial data handling, reading shapefiles and geospatial points. geosphere Calculates distances, great-circle arcs for collaboration lines between countries. rnaturalearth, rnaturalearthdata Provides world map basemaps and country polygons. viridis Color palettes optimized for perceptual uniformity in maps.
countrycode Converts country names/codes to standardized ISO3 codes, necessary for mapping collaborations. Other Utility Packages Package Purpose lubridate Simplifies working with dates and extracting years/decades. stringr String manipulation: cleaning author names, keywords, and affiliations. purrr Functional programming tools: map functions over lists (useful for batch processing decades). 3. Data Import and Preprocessing Step 1: Convert raw files to bibliometrix format dados_scopus <- convert2df(file = paste0(pasta, "scopus.csv"), dbsource = "scopus", format = "csv") dados_wos <- convert2df(file = paste0(pasta, "wos.txt"), dbsource = "wos", format = "plaintext") Explanation: ● convert2df() reads raw exports from Scopus/WoS and converts them to a structured dataframe compatible with bibliometrix. ● Parameters: ○ file: path to export file ○ dbsource: database type ("scopus" or "wos") ○ format: file type ("csv" or "plaintext") Step 2: Merge datasets and remove duplicates dados_total <- mergeDbSources(list(dados_scopus, dados_wos), remove.duplicated = TRUE) write.csv(dados_total, paste0(pasta, "dados_total.csv"), row.names = FALSE, fileEncoding = "UTF-8") Explanation: ● mergeDbSources() combines multiple bibliometric datasets. ● remove.duplicated = TRUE ensures no repeated records remain. ● Output: dados_total.csv – main dataset for all analyses.
4. Bibliometric Analysis resultados <- biblioAnalysis(dados_total, sep = ";") summary(resultados, k = 20) Explanation: ● biblioAnalysis() computes: ○ Total publications ○ Authors, affiliations, countries ○ Most frequent keywords ○ Citation metrics ● summary() produces top-k statistics (k = 20 by default), including: ○ Top authors ○ Top journals ○ Top countries ○ Most cited articles 5. Temporal Analysis ● Plot annual publications, overlaying key events (droughts, El Niño) and policy phases (PPCDAM). ● Code uses ggplot2: ○ geom_bar() for publication count per year ○ geom_vline() for event markers ○ geom_text() for annotation Output: Publications_by_Year_PPCDAM_clean_no2025.png 6. Co-Authorship Networks ● Construct author networks using biblioNetwork(): NetMatrix <- biblioNetwork(dados_total, analysis = "collaboration", network = "authors", sep = ";") graph <- graph_from_adjacency_matrix(NetMatrix, mode = "undirected", weighted = TRUE) ● Visualize with ggraph: ○ Node size = number of publications ○ Edge width = number of co-authored papers
○ Node color = decade of first publication Outputs: Top_Coauthorship_Network.png and Coauthorship_All_Decades.png 7. Thematic Mapping (Keywords) ● Extract keywords from DE (author keywords) and ID (keywords plus) fields. ● Use thematicMap(): ○ Computes centrality (importance) vs. density (development) for each cluster ○ Produces a quadrant-based map: ■ Motor themes (high density, high centrality) ■ Niche themes (high density, low centrality) ■ Emerging/declining (low density, low centrality) ■ Basic/transversal (low density, high centrality) Outputs: Thematic_Map_Bibliometry_Global.png, Thematic_Map_1980s.png, etc. 8. International Collaboration ● Parse affiliation (C1) to extract countries. ● Standardize with countrycode(). ● Create country pairs per publication. ● Use geosphere::gcIntermediate() for great-circle arcs. ● Plot network over world map with sf + ggplot2. Output: Mapa_Collaborations_World.png 9. Methodological Evolution ● Extract methodology-related keywords: Remote Sensing, GIS, Machine Learning, Fire Danger Indices, Statistical Modeling. ● Count occurrences per decade. ● Plot dot plot with ggplot2: ○ X-axis: decade ○ Y-axis: methodology Dot size: frequency of use Dot color: trend over time Output: method_evolution_dotplot.png
10. Top Journals and Publications per Decade ● Count publications by journal: Table_Top20_Journals.csv ● Bar plot by decade: Figure4_Top_Journals.png ● Count publications per decade: Tabela1_Publicacoes_por_Decada.csv 11. Notes ● Ensure UTF-8 encoding to avoid issues with accented characters. ● Large networks may require substantial RAM. ● The script is modular: users can run individual sections (data import, networks, maps) independently.