scieee AI-readable full text Open interactive document viewer

Open Data STAC: Find and access geospatial research datasets easily

Girgin, Serkan; Gohil, Jaykumar Harishbhai

Abstract

In line with Open Science practices and FAIR principles, researchers increasingly publish geospatial datasets through general-purpose research data repositories such as Zenodo and Figshare. However, these repositories often lack mechanisms to effectively utilize the rich spatiotemporal metadata contained in geospatial data files. Metadata is typically entered manually and limited to textual descriptions, reducing data findability, accessibility, and interoperability. Open Data STAC addresses this challenge by automatically detecting geospatial datasets published in major repositories and generating standardized, open-access spatiotemporal metadata catalogs based on the SpatioTemporal Asset Catalog (STAC) specification. Using protocols such as OAI-PMH and repository-specific APIs, datasets are identified through file type analysis and metadata extraction. Extracted spatial and temporal information is combined with repository metadata to create structured STAC items and collections, which are continuously updated in an open catalog. When permitted, cloud-native versions of datasets are also made available to enhance interoperability and accessibility. The project, funded by the Dutch Research Council (NWO) Open Science Fund, is operated by the Centre of Expertise in Big Geodata Science and powered by the Fairly toolset. Through Open Data STAC, geospatial research data becomes more visible, searchable, and usable, bridging the gap between general research repositories and domain-specific geospatial infrastructures.

Full text

OPEN DATA STAC Find and access geospatial research datasets easily! In line with Open Science practices and FAIR principles, researchers are publishing their geospatial research data at research data repositories, such as Zenodo, Figshare. https://opendatastac.org Powered by fairly toolset ∙ Operated by the Centre of Expertise in Big Geodata Science ∙ Development funded by NWO Serkan Girgin <[email protected]>, Jay Gohil Despite containing detailed spatiotemporal metadata, geospatial data files are not effectively utilized by research data repositories. Researchers are required to manually enter spatiotemporal information, typically limited to text descriptions or basic metadata fields. Geospatial research data has been published regularly Geospatial data is largely unfamiliar to data repositories Geospatial metadata can typically be entered only as text Manual metadata entry poses the risk of being done poorly Suggestion by ChatGPT van der Veeren, 2004 Tools to search research data by location is limited Research data repositories often lack effective tools to search by location, such as specifying a geographic extent. Research data publishing bad practices further limit effective data access and interoperability. Easy access and interoperability are limited as well Unfortunately, this is sometimes caused by data repository limitations. Consequently, geospatial research data often becomes "invisible" significantly reducing its findability and accessibility. Geospatial research data becomes practically "invisible" Initiatives exist to facilitate geospatial data discovery On the other hand, there are initiatives that aim to enable access to geospatial data, such as STAC that improves interoperability and data discovery through standard metadata catalogs. Open Data STAC aims to create an open STAC catalog of public research datasets published at major research data repositories. The project "OpenSTAC: an open spatiotemporal catalog to make geospatial research data findable and accessible" with file number OSF23.2.111 of the research programme NWO Open Science Fund 2023 is financed by the Dutch Research Council (NWO) Research data repositories are regularly monitored for newly published datasets using the OAI-PMH protocol and platform‐specific REST APIs. Geospatial datasets are detected by checking for common geospatial file formats through file extensions, file-related metadata analysis, and file type probing. For each identified dataset, spatiotemporal metadata is extracted from the relevant files and combined with the dataset-level metadata from the repository. Related files are grouped using a multi-level strategy that considers both extracted metadata and dataset structure. Corresponding STAC items and collections are generated by using the obtained information to update an open-access, STAC‐based catalog of geospatial research datasets. To improve interoperability, cloud-native versions of the data files are provided, if allowed. doi.org/10.1038/s41597-025-05309-w opendatastac.org/10.1038/s41597-025-05309-w Illustration by Storyset.com Illustration by Storyset.com Illustration by Storyset.com Illustration by Storyset.com Illustration by Storyset.com