scieee AI-readable full text Open interactive document viewer

Capacity Building for Cloud-Native Geospatial Data Use in the Netherlands

Girgin, Serkan

Abstract

Rapid access to and efficient processing of large spatiotemporal datasets remain major challenges in geo-information science and geospatial computing, as data volumes grow rapidly with more diverse sources and higher collection frequencies. Since these datasets are now largely hosted in the Cloud, cloud-native access and processing have become essential skills for public institutions, private organizations, and knowledge institutions. Yet, the inefficient practice of downloading data for local analysis remains common, either because data is not provided in cloud-optimized formats or because users lack the necessary expertise. A similar issue exists in geospatial data publishing, where datasets are often released in formats that hinder efficient cloud access and interoperability. To bridge this gap, a community-driven initiative can effectively promote the use of cloud-native tools and technologies for publishing, accessing, and processing geospatial data by demonstrating benefits through real-life examples and providing targeted training. The CLOUD-NES project addresses this for the Natural and Engineering Sciences (NES) domain in the Netherlands by (a) building a public cloud-native data platform with co-located analysis capabilities, (b) making key datasets such as PDOK and KNMI data available in cloud-optimized formats, (c) demonstrating the efficiency of cloud-native solutions compared to traditional workflows with evidence-based benchmarks, and (d) developing open training materials and organizing training workshops. This will enable stakeholders to gain hands-on experience designing and running cloud-native workflows, as well as creating and publishing datasets following best practices. Project outcomes and lessons learned will be shared with national and international stakeholders through dedicated meetings, together with guidelines and training for building similar infrastructures. This talk presents the community-driven approach to capacity building in cloud-native geospatial data use, introduce the key components of the CLOUD-NES project, and gather ideas and feedback from participants on how to extend its impact to other domains in the Netherlands.

Full text

Capacity Building for Cloud-Native Geospatial Data Use in the Netherlands Dr. Ing. Serkan Girgin MSc1,2 [email protected] https://linkedin.com/in/serkan-girgin/ 1Director, Center of Expertise on Big Geodata Science 2Associate Professor, Department of Geo-information Processing Faculty of Geo-information Science and Earth Observation (ITC) 30 October 2025, Delft Cloudscaping Geo Symposium · Where Cloud Meets Earth Data accessibility plays a crucial role in modern geospatial computing •Large spatial datasets are becoming (publicly) available on the Cloud that are valuable for everyone*. •Quick access and efficient processing of such datasets are necessary to foster successful, timely, economic, and energy efficient activities. •Cloud-native data access and processing are ramping up as modern digital competences that bring computation close to the data to increase efficiency and reduce analysis time. Illustration by Storyset.com Although there are many cloud-native tools and technologies available, the inefficient approach involving data download and local exploration remains mostly as the standard practice. Illustration by Storyset.com Sometimes this is involuntary because the data is not provided in a format* that well suits cloud-based processing. Illustration by Storyset.com (Even if cloud-optimized formats could be used at no additional cost) But it is also not uncommon that the skills necessary for cloud-based data access and processing are lacking. Illustration by Storyset.com Data providers face challenges when supporting cloud-native data formats •Data providers hesitate to change data formats because existing users and software depend on legacy ones. •Decision-makers may not appreciate the gains, such as scalability, reproducibility, and lower egress costs. •While open tools to publish cloud-native geospatial data exist, not all have stable enterprise-ready support. •Poorly tuned chunking, compression, or tiling can make cloudnative formats behave worse than traditional ones. •Questions about where data resides can slow adoption of cloud-native infrastructure. Illustration by Storyset.com The challenge is not just access to cloud-native data, but the ability to use it •Most professionals still lack domain-specific, hands-on training materials that connect these technologies to their actual geospatial workflows. •Tutorials teach how to use COGs, STAC, or Zarr, but not why or when they matter in specific domains. •Cloud-native geospatial data is not systematically taught in geoinformatics or GIS programs. •Public sector and SMEs lack structured reskilling paths. •Training focuses on accessing and analyzing cloud data, not on creating or serving cloud-native datasets. •Cloud-native formats and tools evolve quickly, many tutorials break or use deprecated features. •Most tutorials assume U.S. cloud ecosystems, not European data governance. Illustration by Storyset.com Digital sovereignty is a key pillar of sustainable development •Digital sovereignty is becoming increasingly central to the way communities, institutions and countries think about sustainability, both in terms of data governance and long-term independence. •In the geospatial domain, it means that users and institutions can manage and access their own digital resources, e.g., data, software, computing platforms, without undue dependence on external or commercial actors. •Digital sovereignty is essential for sustainable development because it secures long-term, ethical, and equitable control over the digital foundations of modern life. Illustration by Storyset.com Data access should be coupled with processing capabilities for efficiency •Modern cloud-native data formats enable faster access to large geospatial datasets and efficient processing when combined with co-located computing resources. •Inherent scalability of the cloud infrastructure makes large-scale analysis feasible without significant investment. •This approach allows cost-effective and energy-efficient processing, particularly for resource-intensive modelling and machine learning tasks. •Moreover, cloud-native architecture simplifies infrastructure management for data and compute providers by reducing server components and enabling direct data interaction. Illustration by Storyset.com We are implementing a cloud-native data access and processing infrastructure to support capacity development in the NES domain •The platform consists of a cloud-native object-based storage system, a STAC-compliant data cataloging service, and a JupyterLab-based data analysis environment. •Synergies with international communities (e.g., Pangeo) and other projects (e.g., HPC-DAT) will be explored for the development of the platform. •It will be built on the national research ICT infrastructure provided by SURF, which will enable users to access powerful computing resources*. •To demonstrate on-premise deployment and enable broader applicability across different infrastructures, a twin platform will be established at CRIB. •The platforms will be publicly available without any cost to all interested parties during the project time for activities related to the project. •Following the project, the CRIB platform will be maintained as a core service of the Centre to ensure long-term sustainability. Illustration by Storyset.com Open datasets relevant to the NES domain will be made cloud-native to support benchmarking, capacity development, and self-learning •These will include datasets from the PDOK and KNMI. •Both institutions serve data via OGC-compliant web services and file downloads. Although their infrastructure is on the Cloud, their data is not cloud-native and co-located data analysis capabilities do not exist. •First, data subsets will be ingested by using different cloud-optimized formats and different sets of parameters (e.g., chunk size, compression, data arrangement). •By using preliminary benchmarks, the optimal formats and parameters will be identified for each dataset that minimize data volume, reduce access time, and maximize processing performance. •Then, the complete datasets will be ingested by using the optimal formats and parameters. The original datasets will be stored on the platform to enable reliable comparison with traditional approaches. Illustration by Storyset.com GeoTIFF → COG NetCDF → Zarr Shapefile → Parquet The benefits of cloud-native approach will be demonstrated with quantitative and reproducible evidence •Data access and processing with respect to the original data formats and traditional methods, as well as different cloud-native formats and ingestion parameters, will be benchmarked. •A comprehensive testing framework and reference (cloud-native) datasets will be developed to guarantee transparent and reproducible measurement of storage efficiency, transfer efficiency, access performance, and scalability, as well as efficiency of co-located versus remote analyses. •We will work with the community to design realistic test scenarios that represent typical access and processing patterns. •The results will be shared with research communities, including universities, research institutions, and (inter)national organizations. Illustration by Storyset.com We will develop skills for effective and efficient use of cloud-native data access and processing in the NES domain •Open training materials will be developed on how to use cloud-native data infrastructure efficiently, how to create and publish cloud-native datasets, and how to deploy cloud-native data infrastructure. •Special attention will be made to make the materials (re)usable in different disciplines and scientific domains. •By using the developed training materials, six hands-on training events will be organized at different locations in the Netherlands for researchers, research supporters, and data providers. •These courses will be integrated into the regular training programs of co-applicant institutions for sustainability. •We will also promote the use of the developed training material by training hubs (e.g., ELIXIR, EOSC). Illustration by Storyset.com FAIR practices will enable effective outreach and reuse of project outcomes Illustration by Storyset.com •The scripts to deploy the cloud-native data infrastructure will be published in open-source. The best practices to build a cloud-native data infrastructure will be documented. •The data ingestion pipelines will be made available as open-source software. A guideline for efficient conversion to cloud-optimized formats will be published. •The benchmarking framework will be released as open-source software, along with guidelines on benchmarking cloud-native data access and processing. Benchmarking results will be disseminated through blog posts and open-access publications. •The training materials will be maintained in GitHub repositories to encourage community contributions and will also be archived long-term on Zenodo. •Two symposiums will be organized to engage with leading (inter)national cloud-native initiatives, and share insights and lessons learned from the project with stakeholders. Project activities will be completed in 18 months until March 2027 CLOUD-NES brings together users, data, and infrastructure providers to work efficiently with large geospatial data and advance collective expertise in cloud-native technologies. Illustration by Storyset.com JOIN US! Together we can accelerate cloud-native geospatial adoption across the Netherlands. Illustration by Storyset.com https://ut.onl/cloudnes Contact us if you want to learn more or collaborate! Dr. Ing. Serkan Girgin MSc Head of Department Center of Expertise in Big Geodata Science Associate Professor Department of Geo-information Processing Faculty ITC, University of Twente [email protected] https://linkedin.com/in/serkan-girgin/ Faculty of Geo-information Science and Earth Observation (ITC) Centre of Expertise in Big Geodata Science (CRIB) https://itc.nl