scieee AI-readable full text Open interactive document viewer

VirtualiZarr 2.0 - Create virtual Zarr stores for cloud-friendly access to archival data, using familiar xarray syntax.

Jones, Max

Abstract

ESIP July 2025 Meeting Presentation - Foundational to Emerging Cloud-Native Technology VirtualiZarr allows accessing archival data in a cloud-optimized pattern, without copying or modifying the original data in any way. We will showcase the upcoming major release of VirtualiZarr, which adds easy merging/concatenation with parallelization via Dask, Lithops, or a custom executors via open_virtual_mfdatase; simple, reliable implementation of custom external parsers; direct data loading via Zarr and/or Xarray; and obstore integration for fast, reliable access to the cloud. We demonstrate VirtualiZarr's potential reducing the cost of services and simplifying user workflows.

Full text

Create virtual Zarr stores for cloud-friendly access to archival data, using familiar xarray syntax. VirtualiZarr 2.0 1 2 3 4 Motivation for Virtual Zarr Moving to VirtualiZarr 2.0 Save cost, time, and effort! Improving extensibility, reliability, and performance Virtual Zarr in action What’s next? Speeding up workflows for NASA VEDA Connections with earthaccess, demonstrations, and more! Users and applications want analysis-ready, cloud optimized (ARCO) datasets Figure credit: Sean Harkins Earth observations and modeling results are distributed as files Figure credit: Sean Harkins File-based data solutions do not scale well on the cloud Figure credit: Sean Harkins Zarr (native and virtual) can provide harmonized and cloud-optimized data access on top of data archives Figure credit: Sean Harkins VirtualiZarr enables harmonized, cloud-native access patterns Building cloud-native virtual datacubes requires several steps 1. Parsing native files (e.g., NetCDF, TIFF, FITS, GRIB) into Zarr’s data model Figure credit: HDF and Zarr specs Building cloud-native virtual datacubes requires several steps 2. Combining virtual representations of multiple files into a single dataset Example: Parsing TIFFs with Rust via Async-tiff Slide credit: Kyle Barron VirtualiZarr 2.0 maximizes flexibility and extensibility Example: Loading virtual datasets with Rust via obstore Explore the benchmarks! VirtualiZarr 2.0 improves reliability through Obstore integration Image credits: Kyle Barron Watch Kyle’s Cloud Native Geospatial Talk! Auto-completed AWS parameters Auto-suggested AWS region list VirtualiZarr 2.0 improves performance via Dask, Lithops, Icechunk, Obstore, and more! Image credit: Aimee Barciauskas Massive parallelization minimizes costs and maximizes benefits Credit: Ryan Abernathey and Aimee Barciauskas MUR-SST demonstration from pilot study ●13 total hours of lambda runtime. This includes periodic validation of the dataset. ●9,124 requests, so an average of 5 seconds per request. ●Dataset generation cost estimate $1.23 (21 years) ●Storage cost (note, this includes some native zarr data): $4.44/year Improving time and memory performance for NASA science workflows Also read about the performance benefits and cost savings for services shown by the NASA pilot study with Earthmover on their blog! Credit: Julius Busecke ~40x speedup for opening and >3x speedup for computations in VEDA data story - every time the workflow is run! What’s next? Helping people benefit from this technology through demonstrations, integrations, and collaborations If you’re part of any of these groups, we want to work with you! If not, we still want to work with you! Try it out at part two of the Cloud Native Technologies session! Connect on GitHub and let us know what you think! We improve the tools that improve the planet Washington DC Lisbon