scieee AI-readable full text Open interactive document viewer

A Comprehensive Monitoring Toolkit for Energy Consumption Measurement in Cloud-Based Earth Observation Big Data Processing

Bhawiyuga, Adhitya; Girgin, Serkan; de By, Rolf A.; Zurita-Milla, Raul

Abstract

The processing of earth observation big data (EOBD) in distributed environments has increased significantly, driven by advances in satellite technology and the growing number of earth observation missions. This massive influx of data presents unprecedented opportunities for environmental monitoring, climate change studies, and natural resource management, while simultaneously posing significant computational challenges. Cloud computing has emerged as an enabler for handling such EOBD, offering scalable computational resources, flexible storage solutions, and on-demand processing capabilities through platforms such as Google Earth Engine (GEE), AWS SageMaker, OpenEO, and Pangeo Cloud. While these cloud-based EOBD processing platforms offer varying levels of monitoring capabilities to help users understand their workflow execution, they primarily focus on traditional performance metrics. GEE provides basic performance insights focusing on task execution status, AWS SageMaker offers comprehensive resource utilization metrics through Amazon CloudWatch, and Pangeo Cloud implements the Dask profiler for real-time monitoring of cluster performance. However, a significant gap exists: none of these platforms incorporate energy consumption as a standard monitoring metric. This limitation becomes increasingly critical as the scientific community grows more concerned about the environmental impact of large-scale data processing operations. The absence of energy-related metrics from monitoring may hinder users from understanding the environmental impact associated with their EOBD processing workflows. This knowledge is particularly crucial in the earth observation domain, where the balance between computational requirements and environmental impact directly aligns with the field's core mission of environmental protection. Furthermore, recent green computing initiatives have emphasized the importance of sustainable IT infrastructure, yet the lack of standardized energy consumption metrics in EOBD processing platforms hinders researchers' ability to make informed decisions about computational resource usage. To address this gap, we propose a monitoring toolkit for understanding the energy consumption patterns in distributed EOBD processing. We develop an integrated approach that combines multi-level energy measurements: (1) hardware-level power data collected through RAPL for CPU and DRAM, IPMI for system-level metrics, and external power sensors for overall consumption; (2) software-level resource utilization metrics from the operating system including CPU usage, memory allocation, I/O operations, and network traffic; and (3) application-level profiling through integration with Dask's distributed processing framework. Our methodology employs power ratio modeling to correlate these measurements and estimate process-level energy consumption, enabling fine-grained energy profiling of EOBD workflows. The toolkit generates comprehensive monitoring reports that include energy consumption patterns, resource utilization correlations, and efficiency metrics, allowing users to make informed decisions about their processing strategies. By providing visibility into the energy consumption of computational workflows, this work contributes to the development of more sustainable EOBD processing practices. The toolkit enables users to better evaluate the true environmental cost of their computational workflows and optimize their processing strategies accordingly, supporting the broader goal of environmental protection through more energy-efficient earth observation data processing.

Full text

A Comprehensive Monitoring Toolkit for Energy Consumption Measurement in Cloud-Based Earth Observation Big Data Processing Adhitya Bhawiyuga1, Dr. Serkan Girgin2, Dr. Rolf de By, Prof. Dr. Raul Zurita-Milla ESA Living Planet Symposium 2025 23-27 June 2025, Vienna, Austria Faculty of Geo-information Science and Earth Observation (ITC) Department of Geo-information Processing [email protected] 2 https://linkedin.com/in/serkan-girgin/ [email protected] 1 https://www.linkedin.com/in/bhawiyuga/ •Increasing volume and variety of remote sensing data •Petabyte-scales of continental-level analysis Earth Observation Big Data (EOBD): Increasing Importance Decompose and Distribute Tasks •Spatial-temporal-spectral dimensions •Take benefit of distributed processing •Decompose the data •Perform computation in parallel •Cloud computing for EOBD •Offers scalable, elastic, and ondemand computation resources •Through large pool of computing node with electric power requirement •Example: GEE, Sagemaker, Pangeo •Metrics provided by existing EOBD processing platform •GEE: execution status •AWS Sagemaker: CPU, memory, I/O utilizations through CloudWatch service •Pangeo Cloud: Dask profiler •Challenge: Lack of energy-related metrics to understand the energy cost associated with EOBD computing workflow Metrics on Existing EOBD Processing Platform •460 TW-h data center energy consumption by 2020 •Compared to overall energy consumption of China (366 TW-h) and other Asian countries (236 TW-h) •Green computing initiatives emphasize the importance of sustainable computing inline with EO mission for environmental protection •The energy metrics helps user to understand the energy cost of workflow and make informed decision Why energy metrics matters? Sources: IEA Electricity 2024 IEA Executive Summary •Internal computing nodes •CPU, GPU •Memory •Disk •Power supply •Fan •External •Networking peripherals •External cooling system •Lighting Look into the Cloud Infrastructure Perspective for Energy Consumption Existing Energy-Related Metrics Source Metrics Granularity -level Sampling interval External power sensor Voltage (V), current (A), power (W), energy consumption ( Wh) System -wide (external) 1s IPMI Energy consumption ( Wh) System -wide (internal) 1s RAPL CPU, DRAM, and entire CPU package energy consumption ( Wh ) System -wide, perhardware component 1 ms NVIDIA / AMD SMI GPU Power (W) System -wide 1 s Operating system CPU utilization (%), memory utilization (%), disk I/O (Mb/s), network transmission rate (Mb/s) System -wide, perprocess 4 ms Application profiler Running tasks, CPU utilization (%), memory utilization (%) , disk I/O (Mb/s), network transmission rate (Mb/s) Per -task 1 s Existing Projects on Integrated Energy Measurement Parameters Schapandre Kepler EAR Focus VM Kubernetes (pods) HPC (bare -metal) Node -level No No Yes (IPMI) CPU and memory power Yes (RAPL) Yes (RAPL) Yes (RAPL) GPU power No Yes (NVML) Yes (NVML) I/O power No No No Network power No No No Auxiliary components (cooling, network switch, etc.) No No No VM passthrough Yes (QEMU) No No VM estimate No Yes No Pods -level estimate No Yes No Specific process -level estimate No Yes (pods -level) Yes (jobs level) Often Neglected Gaps in Energy Monitoring Data source variation •Various sensors with different sampling rates •Time synchronization is required Cloud abstraction •Lack of transparency in virtualized environments. •Cloud providers don’t expose detailed energy cost of each running service Hidden costs •Energy from cooling, PSUs, and networking is often hidden. Task Attribution •Challenge in attributing energy to individual task. •Even need hierarchical attribution in virtualized environment Standard benchmark •Lack of standardized frameworks to compare energy efficiency of various EO tasks Icons by https://flaticon.com •Combining multi-level energy-related measurements •Hardware-Level Data: •CPU/DRAM: RAPL •Computer: IPMI •Overall Consumption: External Power Sensors •Software-Level Metrics: •OS metrics: CPU and memory usage, I/O and network activities •Task metrics from application profiler Architecture of Proposed Approach •Modelling relationship between I/O and networking activities to power consumption •Hierarchical power attributions through virtual machine •Test with various benchmarking tasks •Integration with the infrastructure orchestration (e.g. Kubernetes) •Power modelling for public cloud computing instance Future Work Illustration by Storyset.com Adhitya Bhawiyuga PhD Candidate Department of Geoinformation Processing [email protected] https://www.linkedin.com/in/bhawiyuga/ Contact us if you want to learn more or collaborate! Faculty of Geo-information Science and Earth Observation (ITC) Dr. Serkan Girgin Head of Department Center of Expertise in Big Geodata Science Associate Professor Department of Geoinformation Processing [email protected] https://linkedin.com/in/serkan-girgin/ https://itc.nl/big-geodata/