LEAPS-INNOV WP7 D7.3 Report results on promising data reduction pipelines on datasets provided by LEAPS facilities and release as open-source to the community
Abstract
This report presents work on applying modern data reduction pipelines to datasets fromLEAPS facilities, and helping staff at LEAPS facilities to develop pipelines.
Full text
Deliverable no. D7.3 Project information Project full title LEAPS pilot to foster open innovation for accelerator-based light sources in Europe Project acronym LEAPS-INNOV Grant agreement no. 101004728 Instrument Research and Innovation Action (RIA) Duration 01/04/2021 – 31/03/2025 Website https://www.leaps-innov.eu/ Deliverable information Deliverable no. D7.3 Deliverable title Report results on promising data reduction pipelines on datasets provided by LEAPS facilities and release as open-source to the community Deliverable responsible Peter Steinbach (HZDR) Related Work-Package/Task WP7 Data Reduction and Compression Task 7.3: Adopt new algorithms and technologies for lossless and lossy data compression Type (e.g. Report; other) Report Author(s) Peter Steinbach (HZDR) Vincent Favre-Nicolin (ESRF) Francesco Guzzi (ELETTRA) Jérôme Kieffer (ESRF) Zdenek Matej (MAX IV) Alberto Mittone (ALBA-CELLS) Pierre Paléo (ESRF) David Pennicard (DESY) Nicolas Soler (ALBA-CELLS) Thomas Vincent (ESRF) Dissemination level Document Version Date Download page Document information Version no. Date Author(s) Comment 1 20231218 Page 1of 24
Deliverable no. D7.3 Table of contents Project information 1 Deliverable information 1 Document information 1 Abbreviations 3 Introduction 4 Data in Active Use 5 Motivation 5 Micro Data Sheets 6 Collected Datasets for this study 8 The ROFEX dataset 8 The PSI evolving magma dataset 9 Lossless Compression 10 Reproducibility of the Magma Tomography Dataset 10 Lossless Compression with Blosc2 11 Finding good compression parameters with AI guidance 11 Lossy Compression 12 Denoising Tomograms 12 Quantising Tomograms 15 On Metrics and Lossy Compression 17 Minding the gap 19 Summary 21 References 22 Appendix 23 Page 2of 24
Deliverable no. D7.3 D7.3 Applying Modern Data Reduction Pipelines to LEAPS Data Abbreviations zstd zstandard compression codec optimized for speed and high compression ratios blosc high-performance compressor that has been optimized for binary data tiff Tagged Image File Format Page 3of 24
Deliverable no. D7.3 Introduction In recent years, photon science has witnessed a remarkable surge in data rates, as documented in the report of WP7.2. This escalating trend in data generation has brought about new challenges and opportunities that are reverberating throughout the scientific community. Researchers, practitioners, and enthusiasts alike are becoming increasingly aware of the fact that the ongoing growth in data volumes and complexity has pushed the boundaries of conventional data management and analysis methods to their limits. The community anticipates a continuous and exponential increase in data rates, demanding innovative solutions that can not only cope with this deluge of information but also unlock valuable insights that lie within. It is important to emphasize that this rapid increase in data volumes is not confined solely to the realm of photon science; it is a broader trend across many scientific domains and industries. The surge in data rates reflects the progress and advancements in measurement instruments, computational power, and experimental techniques, making it imperative to adapt and evolve our strategies accordingly. This report presents our work on applying modern data reduction pipelines to datasets from LEAPS facilities, and helping staff at our facilities to develop pipelines. This may highlight important findings which may help practitioners and decision makers beyond our community. Page 4of 24
Deliverable no. D7.3 Data in Active Use Motivation In the landscape of modern science, data serves as the lifeblood of research endeavors across diverse academic domains. The realm of photon science is a powerful tool for research in many different fields, and photon science-based experiments, by their very nature, yield copious volumes of data. However, at the inception of our project, we confronted an unfortunate reality: the availability of open datasets, conforming to the principles of Findable, Accessible, Interoperable, and Reusable (FAIR) data sharing, remained in a early stage, with dissemination channels and awareness lagging behind the burgeoning data demand. For this reason, we focussed our efforts in WP7.3 on a small assortment of datasets that stand exemplary for other use cases in photon science. Moreover, we quickly realized a daunting concern extending beyond mere data availability: the provenance of these open datasets. Data provenance refers to the documented history or lineage of a dataset, tracing its origins, transformations, and any changes it has undergone throughout its existence. This information is crucial as it helps researchers and scientists understand the reliability and quality of the data, ensuring that they can trust and use it with confidence in their academic work and scientific inference. Data provenance can be compared to a roadmap that helps navigate the journey of data, making it easier to assess its trustworthiness and suitability for research purposes. Provenance proved to be a crucial aspect for evaluating data quality and reliability, however we observed that it was often absent. In an era where data underpins critical scientific discoveries, the opaqueness surrounding data treatment and processing posed a significant challenge to our project. In response to these systemic gaps, our project took on the task of curating and collecting a small set of representative datasets, engaging in a concerted effort to establish collaborative relationships with data owners who were receptive to our inquiries. This led us to the development of succinct and standardized micro data sheets. These datasheets serve to communicate essential information about the datasets, fostering a shared understanding of their characteristics and provenance, and ultimately contributing to the elevation of data quality and accessibility in photon science and beyond. Page 5of 24
Deliverable no. D7.3 Micro Data Sheets The concept of micro data sheets is inspired by (Gebru 2018), which is a seminal work addressed to a broad audience in machine learning. Although our approach is similar, the list of points covered by the original datasheet article is excessively long and we feared that such a datasheet would not be adopted by data donors due to the time investment required to provide this data manually. In addition, the original micro data sheet in (Gebru 2018) focuses on many aspects which are irrelevant to common use cases in photon science (e.g. civil rights compliance). Thus, we were looking for a simpler, version-control friendly way to describe datasets. The application we had in mind was that such information could accompany a dataset on publicly available platforms like zenodo, dryad and others. While these platforms often comply with the FAIR principles already, the information content relevant to compression was often missing. More concretely, the following points were often relevant for our work: ● What was the motivation behind creating the dataset? (this often hints towards bandwidth and compression ratio requirements) ● How is the dataset composed? (this often hints towards which codec best to use, i.e. is this image or video or 3D like data) ● How was the data collected? (if modern CMOS or CCD sensors were in use, this often implies a concrete noise model for the data) ● How was the dataset preor post-processed? (which algorithms were applied to the data before compression will affect intensity scaling effects and other artefacts) While a standardisation of above mentioned queries would be beneficial, we decided to adopt an agile approach to this and implemented this micro data sheet scheme by virtue of a small lightweight template markdown document. Moreover, our team size was small so that deviations from this template could be handled by communication. An example data sheet is depicted in figure 1. Page 6of 24
Deliverable no. D7.3 Figure 1: Rendering of a markdown based micro data sheet as part of a gitlab hosted version control repository. Easy access, essential information contained. Page 7of 24
Deliverable no. D7.3 Collected Datasets for this study Based on the challenges outlined above, our work culminated on the following datasets. For each of these datasets, we were able to establish contact with personnel who were involved in creating these datasets. The ROFEX dataset The high-performance ROFEX (ROssendorf Fast Electron beam X-ray tomography) imaging technique has been developed at Helmholtz-Zentrum Dresden-Rossendorf for the noninvasive investigation of dynamic processes. Purpose-built for the analysis of multiphase flows in elongated test sections, various research projects out of this field could already benefit from the ROFEX technology. Beyond that it can be applied to diverse other applications such as for instance nondestructive testing. Figure 1: Example frames of a ROFEX multi-phase flow experiment (left) and a single bubble experiment (right). Both images were extracted from a dataset after reconstruction and postprocessing. The dataset depicted in figure 1 left was provided as a hdf5 file as produced by Matlab. Each file consisted of 15000 frames of observations. The dataset itself used in this analysis was encoded as 32 floating point grayscale. The data comprised 15000 frames of WxH = 256x256 pixels each. The dataset of figure 1 right consists 1024 frames of shape WxH = 256x256 pixels each. The grayscale pixel intensities are encoded as 16 bit unsigned integer values. Page 8of 24
Deliverable no. D7.3 The PSI evolving magma dataset In the associated article (Pistone 2021) to these datasets, the authors investigate the effect of lossy compression of original X-ray projections onto the final tomographic reconstructions. The dataset in use was recorded at the TOMCAT experiment at Paul-Scherrer-Institut. The beamline for TOmographic Microscopy and Coherent rAdiology experimentTs (TOMCAT) is operated by the X-ray Tomography Group and offers cutting-edge technology and scientific expertise for exploiting the distinctive features of synchrotron radiation for fast, non-destructive, high resolution, quantitative investigations on a large variety of samples. Figure 2: Example frames of a PSI evolving magma dataset. Both images were extracted from a dataset after reconstruction. The raw dataset contained 751 frames encoded as 16 bit unsigned int. Each frame had a shape of WxHdu = 1900x1008 pixels. The full volume of this dataset comprises 2.743 GB and represents one measurement by the TOMCAT detector. This dataset and many others are available at https://doi.psi.ch/detail/10.16907/05a50450-767f-421d-9832-342b57c201af Page 9of 24
Deliverable no. D7.3 To enhance this approach, we considered a modality by the ROFEX experiment where the science goal is to identify and characterise the position of a single bubble moving through the detector, see Figure 1 right. This dataset was encoded as a float32 grayscale image. As float32 encoding requires 4 bytes to represent a pixel intensity. The outlook of reducing this to 8 bit would offer a reduction in size by 4x. Here, a two-step preprocessing before compression was devised and offered convincing results. Our first step was to threshold the image using Otsu’s method. As this differentiates foreground and background, we quantised the background intensity with 4 color values and the foreground with 252 color values. This quantisation was step 2 and produces a dataset that can be encoded with 8 bits for each pixel. The transformation employed for the quantisation can be stored as a lookup table alongside the data. Eventually, the quantized dataset was compressed using zlib using default parameters. We refer to this approach as TQC,threshold quantize compress. Figure 6: Single Bubble dataset by the ROFEX experiment. Left: Original slice. Right: Same slice made subject to the TQC algorithm and converted back to float32 for visualisation. As figure 6 demonstrates, the TQC algorithm offers visually compelling results when focussing only on the single bubble presented. Moreover, we observed a compression ratio of 18.34x, which in turn opens tremendous perspectives for the ROFEX experiment and its wide reaching scientific portfolio. Page 16 of 24
Deliverable no. D7.3 On Metrics and Lossy Compression The previous section showed how a tailored pipeline of preprocessing steps can help to elevate the compression ratios of modern compression codecs given scientific data from the photon science community. We demonstrated that this can lead to compression ratios above 15x. While the compression ratios were not as high, works like (Pistone 2021) have reported similar conclusions. Lossy compression can be applied to scientific data. Taking this by heart, one central question remains: how much science is obstructed by using lossy compression algorithms. Works like (Pistone 2021) have tried to answer this question by comparing the compressed image with the original and computing image quality metrics to quantify the change of image content. This was then considered an epistemological proxy for how much scientific value is retained despite discarding information during the compression process. As this already indicates this leap of faith bears some imprecision. Moreover, many scientists are rarely trained in the interpretation of image quality metrics. Henceforth, the adoption of lossy compression is not widely spread or even considered. Many scientists are simply afraid to discard information which might turn out to be essential for a yet unknown processing step in a downstream analysis. Within this project, we aspired to change this situation making this fear of information loss concrete. We contacted the dataset owners and asked them if we could get access to a downstream analysis of their data which led to a past publication. To our surprise this enterprise proved fruitless very quickly. While the publication of data has received traction throughout the community, the archival of software which produced this dataset is often discarded or forgotten. For the ROFEX single bubble dataset, we were able to uncover such a downstream analysis. To our luck, the software was still available and working well. For a given ROFEX dataset, this downstream analysis would run calculations on each pixel intensity to obtain velocity fields and the like which describe the liquid and gas which are being observed. This resulted in 15 observable quantities per pixel. We were able to contrast these 15 variables per pixel in the original dataset versus the decompressed dataset from the TQC algorithm. From this difference we computed the normalised root mean square error (NRMSE) value to form a summary statistic. Page 17 of 24
Deliverable no. D7.3 Figure 7: Analysis of the effect of lossy compression on the single bubble ROFEX dataset. The figures compare the NRMSE difference between the original data versus a decompressed version of it. We compare naive quantisation (blue) and the TQC algorithm (green) Top: NRMSE values for slice D009. Bottom: NRMSE values for slice D002. Figure 7 illustrates the impact of our TQC lossy compression scheme on the downstream quality of 15 physics variables. Each variable is computed per pixel of the reconstructed tomograms. For this reason, we report the summary statistics NRMSE describing the difference between the computed value from the original dataset versus the value obtained from the decoded TQC compressed data. The graphs in figure 7 report on two different datasets of the same measurement campaign, i.e. D009 and D002. In figure 7, two observations can be made: first, TQC with thresholding by Otsu’s methods provides lower variations to the original than without; second, all observed differences range below a value 0.01 and thus only exhibit a variation of below 1%. Even more so, the application of the full TQC pipeline (i.e. with thresholding) ranges below or around an NRMSE of 0.002 or 0.2%. Analyses like this can spotlight the effect of lossy compression on the quality of data and the effect on downstream analysis. Especially in the physical sciences, this allows practitioners to grasp the effect quantitatively regarding their scientific output. Moreover, it helps to remove the effect of overestimation of detrimental consequences inherent to the term lossy compression. We hope that this demonstrates an alternative avenue to make lossy compression approachable by people untrained in image analysis. Page 18 of 24
Deliverable no. D7.3 Minding the gap As illustrated in the previous section, the availability of open datasets is rapidly increasing across the scientific community and across photon science facilities. While this is a trend which must be praised extensively, our efforts in this work package have demonstrated a challenge with this. This impediment appears even more pronounced with data modalities which reconstruct the primary scientific output from observed raw data. This challenge originates from the unavailability of post processing software which was used to produce the primary science output of an experiment using the raw data as input. Post processing in itself is almost exclusively a multi-step effort. These pipelines typically comprise filters and transformations from signal processing as well as reconstruction algorithms rooted in the geometry of an experiment. In order to expand the availability of reproducible workflows for such postprocessing, we devised a multi-week effort to intensify the skillsets of beamline scientists to produce and reuse reproducible software in this domain. This effort consisted of a remote hackathon “LEAPS-INNOV Workflow Co-Working Sprint”. This sprint took place from Feb 10, 2023, to March 30, 2023, and stretched over 6 weeks. We invited the community to apply as teams. These teams were matched with experienced research software engineers (RSEs) that provided software support for the team’s projects. This effort is quite new in the photon science community and needed to be advertised through the LEAPS-Innov channels. We had 4 applications from teams based at DESY, ALBA, KIT and HZDR. For the role of mentors, we were able to attract RSEs from HZDR and Seqera Labs (a company supporting the nextflow workflow engine). Page 19 of 24
Deliverable no. D7.3 Figure 8: Agenda of the LEAPS-Innov workflow co-working sprint kickoff meeting. Contributions from two open-source workflow engines (snakemake and nextflow) as well as one standardisation effort (cwl). Four teams presented their plans for the co-working sprint. The hackathon was a full success. Even though the team from KIT dropped out, the process uncovered challenges and low hanging fruit for the participating teams. With respect to technology, two workflow engines were used: nextflow and snakemake. The participants acknowledged the effort invested. To underline this, these testimonies were provided within an anonymous quality assurance survey: “I liked the structure with the first day being an introduction of the tools by experts and then mentors working with teams. I hope there will be co-working events like this in the future.” “Great opportunity to make some concrete big improvements” Page 20 of 24
Deliverable no. D7.3 “I've learned usage of snakemake, liked very personal approach,” We hope that such efforts to co-create software which is bound to be reused extensively during the lifetime of a beamline experiment can be repeated and expanded. We believe that such efforts provide hands-on training and knowledge transfer between RSEs and domain specialists working at a beamline. Summary In the ever-evolving realm of compression, the continuous emergence of new algorithms coincides with the escalating data-generating prowess of our hardware. This convergence necessitates a keen focus on preserving data lineage and ensuring algorithm reproducibility, both pivotal in propelling innovation and establishing benchmarks for cutting-edge compression techniques. It's imperative to construct cohesive workflows that intertwine various algorithms, seamlessly melding post-processing with compression methodologies. Lossless compression, a stalwart in this arena, showcases the potential to achieve compression ratios ranging from 3-4x, surging past speeds of 1GB/s and pushing bandwidths beyond the 20GB/s threshold. The automation of compression codec parameter searches assumes paramount importance, enabling the discovery of configurations finely tuned to match experimental requisites and hardware capabilities. Conversely, the allure of lossy compression lies in its ability to yield notably higher compression ratios, reaching up to 18x, thereby unlocking avenues for broadening scientific data acquisition horizons or curtailing storage expenses. However, navigating the usability evaluation of lossy compression algorithms poses a challenge. The reliability of downstream analysis often remains elusive in terms of reproducibility, rendering image quality metrics as mere indirect surrogates for scientific yield. To comprehensively gauge the efficacy of lossy compression, a quantifiable approach within end-to-end workflows is indispensable, establishing a seamless connection between raw data and tangible scientific outputs such as object segmentations, downstream calculations, and intricate modeling. Notably, novel methodologies, such as collaborative sprint initiatives aimed at fostering software acumen, have surfaced in this work package, garnering positive acclaim from the scientific community for their role in engendering reproducible workflows. Page 21 of 24
Deliverable no. D7.3 References Batson, Joshua, and Loic Royer. 2019. “Noise2Self: Blind Denoising by Self-Supervision.” https://arxiv.org/abs/1901.11365. Gebru, Timnit. 2018. “Datasheets for Datasets.” arXiv. https://arxiv.org/abs/1803.09010. Leonarski, Filip, and Martin Brückner. 2022. “Jungfraujoch: hardware-accelerated data-acquisition system for kilohertz pixel-array X-ray detectors.” Journal of synchrotron radiation 30, no. 1 (10): 227-234. https://doi.org/10.1107/S1600577522010268. Mäkinen, Ymir, Lucio Azzari, and Alessandro Foi. 2020. “Collaborative Filtering of Correlated Noise: Exact Transform-Domain Variance for Improved Shrinkage and Patch Matching.” IEEE Transactions on Image Processing 29, no. n.n. (08): 8339-8354. 10.1109/TIP.2020.3014721. Masui, Kiyoshi. 2015. “A compression scheme for radio data in high performance computing.” arXiv. https://arxiv.org/abs/1503.00638. Pistone, Mittia. 2021. “Seismic Attenuation During Magma Vesiculation: A Combination of Laboratory Constraints and Modeling.” https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2020GL092315. Page 22 of 24
Deliverable no. D7.3 Appendix Use of chatbots for drafting this document Before human editing and writing, this document was drafted using the help of a chatbot (chatGPT at the time of writing). The following lists the prompts which were in use. Note: All content was proofread by a human and henceforth can be considered a human product. - introduction: “I am writing a report to summarize our activities to compress data. Write an introduction with a length of 2 paragraphs which highlights the following points: - as documented in XX, data rates are increasing in photon science - the community in photon science expects this trend to continue - new methods are required - this is a trend not only observed in photon science” - data in active use: “Paraphrase the following points in a positive and academic language! Produce readable prose text that can fill 2 pages. - data is an essential ingredient in modern science - photon science based experiments are by definition among the strong data producers in academia - when project started, availability of open datasets was low when project started (FAIR sharing hardly disseminated) - availability of provenance of open datasets did not exist (if data was available, nobody knew how it was treated) - we collected representative datasets where data owners where responsive to our contact requests - designed short formatted datasheet to communicate about datasets” - summary: “ Q: Paraphrase the following bullet points as if you were a photon scientist which is very proficient in data compression. Produce at least two paragraphs. - new compression algorithms are constantly published, at the same time the capability of hardware to produce data increases - data provenance and algorithm reproducibility are essential to help innovate and benchmark state of the art compression algorithms - workflows to chain algorithms together are important (combine post processing and compression) Page 23 of 24
Deliverable no. D7.3 - lossless compression can achieve compression ratios of 3-4x at more than 1GB/s, reaching bandwidths of 20GB/s and above - automated tests of compression codec parameters is essential to find configurations adapted to the experimental requirements and hardware - lossy compression can achieve higher compression ratios (18x) and thus offers potential to expand scientific data acquisition or lower storage costs - judgement of usability of lossy compression algorithms is tricky as downstream analysis is often not reproducible and image quality metrics can only serve as a indirect proxy for scientific yield - judgement of lossy compression must be quantified in end-to-end workflows which connect raw data and viable scientific outputs (object segmentations, downstream calculations and modelling) - new ways to foster software skills to create reproducible workflows have been established (co-working sprint) and received positive feedback by the community A:” Page 24 of 24