Full text
CDIF-4-XAS: Mappings from Community Standards to CDIF D2: Semantic description of at least two XAS community standards using a CDIF profile (XAS-CDIF) Steven Richard (Committee on Data of the International Science Council [CODATA]), Arofan Gregory (CODATA), Patrick Austin (Scientific Computing Department, Science and Technologies Facilities Council, UKRI), Heike Görzig (Helmholtz-Zentrum Berlin für Materialien und Energie), Simon Hodson (CODATA), Rolf Krahl (Helmholtz-Zentrum Berlin für Materialien und Energie), Markus Kubin (Helmholtz-Zentrum Berlin für Materialien und Energie, Helmholtz Metadata Collaboration), Leandro Liborio (Scientific Computing Department, Science and Technologies Facilities Council, UKRI), Abraham Nieva de la Hidalga (Cardiff University), Deirdre Lungley (UK Data Archive); Flavio Rizzolo (Statistics Canada); Vyacheslav Tykhonov (CODATA)
CDIF-4-XAS project report metadata Project number 01-314 Project acronym CDIF-4-XAS Project name Describing X-ray spectroscopy data for cross domain use Call 1st OSCARS Open Cal Topic OSCARS-PaNOSC Type of action Task 2 Overview of standards, vocabularies (and ontologies), data formats and practices within the XAS area (landscape analysis) Project starting date 01/10/2024 Project duration 24 months REPORTING PERIOD Period covered from 01/10/2024 to 28/02/2025 Reporting period number 2 Periodic report date and version 24/10/2025, version 1 Authors Steven Richard (Committee on Data of the International Science Council [CODATA]), Arofan Gregory (CODATA), Patrick Austin (Scientific Computing Department, Science and Technologies Facilities Council, UKRI), Heike Görzig (Helmholtz-Zentrum Berlin für Materialien und Energie), Simon Hodson (CODATA), Rolf Krahl (Helmholtz-Zentrum Berlin für Materialien und Energie), Markus Kubin (Helmholtz-Zentrum Berlin für Materialien und Energie, Helmholtz Metadata Collaboration), Leandro Liborio (Scientific Computing Department, Science and Technologies Facilities Council, UKRI), Abraham Nieva de la Hidalga (Cardiff University), Deirdre Lungley (UK Data Archive); Flavio Rizzolo (Statistics Canada); Vyacheslav Tykhonov (CODATA) Contributions SR and AG were the lead authors. All other authors contributed editing the draft during revision, and discussed and agreed on the coverage, content, and formatting of the document during the reported period. CHANGE HISTORY VERSION PUBLICATION DATE CHANGE 1.0 24/10/2025 First public version 2
Contents 1. Introduction................................................................................................................... 4 2. Existing Standards and Scenarios of Use......................................................................... 5 3. Inputs and Outputs.........................................................................................................8 3.1 XDI-to-CDIF....................................................................................................................8 3.2 NXxas-to-CDIF................................................................................................................9 3.3 Static Concept Definitions........................................................................................... 11 4. Further Work................................................................................................................ 12 5. References....................................................................................................................13 3
1. Introduction This document and the associated spreadsheet and other materials present an initial attempt to take the major community standards used for X-Ray Absorption Spectroscopy (XAS) and map them to the standards recommended for cross-domain FAIR sharing of data by the Cross Domain Interoperability Framework (CDIF) guidelines. The current standards landscape within the XAS community has been extensively described in the document “Overview of X-Ray Absorption Spectroscopy standards, vocabularies (and ontologies), data formats and practices” [1]. This document builds on the analysis presented in that document as a concrete exploration of how the data and metadata from the two most common XAS standards can be expressed in a CDIF-described package to facilitate use across domain, institutional, and application boundaries. It is felt that a concrete application of the standards and guidelines involved will clearly indicate the next steps for implementation. While the focus of this document is technical, there are also implications for how other activities by domain groups and standards bodies can most effectively be conducted. These implications can be provided as feedback to various groups, and these will be mentioned here. The specific recommendations to be made to such groups do not, however, form part of this deliverable, but will be formulated more completely in other project deliverables in the future. 4
2. Existing Standards and Scenarios of Use The document “Overview of X-Ray Absorption Spectroscopy standards, vocabularies (and ontologies), data formats and practices” [1] clearly identifies two major standards within the XAS community which have become the focus for harmonization of data description and formatting. The first of these is the NeXus format [2, 3] using HDF5 as the file format and the NXxas application profile to hold relevant metadata and description. The second standard is the XAS Data Interchange (XDI) standard [2], which describes a text-based format for both metadata and data (for an overview of these see the “Section 5.1 Community defined standards” in [1]). It should be noted that these two standards are used for different things: NXxas is used to describe the raw data resulting from one or more measurements or scans. While it can have additional application data added to it or linked from it, it builds on the raw data which it describes for these functions. It is used as a means of exchanging the raw data between applications, a task which may involve large amounts of data. Consequently, HDF5 has emerged as the best way to package the data in these scenarios. XDI is much more focused on the analytical aspects of XAS. It is designed to “encapsulate a single spectrum of XAFS along with relevant metadata” [4]. The intention here is to provide this data for exchange between analysis programs, spreadsheets, and data visualization tools. In practical terms, the volume of data to be expressed as XDI is much smaller than that for NXxas, as both the data and metadata are combined in a single text-based format. From the perspective of CDIF, the goal is to allow any potential users of the data to find and assess the data for FAIR purposes, which involves understanding what is in the data sets described. Further, it should be possible to programmatically access the data and transform it into the form necessary for reuse. This “FAIR” functionality is independent of the specific processes supported by the different XAS community standards, but relies on the information contained in those standard formats (as well as some additional information for cataloguing purposes.) These FAIR functions, however, are very much aligned with the stated goals of those who produce the community standards, as we see in their stated goals [5]: ● The benefits of data integration should be not only in data-driven science but also in everyday research. ● The data and metadata should be in as few formats as possible (ideally following an agreed data schema). ● The publication infrastructure should be prepared as a repository with policies for 5
● data utilization, such as the FAIR Principles. ● The database infrastructure should have a search functionality and not just provide online storage. For the purposes of the current mapping exercise, two major scenarios emerge: (1) a data set which is encoded in HDF5 according to the NXxas application profile of NeXus will be transformed into a “cross-domain” CDIF metadata description to accompany the HDF5 file; and (2) an XDI text file will have its data contents described using the “cross-domain” CDIF recommendations for needed metadata. In both cases, the consumer of the data may be ignorant of the source formats – that is, their systems may not be designed to work with the standards employed. Regardless, they should be provided with enough information to maximise the machine-accessibility of these data, and to support search and other FAIR functions. It should be noted that within the XAS community, cataloguing metadata is generally held external to the data files, and this is largely true of both the community standards described here. CDIF combined the cataloguing and data description metadata into a single format for FAIR purposes, and in these cases the XAS data will often be supplemented by cataloguing metadata held in typical “standard” formats such as the DataCite metadata scheme. Additionally, the entire package may need to be supplemented with information regarding access, licensing, and fair use. Such information often comes from institutional policy guidance, and may not be expressed in a standard form. One requirement of CDIF, which is not fully met by existing community standards, is the need for formal definitions of the concepts used in describing the data. These can be exposed as labels for columns of fields and similar descriptors of data. CDIF demands that these be described in a standard fashion, because FAIR users may be unfamiliar with the use of terms in the institution which has produced the data. While there are some formal definitions in the XAS community specifications, these are not expressed in a standard or comprehensive form. They can, however, form the basis of such an expression. Note that CDIF metadata is always expressed in a JSON-LD syntax, optimized for FAIR use/reuse over the Web. It leverages both the familiar Javascript features of JSON, as well as the more powerful RDF features of Linked Data approaches. The standards recommended by CDIF include Schema.org for cataloguing and discovery [6], W3C PROV for provenance information [7], W3C SKOS for controlled vocabularies [8], and the Data Documentation Initiative’s Cross-Domain Integration (DDI-CDI) specification [9]. 6
Each of these specifications are used according to profiles described in the CDIF guidelines, although the provenance guidelines are currently under development. CDIF currently focuses on describing data sets encoded in text-based formats. The need to describe binary formats (including HDF5) has been noted in the initial release of the guidelines, and work in this area is on-going. 7
3. Inputs and Outputs Given the scenarios described above, we can produce high-level schematics of the transformation processes which will be supported. These should make the spreadsheet mappings easier to follow, as there are several different inputs and outputs involved. 3.1 XDI-to-CDIF The schematic in Figure 1 shows the mapping from XDI to the recommended CDIF-4-XAS specification. The inputs include an XDI file, which contains both a metadata section and a data section, and a set of cataloguing metadata (such as a DataCite file). The XDI file may be “passed through” the transformation to act as a data file in the resulting package, or the data may be pulled out and reduced to a simple data-only text file. The metadata is transformed in either case, and placed in different areas of the CDIF package (JSON-LD file and data), based on what type of information it is. The cataloguing of metadata is a fairly straightforward mapping into corresponding Schema.org fields. Figure 1 High level schema of XDI metadata extraction and mapping to the proposed CDIF-4-XAS profile Note that CDIF uses SKOS to describe enumerated codes/values found in the data – this is the set of controlled vocabulary metadata inside the CDIF JSON-LD package. (The dashed 8
line indicates the use of these concepts using links.) The Reference Concepts are a static set of terms drawn from authoritative sources within the XAS community (or from authorities recognised by that community). This is also expressed as SKOS concepts in a SKOS concept scheme, but would be published externally to the CDIF package for each data set (it does not vary from data set to data set, so a single set of definitions can be reused). Note also that the resulting resources are completely “open” in that they are all expressed in open standard formats or as structured text (for the data). No proprietary software is needed to access the files, and all of the standards are Web-based (JSON-LD/RDF, text) so that it is as easy as possible to access and use the resulting metadata. From a FAIR perspective, this is ideal, and very much in line with best practice as endorsed by the CDIF guidelines. The provided metadata is also very rich, giving complete definitions of terms, and indicating the structure and semantics of the data in a (potentially) comprehensive fashion. 3.2 NXxas-to-CDIF The schematic in Figure 2 shows the mapping of NXxas HDF5 files to the recommended CDIF-4-XAS specification. This mapping is very similar to the XDI mapping. The main difference is that the data are always kept in a binary format that of the HDF5 source file. All data are referred to/from the metadata using path expressions which can be resolved by HDF5 software packages. The metadata is all stored in “open” JSON-LD form, in the same type of structure used for the XDI metadata. 9