scieee AI-readable full text Open interactive document viewer

Beyond FAIR: Manifesting marine data synthesis products within the ocean observing system: A biogeochemical essential ocean variables perspective

Lange, Nico-Maurice Dennis

Abstract

Efficient usage of the highly heterogeneous, high-volume biogeochemistry (BGC) essential ocean variables (EOV) data determines the success of BGC observations in supporting the development of evidence-based strategies for climate change mitigation, and adaptation, ecosystem health conservation, and sustainable resource management. Acknowledging the positive impacts of BGC EOV data synthesis products on the BGC data landscape, the overarching goal of this thesis was to further manifest BGC EOV synthesis products as an integral part in the ocean observing system. For this purpose, first, a novel evaluation scheme was developed that enables an objective assessment of the readiness of BGC EOV data synthesis products following the system-engineering approach of the Framework of Ocean Observing. In parallel, over the course of this thesis annual updates of the GLobal Ocean Data Analysis Project (GLODAP) were realized since 2019. The most recent, GLODAPv2.2022, represents the largest and most consistent dataset for carbon-relevant hydrographic cruise data, including data from 1,085 hydrographic cruises with 1,381,248 water samples. Additionally, in pursue of a higher readiness for GLODAP, the vision of a uniform, semi-automatic, and standards-compliant data ingestion system in combination with a modern and versatile data extraction system were developed and outlined. Furthermore, a pilot for the Synthesis Product for Ocean Time Series (SPOTS) was successfully produced during this thesis. Thereby, a template for a sustained living SPOTS was created and the BGC data landscape was expanded by the previously overlooked ship-based time-series programs. For the pilot, a total of 108,332 water samples from 12 ship-based time-series programs were synthesized. Altogether, through this thesis, important steps towards manifesting BGC EOV data synthesis products as an integral part in the ocean observing system were realized.

Full text

Beyond FAIR: Manifesting marine data synthesis products within the ocean observing system – A biogeochemical essential ocean variables perspective Dissertation zur Erlangung des Doktorgrades der Mathematisch-Naturwissenschaftlichen Fakultät der Christian-Albrechts-Universität zu Kiel vorgelegt von Nico-Maurice Dennis Lange Kiel, Juli 2023 II Erster Gutachter: Prof. Dr. Arne Körtzinger Zweiter Gutachter: Tag der mündlichen Prüfung: Prof. Dr. Are Olsen 26.09.2023 III Summary In times of increasing anthropogenic impacts jeopardizing the environmental status of the ocean and the associated services for society, sustained and globally well-coordinated observations of biogeochemistry (BGC) essential ocean variables (EOV) are vital. Efficient usage of the highly heterogeneous, high-volume BGC EOV data determines the success of BGC observations in supporting the development of evidence-based strategies for climate change mitigation, and adaptation, ecosystem health conservation, and sustainable resource management. Recognizing the limiting factor of the fragmented BGC data landscape in this context, and addressing the key data challenges, community-driven BGC EOV data synthesis products constitute key outputs from the Global Ocean Observing System. These products connect BGC EOV observations with societal services by applying customized techniques to combine datasets from multiple sources to provide comparable, Findable, Accessible, Interoperable, Reusable (FAIR), and fit-for-purpose data. However, BGC EOV synthesis products are far from a sustained status, and insufficient BGC data still represent a weak spot in the ocean observing system, resulting in large uncertainties and unresolved issues in marine BGC research. Acknowledging the positive impacts of BGC EOV data synthesis products on the BGC data landscape, the overarching goal of this thesis was to further manifest BGC EOV synthesis products as an integral part in the ocean observing system. For this purpose, first, a novel evaluation scheme was developed that enables an objective assessment of the readiness of BGC EOV data synthesis products following the system-engineering approach of the Framework of Ocean Observing. Its application to four community-driven BGC EOV data synthesis products, in descending readiness level order: The Surface Ocean CO2 Atlas (SOCAT), the Global Ocean Data Analysis Project (GLODAP), the MarinE MethanE and NiTrous Oxide (MEMENTO) data product, and the Global Ocean Oxygen Database and ATlas (GO2DAT) enabled the identification of critical features of BGC synthesis data products. Eventually, the identified features, e.g. traceable and customized Quality Control (QC), and the developed readiness evaluation scheme at large are envisioned to guide new and existing BGC data synthesis products in eliminating their weaknesses and realizing their full potential. In parallel, the goal of annual GLODAP product updates was realized, and over the course of this thesis GLODAPv2.2019, GLODAPv2.2020, GLODAPv2.2021, and GLODAPv2.2022 were released. Combined, a total of 361 new cruises with 381,800 water samples were harmonized, QC’ed, archived, and added. Building upon GLODAPv2, the consistency of the original data could be significantly enhanced. Therefore, the generated GLODAPv2.2022 represents the largest and most consistent dataset for carbon-relevant hydrographic cruise data, including data from 1,085 hydrographic cruises with 1,381,248 water samples, and covering a time period from 1972 until 2021. Several implemented software developments further improved GLODAP’s provenance, and data handling, including the developed Make Ocean merging routine and the applied AtlantOS QC software. Further, in pursue of a higher readiness for GLODAP, the vision of a uniform, semi-automatic, and standards-compliant data ingestion system in combination with a modern and versatile data extraction system were developed and outlined. Furthermore, benefitting from the gained knowledge, and leading the community effort on the Synthesis Product for Ocean Time Series (SPOTS), a pilot was successfully produced during this thesis. Thereby, a template for a sustained living SPOTS was created and the BGC data landscape was expanded by the previously overlooked ship-based time-series programs. For the pilot, a total of 108,332 water samples from 12 ship-based time-series programs, each representative for a different marine environment, and a different program structure, were synthesized. Besides increasing the level of FAIR, implemented i) “Best-Practice” flags, ii) comparisons to GLODAP, and iii) measurement quality IV continuity estimations, further resulted in an increased utility of the data. Moreover, the pilot helped to position ship-based time-series programs for expansion under the United Nations Decade of Ocean Science umbrella as evidenced by the Marine Ecological Time Series network recently obtaining the status of an “Ocean Coordination Group potential new emerging network”. All achievements of this thesis emphasized the value of BGC EOV data synthesis products for the ocean observing system. Particularly, their unique contributions regarding more FAIR, efficient, and utile data were exposed and implemented, guided by Aristotle’s slightly adapted notion “the whole is greater than the sum of its parts”. Thereby, their key aspects for continuous and sustainable success were revealed. Altogether, through the work in this thesis, important steps towards the overarching goal of improving the BGC data landscape through the manifestation of BGC EOV data synthesis products as an integral part in the ocean observing system were realized. However, the thesis also highlighted that more work needs to be done to fully reach this vital goal. V Zusammenfassung In Zeiten zunehmender anthropogener Einflüsse, die den Zustand des Ozeans und die damit verbundenen Leistungen für die Gesellschaft gefährden, sind nachhaltige und global koordinierte Beobachtungen der biogeochemischen (BGC) Essential Ocean Variables (EOV) von entscheidender Bedeutung. Insbesondere bestimmt eine effiziente Nutzung der heterogenen, und großen BGC EOV Datenmengen den Erfolg von marinen BGC Beobachtungen bei der Entwicklung evidenzbasierter Strategien zur Minderung und Anpassung an den Klimawandel, zum Schutz des Ökosystemzustandes und zum nachhaltigem Ressourcenmanagement. In diesem Zusammenhang stellen Communitybasierte BGC EOV Datensynthese Produkte wichtige Outputs des Global Ocean Observing System dar, welche die Herausforderungen einer stark fragmentierten BGC Datenlandschaft erkannt haben und angehen. Diese Produkte verwenden maßgeschneiderte Methoden und Software, um Daten aus verschiedenen Quellen zu kombinieren, um letztendlich vergleichbare, Findable, Accessible, Interoperable, Resuable (FAIR) Daten mit einem hohen Nutz-Faktor zu generieren. Sie stellen so eine direkte Verbindung zwischen BGC EOV Beobachtungen und gesellschaftlichen Dienstleistungen dar. Allerdings haben BGC EOV Datensynthese Produkte noch keinen „sustained“ Zustand im Sinne des Frameworks for Ocean Observing erreicht. Dies ist auch daran zu erkennen, dass unzureichende BGC Daten und Informationsprodukte weiterhin eine Schwachstelle im Ozeanbeobachtungssystem sind, welche folglich zu großen Unsicherheiten und ungelösten Fragen in der marinen BGC Forschung führen. Den Wert von BGC EOV Datensynthese Produkten für die BGC Datenlandschaft anerkennend und darauf aufbauend, war das übergeordnete Ziel dieser Arbeit zu der Manifestierung und Etablierung von BGC EOV Datensynthese Produkte als integralen Bestandteil des Ozeanbeobachtungssystems beizutragen. Für diesen Zweck wurde zunächst ein neuartiges Bewertungsschema entwickelt, das basierend auf dem systemtechnischen Ansatz des Frameworks for Ocean Observing eine objektive Bewertung der „Readiness“ von BGC EOV Datensynthese Produkten ermöglicht. Dessen Anwendung auf vier Community-basierten BGC EOV Datensynthese Produkten; in absteigender Readiness, der Surface Ocean CO2 Atlas (SOCAT), das Global Ocean Data Analysis Project (GLODAP), das MarinE MethanE and NiTrous Oxide (MEMENTO) Datenprodukt und die Global Ocean Oxygen Database and ATlas (GO2DAT) ermöglichte die Identifizierung entscheidender Merkmale von BGC Synthese Datenprodukten. Das entwickelte Bewertungsschema selbst, sowie die identifizierten Merkmale, wie zum Beispiel rückverfolgbare und maßgeschneiderte Qualitätskontrolle, stellen so eine Orientierungshilfe für neue und bestehende BGC Datensynthese Produkte dar. Insbesondere werden so die Erörterung und Beseitigung einzelner Schwächen unterstützt, sodass Produkte ihr volles Potenzial erreichen können. Parallel dazu wurde das Ziel der jährlichen Aktualisierung von GLODAP erreicht, entsprechend wurden im Rahmen dieser Doktorarbeit GLODAPv2.2019, GLODAPv2.2020, GLODAPv2.2021 und GLODAPv2.2022 veröffentlicht. Insgesamt wurden 361 neue Forschungsfahrten mit 381,800 Wasserproben harmonisiert, einer Qualitätskontrolle unterzogen, archiviert und hinzugefügt. Basierend auf GLODAPv2 konnte die Konsistenz der Originaldaten signifikant verbessert werden. So ist das generierte GLODAPv2.2022, mit Daten von 1,085 hydrographischen Forschungsfahrten und 1,381,248 Wasserproben (1972 bis 2021), der größte und konsistenteste Datensatz für kohlenstoffrelevante hydrographische Forschungsfahrtdaten. Mehrere implementierte Softwareentwicklungen haben die Daten Provenanceund Prozesse von GLODAP weiter verbessert. Darunter fallen die entwickelte Make Ocean Merging Routine und die eingesetzte AtlantOS QC Software. Darüber hinaus wurden im Streben nach einer höheren Readiness von GLODAP das Zukunftskonzept eines einheitlichen, halbautomatischen und standardkonformen Daten- VI Aufnahmesystems in Kombination mit einem modernen und flexiblem Datenextraktionssystem entwickelt und skizziert. Des Weiteren, profitierend von dem gewonnenen Wissen und Erfahrungen, wurde im Zuge dieser Doktorarbeit das Synthesis Product for Ocean Time Series (SPOTS) Projekt geführt und dessen Pilotprojekt erfolgreich fertigstellt. Der erstellte Pilot bietet eine Vorlage für einen „sustained“ SPOTS und hat die BGC Datenlandschaft mit den bisher vernachlässigten schiffsbasierten Zeitreihenprogrammen erweitert. Insgesamt wurden hierfür 108,332 Wasserproben aus 12 schiffsbasierten Programmen, die jeweils für eine andere Meeresumgebung und eine andere Programmstruktur repräsentativ sind, synthetisiert. Neben der Verbesserung der Daten-FAIRness, führten implementierte i) „Best-Practice Flags“, ii) Vergleiche mit GLODAP und iii) Berechnungen der Messqualitätskontinuität zu einer erhöhten Nützlichkeit der Zeitreihendaten. Darüber hinaus trug der Pilot dazu bei, die Bedeutung schiffsbasierter Zeitreihenprogramme im Rahmen des Übereinkommens der United Nations Decade of Ocean Science zu verdeutlichen. Der kürzlich erworbene Status des Marine Ecological Time Series Netzwerks als "Ocean Coordination Group potential new emerging network" untermauert dies. Alle Errungenschaften dieser Doktorarbeit haben den Wert von BGC EOV Datensynthese Produkten für das Ozeanbeobachtungssystem hervorgehoben und vertieft. Insbesondere wurden ihre einzigartigen Beiträge zu mehr FAIRness, Effizienz und Nutzbarkeit von BGC Daten ganz im Sinne von Aristoteles These "das Ganze ist mehr als die Summe seiner Teile" herausgestellt und implementiert. Dabei wurden die wichtigsten Elemente für einen kontinuierlichen und nachhaltigen Erfolg von BGC EOV Datensynthese Produkten aufgezeigt. Insgesamt wurden im Rahmen dieser Doktorarbeit wichtige Schritte in Richtung des übergeordneten Ziels, die BGC Datenlandschaft durch die Manifestierung und Etablierung von BGC EOV Datensynthese Produkte als integralen Bestandteil des Ozeanbeobachtungssystems zu verbessern, realisiert. Allerdings hat diese Doktorarbeit auch verdeutlicht, dass weitere Anstrengungen erforderlich sind, um dieses wichtige Ziel vollständig zu erreichen. VII Manuscript overview This thesis contains the following manuscripts which have been prepared in collaboration with other authors: Publication 1 Citation Lange, N., Tanhua, T., Pfeil, B., Bange, H.W., Lauvset, S.K., Grégoire, M., Bakker, D.C.E., Jones, S.D., Fiedler, B., O’Brien, K.M., Körtzinger, A., 2023. A status assessment of selected data synthesis products for ocean biogeochemistry. Front. Mar. Sci. 10, 1078908. https://doi.org/10.3389/fmars.2023.1078908 Status Published Nico Lange’s contribution Conceptualization of study, Coordination of co-author contributions, Development of readiness evaluation scheme, Evaluation of synthesis products, Writing and editing manuscript, Creation of tables and figures Publication 2 Citation Lauvset, S.K., Lange, N., Tanhua, T., Bittig, H.C., Olsen, A., Kozyr, A., Alin, S., Álvarez, M., Azetsu-Scott, K., Barbero, L., Becker, S., Brown, P.J., Carter, B.R., da Cunha, L.C., Feely, R.A., Hoppema, M., Humphreys, M.P., Ishii, M., Jeansson, E., Jiang, L.-Q., Jones, S.D., Lo Monaco, C., Murata, A., Müller, J.D., Pérez, F.F., Pfeil, B., Schirnick, C., Steinfeldt, R., Suzuki, T., Tilbrook, B., Ulfsbo, A., Velo, A., Woosley, R.J., Key, R.M., 2022. GLODAPv2.2022: the latest version of the global interior ocean biogeochemical data product. Earth Syst. Sci. Data 14, 5543–5572. https://doi.org/10.5194/essd-14-5543-2022 Status Published Nico Lange’s contribution Retrieval, formatting, and harmonization of data, QC preparations for reference group meetings, Compilation of product, Discussing QC results, Contributing to the manuscript, Co-Creation of tables and figures Publication 3 Citation Tanhua, T., Lauvset, S.K., Lange, N., Olsen, A., Álvarez, M., Diggs, S., Bittig, H.C., Brown, P.J., Carter, B.R., da Cunha, L.C., Feely, R.A., Hoppema, M., Ishii, M., Jeansson, E., Kozyr, A., Murata, A., Pérez, F.F., Pfeil, B., Schirnick, C., Steinfeldt, R., Telszewski, M., Tilbrook, B., Velo, A., Wanninkhof, R., Burger, E., O’Brien, K., Key, R.M., 2021. A vision for FAIR ocean data products. Commun Earth Environ 2, 136. https://doi.org/10.1038/s43247-021-00209-4 Status Published Nico Lange’s contribution CO-development of (parts of) vision, Co-Creation of figures, Contributing to project and manuscript VIII Publication 4 Citation Lange, N., Fiedler, B., Álvarez, M., Benoit-Cattin, A., Benway, H., Buttigieg, P.L., Coppola, L., Currie, K., Flecha, S., Honda, M., Huertas, I.E., Lauvset, S.K., MullerKarger, F., Körtzinger, A., O’Brien, K.M., Ólafsdóttir, S.R., Pacheco, F.C., RuedaRoa, D., Skjelvan I., Wakita, M., White, A., Tanhua, T., submitted. Synthesis Product for Ocean Time-Series (SPOTS) – A ship-based biogeochemical pilot. Status Submitted to Earth System Science Data Nico Lange’s contribution Conceptualization of time series product, Development of the BGC EOV metadata template and recommended QC Guidelines for BGC EOV time-series programs, Application of additional QC analysis on several original datasets, Execution of the assessments (method evaluation, minimum variability, GLODAP offset), Compilation of the SPOTS pilot and application of the TOATS notebook to the data, Generation of machine-readable metadata (ODIS), Writing and editing publication, Coordination of coauthor contributions, Creation of tables and figures IX List of Figures FIGURE 1: SCHEMATIC ILLUSTRATION OF THE CAPABILITY OF SCIENTISTS TO DISCOVER AND USE LITERATURE (DATA), FROM FERGUSON ET AL. (2014). .................................................................................................................................................... 4 FIGURE 2: THE VALUE CHAIN THAT CONNECTS IN SITU OCEANOGRAPHIC MEASUREMENTS OF CARBON CHEMISTRY TO CLIMATE NEGOTIATIONS. ADAPTED FROM BAKKER ET AL. (2023) BY TANHUA (2023). ................................................................. 6 FIGURE 3: SCHEMATIC ILLUSTRATION OF THESIS STRUCTURE AND ITS POSITION WITHIN THE LARGER BGC DATA LANDSCAPE AND THE OCEAN OBSERVING SYSTEM AT LARGE. ...................................................................................................................... 8 FIGURE 4: FRAMEWORK OF OCEAN OBSERVING (FOO) OCEAN OBSERVING PROCESS DIAGRAM (TANHUA ET AL., 2019A). ............. 9 FIGURE 5: FOO’S THREE MAIN “PHASES” OF AN OBSERVING SYSTEM: “CONCEPT”, “PILOT”, “MATURE” THEIR ATTRIBUTES, AND READINESS (LINDSTROM ET AL., 2012). ................................................................................................................. 10 FIGURE 6: CAPTURED PHENOMENA BY THE EIGHT BGC EOVS FOLLOWING THE BIOGEOCHEMICAL EXPERT PANEL OF GOOS (INTERNATIONAL OCEAN CARBON COORDINATION PROJECT). DARK SQUARES INDICATE THAT A PARTICULAR PHENOMENON IS CAPTURED BY THE RELATED EOV........................................................................................................................... 11 FIGURE 7: THE FAIR GUIDING PRINCIPLES FROM WILKINSON ET AL. (2016). ........................................................................ 13 FIGURE 8: SCHEMATIC ILLUSTRATION DISPLAYING THE SCHEMA.ORG STRUCTURE OF THE SELECTED PROPERTIES LISTED IN TABLE 4. ... 24 FIGURE 9: SCHEMATIC ILLUSTRATION OF CONNECTIONS BETWEEN THE SPOTS DATASET AND RELATED TIME-SERIES EVENTS USING THE SCHEMA.ORG VOCABULARY. GREEN RECTANGLES INDICATE METADATA OF THE TYPE DATASETS, GREY RECTANGLES INDICATE METADATA OF THE TYPE EVENTS, AND YELLOW RECTANGLES METADATA OF THE TYPE SUBEVENTS. THE ARROWS INDICATE THE RELATION (SCHMEA.ORG VOCABULARY) BETWEEN THE DIFFERENT TYPES OF METADATA. CLEAR PROVENANCE, AS WELL AS DIFFERENTIATION BETWEEN METHODS APPLIED DURING DIFFERENT YEARS, AND DIFFERENTIATION BETWEEN PRODUCT GENERATION (DATASET) AND DATA GENERATION (EVENT) IS ENABLED THROUGH THIS STRUCTURE. NOTE THAT THE ILLUSTRATION IS KEPT NON-SPECIFIC AND NON-EXHAUSTING....................................................................................... 25 FIGURE 10: THE APPLIED QC METHODS RELATED TO GLODAP OR SPOTS (OR BOTH). THE COLOR FURTHER INDICATES WHETHER A QC METHOD WAS USED TO INCREASE THE PRECISION (DARK BLUE) OR ACCURACY (LIGHT BLUE) OF THE DATA. FOR GLODAP FLAGGING (PRECISION) AND ADJUSTMENTS (ACCURACY) WERE APPLIED, WHILE FOR SPOTS ONLY FLAGGING WAS APPLIED. .... 27 FIGURE 11: SCHEMATIC ILLUSTRATION OF THE DIFFERENCES BETWEEN PRECISION AND ACCURACY (PORTABLE SPECTRAL SERVICES, 2023). ............................................................................................................................................................ 28 FIGURE 12: SCHEMATIC WORKFLOW OF REGULAR CROSSOVER QC. NOTE THAT THE CROSSOVER RADIUS AND DEPTH SURFACE ARE DEPENDENT ON USER INPUT. THE MOST COMMON ONES ARE SHOWN HERE: 2° AND SIGMA4, RESPECTIVELY. ....................... 31 FIGURE 13: RESULTS OF STATISTICAL OUTLIER QC FOR OXYGEN MEASUREMENTS AT CVOO DURING AUTUMN. EACH SUBPLOT DISPLAYS ONE LAYER/ SLAB, AND RED CIRCLES INDICATE OUTLIERS (TWO-SIGMA). FOR ILLUSTRATION PURPOSES ONLY 5 OF THE 17 “STANDARD CVOO LAYERS” ARE SHOWN. CONCENTRATIONS ARE GIVEN IN µMOL KG-1. ................................................. 33 FIGURE 14: TIMELINE OF SPOTS WITH SELECTED EVENTS HIGHLIGHTED. ............................................................................. 140 2 Introduction 3 1 Introduction Introducing this thesis, this Section first provides the motivation and sets it in context with the present research. The overarching goal, corresponding objectives, and research questions are expressed, and the structure of the thesis is given. Lastly, relevant existing frameworks and Concepts are explained. 1.1 Motivation and Scientific Background Covering approximately 71% of the Earth's surface, the ocean plays a fundamental role in the Earth's system and for our society. Amongst others, the ocean takes part in storing, transporting, and exchanging large amounts of heat, freshwater, and multiple greenhouse gases with the atmosphere (Rhein et al., 2013). One of the key aspects that governs the ocean's functionality is ocean biogeochemistry (Buesseler et al., 2020; Séférian et al., 2020). Understanding its importance is essential, as ocean biogeochemistry is closely linked to climate regulation, ocean acidification, ecosystem health, and sustainable resource management (Doney et al., 2020; IPCC 2021; IPCC, 2022; Jiang et al., 2023). Even more so in times of anthropogenic climate change, as the environmental status of the ocean and the associated services for society are at risk (Cooley et al., 2023). Most prominent, the ocean is the biggest reservoir of carbon in the earth system through both, the physical and biological carbon pump. It currently stores about 25% (Friedligstein et al., 2022) of the anthropogenic carbon emissions of carbon dioxide (CO2) including land-use change, and thereby buffers and mitigates the impacts of climate change (Jiang et al., 2019). The ocean is also a very important component of the global cycles of other greenhouse gases, such as halocarbons, and nitrous oxide (N2O) (Freing et al., 2012, Weber et al., 2019; Yang et al., 2020). In this context, observations of ocean biogeochemistry help to i) elucidate the factors controlling the uptake, storage, and changes of greenhouse gases in the ocean, and are thus essential to close their global budgets, ii) provide insights into the efficiency of the carbon pumps, iii) expand our knowledge regarding the potential feedbacks between oceanic cycling of greenhouse gases and climate change, and iv) enable estimates of ventilation and respiration rates of the ocean (Tanhua et al., 2013; Grégoire et al., 2021; Gruber et al., 2023). Hence, observations of ocean biogeochemistry are vital in supporting the development of informed strategies for climate change mitigation and adaptation. Another reason for observations of ocean biogeochemistry being crucial for society lies in the importance of biogeochemical (BGC) processes for the health and productivity of marine ecosystems. It is the absorption of atmospheric CO2 that leads to the most prominent effect related to ocean health: Ocean acidity increasing by about 30% since the preindustrial period, i.e. ocean acidification (Jiang et al., 2023). Ocean acidification poses large risks to marine organisms and ecosystems with significant effects on all levels of the trophic chain directly impacting future food security (Gattuso et al., 2013). Its importance is also reflected in the Sustainable Development Goal 14.3 “Minimize and address the impacts of ocean acidification, including through enhanced scientific cooperation at all levels”. Furthermore, ocean health is directly influenced by anthropogenic perturbations affecting the elemental cycles of nitrogen and phosphorus (fertilizer over-usage, runoff, atmospheric deposition) (Jickells et al., 2017; Yuan et al., 2018). These perturbations significantly impact ocean chemistry and increasingly lead to eutrophication and hypoxia, as well as harmful algal blooms. The cycles of the underlying essential elements (e.g. nitrogen) directly link to the productivity of phytoplankton, i.e. the growth of primary producers, and the intricate food web dynamics (Maúre et al., 2021). Hence, observations of ocean biogeochemistry are also crucial for gauging ocean health, specifically to better comprehend ecosystem resilience, predict shifts in species distributions, and develop evidence-based strategies to mitigate the impacts of human activities on marine ecosystems. All of which is needed for a sustainable marine resource management beneficial for society. Introduction 4 It is clear that planning and implementing all-encompassing BGC ocean observations are of paramount importance for society (Moltmann et al., 2019, Tanhua et al., 2019a). Regarding observations, it is very important to understand that ocean observations encompass far more than the act of sampling and measuring. In particular, ocean observation encompasses the management of resulting data (Lindstrom et al., 2012). Accordingly, BGC observations and BGC data are closely tight to each other, to the extent that marine BGC observations are only as impactful and useful, as the (resulting) data. This is also reflected in data being an integral part of the Ocean Decade Challenges, emphasizing that BGC data is just as important for society as the other parts of the BGC observation design. This recent shift of the ocean community to put a stronger focus on data also resulted from the observatory-based approach in marine science reaching a global scale following the examples set by the World Ocean Circulation Experiment and the Joint Global Ocean Flux Study (WOCE/JGOFS; JGOFS, 1990) during the 1990s (IOC, 1994), and from ongoing advancements in sensor technologies and autonomous platforms (Tanhua et al., 2019b). Hence, gaining an improved holistic understanding of the climate and the ocean's environmental status in this new “era of marine big data” (Abbott, 2013) strongly depends on the efficient usage of highly heterogeneous, high-volume data originating from a variety of complementing observing platforms. Unfortunately, it is still common practice in academic studies to fail making underlying data publicity available and reusable by either only publishing them as “Supplementary Material”, or not publishing the dataset at all (Starr et al., 2015). The reasons for the lack of data sharing are numerous, ranging from a competitive mindset to having no capacity for data management (Snowden et al., 2019). Especially the latter is very common despite the beneficial cost-reward ratio of good data management significantly improving the return on investment for any ocean observation (Tanhua et al., 2019b). Additionally, ocean observing campaigns are usually funded as research projects and often have very specific research targets. Consequently, for the BGC data that is made publicly available, a multitude of data centers are managing those data with varying extents to which data is further processed (Shepherd 2018; Miguez et al., 2019). Even though, more recently a stronger focus is put into following the guidelines of Findable, Accessible, Interoperable, and Reusable (FAIR) data (Wilkinson et al., 2016), data processing routines of many data centers are restricted to the archival and provision of data by assigning a persistent identifier. In combination with the increasing amount of data, data mining has become increasingly difficult and time-consuming leading to large amounts of “dark data” that remains unused (Figure 1) (Ferguson et al., 2014). Moreover, data users are required to manage a plethora of data versions, file formats, cope with duplicates, and differing levels of documentation and quality. Thus, many parts of the fragmented BGC data landscape are a limiting factor for the value of observations, and scientific progress overall. Figure 1: Schematic illustration of the capability of scientists to discover and use literature (data), from Ferguson et al. (2014). Introduction 5 Community-driven BGC synthesis efforts address these data challenges striving for stream-lined and user-orientated data access (e.g. Bakker et al., 2016; Kock and Bange, 2015; Bowie and Tagliabue, 2018; Lauvset et al., 2022). Focusing on specific societal issues related to marine BGC, BGC synthesis products apply customized techniques to combine datasets from multiple sources to form coherent and consistent data products, going far beyond simply merging individual data. Accordingly, the generation of a synthesis product covers the entire data value chain, from data acquisition to usage (Curry, 2016). More precisely, synthesis includes collection, harmonization, formatting, archival, and Quality Control (QC) of original (meta)data, as well as data integration, product generation, and provision. In particular, the development and application of advanced QC routines represents an incremental part of the synthesis data flow. This combination differentiates synthesis efforts from other data management efforts (Table 1). All data management efforts in turn are (ideally) linked to the International Oceanographic Data and Information Exchange (IODE), the program of the Intergovernmental Oceanographic Commission (IOC) "[...] responsible for enhancing marine research, exploitation and development, by facilitating the exchange of oceanographic data and information between participating Member States, and by meeting the needs of users for data and information products" (IODE, 2022). Table 1: Simplified overview of existing (BGC) Data Landscape. Adapted from Table 1 in Pouliquen et al., 2010. The star (*) denotes the larger category to which synthesis products belong. Who What From To Example Data Providers Collect, measure, and record the data Platform Media Principal Investigator (National) Data Centers Archive, QC, and distribute the (FAIR) (meta)data Media Repository Bundesamt für Seeschifffahrt und Hydrographie Data Assembly Centers (DAC) (Semantic) Harmonization of (FAIR) (meta)data and provision of “data portal” Repository Service Provider Global Argo DAC World Ocean Database Harmonization, enhancement of (FAIR) (meta)data Repository Service Provider Not applicable Thematic Assembly Centers* Harmonization, integration, and enhancement of (FAIR) (meta)data Repository Service Provider Copernicus Marine Service in-Situ TAC Service Providers Providing customized services from data analysis and manipulation (e.g. assessments) Service Provider End-Users Global Carbon Budget Since the first synthesis product release of the Global Ocean Data Analysis Project (GLODAP) in 2004 (Key et al., 2004), BGC data synthesis products continue to gain increasing popularity and demand. While GLODAP was initiated to enable the quantification of the anthropogenic ocean carbon sink (e.g. Sabine et al., 2004), and focuses on carbon-relevant data from repeated hydrography, other prominent BGC synthesis products have different foci. For example, the Surface Ocean CO₂ Atlas (SOCAT; Pfeil et al., 2013, Sabine et al., 2013) focuses on relevant data for the oceanic CO2 uptake and synthesizes insitu surface ocean fCO2 (CO2 fugacity) measurements from multiple platforms. Another example is Introduction 6 given by the MarinE MethanE and NiTtrous Oxide database (MEMENTO; Bange et al., 2009, Kock and Bange 2015) that focuses on hydrographic cruise data relevant for the oceanic distributions, and exchanges of N20 and methane (CH4). Multiple synthesis efforts established to be important components within the ocean observing system, connecting BGC observations with societal services as exemplarily illustrated for measurements of the carbon chemistry in Figure 2. Their impact on research is further evidenced by large numbers of citations, in some cases reaching thousands, as well as their usage in higher-level scientific assessments, e.g. the global carbon budget (Friedligstein et al., 2022). Figure 2: The value chain that connects in situ oceanographic measurements of carbon chemistry to climate negotiations. Adapted from Bakker et al. (2023) by Tanhua (2023). Nevertheless, there are still large uncertainties and many unresolved issues due to insufficient availability of BGC observations, e.g. increasing dissimilarities of the ocean carbon sink estimates (up to 1.1 GtC yr-1 in 2020) between ensemble means of global BGC ocean models and observation-based data products (Friedlingstein et al., 2022); disagreement on the strength and spatial distribution of deoxygenation between models and observation-based products (IPCC, 2019); vastly varying estimated contributions of N2O fluxes from oxygen minimum zones to the global ocean source (4% to 50%; IPCC, 2021). Hence, synthesis efforts are far from complete and the community still works towards having fully comparable, fully FAIR high-quality BGC data, supporting the ocean BGC observing system to reach its full potential and become truly fit-for-purpose. This goal entails: 1. Evaluating BGC Essential Ocean Variable (EOV) synthesis products in support of the elimination of any weaknesses 2. Continuously updating BGC EOV existing living synthesis products with new data 3. Developing and implementing improvements 4. Expanding the BGC EOV synthesis product landscape to previously overlooked observations Introduction 7 1.2 Thesis Objectives and Structure The above-identified objectives, motivated by the identified research gap, all relate to the overarching goal of this thesis: Improving the BGC data landscape through the manifestation of BGC EOV synthesis products as an integral part in the BGC ocean observing system. To better grasp this objective, first, two fundamentally important frameworks and/or concepts are introduced in Section 1.3, (i) the Framework of Ocean Observing (FOO), elaborating on the thesis’ focus on EOVs; and (ii) the concept of FAIR data. Subsequently, Section 2 describes the methodologies used to develop and generate the BGC data synthesis products GLODAP and the Synthesis Product for Ocean Time-Series (SPOTS). Eventually, the publications that relate to the individual goals are included as Sections 3 to 6. Thereby, the chronological order of the individual goals is followed: O1. Evaluating BGC EOV synthesis products in support of the elimination of any weaknesses The first included publication (Section 3), describes the current BGC synthesis product landscape by means of four complementing synthesis products that focus on different EOVs and/or observing platforms. To enable an objective assessment of the maturity of synthesis products, and to guide their further development, a scoring scheme that is based upon FOO is introduced. By exemplarily applying it to the four selected synthesis products, their strengths, weaknesses, and potential improvements are discussed, as is the scoring scheme itself. O2. Continuously updating existing living BGC EOV synthesis products with new data To deepen the understanding of all processes involved in generating and updating a living BGC data synthesis product, the second publication (Section 4) describes GLODAPv2.2022 in detail. This publication is representative for all the work involved in synthesizing global BGC observations from repeated hydrography that mounted in the fourth consecutive annual update of GLODAPv2. O3. Developing and implementing improvements The third publication (Section 5) further emphasizes the importance of GLODAP for climate research. It outlines the vision for FAIR ocean data products and in particular GLODAP. For GLODAP, envisioned improvements for its data handling are depicted accordingly. O4. Expanding the BGC EOV synthesis product landscape to previously overlooked observations The last publication (Section 6) included in this thesis describes the pilot of a new BGC synthesis data product: SPOTS. The generation of this pilot leveraged the knowledge gained from existing synthesis efforts and the results given in the three previous publications. Targeting the gap in temporal resolution of the current BGC data synthesis product landscape, SPOTS complements the existing synthesis products by synthesizing BGC data from fixed location time-series. Section 7 synthesis the contribution of the thesis towards these objectives and discusses their limitation. Eventually, the overall conclusion and outlook of the thesis (Section 8) are guided by two research questions that embody the essence of the overarching goal and the related objectives: • Why are BGC EOV data synthesis products so important for the BGC data landscape, more specifically if all data would be FAIR, will synthesis products become redundant? • What are the key aspects for the sustainable success of BGC EOV data synthesis products? Introduction 8 Figure 3 illustrates the thesis structure, setting the thesis, its overarching goals, the related objectives and the guiding research questions in context with the BGC data landscape and ocean observing system at large. Figure 3: Schematic illustration of thesis structure and its position within the larger BGC data landscape and the ocean observing system at large. Introduction 9 1.3 Frameworks and Concepts 1.3.1 Framework of Ocean Observing (FOO) The Global Ocean Observing System (GOOS, Moltmann et al., 2019), led by IOC of UNESCO, leads and supports the ocean observing community to build an integrated and sustained ocean observing system aiming to deliver maximum impact for society. The GOOS strategy follows the Framework for Ocean Observing (FOO) systems engineering concept (Lindstrom et al., 2012) that encompasses the entire ocean observation value chain 1 , “[…] a chain of processes addressing “why to observe?” (requirement setting process), “what to observe?” (scoping of observational foci), “how to observe?” (coordination of observing elements), and “how to integrate, use and disseminate observational outcomes and understand their impacts?” (Pearlman et al., 2019, p. 2). FOO divides this ocean observing value chain, i.e. the system’s inputs, processes, and outputs, into: “Requirements”, “Observations”, and “Data and Information”. The “Requirements” are the oceanographic information needed to address specific societal issues, the “Observations” the technology and ocean observing networks (platforms) used to collect the required data, and the “Data and Information products 2 ” the resultant data and services (Figure 4). To facilitate the concept, FOO makes use of EOVs, inspired by the success of the Essential Climate Variables. The focus of FOO on EOVs enables to set “essential” requirements for sustained ocean observations. Amongst others, this system approach promotes data standards, broad accessibility, as well as free and open exchange of data and products, following the leading principle of “measure once - use many times” (Lindstrom et al., 2012, p. 5). Figure 4: Framework of Ocean Observing (FOO) Ocean Observing Process Diagram (Tanhua et al., 2019a). The guiding societal issues/drivers for marine biogeochemistry are (Telszewski et al., 2018): 1. The role of ocean biogeochemistry in climate 2. Human impacts on ocean biogeochemistry 3. Ocean ecosystem health 1 a term broadly defined as a set of value-adding activities that one or more communities perform in creating and distributing goods and services (Longhorn and Blakemore, 2008) 2 with its sub-categories: “Oversight and Coordination”, “Data Quality Control”, “Near Real-Time Data Stream delivery”, “Data Repository”, and “Data Products” Introduction 10 The related scientific research questions and applications are: 1.1. How is the ocean carbon content changing? 1.2. How does the ocean influence cycles of non-CO2 greenhouse gases? 2.1. How large are the ocean’s dead zones and how fast are they growing? 2.2. What are the rates and impacts of ocean acidification? 3.1. Is the biomass (production) of the ocean changing? 3.2. How does eutrophication and pollution impact the ocean productivity and water quality? To evaluate the ocean observing system regarding particular societal issues and oceanographic phenomena 3 , e.g. ocean acidification, the FOO concept adopted the technical readiness level, a scheme developed by NASA (National Aeronautics and Space Administration) (Sadin et al., 1989). This holistic approach enables the classification (concept, pilot, mature) of an ocean observing system activity in terms of feasibility, capacity, and impact (Figure 5). Key characteristics of a fit-for-purpose ocean observing system are that i) the resultant “data and information” address the societal issues that determined the requirements, and ii) a feedback loop that links the system’s outcomes to its inputs is established (Figure 4). Hence, for observational networks to become a sustained part of GOOS, networks must first “[…] mature the associated requirements for acceptance, mature their measurement technology for inclusion, and mature their data and information products for appropriate accessibility and application to a range of scientific and societal issues” (Lindstrom et al., 2012). Figure 5: FOO’s three main “phases” of an observing system: “Concept”, “Pilot”, “Mature” their attributes, and readiness (Lindstrom et al., 2012). 3 Defined as: “A phenomenon is an observed process, event, or property, with characteristic spatial and timescale(s), measured or derived from one or a combination of EOVs, and needed to answer at least one of the scientific questions asked in order to address relevant societal need” (Telszewski et al., 2018, p. 138). Introduction 11 Essential Ocean Variables (EOV) The GOOS Expert Panels identify EOVs according to a scoring system based upon the three criteria “Relevance”, “Feasibility”, and “Cost-effectiveness” (Telszewski et al., 2018). • The relevance indicates how effectively a variable addresses the overall GOOS Themes – Climate, Operational Services, and Ocean Health. • Feasibility corresponds to whether deriving the variable, i.e. the actual act of observing and analyzing, has proven to be technically feasible on a global scale. • Cost-effectiveness is defined as the data generation and archiving being affordable (Sloyan et al., 2019). By implementing a scoring system that addresses the above-described criteria and utilizes the outlined marine biogeochemistry research questions, eight BGC EOVs have been defined. The eight BGC EOVs and related phenomena are displayed in Figure 6. Figure 6: Captured phenomena by the eight BGC EOVs following the biogeochemical expert panel of GOOS (International Ocean Carbon Coordination Project). Dark squares indicate that a particular phenomenon is captured by the related EOV. A brief introduction into each BGC EOV following Telszewski et al. (2018) is given in the following. Oxygen Dissolved oxygen (O2) in the ocean is the result of a balance between oxygen supply and consumption. Oxygen supply can be attributed to ocean circulation and ventilation, whereas consumption relates to the process of respiration. Changes in either process strongly influence the absolute amount of oxygen at a given location. The observation of O2 plays a critical role in understanding the, for the most part, strong decline in O2 in the ocean in recent decades. As recent oxygen trends influence our understanding of anthropogenic climate change, O2 should further be regarded as a crucial indicator of climate change. Additionally, O2 observations are vital to interpret water mass ventilation rates; (ii) for calculations of export production and (iii) to interpret repeat hydrographic data, which are of great relevance to document the anthropogenic CO2 inventory in the ocean. 18 Methods 19 2 Methods In this Section, the methods applied in this thesis are briefly presented. A particular focus is given to the applied QC methods. Several methods already existed and are well-established within the BGC community and applied accordingly. Table 2 provides an overview. Table 2: A list of methods that have been used during the course of the thesis, including their main application. Whether the method has been developed as part of this thesis is indicated (Novel), as well. Method Application in Thesis Novel / Pre-existing Engagement with Community Participation in GLODAP Reference Group Establishing SPOTS Core Group Not applicable FOO Readiness Evaluation Scheme Evaluation of: GLODAP, SOCAT, MEMENTO, GO2DAT Novel Data Retrieval and PreProcessing Harmonization of Original Data Pre-existing (Structured) Metadata Templates Collection of Metadata Schema.org compliant SPOTS Metadata Novel AtlantOS QC QC of Oxygen, Inorganic Carbon, Nutrients, Transient Tracer Novel (Consulted) Saturation Plots QC of transient tracer Pre-existing Tracer Ratios QC of transient tracer Pre-existing Crossover Analysis QC of GLODAP’s “core variables” (except tracer) Comparison of SPOTS and GLODAP Pre-existing (Adapted) Comparison to Neural Networks QC of Oxygen, Inorganic Carbon, Nutrients, Transient Tracer Pre-existing Multi-Linear Regressions QC of GLODAP’s “core variables” (except tracer) Pre-existing Statistical Outlier Test(s) QC of ship-based BGC time-series Novel Minimum Variability QC of ship-based BGC time-series Novel Make Ocean Routine Merging of GLODAP Novel Methods 20 2.1 Engaging with the Community Synthesis products are community-driven efforts. Accordingly, the process of updating existing products, as well as developing new ones, relies heavily on engagement and consensus building with and within the community (e.g. Lauvset et al., 2022). 2.1.1 GLODAP – Reference Group The regular (reference group) meetings of GLODAP represent a well-established mechanism for the annual updates of its product. These meetings are mainly related to the organization of a new update, as well as the evaluation of QC results and the determination of adjustment. However, the meetings are also very valuable in terms of: • sharing updates in data handling (e.g. new persistent cruise identifier from OceanOPS 8 ) • obtaining feedback for newly implemented technologies (e.g. data visualization using the Digital Earth Viewer, Python-based crossover tool) • discussing known issues of applied methodologies (e.g. using inorganic carbon interconsistency in the QC) • planning and improving outreach (e.g. defining submission requirements) • identifying and prioritizing problems (e.g. uncertainty calculations) • planning GLODAPv3 Above all, these regular meetings ensure maintaining a strong core group of scientists with complementing expertise (specific region, variable, basin) around GLODAP and ensure that connections to other relevant and closely related communities are established. Examples of the latter include the above-mentioned collaboration with Digital Earth that resulted in the Python-based “Make Ocean” routine (Section 2.6) and the visualization of GLODAP data in the Digital Earth Viewer but also include ongoing communication and data sharing with other data synthesis products (Coastal Ocean Data Analysis Product in North America, CODAP-NA, Jiang et al., 2021; CARbon IN the MEDiterranean Sea, CARIMED, Sanleón-Bartolomé, 2017; GEOTRACES, Schlitzer et al., 2017, Bowie and Tagliabue, 2018) as well as maintaining a strong connection to the International Ocean Carbon Coordination Project (IOCCP). Importantly, the reference group members are routinely exchanged so that new perspectives and insights are continuously incorporated into GLDOAP’s development. In addition to regular reference group meetings, it is also very important to mention the regular contact with data providers, during all stages of the GLODAP workflow (data retrieval, formatting, QC, archival). 2.1.2 SPOTS – Core Group Using the approach of GLODAP for community engagement as a role model, establishing a strong community around SPOTS is vital. To this date, the generation of SPOTS included in total four IOCCPendorsed workshops; the Earth Cube Workshop (Benway et al., 2020), the Marine Ecological Time Series Research Coordination Network (METS-RCN 9 ) informatics meeting, and two SPOTS workshops. The latter two exclusively focused on the development of SPOTS itself. The first SPOTS workshop that was held virtually in November 2020 resulted in the establishment of a general consensus towards SPOTS, including the onset of a concept note that clearly outlines the purpose, benefits, methods, timeline, and participants of its pilot. Moreover, four working groups (Concept/Head; Commonality of methods; Data handling; Data policy) were formed, which, during the course of the following two years, developed the underlying structure and methods of SPOTS (Section 6). During the second SPOTS workshop (virtual, November 2022), the results of the working groups were discussed. Besides, the workshop established consensus on applied “Best-Practice” requirements (Section 6) and led to an 8 https://www.ocean-ops.org/board 9 https://www2.whoi.edu/site/mets-rcn/ Methods 21 accompanying manuscript. The two SPOTS-focused workshops, brought together BGC time-series experts from 15 time-series programs around the globe, as well as numerous experts from other synthesis activities (International Group for Marine Ecological Time Series – IGMETS; O’Brien et al., 2017, GLODAP, SOCAT), and related efforts (Global Ocean Acidification Observing Network - GOA-ON; Newton et al., 2019, Integrated Carbon Observation System – ICOS; Steinhoff et al., 2019). In particular, a strong collaboration with the Ocean Carbon and Biogeochemistry 10 program-led METS-RCN network was established. Through this community engagement synergies were created, and collaborations were fostered (Section 2.1.2). Therefore, at this stage, the SPOTS pilot incarnates one of METS-RCN use-cases i) increasing the outreach of SPOTS (e.g. Ocean Science Townhall), ii) establishing a close cooperation with IODE-led ODIS, resulting in the development of structured metadata (Section 2.4.3), and iii) resulting in the selection of the Biological & Chemical Oceanography Data Management Office 11 as datacenter for SPOTS. 2.2 FOO Readiness Evaluation Scheme for BGC Data Synthesis Products The maturity assessment of BGC EOV data synthesis is carried out using the FOO readiness level scheme for "Data Management and Information Products”, which has adopted the technical readiness level concept developed by NASA (National Aeronautics and Space Administration) (Sadin et al., 1989). Following FOO, the readiness levels are categorized into "Concept," "Pilot," and "Mature" (Lindstrom et al., 2012). Since the corresponding nine FOO readiness levels are quite broad, a customized criteria catalog was developed that refines the existing FOO criteria for each level. This catalog serves as a basis for assessing typical characteristics of data products using a level-by-level (“equal-weighted”) scoring system, with full compliance to the criteria resulting in a 100% score for a given readiness level. A score of 80% or higher is considered a "pass". Although the readiness levels follow a hierarchical structure, a data product can meet some requirements of higher levels before fully complying with all lower levels. To align with the FAIR guidelines, which strongly influence the maturity of a BGC EOV synthesis data product, these guidelines were incorporated into the criteria at multiple readiness levels to varying degrees. Additionally, the criteria catalog considers the degree of being "fit-for-purpose", which is an important requirement within the ocean observing value chain, and the degree of the data flow’s automation, at multiple levels. Given the diverse nature of BGC EOV data, the criteria are intentionally kept as generic as possible. Section 3 provides further insights into the details of this evaluation scheme by exemplarily depicting the criteria and scoring system for readiness level five in detail, and by its application to four selected BGC EOV data synthesis products. 10 https://www.us-ocb.org/ 11 https://www.bco-dmo.org/ Methods 22 2.3 Data Retrieval and Pre-Processing Before data can be retrieved, first awareness about (new) data must be gained. Therefore, GLODAP annually calls for submission of new cruise data. These calls are disseminated through multiple channels (e.g. IOCCP). Besides, through closely collaborating with OCADS and the CLIVAR and Carbon Hydrographic Data Office 12 , the GLODAP team is made aware of new (relevant) additions to their data holdings constantly. Accordingly, for the retrieval of “original data” GLODAP directly works with the data generators (i.e. principal investigators), which provide the “bottle data” directly (via email), or the “bottle data” are retrieved from national data centers or larger repositories (e.g. CCHDO). For the SPOTS pilot on the other hand, data retrieval exclusively works through direct communication with principal investigators of the time-series programs. Given the number of data generators, this is not sustainable nor possible for GLODAP. Once the data is retrieved, each original dataset of GLODAP and SPOTS is formatted into exchange format (Barna et al., 2023). This “harmonization” mainly entails: • Mapping to WOCE ontology, i.e. WOCE variable names (not always trivial, e.g. P(O)M) • Mapping to WOCE flagging scheme (Table 3 exemplarily shows a mapping between three frequently used flagging schemes) • Time, date, and unit conversions (Section 4) • Creating missing required variables (e.g. missing cast number or bottle numbers) • Assigning a cruise-identifier, i.e. an expocode for GLODAP (4-character ship code followed by the date given in yyyymmdd format) • Applying sanity check, i.e. realistic value check (e.g. -100° C water temperature) • Inspecting (linear regression) CTDand bottle fit for salinity and oxygen (Section 4) • Creating a comma separated value file with defined numbers of digits for each variable Eventually, the harmonized original dataset included in GLODAP is archived at OCADS. Further changes resulting from the additional 1st QC (Section 2.5) are forwarded to both, OCADS, and the corresponding national data center. For SPOTS, additionally individual (harmonized) datasets belonging to the same time-series program, but covering different time spans, are merged into one dataset. Table 3: Meaning of primary quality flags in three different often used semantics. Note that other semantics with different flagging schemes exist, see Schlitzer 2023. Flag Ocean Data Viewer (ODV; Schlitzer, 2023) WOCE (Barna et al., 2023) SeaDataNet (L2013) 0 Acceptable Interpolated No quality control 1 Not evaluated / Not calibrated (CTD) Not evaluated / Not calibrated (CTD) Good value 2 Not used Acceptable Probably good value 3 Not used Questionable Probably bad value 4 Questionable Known bad Bad value 5 Not used Not reported Changed value 6 Not used Median of replicates Value below detection 7 Not used Manual chromatographic peak measurement Value in excess 8 Known bad Irregular digital chromatographic peak integration Interpolated value 9 Not used Not measured Missing value 12 https://cchdo.ucsd.edu/ 13 http://seadatanet.maris2.nl/v_bodc_vocab_v2/browse.asp?order=conceptid&formname=search&screen=0&lib= l20 Methods 23 2.4 (Structured) Metadata Metadata is a highly important aspect of datasets regarding FAIR data. However, often an insufficient amount of metadata are provided to data managers and synthesis efforts alike. To enable easier metadata provision, templates to collect rich metadata have been designed and used. Furthermore, not all metadata are machine-readable to the same degree. Consequently, even very rich metadata, when provided in an inferior format, can limit the FAIRness of the data that is described. Using structured metadata, leveraging from existing and well-established structured metadata practices and web standards solves this issue. Structured metadata can be understood as a standardized concept that implements a well-defined metadata scheme and a common vocabulary. Specifically, structured metadata enable efforts like Google’s Data Set Search 14 and GeoCODES 15 to crawl and index the corresponding data into a knowledge graph, enabling the discovery of related datasets, i.e. enabling the discovery and identification of datasets that are merged into synthesis products. 2.4.1 Metadata Collection The collection of metadata is carried out by collaborating with data providers. During the submission process of the synthesis products, data providers are asked to provide additional metadata, ideally by filling out customized metadata templates. Two excel-templates were employed for this thesis: • GLODAP: The Ocean Carbon and Acidification Data System (OCADS) ocean carbon data submission form 16 • SPOTS: SPOTS’ customized metadata template for BGC EOV ship-based time-series programs Both of these templates are developed with ship-based data in mind. If filled out correctly, the templates provide general information about the observation program (e.g. principal investigator, location, and time), and measured variables (e.g. sampling method, and units), as well as about more detailed information (e.g. analytical methods, associated instrumentation, calibration, and QC procedures, and standards). While the OCADS template focuses on the variables of the inorganic carbon system (pH, TA, DIC, pCO2), the SPOTS template additionally focuses on O2, nutrients, DOC, and POM. The focused SPOTS working group “Commonality of methods” (Section 2.1.2) developed the latter template. Using the OCADS template as basis, it implements additional information from the Bermuda Time-Series Workshop report (Lorenzoni and Benway, 2013), GO-SHIP manuals (Langdon et al., 2010; Becker et al., 2019), and results from the Scientific Committee on Research Working Group 147 “Towards comparability of global oceanic nutrient data” (Bakker et al., 2016a; Bakker et al., 2016b; Aoyama et al., 2015). However, note that for GLODAP, many datasets were not directly submitted to its data management team, but rather retrieved from other data centers (e.g. through CCHDO, Section 2.4.1). Accordingly, often submitted metadata do not meet the rich level of the OCADS template. Moreover, presently, only the responsible principal investigator, chief scientist, vessel name and, cruise-date are required metadata for GLODAP. These information are usually attached as header-lines to the data file itself. 2.4.2 Schema.org To derive structured metadata from the above described user-friendly templates, through the collaboration with “Science on Schema” (Shepard et al., 2022) and ODIS, in the course of this thesis, structured metadata (templates) for BGC EOV ship-based time-series datasets were developed for 14 https://datasetsearch.research.google.com/ 15 https://geocodes.earthcube.org/ 16 https://www.ncei.noaa.gov/access/ocean-carbon-acidification-datasystem/support/SubmissionForm_carbon_v1.xlsx Methods 24 SPOTS. For the structure metadata, one of the most common vocabularies in structured metadata, Schema.org, “[…] a collaborative, community activity with a mission to create, maintain, and promote schemas for structured data on the Internet, on web pages, in email messages, and beyond” (Schema.org, 2023), has been applied. Following the recommendations of Google-Search (Google Search Central, 2023), the encoding format used is json-ld. Note that OCADS uses MD Metadata ISO 19115 (ISO, 2014) for GLODAP’s original datasets. Explaining Schema.org entirely is beyond the thesis, however, a few important aspects of the type “Dataset” are explained here to demonstrate the vocabulary and its structure. In Schema.org the type “Dataset” has a fixed list of “Properties” that can be used to describe the dataset, e.g. “Keywords”. These properties in turn are expected to be given as a certain type, e.g. “Text”. Often sub-properties are implemented as well to enable a more sophisticated attribute description, e.g. “DefinedTerm”. For illustration purposes, a non-exhaustive set of example properties is shown in Table 4 and the corresponding structures/connection are displayed in Figure 8. Table 4: Selected properties, their description, and recommended (expected) type for “Datasets” in Schema.org. Property Description Recommended type Name A descriptive name of a dataset text Description A short summary describing a dataset text Url Location of a page describing the dataset url SameAs Other URLs that can be used to access the dataset page url License A license document that applies to this content url isAccessibleForFree Specifying if the dataset is accessible for free Boolean Keywords Keywords summarizing the dataset Defined Term Identifier An identifier for the dataset, such as a DOI PropertyValue VariableMeasured What does the dataset measure? PropertyValue Figure 8: Schematic illustration displaying the Schema.org structure of the selected properties listed in Table 4. Note that to describe the metadata of other schema.org types, e.g. “Events”, another fixed list of expected properties must be used to describe its attributes. The principle remains the same. Methods 25 2.4.3 SPOTS metadata Regarding SPOTS, two different schema.org metadata types are relevant: Datasets and Events. The implementation of both types is beneficial as it allows one to differentiate between methodologies applied to create the dataset and multiple methodologies applied during the actual sample analyzes. This is of particular relevance as the latter commonly vary within one dataset. Hence, the “Dataset” metadata are purely linked to the generated datasets of a time-series or SPOTS itself, while the Event” metadata describe the time-series program with related (sub-)events that describe station visits (e.g. a particular year). Here, information on specific types of measurements (e.g. nutrient sample analysis) that result in the data of datasets are given. Accordingly, clear connections between all related files, as schematically illustrated in Figure 9, are provided. Note that one can focus on any particular event or dataset, i.e. multiple point-of-views are possible. This setup and structure also enable the application of the scheme to all types of time-series measurements, e.g. net measurements, and the degree of granularity is very flexible, enabling a customized approach for each time-series program. However, in some instances, the rather rigid Schema.org construct was extended when necessary (e.g. accepting “DefinedTerm” for the property “variableMeasured”). In addition to the attributes listed in Table 4, the generated metadata files for each time-series program and SPOTS itself provide information on (following the FAIR principle, Section 1.3.2): • Data providersand generators • Funding • Date • Location • Measurement techniques • Cruise reports Figure 9: Schematic illustration of connections between the SPOTS Dataset and related time-series Events using the Schema.org vocabulary. Green rectangles indicate metadata of the type Datasets, Grey rectangles indicate metadata of the type Events, and yellow rectangles metadata of the type subEvents. The arrows indicate the relation (schmea.org vocabulary) between the different types of metadata. Clear provenance, as well as differentiation between methods applied during different years, and differentiation between product generation (Dataset) and data generation (Event) is enabled through this structure. Note that the illustration is kept non-specific and non-exhausting. Methods 26 All structured metadata files are hosted by the METS-RCN GitHub repository 17 , which has been specifically set up for this purpose. From the repository, ODIS obtains access to the metadata and is presently working on a dedicated “Time-Series” source type in its catalog 18 which is based upon these files. 17 https://github.com/earthcube/METS-RCN 18 https://catalogue.odis.org/ Methods 27 2.5 Data Quality Control (QC) Striving towards fully comparable data is a key mission for synthesis products. However, data from multiple heterogenic sources often show large artificial inconsistencies, especially historical data. To combat this and to provide comparable data, a fundamental part of the generation of a synthesis product is the underlying, external QC of the included data. In the following, the concept of QC is introduced and subsequently, the applied QC routines are briefly presented. Methods that are already explained in great detail in one of the included manuscripts are only briefly described, these are: “Neural network comparisons for GLODAP” (Section 4), the “Best-Practice Assessment for SPOTS” (Section 6), and “Minimum variability determination for SPOTS” (Section 6). Figure 10 summarizes which QC methods were applied for which synthesis product. Figure 10: The applied QC methods related to GLODAP or SPOTS (or both). The color further indicates whether a QC method was used to increase the precision (dark blue) or accuracy (light blue) of the data. For GLODAP flagging (precision) and adjustments (accuracy) were applied, while for SPOTS only flagging was applied. 2.5.1 Quality Control (QC), Quality Assurance (QA), and Best-Practices (BP) QC, Quality Assurance (QA), and BPs are often used synonymously. Even though these checks and guides are closely related and all aim at increasing the data quality, it is important to separate them. QA relates to processes that are employed to support the generation of high-quality data during the sampling and analyzing procedures (Bushnell et al., 2019). QC in turn relates to all checks of data quality applied post data generation. BPs are “guides” describing community-accepted methodologies in detail (e.g. Dickson et al., 2007) that have proven to produce the most precise and accurate results relative to other methodologies with the same objective (Pearlman et al., 2019). This distinction is particularly important given that some types of QC rely upon QA “results”, such as comparisons to reference materials, precision estimates from duplicate measurements, or outcomes from interlaboratory calibration exercises (e.g. QUASIMEME, Wells et al., 1997). Further, some types of QC rely upon known BPs to evaluate applied methodologies. During the synthesis of data, the data generation process of the original data is already finished. Moreover, often an internal, i.e. by the data generator, QC of the data has already been applied before data is acquired by synthesis efforts. Hence, the applied checks for synthesis products are restricted to the external QC of the data, but, if necessary, make use of available QA results, and provided BP information. Methods 34 2.5.10 Minimum Variability To assess the consistency of measurement quality within a time-series program, the minimum variability estimation routine was developed. First, for each standard depth layer (+/- 100 dbar) of the time-series program, the coefficient of variation 19 for O2 is calculated. Subsequently, the layer with the least oxygen variability, i.e. the layer with the lowest coefficient of variation in time for O2, is determined. Eventually, for all other BGC EOVs, the minimum variability is calculated on the detected layer with the least oxygen variability (+/- 100 dbar) by means of the coefficient of variation. The choice of using the layer that is closest to an oxygen equilibrium for all calculations relates to oxygen concentrations being linked to (amongst others) variation in ventilation, water mass changes, or changes in consumption and production by biological activity (Sarmiento and Gruber, 2006; Keeling et al., 2010; Stramma and Schmidtko, 2019). This layer thus correlates to (relatively) stable conditions, and the natural variability of other BGC variables is expected to be rather low. For locations that are not characterized by large natural variability (as indicated by high coefficients of variation for oxygen and salinity), a low minimum variability indicates a consistent level of data quality throughout the measurement period for the analyzed variable. However, for locations with high natural variability, high minimum variability estimates do not necessarily relate to inconsistencies in measurement quality. Section 6 gives further insights into this method. 19 Coefficient of Variation = (Standard Deviation / Mean) * 100 Methods 35 2.6 Data Merging Routine: GLODAP’s “Make Ocean” To enable the reproducibility of GLODAP, a Python-based Juypter Notebook that generates the final global and regional GLODAP synthesis products was developed and applied. The notebook also generates consistent unadjusted and adjusted individual cruise files of newly added cruise data. The implemented merging routine follows a strict order of clearly defined steps (based upon Key et al., 2004 and Olsen et al., 2016) and uses the Python libraries numpy, pandas, scipy, shapely, seawater, oct2py, and netCDF4. It is designed for cruise bottle data in exchange format (WOCE semantics, Barna et al., 2023), the officially required submission format of GLODAP. The program processes cruise by cruise and merges the individual consistent cruise datasets in a final step. More details on the specifics of the Python script can be found at https://git.geomar.de/patrick-michaelis/python-for-glodap. The first processing functions import, select, and re-arrange the data and fix small issues of the cruise file. Amongst others, these functions exclude data with WOCE flags 3,4,5, and 8 (Table 3), as well as data without temperature or pressure data, by setting corresponding values to -999 and their WOCE flags to 9. Cruise numbers are also assigned that enable the differentiation between all GLODAPv2 updates. Non-trivial calculations are restricted to missing depths (bottom), and nitrate. Missing bottom depths are assigned by either the maximum sample depth or the extracted bottom depth from ETOPO1 (Amante and Eakins, 2009) – the larger value is used. Missing pressure or depth values are estimated following UNESCO (1981). Lastly, whenever possible, the division of nitrate plus nitrite values (NO2+NO3) into nitrate (NO3) and nitrite (NO2) is executed. If only nitrate plus nitrite values are given, these are renamed to nitrate. The next set of functions start to alter the original data more drastically. Bottle salinityand oxygen (SALNTY, OXYGEN) data is merged with their sensor counterparts (CTDSAL, CTDOXY), following the action (Section 4) implied in the GLODAP adjustment table 20 . Further, the merged salinity, and oxygen data, as well as nutrient data are vertically interpolated to fill data gaps using a quasi-Hermetian piecewise polynomial. Values are only interpolated if the maximum vertical data separation distances (Table 4 in Key et al., 2010) are met. The corresponding WOCE flags are set to 0. The next functions incorporate aspects of GLODAP that result in its high consistency, and usability. Here, the first three functions must be executed first and in succession: 1. The adjustments resulting from the 2nd QC are implemented for the core variables of GLODAP except for pH. Note that all data that have passed the 2nd QC (no adjustment or adjusted) are indicated accordingly through additional 2nd QC flags (Section 4). 2. Whenever necessary, pH is converted to the total scale at 25°C and (pCO2) fCO2 to fCO2 at 20°C and 0 dbar. For the conversion, CO2SYS is employed with TA used as the second inorganic carbon sub-parameter. Missing TA values are approximated as 67 times salinity. The carbonate dissociation constants of Lueker et al. (2000), the bisulfate dissociation constant of Dickson (1990), and the borate-to-salinity ratio of Uppström (1974) are used. Once pH is converted, adjustments to pH are applied as well. 3. If at least two sup-parameters of the inorganic carbon system are available, the missing inorganic carbon sub-parameters are calculated using CO2SYS. Here, some rules were established: • Adjustments and scale conversions have to be applied first • The same constants as for the conversions (pH, fCO2) are used • DIC, TA is the preferred pair to calculate pH and fCO2 20 https://glodap.geomar.de Methods 36 • If either DIC or TA is missing and both pH and fCO2 data existed, pH is preferred • If less than a third of the total number of values is measured, then all values are replaced by calculated values (only for DIC, TA, and pH) The so-calculated inorganic carbon sub-parameters are indicated by a WOCE flag 0 4. Values for potential temperature; potential densities referenced to 0; 1,000; 2,000; 3,000; and 4,000 dbar; neutral density; apparent oxygen utilization are calculated using Fofonoff (1977), Bryden (1973), Sérazin (2011), and Garcia and Gordon (1992) 5. Partial pressures for CFC-11, CFC-12, CFC-113, CCl4, and SF6 are calculated using the solubilities by Warner and Weiss (1985), Bu and Warner (1995), Bullister and Wisegarver (1998), and Bullister et al. (2002) 6. pH in-situ values are obtained following the same method as in 2 The last functions, processing the individual cruise dataset, create the columns DOI, and region, as well as sort the entire dataset according to (hierarchical) station number, pressure, and bottle number. Lastly, an adjusted cruise file that is consistent with the format and semantics of the existing GLODAP updates is saved as a comma-separated value file. Eventually, all created individual consistent adjusted cruise files are appended to the previous GLODAP update to create the GLODAP master file. The regional files are split up using the region information stored in the online adjustment table. 37 3 A status assessment of selected data synthesis products for ocean biogeochemistry 3 A status assessment of selected data synthesis products for ocean biogeochemistry 38 A status assessment of selected data synthesis products for ocean biogeochemistry Nico Lange 1 , Toste Tanhua 1 *, Benjamin Pfeil 2 , Hermann W. Bange 1 , Siv K. Lauvset 3 , Marilaure Gre ´goire 4 , Dorothee C. E. Bakker 5 , Steve D. Jones 2 , Björn Fiedler 1 , Kevin M. O’Brien 6,7 and Arne Körtzinger 1,8 1 GEOMAR, Helmholtz Centre for Ocean Research Kiel, Kiel, Germany, 2 Geophysical Institute, University of Bergen and Bjerknes Centre for Climate Research, Bergen, Norway, 3 NORCE Norwegian Research Centre, Bjerknes Centre for Climate Research, Bergen, Norway, 4 Department of Astrophysics, Geophysics and Oceanography, MAST (Modelling for Aquatic Systems) - FOCUS (Freshwater and OCeanic science Unit of reSearch), University of Liège, Liège, Belgium, 5 Centre for Ocean and Atmospheric Sciences, School of Environmental Sciences, University of East Anglia, Norwich, United Kingdom, 6 Cooperative Institute for Climate, Ocean and Ecosystem Studies, University of Washington, Seattle, WA, United States, 7 Pacific Marine Environmental Laboratory, National Oceanic and Atmospheric Administration, Seattle, WA, United States, 8 Faculty of Mathematics and Natural Sciences, Christian-Albrechts-Universität zu Kiel, Kiel, Germany Ocean data synthesis products for specific biogeochemical essential ocean variables have the potential to facilitate today’s biogeochemical ocean data usage and comply with the Findable Accessible Interoperable and Reusable (FAIR) data principles. The products constitute key outputs from the Global Ocean Observation System, laying the observational foundation for information and services regarding climate and environmental status of the ocean. Using the Framework of Ocean Observing (FOO) readiness level concept, we present an evaluation framework for biogeochemical data synthesis products, which enables a systematic assessment of each product’s maturity. A new criteria catalog provides the foundation for assigning scores to the nine FOO readiness levels. As an example, we apply the assessment to four existing biogeochemical essential ocean variables data products. In descending readiness level order these are: The Surface Ocean CO 2 Atlas (SOCAT); the Global Ocean Data Analysis Project (GLODAP); the MarinE MethanE and NiTrous Oxide (MEMENTO) data product and the Global Ocean Oxygen Database and ATlas (GO 2 DAT). Recognizing that the importance of adequate and comprehensive data from the essential ocean variables will grow, we recommend using this assessment framework to guide the biogeochemical data synthesis activities in their development. Moreover, we envision an overarching cross-platform FAIR biogeochemical data management system that sustainably supports the products individually and creates an integrated biogeochemical essential ocean variables data synthesis product; in short a system that provides truly comparable and FAIR data of the entire biogeochemical essential ocean variables spectrum. KEYWORDS data synthesis product, essential ocean variable, FAIR, technical readiness level, GLODAP, SOCAT, MEMENTO, GO2DAT Frontiers in Marine Science frontiersin.org01 OPEN ACCESS EDITED BY Takafumi Hirata, Hokkaido University, Japan REVIEWED BY Shin-ichiro Nakaoka, National Institute for Environmental Studies (NIES), Japan Piotr Kowalczuk, Polish Academy of Sciences, Poland *CORRESPONDENCE Toste Tanhua [email protected] SPECIALTY SECTION This article was submitted to Ocean Observation, a section of the journal Frontiers in Marine Science RECEIVED 24 October 2022 ACCEPTED 12 April 2023 PUBLISHED 26 April 2023 CITATION Lange N, Tanhua T, Pfeil B, Bange HW, Lauvset SK, Gre ´goire M, Bakker DCE, Jones SD, Fiedler B, O’Brien KM and Körtzinger A (2023) A status assessment of selected data synthesis products for ocean biogeochemistry. Front. Mar. Sci. 10:1078908. doi: 10.3389/fmars.2023.1078908 COPYRIGHT © 2023 Lange, Tanhua, Pfeil, Bange, Lauvset, Gre ´goire, Bakker, Jones, Fiedler, O’Brien and Körtzinger. This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms. TYPE Original Research PUBLISHED 26 April 2023 DOI 10.3389/fmars.2023.1078908 1 Introduction Covering approximately 71% of the Earth’s surface, the ocean’s importance for the earth system and our society is immense. In times of rising carbon dioxide (CO 2 ) and climate change, the environmental status of the ocean and the associated services for society are at risk (Cooley et al., 2022). Even more so as the ocean itself takes a crucial role in “[…] climate by storing and transporting large amounts of heat, freshwater, and carbon, and by exchanging these properties with the atmosphere.”(Rhein et al., 2013). The Global Ocean Observing System (GOOS) has built a structure that coordinates and supports the entire range of ocean observations centered around Essential Ocean Variables (EOVs) (Moltmann et al., 2019;Snowden et al., 2019;Tanhua et al., 2019). Using the Framework of Ocean Observing (FOO) (Lindstrom et al., 2012), the International Ocean Carbon Coordination Project (IOCCP), as the GOOS expert panel for ocean biogeochemistry (BGC), defined the following eight BGC EOVs (IOCCP, 2017): Inorganic carbon, dissolved oxygen (O 2 ), nutrients, particulate matter, dissolved organic carbon, transient tracers and nitrous oxide (N 2 O). A primary objective is to quantify their overall inventories, exchange fluxes and concentration trends. Generally, these quantifications advanced during the past decades, but there are still large uncertainties and many unresolved issues due to insufficient availability of BGC observations. To only mention a few examples, (i) ocean carbon sink estimates from ensemble means of global BGC ocean models and observation-based data products have become increasingly dissimilar with an offset of 1.1 GtC yr -1 in 2020 (Friedlingstein et al., 2022); (ii) models and observation-based products disagree on the strength and spatial distribution of deoxygenation (IPCC, 2019); and (iii) estimated contributions of N 2 Ofluxes from O 2 minimum zones to the global ocean source range from 4% to 50% (IPCC, 2021). To gain an improved holistic understanding of the climate and the ocean’s environmental status, large quantities of easily accessible BGC EOV data –that are spatially and temporally well-resolved, of high quality and from multiple and complementing observing platforms – are required. In particular, it is important to make available observational data FAIR (Findable Accessible Interoperable and Reusable), and enhance the value by proper quality control. Hence, the development of BGC data management systems complying with the FAIR guiding principles for scientificdatamanagementand stewardship has become more important (Wilkinson et al., 2016; Tanhua et al., 2019b). Continuous global efforts aim for more stream-lined and user-orientated data access systems such as the World Ocean Database (Boyer et al., 2018) and the European Marine Observation and Data Network (EMODnet, Miguez et al., 2019). Further user niches are filled in by community-driven synthesis data products that apply (advanced) merging techniques to combine datasets from multiple sources to form a coherent and consistent data product. These synthesis products are either tailored around specific BGC EOVs (e.g. Surface Ocean CO 2 Atlas (SOCAT), the Global Ocean Oxygen Database and Atlas (GO 2 DAT), the MarinE MethanE and NiTrous Oxide (MEMENTO) database) or specific observing platforms [e.g. the Global Ocean Data Analysis Project (GLODAP)]. Generally, these synthesis data products try to solve many obstacles that the current landscape of BGC data has created. Observing campaigns are mostly funded as research projects and often have very specific research questions. Consequently, a multitude of data centers are managing ocean BGC EOV data. These range from local and national data centers (e.g. the Ocean Science Information System at GEOMAR Helmholtz Center for Ocean Research Kiel; the Information and Data Centre at CSIRO National Collections and Marine Infrastructure) to regional infrastructures (e.g. the Integrated Carbon Observing System (ICOS)) or international data centers (e.g. PANGAEA; CCHDO). Hence, data mining has become increasingly difficult and timeconsuming, requiring downloading datasets from different entry points, searching for duplicates, and managing different metadata. Further, BGC EOV data have many users and stakeholders who have highly diverse needs from the data, especially in terms of quality-control (QC). Consequently, a plethora of data versions, file formats and levels of documentation exist (Shepherd, 2018;Miguez et al., 2019;Tanhua et al., 2019). Synthesis data products represent one solution to these data fragmentation issues by the provision of single access points to consistent data and metadata. Nevertheless, some data are collected but not available: for example, many datasets submitted to SOCAT include atmospheric CO 2 measurements that could be useful for air-sea CO 2 flux calculations but are not published as part of the official SOCAT product. Similarly, some ship-based instruments have an O 2 sensor, but the measurements are not processed or archived anywhere. In addition, automated datastreams are uncommon for, in particular, reprocessed or delayed mode data. Such data has passed additional quality control, is characterized by high precision and accuracies and represents data with sufficient quality for climate studies. As a result of the lack of automation, the information exchange between multiple data systems, i.e. interoperability (ISO/IEC/IEEE, 2017), is also limited. These relatively low levels of interoperability hinder data reuse, preservation and integration, and increase associated data management costs (Snowden et al., 2019). The lack of automation also results in large elapsed times from the actual measurement to the provision of the data, i.e. in a high latency. Thus, the many data synthesis efforts are far from complete and in “the era of big data comes to oceanography”(Abbott, 2013) there is a mandate for optimizing fit-for-purpose data synthesis products and their underlying workflows to enhance efficient and interoperable data usage (Tanhua et al., 2019b). The FOO readiness level concept (Lindstrom et al., 2012) becomes useful in this context. Applying it to existing BGC EOV products could guide both existing and new products in their development. Here we introduce such an evaluation framework for four existing BGC EOV data synthesis products: SOCAT, GLODAP, MEMENTO and GO 2 DAT. We first describe the methodology for assessing the products before the four BGC data synthesis products are briefly presented and their maturity is assessed. Finally, we synthesize the Lange et al. 10.3389/fmars.2023.1078908 Frontiers in Marine Science frontiersin.org02 findings and outline our vision of a larger-scale cross-platform BGC EOV data system. 2 Method 2.1 The FOO readiness level concept To assess the maturity of an ocean observing system the Framework of Ocean Observing has adapted the technical readiness level, a scheme developed by NASA (National Aeronotics and Space Administration) (Sadin et al., 1989), and introduced the ocean observing “readiness level”(Lindstrom et al., 2012). Following this framework, ocean observing should be seen as “[…] a chain of processes addressing “why to observe?” (requirement setting process), “what to observe?”(scoping of observational foci), “how to observe?”(coordination of observing elements), and “how to integrate, use and disseminate observational outcomes and understand their impacts?”(data management, analyses and creation and assessment of information products).” (Pearlman et al., 2019). The three pillars of this ocean observing value chain 1 are: “Requirements”,“Observations”and “Data and Information”. For each of these pillars, FOO defined nine readiness levels and grouped these into the categories “Concept”,“Pilot”and “Mature”. A holistic approach enables the evaluation and classification of an entire ocean observing system in terms of feasibility, capacity, and impact. Here we only use the defined readiness levels for “Data Management and Information Products”(Figure 1). We restrict ourselves to climate quality data since these are strongly tied to high-quality BGC EOV synthesis data products, especially to their quality control procedures. The nine readiness levels (Lindstrom et al., 2012) are quite general, so to suit the aim of this work, we have developed a criteria catalog (Appendix 1) which forms an objective basis for the evaluation of the individual data products. Applying the catalog assigns (weighted) scores to typical characteristics of data products on a level-by-level scheme. Full compliance with the criteria yields a 100% score for a given level, with 80% being defined as a “pass”. For example, a product passes readiness level 5 if the data management practices are verified and validated through an existing data policy and archival plan. The criteria catalog (Appendix 1) assigns equally weighted scores to “Policy”,“Archival”and “QC Verification”. These, in turn, are linked to specific data product features, such as having a data usage statement for “Policy”(Figure 2). Note that even though the order of levels is structured hierarchically, a data product can meet some requirements of higher levels before fully complying with all lower levels. Since the maturity of a data product is strongly tied to the FAIR guidelines, we have incorporated the guidelines into the criteria. Following Tanhua et al. (2019), a data product is FAIR if it has a unique persistent identifier with enriched and standardized metadata (findable), enabling access to the machine-readable data and metadata (accessible and interoperable), and can be integrated into other data sources (reusable). The degree of the implementation of the FAIR principles is reflected in the order of the FOO readiness levels. The degree of being “fit-for-purpose”, a requirement of the ocean observing value chain, is also incorporated into the criteria catalog. Given the diverse nature of the data, the criteria have not been further specified and are kept generic on purpose. Workflows and tools used in different products might resemble one another but are tailored toward the specific requirements of the data products. In particular, the data upload (or ingestion) system and quality control methods differ as these are tailored towards the given observing platform, sampling method (continuous or discrete), analysis type, variable (e.g. Johnson et al., 2001;Dickson et al., 2007;Pierrot et al., 2009;Maurer et al., 2021) and stakeholder. Since many research groups and products implement different QC flagging schemes, we have applied a consistent set of quality levels (adapted from ICOS, https://www.icos-cp.eu/data-services/data-collection/data-levelsquality) to describe the data flow and QC of the different products (Table 1). Typical QC examples of the different levels are range tests (level 1), the identification of spikes in space or time (level 2) and the adjustment of known biases (level 3). 3 Synthesis data product assessment In the following, we will briefly describe and evaluate four available BGC EOV data synthesis products for their maturity in terms of FOO readiness. The products were selected based on the goal of covering the entire BGC EOV data synthesis product spectrum. The products cover different BGC EOVs, observing platforms and approaches (cross-platform vs. cross-EOV) and range from products in the planning phase to well-established ones. 3.1 SOCAT The Surface Ocean CO 2 Atlas (Pfeil et al., 2013;Sabine et al., 2013) is an international community-driven effort. It synthesizes insitu surface ocean fCO 2 (fugacity of carbon dioxide) measurements from ships, moored stations, autonomous and drifting surface platforms and yachts with an estimated accuracy better than 10 µatm. SOCAT increases ocean surface fCO 2 data availability and forms the basis of several other data products, such as the SeaFlux data set (Gregor and Fay, 2021) and diverse scientific applications and assessments. The latter range from ocean and climate model and sensor evaluation, regional process studies of surface ocean fCO 2 , the detection and estimation of surface ocean acidification trends (Freeman and Lovenduski, 2015;Lauvset et al., 2015), to the quantification of the ocean carbon sink and its variation (Bakker et al., 2016;Friedlingstein et al., 2022). Thus, SOCAT represents a “[…] key step in the value chain based on in situ inorganic carbon measurements of the oceans, which provides policymakers in climate negotiations with essential information on ocean CO 2 uptake”(Bakker et al., 2020;Guidi et al., 2020). SOCAT’sfirst 1 a term broadly defined as a set of value-adding activities that one or more communities perform in creating and distributing goods and services (Longhorn and Blakemore, 2007) Lange et al. 10.3389/fmars.2023.1078908 Frontiers in Marine Science frontiersin.org03 version (Pfeil et al., 2013;Sabine et al., 2013), was released in 2011 following a call from the international marine carbon community to create a quality-controlled, publicly available synthesis product of surface ocean CO 2 for the global oceans and coastal seas (IOCCP, 2007;Doney et al., 2009). SOCATv2 and SOCATv3 followed in 2013 (Bakker et al., 2014) and 2015 (Bakker et al., 2016), respectively. After the official launch of the SOCAT submission system in September 2015 (SOCAT and SOCOM, 2015), annual product releases have been accomplished. SOCATv2022 includes more than 40 million individual measurements from 1957 to 2021 from more than 100 data contributors (Bakker et al., 2022). The data product consists of 1) the collection of all individual data set files, 2) global and regional synthesis data products, 3) global (monthly, yearly and decadal) gridded products on a 1° latitude by 1° longitude grid and 4) a coastal monthly gridded product on a quarter degree grid. The main synthesis products (2, 3, 4) are based on surface water fCO 2 with an estimated accuracy of better than 5 µatm (33.7 million data points), while fCO 2 values with an accuracy of 5 to 10 µatm are made available separately (6.4 million data points). Recent SOCAT products contain searchable information on the organization where data providers are based, a step towards attributing data sets to funding agencies and countries. While SOCAT synthesis products are made available via ERDDAP (Section 4.1.1.1), metadata of individual data sets in SOCAT are not yet machine-readable. Planned metadata automation will contribute to the initiative led by the Intergovernmental Oceanographic Commission of UNESCO towards a federated data system for the UN Sustainable Development Goal (SDG, UN, 2015)14.3(“Minimize and address the impacts of ocean acidification, including through enhanced scientific cooperation at all levels”). SOCAT also considers to include additional variables to the product, such as atmospheric CO 2 , dissolved inorganic carbon (DIC), total alkalinity (TA), pH, nutrients, methane (CH 4 )andnitrousoxide(N 2 O) concentrations (SOCAT and SOCOM, 2015;Bakker et al., 2016). FIGURE 1 FOO Readiness level for Data Management and Information Products, adapted from Figure 9 in Lindstrom et al. (2012). FIGURE 2 Score assignment scheme for readiness level 5 (Verification). Lange et al. 10.3389/fmars.2023.1078908 Frontiers in Marine Science frontiersin.org04 3.1.1 Software developments 3.1.1.1 ERDDAP The open source software ERDDAP is used as the backbone for SOCAT data quality-control as well as providing access to data and data product. To effectively improve data interoperability, it is not enough to ensure that data are freely and openly available, though both are necessary. To reach a more diverse set of users, including domain and non-domain experts, it is critical to provide effective data services that are easy to use, support multiple data formats, and provide access to humans and machines. One tool that provides all of these capabilities is the open source software ERDDAP. There are several benefits of using ERDDAP as a data server. Among its many features, it (i) supports dozens of popular formats; (ii) provides standards-based metadata and data services and formats; (iii) supports federated access of distributed ERDDAP data services; (iv) supports both human and machine interactions; (v) supports sub-setting of large datasets; (vi) provides improved discovery of datasets through commercial search engines; and (vii) provides support for archival of datasets. The GOOS Observations Coordination Group has adopted ERDDAP as the FAIR-compliant data server of choice for the global ocean networks. Serving data through a tool such as ERDDAP may also help better understand data access patterns. The most accurate method of understanding data usage relies on citations, particularly when using Digital Object Identifiers (DOIs). Using a tool such as ERDDAP also make it possible to gather usage statistics on how data is being accessed, which is a useful additional metric towards a more complete and accurate view of data usage. The usage tracking capabilities of ERDDAP can thus provide a mechanism to track user access, which can largely eliminate the requirements for users to log in. 3.1.1.2 QuinCe The European Research Infrastructure ICOS is developing QuinCe (Steinhoff et al., 2019), as a standardized online tool to ingest, process and QC underway surface ocean fCO 2 measurements from diverse instruments using community-agreed algorithms. While presently QuinCe is only available to a few data providers, in future it will allow data providers to process their data transparently. That includes a record trail that links all applied changes to the original data, i.e. full data provenance is established. QuinCe can automatically export all data in several formats to data centers, near-real-time products, delayed mode products, and the SOCAT data submission system (or dashboard). QuinCe also automatically performs calibrations, data processing, and basic QC of underway instrument data from different platforms (allowing all text formats as input). An interactive user interface with time-series plots, cruise maps and a data table enables the data provider to perform detailed manual QC (Figure 3). The interactive control also enables additional manual scientific1 st QC, i.e. outlier detection, of the level 1 fCO 2 data, which results in level 2 fCO 2 data (World Ocean Circulation Experiment flagging scheme applied). For future traceability, QuinCe records all QC decisions. 3.1.2 FOO readiness SOCAT has implemented a clear concept and management structure “[…] to integrate, use and disseminate observational outcomes and understand their impacts [ …]”(Pearlman et al., FIGURE 3 A screenshot of the main Quality Control page of QuinCe, showing data from sensors in plot and map form together with a table of all sensor and calculated values. Flagged values from automatic and manual QC are highlighted. TABLE 1 Data quality levels. Level Characteristics 0 Uncalibrated 1 Calibrated data with passed automated check (known as ‘sanity check’) 2 Scientific 1st level QC for precision and accuracy has been performed 3 External scientific QC for precision and accuracy has been performed Lange et al. 10.3389/fmars.2023.1078908 Frontiers in Marine Science frontiersin.org05 Research Council of Norway project ICOS Norway and OTC, phase 2 (grant number 296012). DCEB is grateful for support from the UK’s Natural Environment Research Council CUSTARD (Carbon Uptake and Seasonal Traits in Antarctic Remineralisation Depth) project (NE/P02/263/1). Acknowledgments We thank the funding agencies and the data management projects that have made this work possible through dedicated funding for the data management activities and improvements. We thank the many researchers responsible for the collection of data and quality control for their contributions to these data products. The Surface Ocean CO 2 Atlas is an international effort, endorsed by the International Ocean Carbon Coordination Project, the Surface Ocean Lower Atmosphere Study and the Integrated Marine Biosphere Research program, to deliver a uniformly qualitycontrolled surface ocean CO 2 database. This paper contributes to the science plan of the Surface Ocean Lower Atmosphere Study, which is supported by the U.S. National Science Foundation via the Scientific Committee on Oceanic Research. Conflict of interest The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. Publisher’s note All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher. Supplementary material The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fmars.2023.1078908/ full#supplementary-material References Abbott, M. (2013). The era of big data comes to oceanography. Oceanogr. 26, 7–8. doi: 10.5670/oceanog.2013.68 Bakker, D. C. E., Alin, S., Bates, N., Becker, M., Castaño-Primo, R., Cosca, C., et al. (2020) SOCAT version 2020: key in the value chain of surface ocean CO2 measurements. Available at: https://www.socat.info/wp-content/uploads/2020/06/2020_Poster_ SOCATv2020_release.pdf. Bakker, D. C. E., Alin, S. R., Becker, M., Bittig, H., Castaño-Primo, R., Feely, R. A., et al. (2022) SOCAT version 2022 for quantification of ocean CO2 uptake. Available at: https://www.socat.info/wp-content/uploads/2022/06/2022_Poster_SOCATv2022_ release.pdf. Bakker, D. C. E., Lauvset, S. K., Wanninkhof, R., O'Brien, K., Olsen, A., Pfeil, B., et al. (2019) SOCAT version 2019: 26 million in situ surface ocean CO2 observations. Available at: https://www.socat.info/wp-content/uploads/2019/06/2019_Poster_ SOCATv2019_release.pdf. Bakker,D.C.E.,Pfeil,B.,Landa,C.S.,Metzl,N.,O’Brien,K.M.,Olsen,A.,etal.(2016).A multi-decade record of high quality fCO2 data in version 3 of the surface ocean CO 2 atlas (SOCAT). EarthSys.Sci.Data8, 383–413. doi: 10.5194/essd-8-383-2016 Bakker, D. C. E., Pfeil, B., Smith, K., Hankin, S., Olsen, A., Alin, S. R., et al. (2014). An update to the surface ocean CO 2 atlas (SOCAT version 2). Earth Sys. Sci. Data 6, 69–90. doi: 10.5194/essd-6-69-2014 Bange, H. W., Bell, T. G., Cornejo, M., Freing, A., Uher, G., Upstill-Goddard, R. C., et al. (2009). MEMENTO: a proposal to develop a database of marine nitrous oxide and methane measurements. Environ. Chem. 6, 195–197. doi: 10.1071/EN09033 Bittig, H. C., Steinhoff, T., Claustre, H., Fiedler, B., Williams, N. L., Sauzède, R., et al. (2018). An alternative to static climatologies: robust estimation of open ocean CO 2 variables and nutrient concentrations from T, s, and O 2 data using Bayesian neural networks. Front. Mar. Sci. 5. doi: 10.3389/fmars.2018.00328 Boyer,T.P.,Baranova,O.K.,Coleman,C.,Garcia,H.E.,Grodsky,A.,Locarnini,R.A.,etal. (2018). World ocean database 2018 Vol. 87. Ed. A. V. Mishonov (NOAA Atlas NESDIS). Available at: https://www.ncei.noaa.gov/sites/default/files/2020-04/wod_intro_0.pdf. Cooley, S., Schoeman, D., Bopp, L., Boyd, P., Donner, S., Ghebrehiwet, D. Y., et al. (2022). Ocean and coastal ecosystems and their services. in Climate change 2022: impacts, adaptation, and vulnerability. contribution of working group II to the sixth assessment report of the intergovernmental panel on climate change. Chapter 3. IPCC AR6 WGII (Cambridge University Press). https://www.ipcc.ch/report/ar6/wg2/ downloads/report/IPCC_AR6_WGII_FinalDraft_Chapter03.pdf. Diaz, R. J., and Rosenberg, R. (2008). Spreading dead zones and consequences for marine ecosystems. Science 321 (5891), 926–929. doi: 10.1126/science.1156401 Dickson, A. G., Sabine, C. L., and Christian, J. R. (2007). Guide to best practices for ocean CO 2 measurements. PICES Special Publ. 3, 191, 2007.3. doi: 10.25607/OBP-1342 Doney, S. C., Tilbrook, B., Roy, S., Metzl, N., Le Quere, C., Hood, M., et al. (2009). Surface ocean CO 2 variability and vulnerability. Deep-Sea Res. Pt. II 56, 504–511. doi: 10.1016/j.dsr2.2008.12.016 European Marine Board. (2021). Sustaining in situ ocean observations in the age of the digital ocean. EMB Policy Brief 9, 14. doi: 10.5281/zenodo.4836060 Freeman, N. M., and Lovenduski, N. S. (2015). Decreased calcification in the southern ocean over the satellite record. Geophys. Res. Lett. 42, 1834–1840. doi: 10.1002/2014GL062769 Freing, A., and Bange, H. W. (2007). Towards a global database of oceanic nitrous oxide measurements. IMBER Update 8, 3–4. Freing, A., Wallace, D. W. R., and Bange, H. W. (2012). Global oceanic production of nitrous oxide. Phil. Trans. R. Soc. 367 (1593), 1245–1255. doi: 10.1098/rstb.2011.0360 Friedlingstein, P., Jones, M. W., O'Sullivan, M., Andrew, R. M., Bakker, D. C. E., Hauck, J., et al. (2022). Global carbon budget 2021. Earth Sys. Sci. Data 14, 1917–2005. doi: 10.5194/essd-14-1917-2022 Friedlingstein, P., O'Sullivan, M., Jones, M. W., Andrew, R. M., Hauck, J., Olsen, A., et al. (2020). Global carbon budget 2020. Earth Syst. Sci. Data 12, 3269–3340. doi: 10.5194/essd-12-3269-2020 Gregoire, M., Garcon, V., Garcia, H., Breitburg, D., Isensee, K., Oschlies, A., et al. (2021). A global ocean oxygen database and atlas for assessing and predicting deoxygenation and ocean health in the open and coastal ocean. Front. Mar. Sci. 8. doi: 10.3389/fmars.2021.724913 Gregor, L., and Fay, A. R. (2021). SeaFlux data set: air-sea CO 2 fluxes for surface pCO 2 data products using a standardised approach. Zenodo.doi:10.5281/ zenodo.4133802 Gruber, N., Clement, D., Carter, B. R., Feely, R. A., van Heuven, S., Hoppema, M., et al. (2019). The oceanic sink for anthropogenic CO 2 from 1994 to 2007. Science 363, 1193–1199. doi: 10.1126/science.aau5153 Guidi, L., Fernandez Guerra, A., Canchaya, C., Curry, E., Foglini, F., Irisson, J.-O., et al. (2020). Big data in marine science. Eds. B. Alexander, J. J. Heymans, A. Muñiz Piniella, P. Kellett and J. Coopman (Ostend, Belgium: Future Science Brief 6 of the European Marine Board). doi: 10.5281/zenodo.3755793 IOCCP (2007). Surface ocean CO2 variability and vulnerabilities workshop (Paris, France: UNESCO), 11–14. Available at: http://www.ioccp.org/. IOCCP (2017). GOOS Biogeochemistry Essential Ocean Variables (EOVs) - Specification Sheets. Version: 25.08.2017. Available at: http://www.ioccp.org/index. php/foo. accessed on: 18.04.2023 IPCC (2019). IPCC special report on the ocean and cryosphere in a changing climate. Eds. H.-O. Pörtner, D. C. Roberts, V. Masson-Delmotte, P. Zhai, M. Tignor, E. Lange et al. 10.3389/fmars.2023.1078908 Frontiers in Marine Science frontiersin.org12 Poloczanska, K. Mintenbeck, A. Alegrıa, M. Nicolai, A. Okem, J. Petzold, B. Rama and N. M. Weyer Cambridge University Press, Cambridge, UK and New York, NY, USA, pp. 3-35. doi: 10.1017/9781009157964.001 IPCC (2021). “Climate change 2021: the physical science basis,”in Contribution of working group I to the sixth assessment report of the intergovernmental panel on climate change. Eds. V. Masson-Delmotte, P. Zhai, A. Pirani, S. L. Connors, C. Pean, S. Berger, N. Caud, Y. Chen, L. Goldfarb, M. I. Gomis, M. Huang, K. Leitzell, E. Lonnoy, J. B. R. Matthews, T. K. Maycock, T. Waterfield, O. Yelekci, R. Yu and X. X. X. B. Zhou (Cambridge, United Kingdom and New York, NY, USA: Cambridge University Press). doi: 10.1017/9781009157896 ISO/IEC/IEEE (2017). International standard - systems and software engineering– vocabulary SO/IEC/IEEE 24765:2017 (E), 1–541. doi: 10.1109/IEEESTD.2017.8016712 Johnson, G. C., Robbins, P. E., and Hufford, G. E. (2001). Systematic adjustments of hydrographic sections for internal consistency. J. Atmos. Oceanic Technol. 18, 1234– 1244. doi: 10.1175/1520-0426(2001)018 Key, R., Kozyr, A., Chris, S., Lee, K., Wanninkhof, R., Bullister, J., et al. (2004). A global ocean carbon climatology: results from global data analysis project (GLODAP). Global Biogeochem. Cycles 18. doi: 10.1029/2004GB002247 Key, R. M., Olsen, A., van Heuven, S., Lauvset, S. K., Velo, A., Lin, X., et al. (2015). “Global ocean data analysis project, version 2 (GLODAPv2),”in ORNL/CDIAC-162, ND-P093 (Oak Ridge, Tennessee: Carbon Dioxide Information Analysis Center, Oak Ridge National Laboratory, US Department of Energy). doi: 10.3334/CDIAC/ OTG.NDP093_GLODAPv2 Key, R. M., Tanhua, T., Olsen, A., Hoppema, M., Jutterström, S., Schirnick, C., et al. (2010). The CARINA data synthesis project: introduction and overview. Earth Syst. Sci. Data 2, 105–121. doi: 10.5194/essd-2-105-2010 Kock, A., and Bange, H. W. (2015). Counting the ocean's greenhouse gas emissions. Eos Earth Space Sci. News 96, 10–13. doi: 10.1029/2015EO023665 Lauvset, S. K., Currie, K., Metzl, N., Nakaoka, S.-I., Bakker, D. C. E., Sullivan, K., et al. (2018) SOCAT quality control cookbook -for SOCAT version 7. Available at: https:// www.socat.info/wp-content/uploads/2019/01/2018_SOCAT_QC_Cookbook_for_ SOCAT_Version_7.pdf (Accessed 01/11/2020). Lauvset, S., Gruber, N., Landschützer, P., Olsen, A., and Tjiputra, J. (2015). Trends and drivers in global surface ocean pH over the past 3 decades. Biogeosci. 12, 1285– 1298. doi: 10.5194/bg-12-1285-2015 Lauvset, S. K., Lange, N., Tanhua, T., Bittig, H. C., Olsen, A., Kozyr, A., et al. (2021). An updated version of the global interior ocean biogeochemical data product, GLODAPv2.2021. Earth Syst. Sci. Data 13, 5565–5589. doi: 10.5194/essd-13-5565-2021 Lauvset, S., and Tanhua, T. (2015). A toolbox for secondary quality control on ocean chemistry and hydrographic data. Limnol. Oceanogr. Methods, 11601–11608. doi: 10.1002/lom3.10050 Lauvset, S. K., Lange, N., Tanhua, T., Bittig, H. C., Olsen, A., Kozyr, A., et al. (2022). GLODAPv2.2022: the latest version of the global interior ocean biogeochemical data product. Earth Sys. Sci. Data 14 (12), 5543–5572. doi: 10.5194/essd-14-5543-2022 Le Quere, C., Andrew, R. M., Friedlingstein, P., Sitch, S., Hauck, J., Pongratz, J., et al. (2018). Global carbon budget 2018. Earth Syst. Sci. Data 10, 2141–2194. doi: 10.5194/essd-10-2141-2018 Lindstrom,E.,Gunn,J.,Fischer,A.,McCurdy,A.,Glover,L.,Alverson,K.,etal.(2012).“A framework for ocean observing,”in By the task team for an integrated framework for sustained ocean observing (Paris: UNESCO). doi: 10.5270/OceanObs09-FOO Longhorn, R. A., and Blakemore, M. J. (2007). Geographic information - value, pricing, production, and consumption. 1st edition (Boca Raton: CRC Press). Maurer, T. L., Plant, J. N., and Johnson, K. S. (2021). Delayed-mode quality control of oxygen, nitrate, and pH data on SOCCOM biogeochemical profiling floats. Front. Mar. Sci. 8. doi: 10.3389/fmars.2021.683207 Merchant, C. J., Paul, F., Popp, T., Ablain, M., Bontemps, S., Defourny, P., et al. (2017). Uncertainty information in climate data records from earth observation. Earth Syst. Sci. Data 9, 511–527. doi: 10.5194/essd-9-511-2017 Miguez, M., Belen, A., Novellino, A., Vinci, M., Claus, S., Calewaert, J., et al. (2019). The European marine observation and data network (EMODnet): visions and roles of the gateway to marine data in Europe. Front. Mar. Sci. 6. doi: 10.3389/fmars.2019.00313 Moltmann, T., Turton, J., Zhang, H., Nolan, G., Gouldman, C., Griesbauer, L., et al. (2019). A global ocean observing system (GOOS), delivered through enhanced collaboration across regions, communities, and new technologies. Front. Mar. Sci. 6. doi: 10.3389/fmars.2019.00291 NOAA (2022) ERDDAP: easier access to scientific data. Available at: https://www. ncei.noaa.gov/erddap/information.html#:~:text=ERDDAP%20is%20a%20data% 20server,and%20make%20graphs%20and%20maps. Olsen, A., Key, R. M., van Heuven, S., Lauvset, S. K., Velo, A., Lin, X. H., et al. (2016). The global ocean data analysis project version 2 (GLODAPv2) - an internally consistent data product for the world ocean. Earth Syst. Sci. Data 8, 297–323. doi: 10.5194/essd-8297-2016 Olsen, A., Lange, N., Key, R. M., Tanhua, T., A lvarez, M., Becker, S., et al. (2019). GLODAPv2.2019 –an update of GLODAPv2. Earth Sys. Sci. Data 11 (3), 14371461. doi: 10.5194/essd1114372019 Olsen, A., Lange, N., Key, R. M., Tanhua, T., Bittig, H. C., Kozyr, A., et al. (2020). An updated version of the global interior ocean biogeochemical data product, GLODAPv2.2020. Earth Syst. Sci. Data 12, 3653–3678. doi: 10.5194/essd-12-3653-2020 Pearlman, J., Bushnell, M., Coppola, L., Karstensen, J., Buttigieg, P. L., Pearlman, F., et al. (2019). Evolving and sustaining ocean best practices and standards for the next decade. Front. Mar. Sci. 6. doi: 10.3389/fmars.2019.00277 Pfeil, B., Olsen, A., Bakker, D. C. E., Hankin, S., Koyuk, H., Kozyr, A., et al. (2013). A uniform, quality controlled surface ocean CO 2 atlas (SOCAT). Earth Sys. Sci. Data 5, 125–143. doi: 10.5194/essd-5-125-2013 Pierrot, D., Neill, C., Sullivan, K., Castle, R., Wanninkhof, R., Naigur, H., et al. (2009). Recommendations for autonomous underway pCO 2 measuring systems and data-reduction routines. Deep Sea Res. Part II: Topical Stud. Oceanogr. 56, 512–522. doi: 10.1016/j.dsr2.2008.12.005 Rhein, M., Rintoul, S. R., Aoki, S., Campos, E., Chambers, D., Feely, R. A., et al. (2013). “Observations: ocean,”in Climate change 2013: the physical science basis. contribution of working group I to the fifth assessment report of the intergovernmental panel on climate change. Eds. T. F. Stocker, D. Qin, G.-K. Plattner, M. Tignor, S. K. Allen, J. Boschung, A. Nauels, Y. Xia, V. Bex and P. M. Midgley (Cambridge, United Kingdom and New York, NY, USA: Cambridge University Press). Sabine, C., Feely, R., Gruber, N., Key, R., Lee, K., Bullister, J., et al. (2004). The oceanic sink for anthropogenic CO 2 .Sci. (New York N.Y.) 305, 367–371. doi: 10.1126/science.1097403 Sabine, C. L., Hankin, S., Koyuk, H., Bakker, D. C. E., Pfeil, B., Olsen, A., et al. (2013). Surface ocean CO 2 atlas (SOCAT) gridded data products. Earth Sys. Sci. Data 5, 145– 153. doi: 10.5194/essd-5-145-2013 Sadin,S.R.,Frederick,P.P.,andRosen,R. (1989). The NASA technology push towards future space mission systems. Acta Astronautics 20, 73–77. doi: 10.1016/0094-5765(89)90054-4 Shepherd, I. (2018). European Efforts to make marine data more accessible. Ethics Sci. Environ. Polit. 18, 75–81. doi: 10.3354/esep00181 Sloyan, B. M., Wanninkhof, R., Kramp, M., Johnson, G. C., Talley, L. D., Tanhua, T., et al. (2019). The global ocean ship-based hydrographic investigations program (GOSHIP): a platform for integrated multidisciplinary ocean science. Front. Mar. Sci. 6. doi: 10.3389/fmars.2019.00445 Snowden, D., Tsontos, V., Handegard, N. O., Zarate, M., O'Brien, K., Casey, K., et al. (2019). Data interoperability between elements of the global ocean observing system. Front. Mar. Sci. 6. doi: 10.3389/fmars.2019.00442 SOCAT and SOCOM (2015). “SOCAT (Surface ocean CO 2 atlas) and SOCOM (Surface ocean pCO 2 mapping intercomparison) event,”in Report. SOLAS (Surface OceanLower Atmosphere Study) Open Science Conference (University of Kiel, Kiel, Germany) 7 September 2015. Available at: http://www.socat.info/upload/2015_ SOCAT_and_SOCOM_Event_Report.pdf. Steinhoff, T., Gkritzalis, T., Lauvset, S. K., Jones, S., Schuster, U., Olsen, A., et al. (2019). Constraining the oceanic uptake and fluxes of greenhouse gases by building an ocean network of certified stations: the ocean component of the integrated carbon observation system. ICOS-Oceans. Front. Mar. Sci. 6. doi: 10.3389/fmars.2019.00544 Suntharalingam, P., Buitenhuis, E., Le Quere, C., Dentener, F., Nevison, C., Butler, J. H., et al. (2012). Quantifying the impact of anthropogenic nitrogen deposition on oceanic nitrous oxide. Geophys. Res. Lett. 39, Artn L07605. doi: 10.1029/2011gl050778 Suzuki, T., Ishii, M., Aoyama, A., Christian, J. R., Enyo, K., Kawano, T., et al. (2013). “PACIFICA data synthesis project,”in ORNL/CDIAC-159, NDP-092, carbon dioxide information analysis center, oak ridge national laboratory, u. s (Oak Ridge, TN, USA: Department of Energy). Tanhua, T., Lauvset, S., Lange, N., Olsen, A., A lvarez, M., Diggs, S., et al. (2021). A vision for FAIR ocean data products. Commun. Earth Environment 2, 136. doi: 10.1038/s43247-021-00209-4 Tanhua, T., McCurdy, A., Fischer, A., Appeltans, W., Bax, N., Currie, K., et al. (2019). What we have learned from the framework for ocean observing: evolution of the global ocean observing system. Front. Mar. Sci. 6. doi: 10.3389/fmars.2019.00471 Tanhua, T., Pouliquen, S., Hausman, J., O'Brien, K., Bricher, P., de Bruin, T., et al. (2019b). Ocean FAIR data services. Front. Mar. Sci. 6. doi: 10.3389/fmars.2019.00440 Tanhua, T., Steinfeldt, R., Key, R. M., Brown, P., Gruber, N., Wanninkhof, R., et al. (2009). Atlantic Ocean CARINA data: overview and salinity adjustments. Earth Syst. Sci. Data Discuss 2, 241–280. doi: 10.5194/essdd-2-241-2009 Tanhua, T., van Heuven, S., Key, R. M., Velo, A., Olsen, A., and Schirnick, C. (2010). Quality control procedures and methods of the CARINA database. Earth Syst. Sci. Data 2, 205–240. doi: 10.5194/essd-2-35-2010 UN (2015) Draft outcome document of the united nations summit for the adoption of the post2015 development agenda. draft resolution submitted by the President of the general assembly, sixty-ninth session, agenda items 13 (a) and 115, A/69/L.85 (New York: UN) (Accessed 12, 2015). Velo, A., Jesus, C., Fiz, F. P., Toste, T., and Lange, N. (2021). AtlantOS ocean data QC: software packages and best practice manuals and knowledge transfer for sustained quality control of hydrographic sections. Zenodo. doi: 10.5281/zenodo.4532402 Wanninkhof, R., Bakker, D. C. E., Bates, N., Olsen, A., Steinhoff, T., and Sutton, A. J. (2013). Incorporation of alternative sensors in the SOCAT database and adjustments to dataset quality control flags (Oak Ridge, Tennessee: Carbon Dioxide Information Analysis Center, Oak Ridge National Laboratory, US Department of Energy). Available at: http://cdiac.ornl.gov/oceans/Recommendationnewsensors.pdf. doi: 10.3334/CDIAC/OTG.SOCAT_ADQCF Weber, T., Wiseman, N. A., and Kock, A. (2019). Global ocean methane emissions dominated by shallow coastal waters. Nat. Commun. 10, 4584. doi: 10.1038/s41467019-12541-7 Lange et al. 10.3389/fmars.2023.1078908 Frontiers in Marine Science frontiersin.org13 Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J. J., Appleton, G., Axton, M., Baak, A., et al. (2016). The FAIR guiding principles for scientific data management and stewardship. Sci. Data 3, 160018. doi: 10.1038/sdata.2016.18 WMO (2018). WMO greenhouse gas bulletin. the state of greenhouse gases in the atmosphere based on global observations through 2017. Geneva: World Meteorological Organization, No. 14 Yang,S.,Chang,B.X.,Warner,M.J.,Weber,T.S.,Bourbonnais,A.M.,Santoro,A.E.,etal. (2020). Global reconstruction reduces the uncertainty of oceanic nitrous oxide emissions and reveals a vigorous seasonal cycle. Proc. Natl. Acad. Sci.doi:10.1073/pnas.1921914117 Zamora, L. M., Oschlies, A., Bange, H. W., Huebert, K. B., Craig, J. D., Kock, A., et al. (2012). Nitrous oxide dynamics in low oxygen regions of the pacific: insights from the MEMENTO database. Biogeosciences 9, 5007–5022. doi: 10.5194/bg-9-5007-2012 Lange et al. 10.3389/fmars.2023.1078908 Frontiers in Marine Science frontiersin.org14 Glossary ASCII American Standard Code for Information Interchange ASV Autonomous Surface Vehicle AtlantOS Atlantic Ocean Observing Systems AUV Autonomous Underwater Vehicle BGC BioGeoChemical CANYON CArbonate system and Nutrients concentration from hYdrological properties and Oxygen using a Neural-network CARINA CARbon dioxide IN the Atlantic Ocean CSV Comma Separated Value CTD Conductivity, Temperature and Depth DIC Dissolved Inorganic Carbon DOI Digital Object Identifier EMODnet European Marine Observation and Data Network EOV Essential Ocean Variable FAIR Findable Accessible Interoperable Reusable FOO Framework of Ocean Observation FOS Fixed Ocean Station GLODAP Global Ocean Data Analysis Project GO 2 DAT Global Ocean Oxygen Database and Atlas GOOS Global Ocean Observing System ICOS Integrated Carbon Observing System IOCCP International Ocean Carbon Coordination Project IPCC Intergovernmental Panel on Climate Change MEMENTO MarinE MethanE and NiTrous Oxide PACIFICA PACIFic ocean Interior Carbon QC Quality Control QF Quality Flagging RV Research Vessel SCOR Scientific Committee on Oceanic Research SDG Sustainable Development Goal SOCAT Surface Ocean CO 2 Atlas SOCOM Surface Ocean pCO 2 Mapping intercomparison SOOP Ship Of Opportunity Program TA Total Alkalinity UNESCO United Nations Educational, Scientific and Cultural Organization WMO World Meteorological Organization. Lange et al. 10.3389/fmars.2023.1078908 Frontiers in Marine Science frontiersin.org15 54 55 4 GLODAPv2.2022: the latest version of the global interior ocean biogeochemical data product 4 GLODAPv2.2022: the latest version of the global interior ocean biogeochemical data product 56 Earth Syst. Sci. Data, 14, 5543–5572, 2022 https://doi.org/10.5194/essd-14-5543-2022 © Author(s) 2022. This work is distributed under the Creative Commons Attribution 4.0 License. GLODAPv2.2022: the latest version of the global interior ocean biogeochemical data product Siv K. Lauvset1, Nico Lange2, Toste Tanhua2, Henry C. Bittig3, Are Olsen4, Alex Kozyr5, Simone Alin6, Marta Álvarez7, Kumiko Azetsu-Scott8, Leticia Barbero9,10, Susan Becker11, Peter J. Brown12, Brendan R. Carter13,16, Leticia Cotrim da Cunha14, Richard A. Feely6, Mario Hoppema15, Matthew P. Humphreys16, Masao Ishii17, Emil Jeansson1, Li-Qing Jiang18,19, Steve D. Jones4, Claire Lo Monaco20, Akihiko Murata21, Jens Daniel Müller22, Fiz F. Pérez23,24, Benjamin Pfeil4, Carsten Schirnick2, Reiner Steinfeldt25, Toru Suzuki26, Bronte Tilbrook27, Adam Ulfsbo28, Anton Velo23, Ryan J. Woosley29, and Robert M. Key30 1NORCE Norwegian Research Centre, Bjerknes Centre for Climate Research, Bergen, Norway 2GEOMAR Helmholtz Centre for Ocean Research Kiel, Kiel, Germany 3Leibniz Institute for Baltic Sea Research Warnemünde, Rostock, Germany 4Geophysical Institute, University of Bergen and Bjerknes Centre for Climate Research, Bergen, Norway 5NOAA National Centers for Environmental Information, Silver Spring, MD, USA 6Pacific Marine Environmental Laboratory, National Oceanic and Atmospheric Administration, Seattle, Washington, USA 7Instituto Español de Oceanografía, IEO-CSIC, A Coruña, Spain 8Department of Fisheries and Oceans, Bedford Institute of Oceanography, Dartmouth, Nova Scotia, Canada 9Cooperative Institute for Marine and Atmospheric Studies, University of Miami, Miami, Florida, USA 10Atlantic Oceanographic and Meteorological Laboratory, National Oceanic and Atmospheric Administration, Miami, FL, USA 11Scripps Institution of Oceanography, UC San Diego, San Diego, CA 92093, USA 12National Oceanography Centre, Southampton, UK 13Cooperative Institute for Climate Ocean and Ecosystem Studies, University Washington, Seattle, WA, USA 14PPG-Oceanografia, Faculdade de Oceanografia, Universidade do Estado do Rio de Janeiro, Rio de Janeiro (RJ), Brazil 15Alfred Wegener Institute Helmholtz Centre for Polar and Marine Research, Bremerhaven, Germany 16Department of Ocean Systems (OCS), NIOZ Royal Netherlands Institute for Sea Research, Texel, the Netherlands 17Meteorological Research Institute, Japan Meteorological Agency, Tsukuba, Japan 18Cooperative Institute for Satellite Earth System Studies, Earth System Science Interdisciplinary Center, University of Maryland, College Park, MD 20740, USA 19NOAA/NESDIS National Centers for Environmental Information, 1315 East-West Highway, Silver Spring, MD 20910, USA 20LOCEAN, Sorbonne Université, Paris, France 21Research Institute for Global Change, Japan Agency for Marine-Earth Science and Technology, Yokosuka, Japan 22Environmental Physics, Institute of Biogeochemistry and Pollutant Dynamics, ETH Zürich, Zürich, Switzerland 23Instituto de Investigacións Mariñas, IIM – CSIC, Vigo, Spain 24Oceans Department, Stanford University, Stanford, CA 94305, USA 25Institute of Environmental Physics, University of Bremen, Bremen, Germany 26Marine Information Research Center, Japan Hydrographic Association, Tokyo, Japan Published by Copernicus Publications. 5544 S. K. Lauvset et al.: GLODAPv2.2022 27CSIRO Oceans and Atmosphere and Australian Antarctic Program Partnership, University of Tasmania, Hobart, Australia 28Department of Marine Sciences, University of Gothenburg, Gothenburg, Sweden 29Center for Global Change Science, Massachusetts Institute for Technology, Cambridge, MA, USA 30Atmospheric and Oceanic Sciences, Princeton University, Princeton, NJ, 08540, USA Correspondence: Siv K. Lauvset (siv[email protected]) Received: 23 August 2022 – Discussion started: 12 September 2022 Revised: 14 November 2022 – Accepted: 16 November 2022 – Published: 16 December 2022 Abstract. The Global Ocean Data Analysis Project (GLODAP) is a synthesis effort providing regular compilations of surface-to-bottom ocean biogeochemical bottle data, with an emphasis on seawater inorganic carbon chemistry and related variables determined through chemical analysis of seawater samples. GLODAPv2.2022 is an update of the previous version, GLODAPv2.2021 (Lauvset et al., 2021). The major changes are as follows: data from 96 new cruises were added, data coverage was extended until 2021, and for the first time we performed secondary quality control on all sulfur hexafluoride (SF6) data. In addition, a number of changes were made to data included in GLODAPv2.2021. These changes affect specifically the SF6data, which are now subjected to secondary quality control, and carbon data measured on board the RV Knorr in the Indian Ocean in 1994–1995 which are now adjusted using certified reference material (CRM) measurements made at the time. GLODAPv2.2022 includes measurements from almost 1.4 million water samples from the global oceans collected on 1085 cruises. The data for the now 13 GLODAP core variables (salinity, oxygen, nitrate, silicate, phosphate, dissolved inorganic carbon, total alkalinity, pH, chlorofluorocarbon-11 (CFC-11), CFC-12, CFC-113, CCl4, and SF6) have undergone extensive quality control with a focus on systematic evaluation of bias. The data are available in two formats: (i) as submitted by the data originator but converted to World Ocean Circulation Experiment (WOCE) exchange format and (ii) as a merged data product with adjustments applied to minimize bias. For the present annual update, adjustments for the 96 new cruises were derived by comparing those data with the data from the 989 quality-controlled cruises in the GLODAPv2.2021 data product using crossover analysis. SF6data from all cruises were evaluated by comparison with CFC-12 data measured on the same cruises. For nutrients and ocean carbon dioxide (CO2) chemistry comparisons to estimates based on empirical algorithms provided additional context for adjustment decisions. The adjustments that we applied are intended to remove potential biases from errors related to measurement, calibration, and data handling practices without removing known or likely time trends or variations in the variables evaluated. The compiled and adjusted data product is believed to be consistent to better than 0.005 in salinity, 1 % in oxygen, 2 % in nitrate, 2 % in silicate, 2% in phosphate, 4 µmol kg−1in dissolved inorganic carbon, 4µmol kg−1in total alkalinity, 0.01–0.02 in pH (depending on region), and 5% in the halogenated transient tracers. The other variables included in the compilation, such as isotopic tracers and discrete CO2fugacity (fCO2), were not subjected to bias comparison or adjustments. The original data, their documentation, and DOI codes are available at the Ocean Carbon and Acidification Data System of NOAA NCEI (https://www.ncei.noaa.gov/access/ocean-carbon-acidification-data-system/ oceans/GLODAPv2_2022/, last access: 15 August 2022). This site also provides access to the merged data product, which is provided as a single global file and as four regional ones – the Arctic, Atlantic, Indian, and Pacific oceans – under https://doi.org/10.25921/1f4w-0t92 (Lauvset et al., 2022). These bias-adjusted product files also include significant ancillary and approximated data, which were obtained by interpolation of, or calculation from, measured data. This living data update documents the GLODAPv2.2022 methods and provides a broad overview of the secondary quality control procedures and results. Earth Syst. Sci. Data, 14, 5543–5572, 2022 https://doi.org/10.5194/essd-14-5543-2022 S. K. Lauvset et al.: GLODAPv2.2022 5545 1 Introduction The oceans mitigate climate change by absorbing both atmospheric CO2corresponding to a significant fraction of anthropogenic CO2emissions (Friedlingstein et al., 2019; Gruber et al., 2019) and most of the excess heat in the Earth system caused by the enhanced greenhouse effect (Cheng et al., 2017, 2020). The objective of GLODAP (Global Ocean Data Analysis Project; http://www.glodap.info, last access: 27 June 2022) is to provide high-quality and bias-corrected water column bottle data from the ocean surface to the sea floor. These data should be used to document the state and the evolving changes in physical and chemical ocean properties, e.g., the inventory of anthropogenic CO2in the ocean, natural oceanic carbon, ocean acidification, ventilation rates, oxygen levels, and vertical nutrient transports (Tanhua et al., 2021). The core quality-controlled and bias-adjusted variables of GLODAP are salinity, dissolved oxygen, inorganic macronutrients (nitrate, silicate, and phosphate), seawater CO2chemistry variables (dissolved inorganic carbon – TCO2, total alkalinity – TAlk, and pH on the total hydrogen ion, or H+, scale), the halogenated transient tracers chlorofluorocarbon-11 (CFC-11), CFC-12, CFC-113, carbon tetrachloride (CCl4), and sulfur hexafluoride (SF6). Other chemical tracers are measured on many cruises included in GLODAP, such as dissolved organic carbon and nitrogen, as well as stable and radioactive isotope ratios. In many cases, a subset of these data is distributed as part of the GLODAP data product; however, such data have not been extensively quality controlled or checked for measurement biases in this effort. For some of these variables better sources of data exist, for example the product by Jenkins et al. (2019) for helium isotope and tritium data. GLODAP also includes some common derived variables to facilitate interpretation, such as potential density anomalies and apparent oxygen utilization (AOU). A full list of variables included in the data product is provided in Table 1. The oceanographic community largely adheres to principles and practices for ensuring open access to research data, such as the FAIR (Findable, Accessible, Interoperable, Reusable) initiative (Wilkinson et al., 2016), but the plethora of file formats and different levels of documentation, combined with the need to retrieve data on a per cruise basis from different access points, limit the realization of their full scientific potential. In addition, the manual data retrieval is time consuming and prone to data handling errors (Tanhua et al., 2021). For biogeochemical data there is the added complexity of different levels of standardization and calibration and even different units and scales used for the same variable such that the comparability between datasets is often poor. Standard operating procedures have been developed for some variables (Dickson et al., 2007; Hood et al., 2010; Becker et al., 2020), and certified reference materials (CRMs) exist for seawater TCO2and TAlk measurements (Dickson et al., 2003) and reference materials for nutrients in seawater (RMNS, certified based on International Organization for Standardization Guide 34; Aoyama et al., 2012; Ota et al., 2010). Despite all this, biases in data still exist. These can arise from poor sampling and preservation practices, calibration procedures, instrument design and calibration, and inaccurate calculations. The use of CRMs does not by itself ensure accurate measurements of seawater CO2chemistry (Bockmon and Dickson, 2015), and the RMNS have only become available recently and are not universally used. For salinity and oxygen, the lack of calibration of the data from conductivity–temperature– depth (CTD) profiler mounted sensors is an additional and widespread problem, particularly for oxygen (Olsen et al., 2016). For halogenated transient tracers, uncertainties in standard gas composition, extracted water volume, and purge efficiency typically provide the largest sources of uncertainty. In addition to bias, occasional outliers occur. In rare cases poor precision – many multiples worse than that expected with current measurement techniques – can render a set of data of limited use. GLODAP deals with these issues by presenting the data in a uniform format, including any metadata either publicly available or submitted by the data originator, and by subjecting the data to rigorous primary and secondary quality control assessments, focusing on precision and consistency, respectively. The secondary quality control focuses on deep data, in which natural variability is minimal. Adjustments are applied to the data to minimize cases of bias that could be confidently established relative to the measurement precision for the variables and cruises considered. Key metadata are provided in the header of each data file, and original unadjusted data along with full cruise reports submitted by the data providers (where available) are accessible through the GLODAPv2 cruise summary table hosted by the Ocean Carbon and Acidification Data System (OCADS) at the National Oceanographic and Atmospheric Administration (NOAA) National Centers for Environmental Information (NCEI) (https://www.ncei.noaa. gov/access/ocean-carbon-acidification-data-system/oceans/ GLODAPv2_2022/cruise_table_v2022.html, last access: 15 August 2022). This most recent GLODAPv2.2022 data product builds on earlier synthesis efforts for biogeochemical data obtained from research cruises, namely, GLODAPv1.1 (Key et al., 2004; Sabine et al., 2005), Carbon dioxide in the Atlantic Ocean (CARINA) (Key et al., 2010), Pacific Ocean Interior Carbon (PACIFICA) (Suzuki et al., 2013), and notably GLODAPv2 (Olsen et al., 2016). GLODAPv1.1 combined data from 115 cruises with biogeochemical measurements from the global ocean. The vast majority of these were the sections covered during the World Ocean Circulation Experiment and the Joint Global Ocean Flux Study (WOCE/JGOFS) in the 1990s, but data from important “historical” cruises were also included, such as from the Geochemical Ocean Sections Study (GEOSECS), Transient Traces in the Ocean (TTO), and South Atlantic Ventilation Experiment (SAVE). GLOhttps://doi.org/10.5194/essd-14-5543-2022 Earth Syst. Sci. Data, 14, 5543–5572, 2022 5552 S. K. Lauvset et al.: GLODAPv2.2022 Table 4. Summary of salinity and oxygen calibration needs and actions; number of cruises with each of the scenarios identified. Case Description Salinity Oxygen 1 No data are available: no action needed. 0 7 2 No bottle values are available: use CTD values. 58 30 3 No CTD values are available: use bottle values. 0 19 4 Too few data of both types are available for comparison, and >80% of the records have bottle values: use bottle values. 0 0 5 The CTD values do not deviate significantly from bottle values: replace missing bottle values with CTD values. 38 37 6 The CTD values deviate significantly from bottle values: calibrate CTD values using linear fit and replace missing bottle values with calibrated CTD values. 0 1 7 The CTD values deviate significantly from bottle values, and no good linear fit can be obtained for the cruise: use bottle values and discard CTD values. 0 2 Figure 3. Example crossover figure for silicate for cruises 49UF20190207 (blue) and 49RY20110515 (red), as was generated during the crossover analysis. Panel (a) shows all station positions for the two cruises, and (b) shows the specific stations used for the crossover analysis. Panel (d) shows the data of silicate (µmol kg−1) below the upper depth limit (in this case 2000 dbar) versus potential density anomaly referenced to 4000 dbar as points and the interpolated profiles as lines. Non-interpolated data either did not meet minimum depth separation requirements (Table 4 in Key et al., 2010) or are the deepest sampling depth. The interpolation does not extrapolate. Panel (e) shows the mean silicate difference profile (black, dots) with its standard deviation, as well as also the weighted mean offset (straight red lines) and weighted standard deviation. Summary statistics are provided in (c). 3.2.3 pH scale conversion and quality control Altogether 60 of the 96 new cruises included measured, spectrophotometric pH data, and only one required an adjustment (Sect. 4). We also excluded (flag −777) pH on one cruise as a result of the QC work. All except one cruise reported pH data on the total scale and at 25 ◦C. For the one cruise reporting pH on the seawater scale the data were converted following established routines (Olsen et al., 2020). For details on scale and temperature conversions in previous versions of GLODAPv2, we refer to Olsen et al. (2020). In contrast to quality control of pH data in GLODAPv2 (Olsen et al., 2016), the evaluation of the internal consistency of CO2system variEarth Syst. Sci. Data, 14, 5543–5572, 2022 https://doi.org/10.5194/essd-14-5543-2022 S. K. Lauvset et al.: GLODAPv2.2022 5553 Figure 4. Example summary figure for silicate crossovers for 49UF20190207 versus the cruises in GLODAPv2.2021 (with cruise EXPOCODE listed on the xaxis sorted according to the year the cruise was conducted). The black dots and vertical error bars show the weighted mean offset and standard deviation for each crossover (as a ratio). The weighted mean and standard deviation of all these offsets are shown in the red lines and are 1.01 ±0.00. The dashed black lines are the reference line for a ±2 % offset. ables has not been used for the secondary quality control of the pH data in the GLODAPv2 updates of 2020 and onwards. For the 60 new cruises with pH in GLODAPv2.2022 only crossover analysis was used, supplemented by CONTENT and CANYON-B comparisons (Sect. 3.2.4). Recent literature has demonstrated that internal consistency evaluation procedures are subject to errors owing to an incomplete understanding of the thermodynamic constants, major ion contents, measurement biases, and potential contribution of organic compounds or other unknown protolytes to alkalinity. These complications lead to pH-dependent offsets in calculated pH compared with cruise spectrophotometric pH measurements (Álvarez et al., 2020; Carter et al., 2018; Fong and Dickson, 2019; Takeshita et al., 2020). The pH-dependent offsets may be interpreted as biases and generate false corrections (Álvarez et al., 2020; García-Ibáñez et al., 2022). The offsets are particularly strong at pH levels below 7.7, where calculated and measured pH values are different by on average between 0.01 and 0.02. For the North Pacific this is a problem as pH values below 7.7 can occur at the depths used during the QC (>1500 dbar for this region; Olsen et al., 2016). Since any correction, which may be an artifact, would be applied to the full profiles, we use a minimum adjustment of 0.02 for the North Pacific pH data in the merged product files. Elsewhere, the inconsistencies that may have arisen are smaller, since deep pH is typically higher than 7.7 (Lauvset et al., 2020), and at such levels the difference between calculated and measured pH is less than 0.01 on average (Álvarez et al., 2020; Carter et al., 2018). Outside the North Pacific, we believe that the pH data are consistent to within 0.01. Avoiding CO2chemistry internal consistency considerations for these intermediate products helps to reduce the problem, but since the reference dataset (as also used for the generation of the CANYON-B and CONTENT algorithms) may have these issues, a future full re-evaluation, envisioned for GLODAPv3, is needed to address the problem completely. 3.2.4 CANYON-B and CONTENT analyses CANYON-B and CONTENT (Bittig et al., 2018) were used to support decisions regarding the application of adjustments (or not). CANYON-B is a neural network for estimating nutrients and seawater CO2chemistry variables from temperature, salinity, and oxygen content. CONTENT additionally considers the consistency among the estimated CO2chemistry variables to further refine them. These approaches were developed using the data included in the GLODAPv2 data product (i.e., the 2016 version without any more recent updates). Their advantage compared to crossover analyses for evaluating consistency among cruise data is that effects of water mass changes on ocean properties are represented in the nonlinear relationships in the underlying neural network. For example, if elevated nutrient values measured on a cruise are not due to a measurement bias but actual aging of the water masses that have been sampled and as such accompanied by a decrease in oxygen content, the measured values and the CANYON-B estimates are likely to be similar. Vice https://doi.org/10.5194/essd-14-5543-2022 Earth Syst. Sci. Data, 14, 5543–5572, 2022 5554 S. K. Lauvset et al.: GLODAPv2.2022 versa, if the nutrient values are biased, the measured values and CANYON-B predictions will be dissimilar. Used in the correct way and with caution this tool is a powerful supplement to the traditional crossover analyses which form the basis of our analyses. Specifically, we gave no weight to comparisons in which the crossover analyses had suggested that the salinity and/or O2data were biased, as this would lead to error in the predicted values. We also considered the uncertainties of the CANYON-B and CONTENT estimates. These uncertainties are determined for each predicted value, and for each comparison the ratio of the difference (between measured and predicted values) to the local uncertainty was used to gauge the comparability. As an example, the CANYON-B and CONTENT analyses of the data obtained for 49UF20190207 are presented in Fig. 5. The CANYON-B and CONTENT results confirmed the crossover comparisons for silicate discussed in Sect. 3.2.2 showing an inconsistency of 1.01. For the other variables, the inconsistencies are low and agree with the crossover results (not shown here but results can be accessed through the adjustment table). Another advantage of the CANYON-B and CONTENT comparisons is that these procedures provide estimates at the level of individual data points; e.g., pH values are determined for every sampling location and depth where temperature, salinity, and O2data are available. Cases of strong differences between measured and estimated values are always examined. This has helped us to identify primary QC issues for some cruises and variables, for example a case of an inverted pH profile on cruise 32PO20130829, which was identified and amended in GLODAPv2.2020. 3.2.5 Halogenated transient tracers and SF6 For the halogenated transient tracers (CFC-11, CFC-12, CFC-113, and CCl4; CFCs for short), an inspection of surface saturation levels and an evaluation of relationships between the tracers for each cruise were used to identify biases rather than crossover analyses. Crossover analysis is of limited value for these variables given their transient nature and low contents at depth. As for GLODAPv2, the procedures were the same as those applied for CARINA (Jeansson et al., 2010; Steinfeldt et al., 2010). Beginning with GLODAPv2.2022, we have performed secondary quality control for SF6data, as this tracer is increasingly being measured and has proven a valuable addition to CFCs. The procedure is mainly based on comparisons with the quality-controlled CFC-12 data, which are available for all cruises with SF6measurements. We compare the surface saturation of SF6with that of CFC-12 and also consider the correlation between SF6and CFC-12 in the ocean interior. Typically, this relation shows some scatter and does not follow a distinct curve (Fig. 6). However, for a given CFC12 value the SF6content should fall into a certain range, and this range can be estimated by the transit time distribution (TTD; Hall et al., 2002) method. Note that we are not trying to adjust SF6to perfectly correlate with CFC-12 as that would severely decrease the value of SF6as an independent constraint on ocean circulation. We merely confirm that the SF6content is within an allowable range and only apply adjustments if all lines of evidence suggest it is warranted. In GLODAPv2.2022 no adjustment smaller than 10 % has been applied. As TTD, we use an inverse Gaussian function, which can be described by two parameters: the mean age (0) and the width (1) (Hall et al., 2002). Typically, the ratios of 1/ 0 are chosen as a fixed parameter, and 0is varied. Here, we use a range of 0between 0 and 2000 years and two values for 1/ 0: 0.5 and 2. This range of TTD parameters reproduces simultaneous observation of different tracers, like CFC-12 and SF6, when calculating the tracer contents from the TTD and the atmospheric mixing ratio (Steinfeldt et al., 2009). Typically, for the same CFC-12 value derived from the TTD, the corresponding SF6value increases with the 1/ 0 ratio of the TTD, and it also increases with decreasing saturation (α). As range for the expected SF6to CFC-12 relation we use the TTD with 1/ 0 =0.5 and α=1 as the lower boundary and the TTD with 1/ 0 =0.5 and 80% saturation as the upper boundary. In some cases, like deep water formation or an ice-covered region, the tracer saturation might be lower, as the minimum of 65 % from Steinfeldt et al. (2009) indicates, but the majority of the data is actually located between our assumed lower and upper boundaries (see results for cruise 096U20160426 in Fig. 6). A few exceptions are found for cruises in the Southern Ocean, as has already been shown in Stöven et al. (2015). Note that in 1996, a SF6release experiment was performed in the Greenland Sea (Watson et al., 1999). This leads to a large excess of SF6compared to CFC12 in the Nordic Seas, which is clearly visible in our analyses and hampers the quality control of the SF6data in this region. 3.3 Merged product generation The merged product file for GLODAPv2.2022 was created by updating cruises and correcting known issues in the GLODAPv2.2021 merged file and then appending a merged and bias-corrected file containing the 96 new cruises – sorted according to EXPOCODE, station, and pressure – to this updated GLODAPv2.2021 file. GLODAP cruise numbers were assigned consecutively, starting from 4001, so they can be distinguished from the GLODAPv2.2021 cruises, which ended at 3043. The merging was otherwise performed following the procedures used for previous GLODAP versions (Olsen et al., 2019, 2020; Lauvset et al., 2021). 3.3.1 Updates and corrections for GLODAPv2.2021 For GLODAPv2.2022 we made several updates to cruises included in GLODAPv2.2021 (and earlier versions). The major updates were (i) to perform secondary quality control on all Earth Syst. Sci. Data, 14, 5543–5572, 2022 https://doi.org/10.5194/essd-14-5543-2022 S. K. Lauvset et al.: GLODAPv2.2022 5555 Figure 5. Example summary figure for CANYON-B and CONTENT analyses for 49UF20190207. Any data from regions where CONTENT and CANYON-B were not trained are excluded. The top row shows the nutrients and the bottom row the seawater CO2chemistry variables. All are shown versus sampling pressure (dbar), and the unit is micromoles per kilogram (µmolkg−1) for all except pH, which is on the total scale at in situ temperature and pressure. Black dots (which to a large extent are hidden by the predicted estimates) are the measured data, blue dots are CANYON-B estimates, and red dots are the CONTENT estimates. Each variable has two figure panels. The left shows the depth profile, while the right shows the absolute difference between measured and estimated values divided by the CANYON-B and CONTENT uncertainty estimate, which is determined for each estimated value. These values are used to gauge the comparability; a value below 1 indicates a good match, as it means that the difference between measured and estimated values is less than the uncertainty of the latter. The statistics in each panel are for all data deeper than 500dbar, and N is the number of samples considered. A multiplicative adjustment and its interquartile range are given for the nutrients. For the seawater CO2chemistry variables the numbers in each panel are the median difference between measured and predicted values for CANYON-B (upper) and CONTENT (lower). Both are given with their interquartile range. Table 5. Possible outcomes of the secondary QC and their codes in the online adjustment table. Secondary QC result Code The data are of good quality, are consistent with the rest of the dataset, and should not be adjusted 0/1∗ The data are of good quality but are biased: adjust by adding (for salinity, TCO2, TAlk, pH) or by multiplying (for oxygen, nutrients, CFCs) the adjustment value Adjustment value The data have not been quality controlled, are of uncertain quality, and are suspended until full secondary QC has been carried out −666 The data are of poor quality and excluded from the data product −777 The data appear of good quality, but their nature, being from shallow depths and coastal regions without crossovers or similar, prohibits full secondary QC −888 No data exist for this variable for the cruise in question −999 ∗The value of 0 is used for variables with additive adjustments (salinity, TCO2, TAlk, pH) and 1 for variables with multiplicative adjustments (for oxygen, nutrients, CFCs). This is mathematically equivalent to “no adjustment” in both cases. https://doi.org/10.5194/essd-14-5543-2022 Earth Syst. Sci. Data, 14, 5543–5572, 2022 5556 S. K. Lauvset et al.: GLODAPv2.2022 Figure 6. Example of plots used as basis for the SF6QC procedure. Shown are results for cruises 096U20160426 (left) and 320620170703 (right). (a, e) CFC-12 versus pressure for the specific cruise (red), together with all data from the corresponding GLODAP region (Pacific in this case, grey). (b, f) Same as upper row but for SF6.(c, g) CFC-12 versus SF6(red dots), here the measured contents have been converted into atmospheric mixing ratios. Solid black line: atmospheric time history of CFC-12 versus that of SF6. Dotted lines: CFC-12 versus SF6derived from the TTD method for two different sets of TTD parameters. (d, h) CFC-12 versus SF6saturation for the surface layer (P < 20 dbar), where the numbers give the mean saturation. Table 6. Summary of secondary QC results for the 96 new cruises, in number of cruises per result and per variable. Sal. Oxy. NO3Si PO4TCO2TAlk pH CFC-11 CFC-12 CFC-113 CCl4SF6 With data 96 90 91 92 93 93 94 60 5 6 1 0 2 No data 0 6 5 4 3 3 2 36 91 90 95 96 94 Unadjusteda35 33 33 5 33 35 34 28 3 4 1 0 2 Adjustedb0 2 0 29 1 0 1 1 1 1 0 0 0 −888c61 55 58 58 58 58 59 30 1 1 0 0 0 −666d0 0 0 0 1 0 0 0 0 0 0 0 0 −777e0 0 0 0 0 0 0 1 0 0 0 0 0 aThe data are included in the data product file as is, with a secondary QC flag of 1. bThe adjusted data are included in the data product file with a secondary QC flag of 1. cData appear of good quality but have not been subjected to full secondary QC. They are included in data product with a secondary QC flag of 0. dData are of uncertain quality and suspended until full secondary QC has been carried out; they are excluded from the data product. eData are of poor quality and excluded from the data product. Earth Syst. Sci. Data, 14, 5543–5572, 2022 https://doi.org/10.5194/essd-14-5543-2022 S. K. Lauvset et al.: GLODAPv2.2022 5557 Figure 7. Distribution of applied adjustments for each core variable that received secondary QC, in micromoles per kilogram (µmolkg−1) for TCO2and TAlk and unitless for salinity and pH (but multiplied by 1000 in both cases so a common xaxis can be used), while for the other properties adjustments are given in percent ((adjustment ratio −1)×100). Grey areas depict the initial minimum adjustment limits. The figure includes numbers for data subjected to secondary quality control only. Note also that the y-axis scale is set to render the number of adjustments visible, so the bar showing zero offset (the 0 bar) for each variable is cut off (see Table 6 for these numbers). SF6data (see Sect. 3.2.5) and (ii) to apply small adjustments to TCO2and TAlk data measured on board the RV Knorr in 1994–1995 (EXPOCODES 316N199*; Table A2). These adjustments are derived from offsets in the CRM measurements which were previously reported but never applied to the seawater measurements (Christopher Sabine and Douglas Wallace, personal communication, 2022; Johnson et al., 2002). These offsets are lower than the minimum adjustment limits defined for GLODAP. Applying these adjustments achieves procedural consistency with other CO2chemistry data that are usually corrected for CRM offsets before being subjected to secondary QC. For TAlk the original CRM offsets were derived from Table 2 in Millero et al. (1998), who reported repeated CRM measurements on different titration cells for each cruise. The mean measured CRM value across all cells was calculated and compared to the published reference value for the same batch, and, if necessary, the offsets obtained from multiple CRM batches measured on one cruise were averaged. For TCO2the original CRM offsets were calculated from Table 3 in Johnson et al. (1998), who reported offsets for two measurement systems, which were here averaged. Johnson et al. (2002) report that their TCO2measurements were affected by changes in pipette volumes, which they were able to correct for in the CRM measurements. However, these volume corrections were most likely not applied to the seawater measurements (Douglas Wallace, personal communication, 2022; Johnson et al., 2002), and we therefore use the CRM offsets reported before correcting for the changes in pipette volume. For both TAlk and TCO2we calculate and use the mean CRM offset across all Indian Ocean cruises on the RV Knorr from 1994–1995 (−3.5 µmol kg−1for TAlk and 1.7 µmol kg−1for TCO2) as a bulk adjustment value for the seawater measurements on these cruises. The GLODAP policy for avoiding small adjustments does not apply in this instance because there is a documented reason for the adjustment beyond improving internal consistency of the GLODAPv2 data product. Encouragingly, we also note that applying these adjustments improves the consistency with more recent (post-2000) Indian ocean data in GLODAPv2: for TAlk the mean absolute offset decreased from 2.8µmolkg−1for the unadjusted data to −0.7 µmol kg−1for the adjusted data, while for TCO2the mean absolute offset decreased from −2.3 µmol kg−1for the unadjusted data to −0.6 µmol kg−1 for the adjusted data, respectively. Table A2 in the Appendix shows a list of the cruises that have been updated, as well as what the update consists of. https://doi.org/10.5194/essd-14-5543-2022 Earth Syst. Sci. Data, 14, 5543–5572, 2022 5558 S. K. Lauvset et al.: GLODAPv2.2022 In addition, several minor omissions and errors have been identified and corrected. –An error was corrected in the QC flagging of calculated CO2chemistry variables when fCO2was used as one of the inputs (changed from 1 to 0). –CFC-12 data were added to cruise 06M320150501. –Missing bottle number were added to cruises 29AH20160617 and 29HE20190406. –For cruise 316N19831007 the WOCE flag on TAlk was changed from 2 to 0. –Oxygen concentrations of 49UP19970912 have been adjusted 1.5 % upward. –pH values of 49HG19960807 have been adjusted downward by 0.05. –The time series from Weather Station M in the Norwegian Sea was updated with data from 2008–2021. –In addition to DOIs for all original data files, DOIs for the included data products (CODAP-NA and GEOTRACES) have been added to the product files. –An extra column “G2expocode” has been added, listing the EXPOCODE for each entry. 4 Secondary quality control results and adjustments The secondary QC has five possible outcomes which are summarized in Table 5, along with the corresponding codes that appear in the online adjustment table and that are also occasionally used as shorthand for decisions in the text below. Some cruises were not applicable for full secondary QC. Specifically, in some cases data were too shallow or geographically too isolated for full and conclusive consistency analyses. In other cases, the results of these analyses were inconclusive, but we have no reason to believe that the data in question are of poor quality. A secondary QC flag has been included in the merged product files to enable their identification, with “0” used for variables and cruises not subjected to full secondary QC (corresponding to code −888 in Table 5) and “1” for variables and cruises that were subjected to full secondary QC. The secondary QC flags are assigned per cruise and variable, not for individual data points, and are independent of – and included in addition to – the primary (WOCE) QC flag on individual measurements. For example, interpolated (salinity, oxygen, nutrients) or calculated (TCO2, TAlk, pH) values, which have a primary QC flag of 0, may have a secondary QC flag of 1 if the measured data these values are based on have been subjected to full secondary QC. Conversely, individual data points may have a secondary QC flag of 0 even if their primary QC flag is 2 (good data). Prominent examples for this version are the CODAP-NA data (Jiang et al., 2021), which as a primarily coastal dataset typically has quite shallow sampling depths that rendered conclusive secondary QC impossible. As a consequence, most, but not all, of these data are included with a secondary QC flag of 0. The secondary QC actions for the 13 core variables and the distribution of adjustments applied on the 96 new cruises are summarized in Table 6 and Fig. 7, respectively. For most variables only a small fraction of the data were adjusted: no salinity, TCO2, or nitrate data, 1.1 % TAlk data and phosphate data, 2.2 % of oxygen data, and 31 % of silicate data. The large percentage of silicate data requiring adjustment in this version is due to a consistent 1% offset in the silicate data from the Japan Meteorological Agency (JMA) after 2018 (compared to older data from JMA). This offset has been traced to a change in the batch of Merck silicate standard solution used. In GLODAPv2.2022 this offset has been corrected by adjusting the new data (after 2018) to be consistent with the older data. For the CFCs, CFC-11 required adjustment for one out of the five new cruises and CFC12 required adjustment on one out of six new cruises. For the total of 82 cruises with SF6data in GLODAPv2.2022, two cruises (06MT20060712 and 325020080826) could not be subjected to secondary quality control (−888), and five cruises received an upward adjustment (see example for cruise 320620170703 in Fig. 6). The magnitude of the adjustment was calculated using the saturation of CFC-12 as a benchmark. Additionally, for two cruises (49K619990523 and 58GS20090528), the SF6values are out of the TTDderived range, as are the surface saturations. In these cases, the SF6data are discarded (QC flag −777). Of the 96 new cruises in GLODAPv2.2022 only two include SF6, and neither required an adjustment. Overall, the magnitudes of the various adjustments applied are small, and the tendency observed during the production of the three previous updates remains, namely that the large majority of recent cruises are consistent with earlier releases of the GLODAP data product. A total of 60 out of the 96 new cruises included measured pH data, but only one received an adjustment (and one was flagged −777). However, the new crossover and inversion analysis of all pH data in the northwestern Pacific that was planned following the release of GLODAPv2.2020 has not yet been performed. Such an analysis is planned for the next full update of GLODAP, i.e., GLODAPv3. Therefore, the conclusion from GLODAPv2.2020 remains that some caution should be exercised if looking at trends in ocean pH in the northwestern Pacific using GLODAPv2.2022 or earlier versions. For the nutrients, adjustments were applied to maintain consistency with data included in GLODAPv2.2021 and earlier versions. An alternative goal for the adjustments would be maintaining consistency with data from cruises that employed reference materials (RMNS) to ensure accuracy of nutrient analyses. Such a strategy was adopted Earth Syst. Sci. Data, 14, 5543–5572, 2022 https://doi.org/10.5194/essd-14-5543-2022 S. K. Lauvset et al.: GLODAPv2.2022 5559 Figure 8. Magnitude of applied adjustments relative to minimum adjustment limits (Table 3) per decade for the 1085 cruises included in GLODAPv2.2022. by Aoyama (2020) for preparation of the Global Nutrients Dataset 2013 (GND13) and is being considered for GLODAP as well. However, as this would require a re-evaluation of the entire dataset, this will not occur until the next full update of GLODAP. For now, we note the overall agreement between the adjustments applied in these two efforts (Aoyama, 2020) and that most disagreements appear to be related to cases where no adjustments were applied in GLODAP. The improvement in data consistency resulting from the secondary QC process is evaluated by comparing the weighted mean of the absolute offsets for all crossovers before and after the adjustments have been applied. This “consistency improvement” for core variables is presented in Table 7. The data for CFCs were omitted from these analyses for previously discussed reasons (Sect. 3.2.5). Globally, the improvement is modest. Considering the initial data quality, this result was expected. However, this does not imply that the data initially were consistent everywhere. Rather, for some regions and variables there are substantial improvements when the adjustments are applied. For example, oxygen, silicate, and phosphate in the Atlantic Ocean all show a considerable improvement. Table 7. Improvements resulting from quality control of the 96 new cruises per basin and for the global dataset. The values in the table are the weighted mean of the absolute offset of unadjusted and adjusted data versus GLODAPv2.2021. The total number of valid crossovers in the global ocean for the variable in question is n. The values in this table represent the inter-cruise consistency in the GLODAPv2.2022 product. Arctic Atlantic Indian Pacific Global Unadj. Adj. Unadj. Adj. Unadj. Adj. Unadj. Adj. Unadj. Adj. n(global) Sal (×1000) NA ⇒NA 4.6 ⇒4.6 0.7 ⇒0.7 1.2 ⇒1.2 1.3 ⇒1.3 1105 Oxy (%) NA ⇒NA 1.5 ⇒0.8 0.5 ⇒0.5 0.4 ⇒0.4 0.5 ⇒0.4 1064 NO3(%) NA ⇒NA 1.7 ⇒1.7 0.7 ⇒0.7 0.4 ⇒0.4 0.4 ⇒0.4 940 Si (%) NA ⇒NA 3.0 ⇒2.6 0.9 ⇒0.9 1.4 ⇒0.6 1.4 ⇒0.6 916 PO4(%) NA ⇒NA 2.0 ⇒1.1 0.7 ⇒0.7 0.7 ⇒0.7 0.7 ⇒0.7 936 TCO2(µmol kg−1) NA ⇒NA 7.3 ⇒7.3 2.0 ⇒2.0 1.8 ⇒1.8 2.4 ⇒2.4 544 TAlk (µmol kg−1)) NA ⇒NA 4.5 ⇒3.1 5.2 ⇒5.2 1.8 ⇒1.8 1.9 ⇒1.8 515 pH (×1000) NA ⇒NA 11.6 ⇒11.6 NA ⇒NA 5.5 ⇒5.3 5.5 ⇒5.4 462 NA – not available https://doi.org/10.5194/essd-14-5543-2022 Earth Syst. Sci. Data, 14, 5543–5572, 2022 5560 S. K. Lauvset et al.: GLODAPv2.2022 Figure 9. Locations of stations included in the (a) Arctic, (b) Atlantic, (c) Indian, and (d) Pacific ocean product files for the complete GLODAPv2.2022 dataset. The various iterations of GLODAP provide insight into initial data quality covering more than 4 decades. Figure 8 summarizes the applied absolute adjustment magnitude per decade. These distributions are broadly unchanged compared to GLODAPv2.2021 (Fig. 7 in Lauvset et al., 2021). Most TCO2and TAlk data from the 1970s needed an adjustment, but this fraction steadily declines until only a small percentage is adjusted in recent years. This is encouraging and demonstrates the value of standardizing sampling and measurement practices (Dickson et al., 2007), the widespread use of CRMs (Dickson et al., 2003), and instrument automation. The pH adjustment frequency also has a downward trend; however, there remain issues with the pH adjustments, and this is a topic for future development in GLODAP, with the support from the Ocean Carbon & Biogeochemistry (OCB) Ocean Carbonate System Intercomparison Forum (OCSIF, https://www.us-ocb.org/ ocean-carbonate-system-intercomparison-forum/, last access: 27 June 2022) working group (Álvarez et al., 2020). For the nutrients and oxygen, only the phosphate adjustment frequency decreases from decade to decade. However, we do note that the more recent data from the 2010s receive the fewest adjustments. This may reflect recent increased attention that seawater nutrient measurements have received through an operation manual (Becker et al., 2020; Hydes et al., 2010), availability of RMNS (Aoyama et al., 2012; Ota et al., 2010), and the Scientific Committee on Oceanic Research (SCOR) working group no. 147 towards comparability of global oceanic nutrient data (COMPONUT). For silicate, the fraction of cruises receiving adjustments peaks in the 1990s and 2000s. This is related to the 2% offset between US and Japanese cruises in the Pacific Ocean that was revealed during production of GLODAPv2 and discussed in Olsen et al. (2016). For salinity and the halogenated transient tracers, the number of adjusted cruises is small in every decade. 5 Data availability The GLODAPv2.2022 merged and adjusted data product is archived at the OCADS of NOAA NCEI (https://doi.org/10.25921/1f4w-0t92, Lauvset et al., 2022). These data and ancillary information are also available via our web pages and https://www.ncei.noaa.gov/ access/ocean-carbon-acidification-data-system/oceans/ GLODAPv2_2022/ (last access: 15 August 2022). The data are available as comma-separated ascii files (*.csv) and as binary MATLAB files (*.mat) that use the open-source Hierarchical Data Format version 5 (HDF5). The data product is also made available as an Ocean Data View (ODV) file which can be easily explored using the “webODV Explore” online data service (https://explore.webodv.awi.de/, webODV Explore, 2022). Regional subsets are available for the Arctic, Atlantic, Pacific, and Indian oceans. There are no data overlaps between regional subsets, and each cruise exists in only one basin file even if data from that cruise Earth Syst. Sci. Data, 14, 5543–5572, 2022 https://doi.org/10.5194/essd-14-5543-2022 S. K. Lauvset et al.: GLODAPv2.2022 5561 Table 8. Table listing the number of data points in GLODAPv2.2022, as well as the number of data with various combinations of variables. Variables Number of records All core (salinity, oxygen, nitrate, silicate, phosphate, TCO2, TAlk, pH, CFC-11, CFC-12, CFC-113, CCl4, and SF6) 174 All core except SF62029 Salinity, oxygen, nitrate, silicate, phosphate, CFC-11, CFC-12, CFC-113, CCl4, and SF6plus two of TCO2, TAlk, and pH 636 Salinity, oxygen, nitrate, silicate, phosphate, TCO2, TAlk, and pH 168 330 CFC-11, CFC-12, CFC-113, CCl4, and SF6926 At least one transient tracer species or SF6427 913 SF698 951 Two out of the three CO2chemistry core variables (TCO2, TAlk, pH) 448 024 Measured fCO233 844 Salinity, oxygen, nitrate, silicate, and phosphate 861 650 Salinity and oxygen 1 165 389 No salinity 27 906 Total in GLODAPv2.2022 1 381 248 cross basin boundaries. The station locations in each basin file are shown in Fig. 9. The product file variables are listed in Table 1. As well as being included in the .csv and .mat files, lookup tables for matching the EXPOCODE and DOI of a cruise with GLODAP cruise number are provided with the data files. A “known issues document” accompanies the data files and provides an overview of known errors and omissions in the data product files. It is regularly updated, and users are encouraged to inform us whenever any new issues are identified. It is critical that users consult this document whenever the data products are used. All material produced during the secondary QC is available via the online GLODAP adjustment table hosted by GEOMAR, Kiel, Germany, at https://glodapv2-2022.geomar.de/ (GLODAP, 2022a) and can also be accessed through http:// www.glodap.info (GLODAP, 2022b). This is similar in form and function to the GLODAPv2 adjustment table (Olsen et al., 2016) and includes a brief written justification for any adjustments applied. The original cruise files, with updated flags determined during additional primary GLODAP QC, are available through the GLODAPv2.2022 cruise summary table (CST) hosted by OCADS: https://www.ncei.noaa.gov/ access/ocean-carbon-acidification-data-system/oceans/ GLODAPv2_2022/cruise_table_v2022.html (GLODAP, 2022c). Each of these files has been assigned a DOI, which is included in the data product files but not listed here. The CST also provides brief information on each cruise and access to metadata, cruise reports, and its adjustment table entry. While GLODAPv2.2022 is made available without any restrictions, users of the data should adhere to the fair data use principles: for investigations that rely on a particular (set of) cruise(s), recognize the contribution of GLODAP data contributors by at least citing both the cruise DOI and any articles where the data are described, as well as, preferably, contacting principal investigators to explore opportunities for collaboration and co-authorship. To this end, DOIs are provided in the product files, as well as relevant articles and principal investigator names in the cruise summary table. Contacting principal investigators comes with the additional benefit that the principal investigators often possess expert insight into the data and/or specific region under investigation. This can improve scientific quality and promote data sharing. This paper should be cited in any scientific publications that result from usage of the product. Citations provide the most efficient means to track use, which is important for attracting funding to enable the preparation of future updates. 6 Summary GLODAPv2.2022 is an update of GLODAPv2.2021. Data from 96 new cruises have been added to supplement the earlier release and extend temporal coverage by 1 year. GLODAP now includes 48 years, 1972–2021, of global interior ocean biogeochemical data from 1085 cruises. The total number of data records is 1 381 248 (Table 8). Records with measurements for all 13 core variables (salinity, oxygen, nitrate, silicate, phosphate, TCO2, TAlk, pH, CFC-11, CFC-12, CFC-113, CCl4, and SF6) are very rare (174), and requiring only two out of the three core seawater CO2chemistry variables, in addition to all the other core variables, is still very rare with only 636 records (Table 8). A major limiting factor to having all core variables is the simultaneous availability of data for all four transient tracer species and SF6. In GLODAPv2.2022 there are 98 951 records with SF6 data and 427 913 records with at least one transient tracer or SF6. A total of 2 % (27 906) of all data records do not have salinity. There are several reasons for this, the main one being the inability to vertically interpolate due to a separation that is too large between measured samples. Other reasons for missing salinity include salinity not being reported and missing depth or pressure. As for previous versions there is a bias toward summertime in the data in both hemispheres; most data are collected during April through November in the Northern Hemisphere, https://doi.org/10.5194/essd-14-5543-2022 Earth Syst. Sci. Data, 14, 5543–5572, 2022 5568 S. K. Lauvset et al.: GLODAPv2.2022 Note on former version. Former versions of this article were published on 15 August 2016, 25 September 2019, 23 December 2020, and 3 December 2021 and are available at https://doi.org/10.5194/essd8-297-2016, https://doi.org/10.5194/essd-11-14372019, https://doi.org/10.5194/essd-12-3653-2020, and https://doi.org/10.5194/essd-13-5565-2021. Supplement. The supplement related to this article is available online at: https://doi.org/10.5194/essd-14-5543-2022-supplement. Author contributions. SKL and TT led the team that produced this update. RMK, AK, BP, and SDJ compiled the original data files. NL conducted the primary and secondary QC analyses. HCB conducted the CANYON-B and CONTENT analyses. CS manages the adjustment table e-infrastructure. AK maintains the GLODAPv2 web pages at NCEI/OCADS. JDM was responsible for identifying the small offsets in the historical Indian Ocean data. LQJ, RAF, BRC, SRA, and LB conducted CODAP-NA QC efforts prior to ingestion into GLODAP. TT, RS, and EJ performed the secondary QC on all transient tracers. All authors contributed to the interpretation of the secondary QC results and made decisions on whether to apply adjustments. Many conducted ancillary QC analyses. SKL updated the living data manuscript with contributions from all authors. Competing interests. At least one of the (co-)authors is a member of the editorial board of Earth System Science Data. The peerreview process was guided by an independent editor, and the authors also have no other competing interests to declare. Disclaimer. Publisher’s note: Copernicus Publications remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Acknowledgements. GLODAPv2.2022 would not have been possible without the effort of the many scientists who secured funding, dedicated time to collect data, and shared the data that are included. Chief scientists at the various cruises and principal investigators for specific variables are listed in the online cruise summary table. The author team also want to thank the large GLODAP user community for useful input and notification about potential issues in the data products. Such input is invaluable and helps ensure that GLODAP maintains its high quality and consistency over time. This is CICOES and PMEL contribution numbers 2022-1223 and 5414, respectively. This activity is supported by the International Ocean Carbon Coordination Project (IOCCP). The authors thank Christopher Sabine, Douglas Wallace, Ernie Lewis, and Kenneth M. Johnson for advising the author team with respect to additional corrections for the 1994–1995 Indian Ocean data from the RV Knorr. The authors thank the CODAP-NA team, including Dana Greeley, Denis Pierrot, Charles Featherstone, James Hooper, Chris Melrose, Natalie Monacci, Jonathan Sharp, Shawn Shellito, Yuan-Yuan Xu, Alex Kozyr, Robert H. Byrne, Wei-Jun Cai, Jessica Cross, Gregory C. Johnson, Burke Hales, Chris Langdon, Jeremy Mathis, Joe Salisbury, and David W. Townsend for contributing cruise data and participating in the quality control efforts of CODAP-NA and for providing advice on how to perform secondary QC on these data. The authors thank the GEOTRACES data management team for help in identifying and retrieving the data files relevant for GLODAP. Financial support. Nico Lange was funded by EU Horizon 2020 through the EuroSea action (grant agreement 862626). Siv K. Lauvset acknowledges internal strategic funding from NORCE Climate. Leticia Cotrim da Cunha was supported by Prociencia/UERJ 20222024 and CNPq/PQ2 309708/2021-4 grants. Marta Álvarez was supported by IEO RADPROF project. Peter J. Brown was partly funded by the UK Climate Linked Atlantic Sector Science (CLASS) NERC National Capability Long-term Single Centre Science Programme (grant NE/R015953/1). Anton Velo and Fiz F. Pérez were supported by BOCATS2 (PID2019-104279GB-C21) project funded by MCIN/AEI/10.13039/501100011033 and contributing to WATER:iOS CSIC PTI. Funding for Li-Qing Jiang and the CODAPNA development team (Simone R. Alin, Leticia Barbero, Richard A. Feely, Brendan R. Carter) comes from the NOAA Ocean Acidification Program (OAP, project number: OAP 1903-1903) and NOAA National Centers for Environmental Information (NCEI). Brendan R. Carter thanks the Global Ocean Monitoring and Observing (GOMO) program of the National Oceanic and Atmospheric Administration (NOAA) for funding their contributions (project no. 100007298) through the Cooperative Institute for Climate, Ocean, & Ecosystem Studies (CIOCES) under NOAA Cooperative Agreement NA20OAR4320271, contribution no. 2022-2012. Richard A. Feely and Simone R. Alin acknowledge the NOAA GOMO (project no. 100007298) and the NOAA Pacific Marine Environmental Laboratory. Henry C. Bittig gratefully acknowledges financial support by the BONUS INTEGRAL project (grant no. 03F0773A). Bronte Tilbrook was supported through the Australian Antarctic Program Partnership and the Integrated Marine Observing System. Matthew P. Humphreys acknowledges EU Horizon 2020 action SO-CHIC (grant no. 821001). Adam Ulfsbo was supported by the Swedish Research Council FORMAS (grant no. 2018-01398). Jens Daniel Müller acknowledges support from the European Union’s Horizon 2020 research and innovation program under grant agreement no. 821003 (project 4C). Alex Kozyr and Li-Qing Jiang were supported by NOAA grant NA19NES4320002 (Cooperative Institute for Satellite Earth System Studies – CISESS) at the University of Maryland/ESSIC. GLODAP also acknowledge funding from the Initiative and Networking Fund of the Helmholtz Association through the project “Digital Earth” (ZT-0025) and from the United States National Science Foundation grant OCE-2140395 to the Scientific Committee on Oceanic Research (SCOR, United States) for International Ocean Carbon Coordination Project. The contribution of Leticia Barbero was carried out under the auspices of CIMAS and NOAA, cooperative agreement no. NA20OAR4320472. Review statement. This paper was edited by Giuseppe M. R. Manzella and reviewed by two anonymous referees. Earth Syst. Sci. Data, 14, 5543–5572, 2022 https://doi.org/10.5194/essd-14-5543-2022 S. K. Lauvset et al.: GLODAPv2.2022 5569 References Álvarez, M., Fajar, N. M., Carter, B. R., Guallart, E. F., Pérez, F. F., Woosley, R. J., and Murata, A.: Global ocean spectrophotometric pH assessment: consistent inconsistencies, Environ. Sci. Technol., 54, 10977–10988, https://doi.org/10.1021/acs.est.9b06932, 2020. Aoyama, M.: Global certified-reference-materialor referencematerial-scaled nutrient gridded dataset GND13, Earth Syst. Sci. Data, 12, 487–499, https://doi.org/10.5194/essd-12-487-2020, 2020. Aoyama, M., Ota, H., Kimura, M., Kitao, T., Mitsuda, H., Murata, A., and Sato, K.: Current status of homogeneity and stability of the reference materials for nutrients in Seawater, Anal. Sci., 28, 911–916, https://doi.org/10.2116/analsci.28.911, 2012. Becker, M., Andersen, N., Erlenkeuser, H., Humphreys, M. P., Tanhua, T., and Körtzinger, A.: An internally consistent dataset of δ13C-DIC in the North Atlantic Ocean – NAC13v1, Earth Syst. Sci. Data, 8, 559–570, https://doi.org/10.5194/essd-8-559-2016, 2016. Becker, S., Aoyama, M., Woodward, E. M. S., Bakker, K., Coverly, S., Mahaffey, C., and Tanhua, T.: GO-SHIP Repeat Hydrography Nutrient Manual: The Precise and Accurate Determination of Dissolved Inorganic Nutrients in Seawater, Using Continuous Flow Analysis Methods, Front. Mar. Sci., 7, 90 pp., https://doi.org/10.3389/fmars.2020.581790, 2020. Bittig, H. C., Steinhoff, T., Claustre, H., Fiedler, B., Williams, N. L., Sauzède, R., Körtzinger, A., and Gattuso, J.-P.: An alternative to static climatologies: Robust estimation of open ocean CO2variables and nutrient concentrations from T, S, and O2 data using Bayesian Neural Networks, Front. Mar. Sci., 5, 328, https://doi.org/10.3389/fmars.2018.00328, 2018. Bockmon, E. E. and Dickson, A. G.: An inter-laboratory comparison assessing the quality of seawater carbon dioxide measurements, Mar. Chem., 171, 36–43, https://doi.org/10.1016/j.marchem.2015.02.002, 2015. Brakstad, A., Våge, K., Håvik, L., and Moore, G. W. K.: Water Mass Transformation in the Greenland Sea during the Period 1986–2016, J. Phys. Oceanogr., 49, 121–140, https://doi.org/10.1175/JPO-D-17-0273.1, 2019. Carter, B. R., Feely, R. A., Williams, N. L., Dickson, A. G., Fong, M. B., and Takeshita, Y.: Updated methods for global locally interpolated estimation of alkalinity, pH, and nitrate, Limnol. Oceanogr.-Meth., 16, 119–131, https://doi.org/10.1002/lom3.10232, 2018. Cheng, L. J., Trenberth, K. E., Fasullo, J., Boyer, T., Abraham, J., and Zhu, J.: Improved estimates of ocean heat content from 1960 to 2015, Sci. Adv., 3, e1601545, https://doi.org/10.1126/sciadv.1601545, 2017. Cheng, L. J., Abraham, J., Zhu, J., Trenberth, K. E., Fasullo, J., Boyer, T., Locarnini, R., Zhang, B., Yu, F. J., Wan, L. Y., Chen, X. R., Song, X. Z., Liu, Y. L., and Mann, M. E.: Record-setting ocean warmth continued in 2019, Adv. Atmos. Sci, 37, 137–142, https://doi.org/10.1007/s00376-020-9283-7, 2020. Dickson, A. G., Afghan, J. D., and Anderson, G. C.: Reference materials for oceanic CO2analysis: a method for the certification of total alkalinity, Mar. Chem., 80, 185–197, https://doi.org/10.1016/S0304-4203(02)00133-0, 2003. Dickson, A. G., Sabine, C. L., and Christian, J. R.: Guide to Best Practices for Ocean CO2measurements, PICES Special Publication 3, North Pacific Marine Science Organization, 191 pp., 2007. Falck, E. and Olsen, A.: Nordic Seas dissolved oxygen data in CARINA, Earth Syst. Sci. Data, 2, 123–131, https://doi.org/10.5194/essd-2-123-2010, 2010. Fong, M. B., and Dickson, A. G.: Insights from GO-SHIP hydrography data into the thermodynamic consistency of CO2 system measurements in seawater, Mar. Chem., 211, 52–63, https://doi.org/10.1016/j.marchem.2019.03.006, 2019. Friedlingstein, P., Jones, M. W., O’Sullivan, M., Andrew, R. M., Hauck, J., Peters, G. P., Peters, W., Pongratz, J., Sitch, S., Le Quéré, C., Bakker, D. C. E., Canadell, J. G., Ciais, P., Jackson, R. B., Anthoni, P., Barbero, L., Bastos, A., Bastrikov, V., Becker, M., Bopp, L., Buitenhuis, E., Chandra, N., Chevallier, F., Chini, L. P., Currie, K. I., Feely, R. A., Gehlen, M., Gilfillan, D., Gkritzalis, T., Goll, D. S., Gruber, N., Gutekunst, S., Harris, I., Haverd, V., Houghton, R. A., Hurtt, G., Ilyina, T., Jain, A. K., Joetzjer, E., Kaplan, J. O., Kato, E., Klein Goldewijk, K., Korsbakken, J. I., Landschützer, P., Lauvset, S. K., Lefèvre, N., Lenton, A., Lienert, S., Lombardozzi, D., Marland, G., McGuire, P. C., Melton, J. R., Metzl, N., Munro, D. R., Nabel, J. E. M. S., Nakaoka, S.-I., Neill, C., Omar, A. M., Ono, T., Peregon, A., Pierrot, D., Poulter, B., Rehder, G., Resplandy, L., Robertson, E., Rödenbeck, C., Séférian, R., Schwinger, J., Smith, N., Tans, P. P., Tian, H., Tilbrook, B., Tubiello, F. N., van der Werf, G. R., Wiltshire, A. J., and Zaehle, S.: Global Carbon Budget 2019, Earth Syst. Sci. Data, 11, 1783–1838, https://doi.org/10.5194/essd-111783-2019, 2019. Fröb, F., Olsen, A., Våge, K., Moore, G. W. K., Yashayaev, I., Jeansson, E., and Rajasakaren, B.: Irminger Sea deep convection injects oxygen and anthropogenic carbon to the ocean interior, Nat. Commun., 7, 13244, https://doi.org/10.1038/ncomms13244, 2016. García-Ibáñez, M. I., Takeshita, Y., Guallart, E. F., Fajar, N. M., Pierrot, D., Pérez, F. F., Cai, W.-J., and Álvarez, M.: Gaining insights into the seawater carbonate system using discrete fCO2measurements, Mar. Chem., 245, 104150, https://doi.org/10.1016/j.marchem.2022.104150, 2022. GLODAP: GLODAPv2.2022 Adjustments, https://glodapv2-2022. geomar.de/, last access: 9 December 2022a. GLODAP: A uniformly calibrated open ocean data product of inorganic and carbon-relevant variables, http://www.glodap.info, last access: 9 December 2022b. GLODAP: Original Cruise Information and Data Table for GLODAPv2.2022, https://www.ncei.noaa.gov/access/ ocean-carbon-acidification-data-system/oceans/GLODAPv2_ 2022/cruise_table_v2022.html, last access: 9 December 2022c. Gordon, A. L.: Deep Antarctic covection west of Maud Rise, J. Phys. Oceanogr., 8, 600–612, https://doi.org/10.1175/15200485(1978)008<0600:DACWOM>2.0.CO;2, 1978. Gruber, N., Clement, D., Carter, B. R., Feely, R. A., van Heuven, S., Hoppema, M., Ishii, M., Key, R. M., Kozyr, A., Lauvset, S. K., Lo Monaco, C., Mathis, J. T., Murata, A., Olsen, A., Perez, F. F., Sabine, C. L., Tanhua, T., and Wanninkhof, R.: The oceanic sink for anthropogenic CO2from 1994 to 2007, Science, 363, 1193-1199, https://doi.org/10.1126/science.aau5153, 2019. https://doi.org/10.5194/essd-14-5543-2022 Earth Syst. Sci. Data, 14, 5543–5572, 2022 5570 S. K. Lauvset et al.: GLODAPv2.2022 Hall, T. M., Haine, T. W. N., and Waugh, D. W.: Inferring the concentration of anthropogenic carbon in the ocean from tracers, Global Biogeochem. Cy., 16, GB1131, https://doi.org/10.1029/2001gb001835, 2002. Hood, E. M., Sabine, C. L., and Sloyan, B. M. (Eds.): The GO-SHIP hydrography manual: A collection of expert reports and guidelines, IOCCP Report Number 14, ICPO Publication Series Number 134, http://www.go-ship.org/HydroMan.html (last access: 1 July 2022), 2010. Hydes, D. J., Aoyama, A., Aminot, A., Bakker, K., Becker, S., Coverly, S., Daniel, A., Dickson, A. G., Grosso, O., Kerouel, R., van Ooijen, J., Sato, K., Tanhua, T., Woodward, E. M. S., and Zhang, J.-Z.: Determination of dissolved nutrients in seawater with high precision and intercomparability using gas-segmented continuous flow analysers, in: The GO SHIP Repeat Hydrography Manual: A Collection of Expert Reports and Guidelines, edited by: Hood, E. M., Sabine, C., and Sloyan, B. M., IOCCP Report Number 14, ICPO Publication Series Number 134, ICPO, http: //www.go-ship.org/HydroMan.html (last access: 1 July 2022), 2010. Jeansson, E., Olsson, K. A., Tanhua, T., and Bullister, J. L.: Nordic Seas and Arctic Ocean CFC data in CARINA, Earth Syst. Sci. Data, 2, 79–97, https://doi.org/10.5194/essd-2-79-2010, 2010. Jenkins, W. J., Doney, S. C., Fendrock, M., Fine, R., Gamo, T., Jean-Baptiste, P., Key, R., Klein, B., Lupton, J. E., Newton, R., Rhein, M., Roether, W., Sano, Y. J., Schlitzer, R., Schlosser, P., and Swift, J.: A comprehensive global oceanic dataset of helium isotope and tritium measurements, Earth Syst. Sci. Data, 11, 441454, https://doi.org/10.5194/essd-11-441-2019, 2019. Jiang, L.-Q., Feely, R. A., Wanninkhof, R., Greeley, D., Barbero, L., Alin, S., Carter, B. R., Pierrot, D., Featherstone, C., Hooper, J., Melrose, C., Monacci, N., Sharp, J. D., Shellito, S., Xu, Y.-Y., Kozyr, A., Byrne, R. H., Cai, W.-J., Cross, J., Johnson, G. C., Hales, B., Langdon, C., Mathis, J., Salisbury, J., and Townsend, D. W.: Coastal Ocean Data Analysis Product in North America (CODAP-NA) – an internally consistent data product for discrete inorganic carbon, oxygen, and nutrients on the North American ocean margins, Earth Syst. Sci. Data, 13, 2777–2799, https://doi.org/10.5194/essd-13-2777-2021, 2021. Jiang, L.-Q., Pierrot, D., Wanninkhof, R., Feely, R. A., Tilbrook, B., Alin, S., Barbero, L., Byrne, R. H., Carter, B. R., Dickson, A. G., Gattuso, J.-P., Greeley, D., Hoppema, M., Humphreys, M. P., Karstensen, J., Lange, N., Lauvset, S. K., Lewis, E. R., Olsen, A., Pérez, F. F., Sabine, C., Sharp, J. D., Tanhua, T., Trull, T. W., Velo, A., Allegra, A. J., Barker, P., Burger, E., Cai, W.-J., Chen, C.-T. A., Cross, J., Garcia, H., Hernandez-Ayon, J. M., Hu, X., Kozyr, A., Langdon, C., Lee, K., Salisbury, J., Wang, Z. A., and Xue, L.: Best Practice Data Standards for Discrete Chemical Oceanographic Observations, Front. Mar. Sci., 8, https://doi.org/10.3389/fmars.2021.705638, 2022. Johnson, K. M., Dickson, A. G., Eischeid, G., Goyet, C., Guenther, P., Key, R. M., Millero, F. J., Purkerson, D., Sabine, C. L., Schottle, R. G., Wallace, D. W. R., Wilke, R. J., and Winn, C. D.: Coulometric total carbon dioxide analysis for marine studies: assessment of the quality of total inorganic carbon measurements made during the US Indian Ocean CO2Survey 1994– 1996, Mar. Chem., 63, 21–37, https://doi.org/10.1016/S03044203(98)00048-6, 1998. Johnson, K. M., Dickson, A. G., Eischeid, G., Goyet, C., Guenther, P. R., Key, R. M., Lee, K., Lewis, E. R., Millero, F. J., Purkerson, D., Sabine, C. L., Schottle, R. G.,l, Wallace, D. W. R., Wilke, R. J., and Winn, C. D.: Carbon Dioxide, Hydrographic and Chemical Data Obtained During the Nine RIV Knorr Cruises Comprising the Indian Ocean CO2Survey (WOCE Sections I8SI9S, I9N, I8NI5E, /3, I5WI4, I7N, II, lIO, and 12, 1 December, I994– January 22, 1996), edited by: Kozyr, A., ORNUCDIAC-138, NDP-080, Carbon Dioxide Information Analysis Center, Oak Ridge National Laboratory, U.S. Department of Energy, Oak Ridge, Tennessee, 59 pp., 2002. Joyce, T. and Corry, C.: Chapter 4. Hydrographic Data Formats, in Requirements for WOCE Hydrographic Programme Data Reporting, WOCE Hydrographic Programme Office. Woods Hole, MA: Woods Hole Oceanographic Institution, 1994. Jutterström, S., Anderson, L. G., Bates, N. R., Bellerby, R., Johannessen, T., Jones, E. P., Key, R. M., Lin, X., Olsen, A., and Omar, A. M.: Arctic Ocean data in CARINA, Earth Syst. Sci. Data, 2, 71–78, https://doi.org/10.5194/essd-2-71-2010, 2010. Key, R. M., Kozyr, A., Sabine, C. L., Lee, K., Wanninkhof, R., Bullister, J. L., Feely, R. A., Millero, F. J., Mordy, C., and Peng, T. H.: A global ocean carbon climatology: Results from Global Data Analysis Project (GLODAP), Global Biogeochem. Cy., 18, GB4031, https://doi.org/10.1029/2004GB002247, 2004. Key, R. M., Tanhua, T., Olsen, A., Hoppema, M., Jutterström, S., Schirnick, C., van Heuven, S., Kozyr, A., Lin, X., Velo, A., Wallace, D. W. R., and Mintrop, L.: The CARINA data synthesis project: introduction and overview, Earth Syst. Sci. Data, 2, 105– 121, https://doi.org/10.5194/essd-2-105-2010, 2010. Lauvset, S. K. and Tanhua, T.: A toolbox for secondary quality control on ocean chemistry and hydrographic data, Limnol. Oceanogr.-Meth., 13, 601–608, https://doi.org/10.1002/lom3.10050, 2015. Lauvset, S. K., Key, R. M., Olsen, A., van Heuven, S., Velo, A., Lin, X., Schirnick, C., Kozyr, A., Tanhua, T., Hoppema, M., Jutterström, S., Steinfeldt, R., Jeansson, E., Ishii, M., Perez, F. F., Suzuki, T., and Watelet, S.: A new global interior ocean mapped climatology: the 1° × 1° GLODAP version 2, Earth Syst. Sci. Data, 8, 325–340, https://doi.org/10.5194/essd-8-325-2016, 2016. Lauvset, S. K., Carter, B. R., Perez, F. F., Jiang, L.-Q., Feely, R. A., Velo, A., and Olsen, A.: Processes Driving Global Interior Ocean pH Distribution, Global Biogeochem. Cy., 34, e2019GB006229, https://doi.org/10.1029/2019gb006229, 2020. Lauvset, S. K., Lange, N., Tanhua, T., Bittig, H. C., Olsen, A., Kozyr, A., Álvarez, M., Becker, S., Brown, P. J., Carter, B. R., Cotrim da Cunha, L., Feely, R. A., van Heuven, S., Hoppema, M., Ishii, M., Jeansson, E., Jutterström, S., Jones, S. D., Karlsen, M. K., Lo Monaco, C., Michaelis, P., Murata, A., Pérez, F. F., Pfeil, B., Schirnick, C., Steinfeldt, R., Suzuki, T., Tilbrook, B., Velo, A., Wanninkhof, R., Woosley, R. J., and Key, R. M.: An updated version of the global interior ocean biogeochemical data product, GLODAPv2.2021, Earth Syst. Sci. Data, 13, 5565–5589, https://doi.org/10.5194/essd-13-5565-2021, 2021. Lauvset, S. K., Lange, N., Tanhua, T., Bittig, H. C., Olsen, A., Kozyr, A., Alin, S. R., Álvarez, M., Azetsu-Scott, K., Barbero, L., Becker, S., Brown, P. J., Carter, B. R., Cotrim da Cunha, L., Feely, R. A., Hoppema, M., Humphreys, M. P., Ishii, M., Jeansson, E., Jiang, L.-Q., Jones, S. D., Lo Monaco, C., MuEarth Syst. Sci. Data, 14, 5543–5572, 2022 https://doi.org/10.5194/essd-14-5543-2022 S. K. Lauvset et al.: GLODAPv2.2022 5571 rata, A., Müller, J. D., Pérez, F. F., Pfeil, B., Schirnick, C., Steinfeldt, R., Suzuki, T., Tilbrook, B., Ulfsbo, A., Velo, A., Woosley, R. J., and Key, R. M.: Global Ocean Data Analysis Project version 2.2022 (GLODAPv2.2022) (NCEI Accession 0257247), NOAA National Centers for Environmental Information [data set], https://doi.org/10.25921/1f4w-0t92, 2022. Millero, F. J., Dickson, A. G., Eischeid, G., Goyet, C., Guenther, P., Johnson, K. M., Key, R. M., Lee, K., Purkerson, D., Sabine, C. L., Schottle, R. G., Wallace, D. W. R., Lewis, E., and Winn, C. D.: Assessment of the quality of the shipboard measurements of total alkalinity on the WOCE Hydrographic Program Indian Ocean CO2survey cruises 1994–1996, Mar. Chem., 63, 9–20, https://doi.org/10.1016/S0304-4203(98)00043-7, 1998. National Geophysical Data Center/NESDIS/NOAA/U.S. Department of Commerce: ETOPO2, Global 2 Arc-minute Ocean Depth and Land Elevation from the US National Geophysical Data Center (NGDC), Research Data Archive at the National Center for Atmospheric Research, Computational and Information Systems Laboratory [data set], https://doi.org/10.5065/D6668B75, 2006. Olsen, A., Key, R. M., van Heuven, S., Lauvset, S. K., Velo, A., Lin, X., Schirnick, C., Kozyr, A., Tanhua, T., Hoppema, M., Jutterström, S., Steinfeldt, R., Jeansson, E., Ishii, M., Pérez, F. F., and Suzuki, T.: The Global Ocean Data Analysis Project version 2 (GLODAPv2) – an internally consistent data product for the world ocean, Earth Syst. Sci. Data, 8, 297–323, https://doi.org/10.5194/essd-8-297-2016, 2016. Olsen, A., Lange, N., Key, R. M., Tanhua, T., Álvarez, M., Becker, S., Bittig, H. C., Carter, B. R., Cotrim da Cunha, L., Feely, R. A., van Heuven, S., Hoppema, M., Ishii, M., Jeansson, E., Jones, S. D., Jutterström, S., Karlsen, M. K., Kozyr, A., Lauvset, S. K., Lo Monaco, C., Murata, A., Pérez, F. F., Pfeil, B., Schirnick, C., Steinfeldt, R., Suzuki, T., Telszewski, M., Tilbrook, B., Velo, A., and Wanninkhof, R.: GLODAPv2.2019 – an update of GLODAPv2, Earth Syst. Sci. Data, 11, 1437–1461, https://doi.org/10.5194/essd-11-1437-2019, 2019. Olsen, A., Lange, N., Key, R. M., Tanhua, T., Bittig, H. C., Kozyr, A., Álvarez, M., Azetsu-Scott, K., Becker, S., Brown, P. J., Carter, B. R., Cotrim da Cunha, L., Feely, R. A., van Heuven, S., Hoppema, M., Ishii, M., Jeansson, E., Jutterström, S., Landa, C. S., Lauvset, S. K., Michaelis, P., Murata, A., Pérez, F. F., Pfeil, B., Schirnick, C., Steinfeldt, R., Suzuki, T., Tilbrook, B., Velo, A., Wanninkhof, R., and Woosley, R. J.: An updated version of the global interior ocean biogeochemical data product, GLODAPv2.2020, Earth Syst. Sci. Data, 12, 3653–3678, https://doi.org/10.5194/essd-12-3653-2020, 2020. Oka, E., Katsura, S., Inoue, H., Kojima, A., Kitamoto, M., Nakano, T., and Suga, T.: Long-term change and variation of salinity in the western North Pacific subtropical gyre revealed by 50-year long observations along 137 degrees E, J. Oceanogr., 73, 479– 490, https://doi.org/10.1007/s10872-017-0416-2, 2017. Oka, E., Ishii, M., Nakano, T., Suga, T., Kouketsu, S., Miyamoto, M., Nakano, H., Qiu, B., Sugimoto, S., and Takatani, Y.: Fifty years of the 137A degrees E repeat hydrographic section in the western North Pacific Ocean, J. Oceanogr., 74, 115–145, https://doi.org/10.1007/s10872-017-0461-x, 2018. Ota, H., Mitsuda, H., Kimura, M., and Kitao, T.: Reference materials for nutrients in seawater: Their development and present homogenity and stability, in: Comparability of nutrients in the world’s oceans, edited by: Aoyama, A., Dickson, A. G., Hydes, D. J., Murata, A., Oh, J. R., Roose, P., and Woodward, E. M. S., Mother Tank, Tsukuba, Japan, 2010. Sabine, C., Key, R. M., Kozyr, A., Feely, R. A., Wanninkhof, R., Millero, F. J., Peng, T.-H., Bullister, J. L., and Lee, K.: Global Ocean Data Analysis Project (GLODAP): Results and Data, ORNL/CDIAC-145, NDP-083, Carbon Dioxide Information Analysis Center, Oak Ridge National Laboratory, U.S. Department of Energy, Oak Ridge, TN, USA, 2005. Sloyan, B. M., Wanninkhof, R., Kramp, M., Johnson, G. C., Talley, L. D., Tanhua, T., McDonagh, E., Cusack, C., O’Rourke, E., McGovern, E., Katsumata, K., Diggs, S., Hummon, J., Ishii, M., Azetsu-Scott, K., Boss, E., Ansorge, I., Perez, F. F., Mercier, H., Williams, M. J. M., Anderson, L., Lee, J. H., Murata, A., Kouketsu, S., Jeansson, E., Hoppema, M., and Campos, E.: The Global Ocean Ship-Based Hydrographic Investigations Program (GO-SHIP): A Platform for Integrated Multidisciplinary Ocean Science, Front. Mar. Sci., 6, https://doi.org/10.3389/fmars.2019.00445, 2019. Steinfeldt, R., Rhein, M., Bullister, J. L., and Tanhua, T.: Inventory changes in anthropogenic carbon from 1997-2003 in the Atlantic Ocean between 20◦S and 65◦N, Global Biogeochem. Cy., 23, GB3010, 10.1029/2008GB003311, 2009. Steinfeldt, R., Tanhua, T., Bullister, J. L., Key, R. M., Rhein, M., and Köhler, J.: Atlantic CFC data in CARINA, Earth Syst. Sci. Data, 2, 1–15, https://doi.org/10.5194/essd-2-1-2010, 2010. Stöven, T., Tanhua, T., Hoppema, M., and Bullister, J. L.: Perspectives of transient tracer applications and limiting cases, Ocean Sci., 11, 699–718, https://doi.org/10.5194/os-11-6992015, 2015. Suzuki, T., Ishii, M., Aoyama, A., Christian, J. R., Enyo, K., Kawano, T., Key, R. M., Kosugi, N., Kozyr, A., Miller, L. A., Murata, A., Nakano, T., Ono, T., Saino, T., Sasaki, K., Sasano, D., Takatani, Y., Wakita, M., and Sabine, C.: PACIFICA Data Synthesis Project, ORNL/CDIAC-159, NDP-092, Carbon Dioxide Information Analysis Center, Oak Ridge National Laboratory, U.S. Department of Energy, Oak Ridge, TN, USA, https://doi.org/10.3334/CDIAC/OTG.PACIFICA_NDP092, 2013. Swift, J.: Reference-quality water sample data: Notes on aquisition, record keeping, and evaluation, in: The GO-SHIP Repeat Hydrography Manual: A Collection of Expert Reports and Guidelines, edited by: Hood, E. M., Sabine, C., and Sloyan, B. M., IOCCP Report Number 14, ICPO Publication Series Number 134, 2010. Swift, J. and Diggs, S. C.: Description of WHP exchange format for CTD/Hydrographic data, CLIVAR and Carbon Hydrographic Data Office, UCSD Scripps Institution of Oceanography, San Diego, Ca, US, 2008. Takeshita, Y., Johnson, K. S., Coletti, L. J., Jannasch, H. W., Walz, P. M., and Warren, J. K.: Assessment of pH dependent errors in spectrophotometric pH measurements of seawater, Mar. Chem., 223, 103801, https://doi.org/10.1016/j.marchem.2020.103801, 2020. Talley, L. D., Feely, R. A., Sloyan, B. M., Wanninkhof, R., Baringer, M. O., Bullister, J. L., Carlson, C. A., Doney, S. C., Fine, R. A., Firing, E., Gruber, N., Hansell, D. A., Ishii, M., Johnson, G. C., Katsumata, K., Key, R. M., Kramp, M., Langdon, C., Macdonald, A. M., Mathis, J. T., McDonagh, E. L., Mecking, S., https://doi.org/10.5194/essd-14-5543-2022 Earth Syst. Sci. Data, 14, 5543–5572, 2022 5572 S. K. Lauvset et al.: GLODAPv2.2022 Millero, F. J., Mordy, C. W., Nakano, T., Sabine, C. L., Smethie, W. M., Swift, J. H., Tanhua, T., Thurnherr, A. M., Warner, M. J., and Zhang, J. Z.: Changes in ocean heat, carbon content, and ventilation: A review of the first decade of GO-SHIP global repeat hydrography, Annu. Rev. Mar. Sci., 8, 185–215, https://doi.org/10.1146/annurev-marine-052915-100829, 2016. Tanhua, T., van Heuven, S., Key, R. M., Velo, A., Olsen, A., and Schirnick, C.: Quality control procedures and methods of the CARINA database, Earth Syst. Sci. Data, 2, 35–49, https://doi.org/10.5194/essd-2-35-2010, 2010. Tanhua, T., Lauvset, S. K., Lange, N., Olsen, A., Álvarez, M., Diggs, S., Bittig, H. C., Brown, P. J., Carter, B. R., da Cunha, L. C., Feely, R. A., Hoppema, M., Ishii, M., Jeansson, E., Kozyr, A., Murata, A., Pérez, F. F., Pfeil, B., Schirnick, C., Steinfeldt, R., Telszewski, M., Tilbrook, B., Velo, A., Wanninkhof, R., Burger, E., O’Brien, K., and Key, R. M.: A vision for FAIR ocean data products, Commun. Earth Environ., 2, 136, https://doi.org/10.1038/s43247-021-00209-4, 2021. Velo, A., Cacabelos, J., Lange, N., Perez, F. F., and Tanhua, T.: Ocean Data QC: Software package for quality control of hydrographic sections (v1.4.0). Zenodo [code], https://doi.org/10.5281/zenodo.4532402, 2021. Watson, A. J., Messias, M. J., Fogelqvist, E., Van Scoy, K. A., Johannessen, T., Oliver, K. I. C., Stevens, D. P., Rey, F., Tanhua, T., and Olsson, K. A.: Mixing and convection in the Greenland Sea from a tracer-release experiment, Nature, 401, 902–904, https://doi.org/10.1038/44807, 1999. Weatherall, P., Marks, K. M., Jakobsson, M., Schmitt, T., Tani, S., Arndt, J. E., Rovere, M., Chayes, D., Ferrini, V., and Wigley, R.: A new digital bathymetric model of the world’s oceans, Earth Space Sci., 2, 331–345, https://doi.org/10.1002/2015EA000107, 2015. webODV Explore: https://explore.webodv.awi.de/, last access: 9 December 2022. Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J.-W., da Silva Santos, L. B., Bourne, P. E., Bouwman, J., Brookes, A. J., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., Evelo, C. T., Finkers, R., Gonzalez-Beltran, A., Gray, A. J. G., Groth, P., Goble, C., Grethe, J. S., Heringa, J., ’t Hoen, P. A. C., Hooft, R., Kuhn, T., Kok, R., Kok, J., Lusher, S. J., Martone, M. E., Mons, A., Packer, A. L., Persson, B., Rocca-Serra, P., Roos, M., van Schaik, R., Sansone, S.-A., Schultes, E., Sengstag, T., Slater, T., Strawn, G., Swertz, M. A., Thompson, M., van der Lei, J., van Mulligen, E., Velterop, J., Waagmeester, A., Wittenburg, P., Wolstencroft, K., Zhao, J., and Mons, B.: The FAIR Guiding Principles for scientific data management and stewardship, Sci. Data, 3, 160018, https://doi.org/10.1038/sdata.2016.18, 2016. Yashayaev, I. and Loder, J. W.: Further intensification of deep convection in the Labrador Sea in 2017, Geophys. Res. Lett., 44, 1429–1438, https://doi.org/10.1002/2016GL071668, 2017. Earth Syst. Sci. Data, 14, 5543–5572, 2022 https://doi.org/10.5194/essd-14-5543-2022 87 5 A vision for FAIR ocean data products 5 A vision for FAIR ocean data products 88 COMMENT A vision for FAIR ocean data products Toste Tanhua 1✉, Siv K. Lauvset2, Nico Lange1, Are Olsen3, Marta Álvarez 4, Stephen Diggs5, Henry C. Bittig 6, Peter J. Brown 7, Brendan R. Carter8,9, Leticia Cotrim da Cunha 10, Richard A. Feely9, Mario Hoppema 11, Masao Ishii 12, Emil Jeansson 2, Alex Kozyr13, Akihiko Murata14, Fiz F. Pérez 15, Benjamin Pfeil3, Carsten Schirnick 1, Reiner Steinfeldt 16, Maciej Telszewski17, Bronte Tilbrook 18, Anton Velo 15, Rik Wanninkhof19, Eugene Burger9, Kevin O’Brien8,9 & Robert M. Key20 The ocean is mitigating global warming by absorbing large amounts of excess carbon dioxide from human activities. To quantify and monitor the ocean carbon sink, we need a state-of-the-art data resource that makes data submission and retrieval machine-compatible and efficient. Human activities such as combustion of fossil fuel, land use change, and cement production increased the atmospheric carbon dioxide (CO 2 ) concentration to 418 ppm in April 2021. This level is almost 50% higher than at the beginning of the industrial age. The greenhouse effect of atmospheric CO 2 and other gases has led to significant warming and increased stratification in the ocean, and has consequences for ecosystems and marine ecosystem services. Notably, atmospheric CO 2 concentrations would now be around another 76 ppm higher than current levels1if the ocean had not taken up a significant fraction of our emissions from the atmosphere2. The ocean is one of the largest carbon pools on the planet, second only to the Earth’s crust. The ocean contains about 38,000 Gigatonnes of carbon and thereby dwarfs the cumulative emissions of fossil CO 2 since the Industrial Revolution from fossil fuel combustion (about 440 GtC to 2019) and land-use change (about 210 GtC)1. As such, the accumulation rate of carbon in the surface ocean of about 1 µmol kg−1year−1driven by anthropogenic CO 2 emissions is much smaller than the natural variations in dissolved inorganic carbon content, over a range of 500 µmol kg−1regionally and 100 µmol kg−1seasonally3. Thus, any emission-driven trends in ocean carbon concentrations or changes in biogeochemical cycles are expressed amid large https://doi.org/10.1038/s43247-021-00209-4 OPEN 1GEOMAR Helmholtz Centre for Ocean Research Kiel, Kiel, Germany. 2NORCE Norwegian Research Centre, Bjerknes Centre for Climate Research, Bergen, Norway. 3Geophysical Institute, University of Bergen and Bjerknes Centre for Climate Research, Bergen, Norway. 4Instituto Español de Oceanografía, A Coruña, Spain. 5UC San Diego, Scripps Institution of Oceanographyi, San Diego, CA, USA. 6Leibniz Institute for Baltic Sea Research Warnemunde, Rostock, Germany. 7National Oceanography Centre, Southampton, UK. 8Cooperative Institute for Climate, Ocean, and Ecosystem Studies, University Washington, Seattle, WA, USA. 9Pacific Marine Environmental Laboratory, National Oceanic and Atmospheric Administration, Seattle, WA, USA. 10 Faculdade de Oceanografia/PPG-Oceanografia, Universidade do Estado do Rio de Janeiro, Rio de Janeiro (RJ), Brazil. 11 Alfred Wegener Institute Helmholtz Centre for Polar and Marine Research, Bremerhaven, Germany. 12 Meteorological Research Institute, Japan Meteorological Agency, Tsukuba, Japan. 13 NOAA National Centers for Environmental Information, Silver Spring, MD, USA. 14 Research Institute for Global Change, Japan Agency for Marine-Earth Science and Technology, Yokosuka, Japan. 15 Instituto de Investigaciones Marinas, IIM –CSIC, Vigo, Spain. 16 University of Bremen, Institute of Environmental Physics, Bremen, Germany. 17 International Ocean Carbon Coordination Project, Institute of Oceanology of Polish Academy of Sciences, Sopot, Poland. 18 CSIRO Oceans and Atmosphere and Australian Antarctic Program Partnership, Hobart, Australia. 19 Atlantic Oceanographic and Meteorological Laboratory, National Oceanic and Atmospheric Administration, Miami, FL, USA. 20 Atmospheric and Oceanic Sciences, Princeton University, Princeton, NJ, USA. ✉email: [email protected] COMMUNICATIONS EARTH & ENVIRONMENT | (2021) 2:136 | https://doi.org/10.1038/s43247-021-00209-4 | www.nature.com/commsenv 1 1234567890():,; natural variability in these seawater properties across a range of spatial and temporal scales. Accurately quantifying a small change against a large and variable background requires precise and accurate measurements made over decades. The GLobal Ocean Data Analysis Project (GLODAP)4,5, initiated in 2004 and subsequently updated6–8, has been instrumental in delivering carbon-relevant interior ocean data that support well-quantified estimates of the ocean carbon sink. The project delivers near-global data coverage; standardized quality control procedures; a high degree of internal consistency; common data formats; and open and free access to the available data. Compared to its first version, the GLODAP data inventory has more than tripled in size (Fig. 1). In order to continue to serve its purpose, GLODAP needs to advance both its data ingestion systems and its data extraction systems to become more streamlined and automated. In order to decrease the amount of routine manual work as well as the potential for errors, data submission workflows must become uniform, semi-automated, and compatible with machine-learning techniques for quality control. The data extraction system also needs to accommodate a wider range of filtering to fine-tune requests from users. Global ocean carbon data Faced with the challenge of quantifying the ocean’s storage of anthropogenic carbon, the ocean community began to systematically measure marine inorganic carbon concentrations in the 1970 and 1980’s4. These efforts ramped up significantly during the World Ocean Circulation Experiment and the Joint Global Ocean Flux Study (WOCE/JGOFS) during the 1990’s, and have later been continued along selected WOCE lines in the repeat hydrographic programs including the Global Ocean Ship-based Hydrographic Investigations Program (GO-SHIP)9. The primary focus of GLODAP is synthesizing seawater inorganic carbon chemistry data from these global cruise campaigns. However, data for ocean hydrography, dissolved oxygen, transient tracers, inorganic nutrients, and a range of other variables are included to facilitate interpretation. A unique feature of GLODAP is the addition of several layers of quality control and adjustments conducted to minimize inconsistencies and biases in the data10 using a range of tools such as comparison of deep water values at nearby locations. GLODAP offers uniform data at three levels; (1) data from individual cruises in a uniform format with coherent quality control and unit conversion applied, (2) a bias adjusted data product, and 3) a global 1° × 1° mapped climatology11. The GLODAP data product has supported more than 2000 articles (and counting) since the year 2000, evidencing its extensive use by the scientific community and the trust placed in it. Seminal contributions on the oceanic anthropogenic carbon content and temporal evolution would not have been possible without GLODAP2,12,13. The knowledge from these studies informs, for instance, the Intergovernmental Panel for Climate Change (IPCC) assessments, and the Sustainable Development Goals (SDG) of the UN Agenda 2030 and the Global Climate Observing System (GCOS) indicators on ocean acidification. GLODAP is also an essential reference data set for autonomous observing networks, such as Biogeochemical-Argo: “The longterm success of a global chemical sensor observing system will depend on support from an ongoing, shipboard hydrographic program to produce a high-quality data set for deep waters at the global scale.”14. With the growth in the amount of data, the ongoing need to provide information on the ocean carbon sink to inform global carbon emission-reduction efforts, and the emerging need to monitor impacts of initiatives in geoengineering and sustainable use of the oceans, the importance of GLODAP will only increase. However, despite receiving short-term funds from a range of projects, GLODAP is a largely unfunded community effort organized and executed by the GLODAP team. Such a situation is unsustainable, and there is significant risk that the effort will diminish or disappear in the next few years. The building and supporting of infrastructure will be critical to ensure that GLODAP continues to provide a valuable service to the global community. Fig. 1 Key outputs and metrics of GLODAP. a Interior ocean concentration of anthropogenic carbon along a section indicated with a black line in panel (b). bIntegrated column inventory of anthropogenic carbon22. Both panels used transient tracer data and the Transit Time Distribution method to calculate anthropogenic carbon23 content. cCumulative number of samples in GLODAPv2.2020 over time. COMMENT COMMUNICATIONS EARTH & ENVIRONMENT | https://doi.org/10.1038/s43247-021-00209-4 2COMMUNICATIONS EARTH & ENVIRONMENT | (2021) 2:136 |https://doi.org/10.1038 /s43247-021-00209-4 | www.nature.com/commsenv Improved efficiency and service The current GLODAP workflow requires substantial manual work that necessitates dedicated time from, and funding for, data experts, and that introduces opportunities for data handling errors. GLODAP has matured over the last decade with a set of well-documented protocols and development of dedicated software, as well as a backbone of data management support. However, the GLODAP team now strives for advancements on both data input and output, toward a semi-automated system that will reduce the manual work intensity and associated errors. First, the team aims to implement a uniform, semi-automatic, and standards-compliant data ingestion system that will facilitate the data submission and quality control procedures. This will enable direct interaction with data providers, leading to improvements in data handling, data quality control, and documentation. The envisaged changes will also enable rapid application of novel quality control approaches using machinelearning techniques. Second, we want to upgrade to a versatile data extraction system. Such a system will provide more flexibility and options to users, such as requesting output with originally submitted data (without adjustments), or only sub-sets of the data in various formats. These upgrades will streamline repository workflows to insure the data products are FAIR (findable, accessible, interoperable, and reusable)15, while reducing the burden of data management on scientists. Nevertheless, there will remain a need for experts to spend time on quality control and internal consistency adjustments. Branch out to keep data accessible We expect that the improvements will encourage submission of data through building a community of data providers, and will simplify and streamline the process of providing regular updates of the GLODAP products. At the same time, access to GLODAP data will increase. Workflow improvements would allow for enhanced data access systems supporting machine-to-machine services, and better integrated data visualization products16,17. The GO-SHIP repeat hydrography effort currently provides the backbone of GLODAP thanks to its high data quality and rapid availability. However, many other datasets reach GLODAP through the extensive network of the GLODAP team; some of these datasets will be functionally lost if not collated by GLODAP. An automated system can aid rescue these data for reuse, by providing a streamlined process for scientists to submit data and metadata, and for users to access and visualize the data. Upgrades of GLODAP will benefit from the data system that has already been developed for the Surface Ocean CO 2 Atlas (SOCAT)18. SOCAT successfully streamlined data submission, quality control, and release of an annual synthesis product, but faces the same resourcing challenges as GLODAP to sustain regular updates. Leveraging an existing, and proven, workflow translates to a significant reduction in both cost and labor of developing a similar system for GLODAP. An investment for the planet’s future GLODAP needs continued support from the scientific community, but also needs support from funding agencies and stakeholders. Without the updated infrastructure and adequate sustained resourcing in place, GLODAP services may not be able to be maintained on a regular basis. While the ocean currently takes up about 2.6 Gt of anthropogenic carbon annually, we must understand the evolution, efficiency, and regional patterns of the ocean carbon sink if we want to be able to predict the climate effect of future emissions, as well as to quantify and assess mitigation efforts. Furthermore, human activities affect ocean biogeochemistry in other ways as well, such as de-oxygenation19, changes in nutrient supply20, and ocean acidification21, issues that all need high quality, consistent ocean biogeochemical data to quantify trends, and variability. Co-located high-quality measurements of physical and biogeochemical parameters that allow for the separation of natural variability from anthropogenic changes—as delivered by GLODAP—are a key component to monitoring, understanding, and mitigating the human influence on the Earth’s climate. Received: 8 May 2021; Accepted: 9 June 2021; References 1. Friedlingstein, P. et al. Global carbon budget 2020. Earth Syst. Sci. Data 12, 3269–3340 (2020). 2. Gruber, N. et al. The oceanic sink for anthropogenic CO 2 from 1994 to 2007. Science 363, 1193–1199 (2019). 3. Broullón, D. et al. A global monthly climatology of oceanic total dissolved inorganic carbon: a neural network approach. Earth Syst. Sci. Data 12, 1725–1743 (2020). 4. Key, R. M. et al. A global ocean carbon climatology: results from Global Data Analysis Project (GLODAP). Global Biogeochem. Cycle 18, GB4031 (2004). 5. GLODAP. https://www.glodap.info/ (2021). 6. Olsen, A. et al. The Global Ocean Data Analysis Project version 2 (GLODAPv2)—an internally consistent data product for the world ocean. Earth Syst. Sci. Data 8, 297–323 (2016). 7. Olsen, A. et al. GLODAPv2.2019—an update of GLODAPv2. Earth Syst. Sci. Data 11, 1437–1461 (2019). 8. Olsen, A. et al. An updated version of the global interior ocean biogeochemical data product, GLODAPv2.2020. Earth Syst. Sci. Data 12, 3653–3678 (2020). 9. Sloyan, B. M. et al. The global ocean ship-based hydrographic investigations program (GO-SHIP): a platform for integrated multidisciplinary ocean science. Front. Mar. Sci. 6,https://doi.org/10.3389/fmars.2019.00445 (2019). 10. Tanhua, T. et al. Quality control procedures and methods of the CARINA database. Earth Syst. Sci. Data 2,35–49 (2010). 11. Lauvset, S. K. et al. A new global interior ocean mapped climatology: the 1°×1° GLODAP version 2. Earth Syst. Sci. Data 8, 325–340 (2016). 12. Sabine, C. L. et al. The Oceanic sink for Anthropogenic CO 2 .Science 305, 367–371 (2004). 13. Khatiwala, S. et al. Global storage of anthropogenic carbon. Biogeosciences 10, 2169–2191 (2013). 14. Johnson, K. S. et al. Biogeochemical sensor performance in the SOCCOM profiling float array. J. Geophys. Res. Oceans 122, 6416–6436 (2017). 15. Wilkinson, M. D. et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci. Data 3, 160018 (2016). 16. webODV. https://hdl.handle.net/20.500.12085/webodv-glodap (2021). 17. digitalearthviewer-glodap. https://hdl.handle.net/20.500.12085/ digitalearthviewer-glodap (2021). 18. Bakker, D. C. E. et al. A multi-decade record of high-quality fCO 2 data in version 3 of the Surface Ocean CO 2 Atlas (SOCAT). Earth Syst. Sci. Data 8, 383–413 (2016). 19. Schmidtko, S., Stramma, L. & Visbeck, M. Decline in global oceanic oxygen content during the past five decades. Nature 542, 335–339 (2017). 20. Moon, J.-Y., Lee, K., Tanhua, T., Kress, N. & Kim, I.-N. Temporal nutrient dynamics in the Mediterranean Sea in response to anthropogenic inputs. Geophys. Res. Lett. 43, 5243–5251 (2016). 21. Feely, R. A. et al. Impact of anthropogenic CO 2 on the CaCO 3 system in the oceans. Science 305, 362–366 (2004). 22. Lauvset, S. K. et al. Processes driving global interior ocean pH distribution. Global Biogeochem. Cycle 34, e2019GB006229 (2020). 23. Waugh, D. W., Hall, T. M., McNeil, B. I., Key, R. & Matear, R. J. Anthropogenic CO 2 in the Oceans estimated using transit-time distributions. Tellus 58B, 376–389 (2006). Acknowledgements We acknowledge the efforts of scientists that secured funding, dedicated time to collect and share the data. This effort has been supported by EU Horizon 2020 through the EuroSea action (grant no. 862626), the Helmholtz Association through “Digital Earth” (grant no. ZT-0025), L.C.C. acknowledges the UERJ/Prociencia grant (2018-2021), M.H. acknowledge EU Horizon 2020 action SO-CHIC (grant N°821001). B.C., R.A.F., E.B., COMMUNICATIONS EARTH & ENVIRONMENT | https://doi.org/10.1038/s43247-021-00209-4 COMMENT COMMUNICATIONS EARTH & ENVIRONMENT | (2021) 2:136 | https://doi.org/10.1038/s43247-021-00209-4 | www.nature.com/commsenv 3 4 2. Data Sources The SPOTS pilot includes data from 12 fixed ship-based time-series programs (Fig. 1), all of which routinely measure BGC EOVs. All major climate zones are covered, although not all ocean biogeochemical zones are 125 (Reygondeau et al., 2013). Existing datasets were extended whenever possible by publicly available and more recent data (Table S1.). In addition to capturing different marine environments (Sect. 2.1), the characteristics of the time-series programs also differ in terms of the station visit frequency, i.e. temporal resolution (monthly, seasonal, or irregular), the time range of the observational period, the bottom depth and whether a dedicated research vessel is used (Table 1). If a time-series program consists of two or more related stations, usually the 130 deepest station was selected. The included data from GIFT and RADCOR display exceptions to this rule as for both sites data from three related stations were selected. Table 1: Key metadata of participating time-series programs. Colors indicate ocean basins: Green: Pacific; Light blue: Atlantic; Orange: Marginal Seas; Dark Blue: Nordic Seas. S=Salinity (either bottle or CTD-data); O2=Oxygen (either bottle or CTD-135 data); NO3=Dissolved inorganic nitrate; NO2=Dissolved nitrite; PO4=Dissolved phosphate; SiOH4=Dissolved silicate; NH4=Dissolved ammonium; DIC=Dissolved inorganic carbon; TA=Total alkalinity; pCO2=Partial pressure of carbon dioxide; POC=Particulate organic carbon; PON=Particulate organic nitrogen; POP=Particulate organic phosphorus; DOC=Dissolved organic carbon. Time-Series Site Location Time Range Temporal Resolution Bottom Depth # of Visits Dedicated Vessel Variables Original DOI(s) KNOT 44.0°N 155.0°E 1997– 2020 1-3 cruises yr-1 6000 m 21 No S,O2,NO3,NO2,SiOH4, PO4,NH4,DIC,TA,pH 10.25921/tarq-6v91 K2 47.0°N 160.0°E 1999– 2020 1-3 cruises yr-1 6000 m 49 No S,O2,NO3,NO2,SiOH4, PO4,NH4,DIC,TA,pH, DOC 10.25921/mpfz-sv16 ALOHA 22.8°N 158.0°W 1988– 2019 Monthly 4750 m 311 Yes S,O2,NO3,NO2,SiOH4, PO4,DIC,TA,pH, POC,PON,POP,DOC 10.1575/1912/bcodmo.3773.1 Munida 45.8°S 171.5°E 1998– 2019 6 cruises yr-1 1000 m 80 Yes S,NO3,SiOH4,PO4,DIC ,TA NA GIFT 36.9°N 6.0°W 2005– 2015 Seasonal 315 m – 842 m 26 Yes S,O2,NO3,SiOH4,PO4, TA,pH, DOC 10.20350/digitalCSI C/10549 CVOO 17.6°N 24.3°W 2006– 2019 1-3 cruises yr-1 3600 m 42 Partly S,O2,NO3,NO2,SiOH4, PO4,NH4,DIC,TA, POC,PON,POP 10.1594/PANGAEA .958597 RADCOR 43.4°N 8.4°E 2013– 2020 Monthly 15 m – 80 m 80 Yes S,O2,NO3,NO2,SiOH4, PO4,DIC,TA,pH NA CARIACO 10.5°N 64.7°W 1995– 2017 Monthly 1300 m 230 Yes S,O2,NO3,NO2,SiOH4, PO4,NH4,TA,pH, POC,PON,POP,DOC 10.1575/1912/bcodmo.3093.1 DYFAMED 42.3°N 7.5°E 1991– 2017 Monthly 2400 m 190 No S,O2,NO3,NO2,SiOH4, PO4,,DIC,TA,pH 10.17882/43749 IrmingerSea 64.3°N 28.0°W 1983– 2019 Seasonal 1000 m 131 Yes S,O2,NO3,SiOH4,PO4 , DIC,TA,pCO2 10.3334/cdiac/otg.ca rina_irmingersea_v2 ; 10.25921/vjmy8h90 IcelandSea 68.0°N 12.7°W 1983– 2019 Seasonal 1850 m 146 Yes S,O2,NO3,SiOH4,PO4, DIC,TA,pCO2 10.3334/cdiac/otg.ca rina_icelandsea; 10.25921/qhed-3h84 OWSM 66.0°N 2.0°E 2001– 2021 4-12 cruises yr-1 2100 m 147 Until 2009 S,O2,NO3,SiOH4,PO4, DIC,TA 10.3334/cdiac/otg_ts m_ows_m_66n_2e 140 5 Figure 1: Locations of participating ship-based time-series stations. 2.1. Marine Environment of Time-Series Sites 2.1.1. A Long-term Oligotrophic Habitat Assessment (ALOHA) The deep water (~4750 m) time-series station of the Hawaii Ocean Time-Series program (HOT), ALOHA (Karl 145 and Church, 2019), is located 100 km north of Oahu, Hawaii, more than one Rossby radius (50 km) away from the steep topography associated with the Hawaiian Ridge. ALOHA serves as an open ocean benchmark and its research goals are aligned with the main objectives of the JGOFS and the World Ocean Circulation Experiment (WOCE). One of the principals of the HOT program is to observe seasonal and interannual variations in water mass characteristics and BGC variables. The monthly measurements since 1988 are representative of the 150 oligotrophic North Pacific eastern subtropical gyre with Station ALOHA lying in the center of the North Pacific and North Equatorial Current. Typically, the site is characterized by a relatively deep permanent pycnocline (and nutricline), and a shallow mixed-layer depth. Intermittent local wind forcing caused by extratropical cyclones' cold fronts impacts the annual cycle of the surface waters (Karl et al., 1996). 155 2.1.2. CArbon Retention In A Colored Ocean (CARIACO) The station of the CARIACO Oceanographic Time-Series Program (Muller-Karger et al., 2019) is located in the Cariaco Basin, a semi-enclosed tectonic depression located on the continental shelf off northern Venezuela in the southern Caribbean Sea. The Cariaco Basin is composed of two approximately 1400 m deep sub-basins that are connected to the Caribbean Sea by two shallow (140 m deep) channels. These channels allow for the open 160 exchange of near-surface water. The restricted circulation below the 140 m sills, coupled with highly productive surface waters due to seasonal wind-driven coastal upwelling (around 450 g C m-2 y-1; Muller-Karger et al., 2010), has led to sustained anoxia below around 250 m. The goal of the near-monthly measurements at CARIACO between 1995 and 2017 was to observe linkages between oceanographic processes and the production, remineralization, and sinking flux of particulate matter in the Cariaco Basin, and how these change over time. It 165 also aimed at understanding climatic changes in the region. 2.1.3. Cape Verde Ocean Observatory (CVOO) CVOO is located in the eastern tropical North Atlantic about 800 km from the west coast of Africa, which is influenced by the seasonal eastern boundary upwelling system, high Saharan dust deposition rates, and frequently 170 passing eddies (Schütte et al., 2016). It is part of the Cape Verde Observatory, which also includes an operational atmospheric monitoring site. The combined observations aim at investigating long-term changes of greenhouse gas concentrations in the atmosphere and in the ocean in a key region for air-sea interaction. The irregular measurements of BGC variables at CVOO started in 2006 and are still ongoing, and the project strives for more regular measurements in the future by having a dedicated vessel available. The station has a bottom depth of 3600 175 m and lies in the center of the Cape Verde Fontal Zone, resulting in large variations of the present oligotrophic water masses. The frontal zone separates most of the eastern tropical North Atlantic from the anticyclonic subtropical gyre system in the North Atlantic (Stramma et al., 2005). This further results in an ocean shadow zone 6 and an oxygen-poor layer between 400 m to 500 m (Stramma et al., 2008), which is being sampled at CVOO. Below the mixed layer, subtropical underwater from the subtropical gyre system, as well as North Atlantic Central 180 Water and South Atlantic Central Water can be present (Tomczak 1981; Pastor et al., 2008). 2.1.4. DYFAMED DYFAMED is located in the central part of the Ligurian Sea, about 50 km off Nice, on the Nice Corsica transect, and is representative of open sea western Mediterranean basin waters. Ongoing multidisciplinary monthly 185 measurements at DYFAMED have been performed since 1991 observing: i) the evolution of the water mass properties, ii) the carbon export change, and iii) the variability of the biological species relative to climate forcing. The water column can be divided into three principal layers: deep, intermediate, and surface. The latter, typically for the Mediterranean trophic environment, experiences large seasonal variability. Further, the Northern Current front acts as a barrier to exchanges with the coastal zone of the Ligurian Sea and prevents DYFAMED from lateral 190 inputs (Vescovali et al., 1998). Consequently, the primary production depends on inputs of nutrients from deeper waters and atmospheric inputs of nitrogen and some trace metals, particularly during summer (Miquel, 2011). The DYFAMED site is characterized by intermediate water (300-400m) that is lower in oxygen concentrations (Levantine Intermediate Water) and deep water that is richer in oxygen, primarily induced by vertical mixing occurring in winter during intense and cold winds (convection processes; Coppola et al., 2018). 195 2.1.5. Gibraltar Fixed Time series (GIFT) Seasonal measurements at GIFT were established in 2005 to quantify the exchange of carbon between the Mediterranean Sea and the adjacent Atlantic Ocean and assess the temporal evolution of BGC fluxes. The three GIFT time-series stations (Flecha et al., 2019) are located along the longitudinal axis of the Strait of Gibraltar, 200 which connects the two basins. The Strait is surrounded by the Gulf of Cadiz (west) and the Alboran Sea (east). Water circulation in the channel can be described as a bi-layer system characterized by an inward (eastward) flow of the North Atlantic Central Water in the upper layer and an outward (westward) flow of Mediterranean waters (predominantly formed by a mixture of the Levantine Intermediate Water and the Western Mediterranean Deep Water) at the bottom layer. The depth and thickness of each water mass vary along the Strait, due to topography 205 in the channel and the influence of physical mechanisms. In particular, the Espartel sill (358 m depth) and the Camarinal sill (285 m depth) lead to large variability in the proportion of water flows’ position. Therefore, sampling depths vary from one campaign to another due to the instant position of the incoming and outcoming flows that are identified by their thermohaline properties through the CTD casts. 210 2.1.6. Irminger Sea station (IRM-TS) and Iceland Sea station (IC-TS) In 1983, seasonal measurements at the IRM-TS and the IC-TS (Olafsson et al., 2010) were initiated to observe the seasonal variability of carbon-nutrient chemistry in the North Atlantic off the Iceland shelf. The stations are located in two hydrographically different regions north and southwest of Iceland (Takahashi et al., 1985; Peng et al., 1987). The station in the northern Irminger Sea (IRM-TS) is characterized by relatively warm and saline (S > 35) Modified 215 North Atlantic Water derived from the North Atlantic Drift. Winter mixing is induced by strong winds and loss of heat to the atmosphere. This location may also be described as representing the subpolar gyre (Hatún et al., 2005). The IS-TS is located in the central Iceland Sea north of the Greenland-Scotland Ridge. At the IC-TS cold Arctic Intermediate Water, formed from Atlantic Water and low salinity Polar Water, usually predominates and overlays Arctic Deep Water (Olafsson et al., 2009). The Polar Water influence in the surface layers is variable (Stefansson, 220 1962; Hansen and Østerhus, 2000). Both regions are important sources of North Atlantic Deep Water. 2.1.7. K2 and KNOT The K2 and KNOT stations (Wakita et al., 2017) are located approximately 400 km northeast of Hokkaido Island, Japan in the subarctic western North Pacific. Since 2001 and 1997, respectively, irregular field observations have 225 been conducted at these stations to investigate the inorganic carbon system dynamics in response to variations in hydrography and biological processes. The overarching goal is to investigate the response of the biological pump to climate forcing in the western subarctic Pacific gyre. The region is characterized by high primary productivity, abundant marine resources (FAO, 2016) and might be the first region of the ocean to become undersaturated with respect to calcium carbonate during winter (Orr et al., 2005). The sites are representative of the southwestern 230 subarctic gyre with both stations lying offshore of the Oyashio Current and just north of the subpolar front. Seasonal cycles are present (e.g., Takahashi et al., 2006; Tsurushima et al., 2002; Wakita et al., 2013) with a highly productive biological pump from spring to fall and strong vertical mixing of deep waters that are rich in dissolved inorganic carbon (DIC) in winter. 7 235 2.1.8. Munida This deep-water station is located in the Southwest Pacific Ocean 65 km off the southeast coast of New Zealand and is part the Munida Time Series Transect, which is sampled every two months. Measurements at Munida were established in 1998 to study the role of these waters in the uptake of atmospheric carbon dioxide, and the seasonal, interannual, and long-term changes of the carbonate chemistry. The subantarctic waters are a sink for atmospheric 240 carbon dioxide (Currie et al., 2011), and the seasonal cycles of DIC are primarily driven by net community production (Brix et al, 2013; Jones et al., 2013) with modification by the annual cycle of sea surface temperature. 2.1.9. Ocean Weather Station Mike (OWSM) OWSM is located in the Norwegian Sea at the western baroclinic branch of the northwards-flowing Norwegian 245 Atlantic Current where the water depth is 2100 m (Skjelvan et al., 2008; 2022). Hydrographic measurements date back to 1948 while carbonate chemistry measurements started in 2001 to monitor long-term changes in the biogeochemistry. Between 2001 and 2009, the station was sampled monthly, and since 2010, the sampling frequency has been four to six times per year. The site encompasses the cold Norwegian Sea Deep Water and the Arctic Intermediate Water in addition to the relatively warm and saline Atlantic Water. Occasionally during late 250 summer, fresh Norwegian Coastal Current Water meanders all the way out to OWSM, influencing the surface water at the station. Seasonal variability is observed in the uppermost ~200 m, and long-term trends of carbonate variables are observed at all water depths. Over time, the surface water CO2 content at OWSM has increased at a faster rate than atmospheric pCO2 at this site (Skjelvan et al., 2022). 255 2.1.10. A Coruña RADIALES (RADCOR) The RADIALES program started in 1989 aiming to obtain reliable baselines for long-term studies on climate change and ecosystem dynamics in times of increasing anthropogenic disturbances along the northern and northwestern Spanish coasts (Valdés et al., 2021). The program consists of monthly multidisciplinary perpendicular sections covering the Cantabrian Sea and northwest coastal and neritic Spanish ocean. The A Coruña 260 (NW Galician coast) section (RADCOR) started in 1990 (Bode et al., 2020) and CO2 variables have been incorporated since 2013 in two stations, E2CO and E4CO. RADCOR is located on the northern edge of the Iberian Upwelling Region. Here, the classical pattern of seasonal stratification of the water column in temperate regions is masked by upwelling events from May to September. These upwelling events provide nutrients to support both primary and secondary production in summer. Nevertheless, upwelling is highly variable in intensity and 265 frequency, demonstrating substantial interannual variability, mostly affecting the E2CO station (80 m), while the station closest to shore, E4CO (15 m), is more impacted by estuarine and benthic processes. 8 3. Methods The data flow of the SPOTS pilot depicting the main steps of the synthesis is schematically illustrated in Fig. 2. In 270 the following, the individual components of this data flow are described in detail. Figure 2: Schematic data flow of the SPOTS pilot 3.1. Data Collection 275 The data from the 12 participating time-series programs were retrieved from data centers or directly obtained from the responsible principal investigator (Table S1). In the latter case, merging, formatting, additional quality-control (QC), and archiving of existing data were carried out. Only bottle data for BGC EOVs, that had been measured by at least two of the participating programs were included in the pilot project, along with accompanying ancillary pressure, salinity, and temperature data. We have also developed a metadata template for BGC EOV ship-based 280 time-series data (Table S2). The template has subsequently been used to collect all relevant metadata information from each participating time-series program. The collected metadata includes general information about the program, such as information about the principal investigator and the location and timeframe of related station(s). It also includes detailed information on the measured variables - e.g., units; sampling and analytical methods and associated instrumentation; calculation, calibration, and quality control procedures; and standards or (certified) 285 reference materials used. The latter not only vary among the time-series programs, but can also vary within a timeseries program over time. 3.2. Data Assembly The SPOTS pilot was created by standardizing data format, units, header names, primary QC flags, times, locations, and fill values and subsequently merging the individual datasets of each time-series program into one 290 file. Only data that received a WOCE quality flag 2 (Table S3) were included in the product. Existing data were altered as little as possible without interpolation or calculation of “missing” variables. Similarly, original station-, castand bottle numbers were kept or created artificially if non-existent to ensure consistency. The headers, units, and flags of the individual time-series datasets were standardized (Table S4) to conform with the WOCE exchange bottle data format (Swift and Diggs, 2008), a comma-delimited ASCII format for bottle data from hydrographic 295 cruises. To enable an automated mapping to other existing vocabularies, we also mapped the WOCE headers to the Natural Environment Research Council (NERC) British Oceanographic Data Centre P01 vocabulary collection, as well as to the newly proposed BGC bottle standard by Liqing et al. (2022). We did not use the latter as “central” semantics due to restrictions of existing QC-tools (e.g., AtlantOS QC (Velo et al., 2022) and the crossover toolbox (Tanhua et al., 2010; Lauvset and Tanhua, 2015)) to WOCE semantics. 300 The standardization process also entailed unit conversions, most frequently from micromoles per liter (µmol L-1; nutrients and dissolved organic carbon (DOC)) or from micrograms per kilogram (µg kg-1; particulate matter) to micromoles per kilogram of seawater (µmol kg-1). The default procedure to convert from volumetric to gravimetric 9 units was to use seawater density at in-situ salinity, reported laboratory temperature (otherwise assuming 20°C as laboratory conditions), and pressure of 1 atm (following recommendations from Liqing et al., 2022). For some 305 time-series datasets, the combined concentration of nitrate and nitrite was reported (Table S4). If explicit nitrite concentrations were provided, these were subtracted to obtain the nitrate values. If not, the combined concentration was renamed to nitrate assuming that the relative nitrite amount is negligible. For the HOT program specifically, low-level, high-sensitivity measurements of macronutrients (phosphate and nitrate) were available but not included in the pilot product. Particulate organic matter was derived by subtracting the particulate inorganic matter from 310 the total particulate matter, if available. For particulate organic carbon and particulate organic nitrogen, the factors 1/12.01 and 1/14.01 (inverse standard atomic masses) were used, respectively, for the unit conversion to micromoles per kilogram. For the HOT program, particulate carbon and nitrogen measurements correspond to total particle concentrations (PC and PN), but are here assumed to approximate POC and PON. If neither temperature nor pressure was provided, all corresponding data entries were excluded from the product. The 315 potential density anomaly1 is the only calculated variable. Missing and excluded values were set to -999. 3.3. Qualitative Assessment of Data 3.3.1. Internally Applied Quality-Control (QC) The majority of the programs have established their own routines for QC and correspondingly flag their data using different flagging schemes. We did not double-check the applied flags, nor did we run additional QC checks. The 320 applied QC on the collected stations include statistical outlier checks on routinely measured pressure intervals using either a twoor three- (seasonal) sigma criteria, visual inspections of property-property plots (PPP), and application of crossovers using reference layers (Table S5). For example, K2 and KNOT used North Pacific Deep Water (NPDW), defined as the water mass between 27.69 σθ (around 2000 dbar) and 27.77 σθ (around 3500 dbar) (Wakita et al., 2017), as the reference layer for their internal crossover checks. For CVOO and Munida, we 325 performed QC by applying a seasonal two-sigma criterium to the data, and for CVOO, we made additional use of comparisons to CANYON-B (Bittig et al., 2018) and crossovers. Since the QC procedures differ from program to program, we have provided recommendations for the QC of future data, so that the flags are applied more consistently across different programs (Sect. 6.3). Further, the standardization of the SPOTS pilot also entailed mapping to a central flagging scheme. We chose the WOCE bottle flag scheme (Table S3). Flags indicating 330 replicate measurements (WOCE flag of 6) were set to 2, whereas all other flags were set to 9 and the corresponding values to -999. 3.3.2. Best-Practices (BP) Assessment Given the inconsistencies in the applied internal quality checks and the fact that bias corrections following 335 crossovers analyses are presently impossible to apply to all included time-series datasets2, the comparability of the data for the SPOTS pilot was qualitatively assessed. The information on the applied methods of each time-series program, as provided through the metadata collection, was evaluated against, ideally, published Best Practices (BPs), and otherwise known standard operating procedures (SOPs). “BP Flags” were assigned accordingly to each cruise of a time-series program (Table 2). 340 Table 2: Meaning of assigned BP Flags. Flag Definition 0 No data 1 Methods meet all BP requirements (including “desired”) 2 Methods only meet “required” BP requirements 3 Methods do not meet the BP requirements (or no metadata given) The majority of the defined “BP requirements” used for the evaluation are based on the Bermuda Time-Series Workshop report (Lorenzoni and Benway, 2013), with additional implementation of: GO-SHIP manuals (Langdon 345 et al., 2010; Becker et al., 2019); the CARIACO Methods Manual (Astor et al., 2013); HOT analytical methods (https://hahana.soest.hawaii.edu/hot/protocols/protocols.html), which are based on the Joint Global Ocean Flux 1 Calculated using the Matlab seawater toolbox (Morgan, 1994) 2 Crossover require a “constant” reference layer over the entire span of measurements. Especially in coastal and shallow water formation regions this layer is nonexistent. Detrending might make this criterion redundant. However, detrending techniques rely on regular measurement intervals, which is not the case for most ship-based time-series sites. 10 Study protocols (IOC, 1994); the guide to BPs for ocean CO2 measurements (Dickson et al., 2007); results from the Scientific Committee on Research Working Group 147 “Towards comparability of global oceanic nutrient data” (Bakker et al., 2016; Aoyama et al., 2015); and studies on preservation techniques for nutrients (e.g. Dore et 350 al., 1996). The requirements were grouped into “Required” and “Desired” BP, see Table 3. To fulfill all requirements, i.e. receive a BP flag of 1, the metadata must show that the methods also met the corresponding “Desired” requirements. Only time-series programs that provided granular metadata, i.e. metadata differentiating between different methods applied in time, could obtain a BP flag of 1. 355 Table 3: BP requirements used for the method evaluation. Variable Required Desired Salinity AutoSal (Sub - ) standard used regularly Temperature constant Glass bottles Dissolved Oxygen Winkler Draw temperature used for mass calculation if difference to in situ temperature > 2.5 ° C Titration reagent assessed using CSK/OSIL primary standard Nutrients All Autoanalyser; If stored: Frozen upright Carrier Solution documented Calibrated against Reference Material Silicate Autoanalyser; If stored and concentrations are above 40 µmol L-1: Poisoned and kept cold Carrier Solution documented Calibrated against Reference Material Dissolved Inorganic Carbon Coulometry Calibrated against Certified Reference Material (Andrew Dickson, SIO) If stored: Poisoned, kept in dark and cool location Potentiometric (closed-cell); Calibrated against Certified Reference Material (Andrew Dickson, SIO); If stored: Poisoned and kept in a cold, dark location 3 Not applicable Total Alkalinity Potentiometric Titration (multi-step) Open Cell or curve fitting method documented Calibrated against Certified Reference Material (Andrew Dickson, SIO) If stored: Poisoned, kept in dark and cool location Spectrophotometric Indicator dye: bromocresol green Calibrated against Certified Reference Material (Andrew Dickson, SIO) If stored: Poisoned, kept in dark and cool location pH Spectrophotometric with scale and temperature reported Indicator dye: m-cresol purple Indicator dye: Purified If dye is not purified: Correction applied to impurities Partial pressure of CO2 Gas-chromatography Temperature and standard reported If stored: Poisoned, kept in dark and cool location Infrared-based system Temperature and standard reported If stored: Poisoned, kept in dark and cool location Particulate Matter Carbon and Nitrate High temperature combustion with reported filter volume and pore size Dried filters (60°C) Standards reported 3 Capped at a BP flag 2. 11 Phosphorus Ash hydrolysis with reported filter volume and pore size Dried filters (60°C) Standards reported Dissolved Organic Carbon High temperature combustion Filtered If stored: Frozen or acidified to and refrigerated Calibrated against Reference Material (Dennis Hansell, University of Miami) 3.4. Quantitative Assessment of Data In addition to the qualitative BP assessment (Sect. 3.3), the bottle data of the time-series are described by our quantitative descriptors: 1) precision, 2) accuracy, 3) variability on the most consistent depth layer, and 4) 360 consistency with GLODAP (Lauvset et al., 2022). Precision and accuracy were included in the SPOTS pilot dataset file, the latter two were not included and are only described here. 3.4.1. Precision and Accuracy Precision and accuracy estimates, as provided by each time-series program’s primary quality-assurance procedure, 365 were assigned to the bottle data. The temporal resolution of these estimates varies from estimates given for each cruise, i.e. on a cruise-to-cruise basis, to estimates given for longer time periods (covering multiple cruises) without recorded changes in applied methodology (Table S6), depending on the individual time-series’ internal procedure. If only one estimate was given for a variable for the entire time-series, that estimate was only assigned to the most recently applied method. The units correspond to the units of the respective variable. 370 Precision estimates are based on replicate samples and expressed as one standard deviation of the replicate measurements4. For the carbon variables, the assigned accuracy estimates represent the deviation from certified reference materials from the A. Dickson Laboratory (Scripps Institution of Oceanography). The pH accuracies of RADCOR are an exception, representing the difference from the theoretical TRIS buffer value at 25°C. For oxygen concentrations, the assigned accuracy estimates represent the accuracy of the KIO3 primary standard normality 375 assessed using a certified reference standard from either Ocean Scientific International Ltd (OSIL) or Wako Pure Chemical Industries (WAKO). For nutrient concentrations, the assigned accuracy estimates represent the deviation from reference material from either OSIL, WAKO, or QUASIMEME (Wells et al., 1997) or from certified reference material from Kanso Technos Co., Ltd. (KANSO). For particulate phosphorus concentrations, the assigned accuracy estimates represent deviations from National Institute of Science and Technology (NIST) apple 380 leaves (0.159% P by weight). For DOC, the accuracy estimates represent deviations from deep seawater reference material from D. Hansell (RSMAS, University of Miami). The exact calculations to express the above deviations from reference materials differ slightly across the time-series programs (Table S7), thereby preventing combined precision and accuracy estimates to calculate a total uncertainty in a consistent manner. The estimates should not be confused with values provided by instrument manufacturers, which are ideal values and are usually well below 385 real-world uncertainties. 3.4.2. Minimum Variability To provide an internal consistency measure of measurement quality, we determined the minimum variability of each BGC variable for each time-series station on the pressure surface (+/- 100 dbar) with the least oxygen 390 variability, i.e. the layer on which oxygen has the lowest coefficient of variation. We chose oxygen as natural variability in oxygen can be linked to either variation in ventilation, water mass changes, or changes in consumption and production by biological activity5 (Sarmiento and Gruber, 2006; Keeling et al., 2010; Stramma and Schmidtko, 2019). As these natural oxygen changes are likely to be accompanied by changes in other BGC variables, we used the layer that is closest to an oxygen equilibrium as an approximation for the least natural 395 variability in ocean BGC. In addition, this choice allowed us to use the salinity variability as an independent indicator of natural variability. For i) CARIACO, ii) GIFT, iii) Munida, and iv) RADCOR, this layer could not be determined properly, respectively, due to i) anoxic water masses below the mixed layer, ii) varying measurement depths, iii) no oxygen data and iv) a shallow bottom depth of 80 m. The minimum variabilities of the other variables 4 Exception: IRMand IC TS using Vdub*Cmean (following OSPAR, 2011), where Vdub is coefficient of variation calculated from dublicates and Cmean is the mean of the concentration measured. 5 Not represented in the variability of salinity 12 were subsequently determined by calculating the coefficient of variation of all samples on the identified pressure 400 surface. A minimum of 10 samples on the pressure surface was required. 3.4.3. Comparisons to GLODAP The final quantitative descriptor indicates how well the time-series data compares to the GLODAP dataset (GLODAPv2.2021, Lauvset et al., 2021) and vice-versa, with no a prioi assumption of which is ‘correct’. To this 405 end, we applied an adapted version of the GLODAP crossover routine to all individual cruises of the time-series programs. Generally, the crossover routine calculates a depth-independent offset between a cruise and a reference dataset based on multiple crossing cruises, i.e. “crossover-pairs”. The secondary quality control of GLODAP depends heavily on this routine to determine and correct for biases of new cruises, which results in the high internal consistency of the core GLODAP variables. In the following, we first describe the crossover routine of GLODAP 410 in detail and subsequently highlight the modifications applied to the routine so that it fits our pilot product needs. For a given variable the depth-independent offset of a new cruise against GLODAP is calculated using the following steps: Step 1) Detect all GLODAP cruises that cross the to-be-compared cruise (denoted as Cruise A in the following), 415 i.e. find all “crossover-pairs” of Cruise A in GLODAP. In the 2nd QC of GLODAP, a “crossover-pair” is defined by two cruises that have (at least) three stations within a 2° radius that include (at least) three samples below a minimum of 1500 dbar. These requirements ensure that the influence of natural signals on the calculated offsets is limited. That becomes especially important if the time period between Cruise A and a crossing GLODAP cruise (denoted as Cruise B in the following) is large. 420 Step 2) Interpolate the samples of Cruise A and Cruise B to the same standard depths. Usually, the concentrations are compared on sigma-4 surfaces6. Samples above the chosen minimum depth are ignored to exclude layers that are influenced by daily to interannual variability. 425 Step 3) Compare all existing samples of Cruise A and B that are on the same depth surface and from stations within 2°. For each depth surface, the individual offsets are averaged to obtain depth-dependent mean offsets and standard deviations. For nutrients and oxygen, the offsets are multiplicative, and for the carbon variables and salinity, the offsets are additive. 430 Step 4) Calculate the constant offset of Cruise A against Cruise B by inverse variance weighting all depthdependent offsets. The resultant depth-independent offset is also known as the crossover-pair offset. Step 5) Calculate the standard deviation of the crossover-pair offset by inverse variance weighting all depthdependent standard deviations. This crossover-pair standard deviation reflects the similarity of the offsets within 435 one depth surface and across all depth surfaces. The lower it is, the higher the confidence in the crossover-pair offset. Step 6) Repeat Steps 2) to 5) for all identified crossover-pairs. 440 Step 7) Calculate the total offset of Cruise A against GLODAP by inverse variance weighting all calculated crossover-pair offsets. The resultant standard deviation describes the overall uncertainty in the total offset. For our purposes, we applied an adapted version of the above-described crossover routine using GLODAP as the “reference dataset” against which each time-series station is compared. The term “reference dataset” does not 445 imply that the quality of GLODAP is higher than the quality of the time-series programs, only that it represents a dataset with known consistency in time and space. Each cruise of a time-series station, i.e. station visit, represents another Cruise A in the above-outlined crossover steps. For a given time-series station and variable, our adapted crossover routine starts with the identification of crossover-pairs for each station visit, similar to Step 1. However, since multiple time-series cruises only take one 450 profile with fewer than three samples below 1500 dbar, we could not apply the same crossover-pair requirements. We kept the distance requirement of 2° and added a new temporal requirement, that only crossover cruises within +/- 45 days were included in the routine. That permitted relaxing the minimum depth requirement and dropping 6 In regions with a high probability of internal waves, in upwelling and water formation regions the offsets are calculated on pressure surfaces. 13 the requirement of the minimum number of profiles. Note that we excluded crossover-pairs of cruises that are included in both products (parts of: IC-TS, IRM-TS, and OWSM). Steps 2 to 6 of the routine are identical and 455 repeated for all time-series station cruises. In the next step, all crossover pair offsets against the same GLODAP cruise, i.e. a particular Cruise B, are averaged. This step was necessary when multiple time-series cruises took place within 90 days and all were compared to the same Cruise B. Consequently, we obtained one depthindependent offset (and standard deviation) of the time-series station against each GLODAP cruise that meets the crossing requirements. In a final calculation, we determine the total offset of the time-series station against 460 GLODAP by inverse variance weighting all obtained time-series station offsets. If the standard deviation of the time-series station offset against a particular cruise B was below the consistency estimates of GLODAPv2 (see Table 11 in Olsen et al., 2016), the latter ones were used as standard deviations (e.g. only one crossover pair exits between the entire time-series and a particular Cruise B). The routine was only applied to variables defined as core variables7 in GLODAP. Negative (or lower than unity) offsets indicate lower values compared to GLODAP and 465 vice versa. 7 Salinity, oxygen, nitrate, phosphate, silicate, DIC, total alkalinity and pH 20 4.3.3. Iceland Sea The crossover offsets of the IC-TS of salinity, oxygen, nitrate, and DIC against GLODAP are within the consistency limits of GLODAP, i.e. no significant offset is remarkable between the two products. For nitrate, the variability between the individual offsets is large, which reduces confidence in the analysis. For phosphate, the 675 SPOTS pilot has 6% lower concentrations than GLODAP based upon four cruises from the IC-TS (B17-94, B996, B12-96, and B5-2002) and three GLODAP cruises (58JH19941028, 58JH19961030 and 316N20020530), which all passed GLODAP’s 2nd QC. This large offset mainly originates from the 2002 cruise, while cruises from 1996 indicate a good fit. The same cruises show a -4% offset for silicate, and the underlying data show a similar pattern. However, the relatively large minimum variability of salinity (Sect. 4.2) demonstrates that the Iceland Sea 680 is a dynamically active region with deep open ocean convection and complex seasonally varying currents; this high natural variability reduces confidence in the crossover analysis for the Iceland Sea region. 4.3.4. Irminger Sea All crossover offsets of the IRM-TS against GLODAP are above GLODAP’s consistency limits except for 685 phosphate. However, given i) that the minimum depth had to be set to only 500 m in a deep water formation area and ii) the relatively large minimum variability of salinity (Sect. 4.2), the larger offsets were expected and are likely attributable to the inherent natural variability of this region. Further, the relatively small number of crossovers does not allow for a more in-depth investigation of the offsets. 690 4.3.5. KNOT Crossover offsets could be calculated for all GLODAP core variables. The calculations were performed on the NPDW, which has a residence time of about 500 years (Stuiver et al., 1983). Following the definition from Wakita et al. (2010), we used 27.69 σ (around 2000 dbar) and 27.77 σ (around 3500 dbar) as limits. All of the so-calculated offsets of KNOT against GLODAP are clearly within the consistency limits except for total alkalinity (-5 µmol 695 kg-1). Confidence in the analysis is provided through a large number of crossover cruises and consistency of calculated offsets. Data from a few cruises are present in both products. 4.3.6. K2 Crossover offsets could be calculated for all GLODAP core variables. The calculations were again performed on 700 the NPDW using the identical limits as those of KNOT. All of the so-calculated offsets of K2 against GLODAP are clearly within the consistency limits. Confidence in the analysis is provided through a large number of crossover cruises and consistency of calculated offsets, as exemplarily shown for nitrate (Fig. 6). Data from a few cruises are present in both products. 705 Figure 6: Total weighted offset of the SPOTS pilot nitrate data against GLODAPv2.2021 at station K2 in the North Pacific Deep Water (NPDW) layer. The total weighted offset is multiplicative and illustrated by the red line. The dashed red lines are 21 the corresponding standard deviation. The black dots display the weighted offsets of individual K2 cruises against GLODAP cruises with the corresponding error bars displaying their standard deviation. If the calculated standard deviation of the 710 individual cruises is lower than GLODAP’s nitrate consistency limit (2%) it is set to the latter. The summary figure indicates a very good fit between the SPOTS pilot product and GLODAP at the K2 station for nitrate with a total weighted offset of 0.0%. 4.3.7. OWSM 715 Crossover offsets at OWSM indicate slight mismatches between the nitrate, phosphate, and DIC concentrations of the SPOTS pilot vs. GLODAP. The total weighted mean offsets are -3%, -3%, and 6 µmol kg-1, respectively. The former two offsets are only based upon a comparison between the OWSM cruise from 20020415 (no CRUISE ID present) and 316N20020530. Three more recent OWSM cruises from 2019 are additionally checked against 58JH20190515. Both GLODAP cruises passed GLODAP’s 2nd QC. However, the DIC offsets are very dependent 720 on the crossover pair and the final offset should be treated with caution. The small number of crossovers does not allow for a more in-depth investigation of the relatively small offsets. 22 5. Product File Description The product file variable names are described in Table S8. Each fixed-location time-series station is identified by the entry under “TimeSeriesSite”, and individual cruises are identified by “CRUISE”. Station, cast, and bottle 725 numbers are linked to the original cruise campaign numbering (if provided). In some cases, station number duplicates within the same time-series program exist as the data originates from different research vessels of opportunity (Table 1). Nitrate values can contain nitrite concentrations (Table S4). Similarly, ALOHA’s particulate organic matter includes particulate inorganic components. Since all pH values were reported on the total scale at 25°C, no additional pH temperature entry is provided. Conversely, for pCO2 corresponding temperature 730 measurements are given. In addition to the WOCE flags, each bottle variable is further accompanied by the assigned BP Flag (Sect. 4.1) and by the provided precision and accuracy estimates (Sect. 3.4). The last column lists the digital object identifier (DOI) of the original dataset. All missing entries are indicated by -999. A total of 108,332 water samples are included in the product. Bottle salinity with 75,654 measurements is the variable with the most abundant data (Table 6). The number of bottle salinity samples is about twice the number 735 of bottle oxygen and nutrient (excluding ammonium and nitrite) samples and almost five times the number of included DIC and total alkalinity samples. pH and nitrite have around 10,000 samples and the product includes between 4,900 and 7,600 samples of particulate matter, DOC, and ammonium. With 1,898 samples from the IRMTS and the IC-TS, pCO2 is the variable with the fewest measurements. Silicate, nitrate, and total alkalinity are the only variables measured at all sites. Around 56% of all bottle data values originate from ALOHA (Table 6) and 740 14% from CARIACO. The remaining 25% are distributed rather equally across the different programs. ALOHA’s large percentage can be explained by measurements at ALOHA i) having taken place consistently on a monthly basis for >30 years; ii) including up to 30 hydrocasts per station visit; and iii) including all but two of the product’s bottle variables. The dominance of ALOHA’s measurements is most pronounced for salinity, particulate phosphate (inorganic and organic), and DOC (around 70% - 80% of the samples are measured at ALOHA). For oxygen and 745 nutrients, ALOHA’s samples represent around 52% of all samples, and for the inorganic carbon variables (DIC, total alkalinity, and pH) between 32% - 42%. Table 6: Summary statistics showing the total number of samples per variable included in the SPOTS pilot of each timeseries site. Percentages in brackets show fractions in comparison to the total number per variable except for the last column. 750 Percentages are rounded; thus, the sum is not always equal to exactly 100%. Variable abbreviations are identical to Table 1. S O2 NO3 NO2 PO4 SiOH4 NH4 DIC TA pH pCO2 POC PON POP DOC Total ALOHA 63334 (84%) 21937 (57%) 18130 (52%) 750 (6%) 17648 (53%) 17656 (52%) 0 5911 (35%) 5780 (32%) 4124 (42%) 0 3659 (48%) 3637 (49%) 3675 (75%) 4778 (67%) 171019 (56%) CARIACO 4026 (5%) 3528 (9%) 3705 (11%) 3768 (32%) 3724 (11%) 3691 (11%) 3680 (69%) 0 3687 (21%) 3760 (39%) 0 3870 (51%) 3804 (51%) 1221 (25%) 975 (14%) 43439 (14%) CVOO 345 (<1%) 534 (1%) 451 (1%) 507 (4%) 451 (1%) 411 (1%) 73 (1%) 346 (2%) 304 (2%) 0 0 39 (1%) 39 (1%) 24 (<1%) 0 3524 (1%) DYFAMED 0 2328 (6%) 1525 (4%) 1670 (14%) 1611 (5%) 1482 (4%) 0 1086 (6%) 1114 (6%) 56 (1%) 0 0 0 0 0 10872 (4%) GIFT 0 480 (1%) 479 (1%) 0 0 477 (1%) 0 0 470 (3%) 463 (5%) 0 0 0 0 199 (3%) 2568 (1%) IcelandSea 2214 (3%) 2111 (5%) 2070 (6%) 0 2087 (6%) 2101 (6%) 0 1824 (11%) 280 (2%) 0 1101 (58%) 0 0 0 0 13788 (4%) IrmingerSea 1901 (3%) 1792 (5%) 1774 (5%) 0 1767 (5%) 1784 (5%) 0 1477 (9%) 209 (1%) 0 797 (42%) 0 0 0 0 11501 (4%) K2 1921 (3%) 1904 (5%) 1996 (6%) 1997 (17%) 1994 (6%) 1983 (6%) 1188 (22%) 1897 (11%) 1805 (10%) 509 (5%) 0 0 0 0 1129 (16%) 18323 (6%) KNOT 1864 (2%) 1997 (5%) 1859 (5%) 1893 (16%) 1851 (6%) 1862 (5%) 376 (7%) 1821 (11%) 1802 (10%) 174 (2%) 0 0 0 0 0 15445 (5%) Munida 0 0 285 (1%) 0 285 (1%) 280 (1%) 0 220 (1%) 298 (2%) 0 0 0 0 0 0 1368 (<1%) OWSM 49 (<1%) 905 (2%) 1004 (3%) 0 911 (3%) 1004 (3%) 0 2053 (12%) 1320 (7%) 0 0 0 0 0 0 7246 (2%) RADCOR 0 1215 (3%) 1270 (4%) 1279 (11%) 1268 (4%) 1284 (4%) 0 190 (1%) 739 (4%) 678 (7%) 0 0 0 0 0 7923 (3%) Total 75654 38731 34548 11810 33597 34015 5317 16825 17808 9764 1898 7568 7480 4920 7081 307016 23 6. Stakeholders The main stakeholder groups of SPOTS are the data providers on the upstream-end, i.e. the individual time-series programs (Sect. 2), and users of time-series data on the downstream-end. Regarding the latter, the SPOTS pilot is 755 intended to be applied in different ocean BGC fields: evaluations of ocean BGC, neural networks such as CANYON-B (Bittig et al., 2018), CANYON-MED (Fourrier et al. 2020), or ESPER (Carter et al., 2021), regional ocean BGC models, (e.g., models participating in RECCAP such as Ishii et al., 2015), 1D model applications (e.g., Mamnun et al., 2022 using REcoM2), global ocean BGC models participating in model intercomparison projects (e.g., Coupled Model Intercomparison Project - Orr et al., 2016); evaluations of autonomous BGC observing 760 networks such as BGC Argo (Bittig et al., 2019); global scientific assessments such as the Global Carbon Budget (Friedlingstein et al., 2022); or multi time-series studies and analyses (e.g., Bates et al., 2014; O’Brien et al., 2017). These time-series can also contribute ocean carbonate chemistry data to the United Nations Sustainable Development Goals, especially target 14.3 to minimize and address the impacts of ocean acidification. 6.1. Benefits 765 The main goal of SPOTS is that both stakeholder groups benefit from the product. Through a use-case, the benefits for the users are implicitly demonstrated in Sect. 6.2. On the upstream end, data providers benefit from the product in several ways. First of all, the product increases the impact of individual ship-based time-series programs. For smaller and less well-known time-series programs, the impact is particularly improved by increasing their visibility and discoverability. Here, two “pull factors” 770 contribute: i) the popularity and success of the included larger time-series programs and ii) being exposed on the Ocean Data and Information System (ODIS) catalog (https://book.oceaninfohub.org) in a schema.org-friendly way (Sect. 7.2). The larger sites also benefit from the latter, but the impact of larger time-series programs is in particular increased through enhanced usability of their data. Here, the proverb “the whole is greater than the sum of its parts” perfectly describes the benefits of SPOTS. The envisioned (non-exhaustive) list of users underscores the 775 idea that consistent and inter-comparable data from multiple time-series programs (i.e. the “whole”) leads to an extended range of applications relative to those of a single time-series program. The data being automatically uploaded to ERDDAP, which increases the accessibility, interoperability, and machine-readability (Sect. 7.2), also becomes important in broadening users and applications of data from these time-series programs. Further, participating time-series programs benefit from optional data management support for formatting, QC, 780 and data archival. This support aims at reducing the data management workload of individual programs and being directly ascribed to the FAIR data practices. Regarding guidelines and BPs, the participating time-series programs also benefit from the product fostering collaborations across several programs, which is especially relevant for emerging time-series programs. Ship-based time-series programs represent one of our most powerful tools for monitoring marine ecosystem 785 changes. The product contributes to the development of a sustained, globally distributed network of time-series observatories that sample a core set of biogeochemical and ecological variables guided by common best practices (methodological, FAIR data, etc.). These are required attributes of a GOOS observing network, and achieving this status would ultimately help position ship-based time-series programs for expansion under the United Nations Decade of Ocean Science umbrella. In addition, the product links individual time-series efforts to larger policy 790 directives such as the Marine Strategy Directive Framework in Europe with respect to e.g., ocean monitoring indicators. 6.2. Use-Case As an example to demonstrate both the utility and potential misuses of the SPOTS pilot, we applied the recently developed Trends of Ocean Acidification Time Series software (TOATS, https://github.com/NOAA-795 PMEL/TOATS) to the mixed layer total alkalinity data included in the product (Fig. 7). The TOATS software is a supplement to the recently published best practices for assessing trends of ocean acidification time-series and provides a python based Jupyter Notebook to compare trends across different (BGC) time-series data sets (Sutton et al., 2022). It was developed based on several published trend analysis techniques to standardize estimating and reporting trends from ocean carbon time-series data sets. Following a strict sequence of approaches8, TOATS 800 8 1. assess data gaps in the time-series; 2. remove periodic signals (i.e. normally occurring variations due to predictable cycles) from the time-series;3. assess a linear fit to the data with the periodic signal(s) removed; 4. estimate whether a statistically significant trend can be detected from the time-series; 5. consider uncertainty in the measurements and reported trends; and 6. present trend analysis results in the context of natural variability and uncertainty. 24 estimates i) the linear trend, ii) its uncertainties, and iii) the trend detection time of the assessed time-series data. The latter indicates the minimum observational period needed to statistically distinguish between natural variability (noise) and anthropogenic forcing. This method requires time-series with sub-seasonal sampling frequency to constrain seasonal variability of surface ocean carbonate chemistry; however, for the purpose of this example, we assessed all time-series programs rather than restricting the assessment to time-series datasets with 805 regular monthly measurements. The only non-trivial calculation step we applied before running TOATS was to calculate the surface mixed layer depth for each cruise (defined using a 0.3 potential density anomaly criteria following de Boyer Montégut et al. (2004)) and to average total alkalinity concentrations within the estimated mixed layers. The results of our use-case (Fig. 7) show trends in alkalinity for all time-series (seven of them with significant trends). 810 The ease of use in applying TOATS to multiple time-series demonstrates the main benefits and potential misuse of the SPOTS pilot at the same time. Concerning the benefits, the combination of the SPOTS pilot and TOATS enables any user to perform joint time-series studies that follow published BPs without requiring any in-depth programming knowledge. The need to, a priori, know about existing time-series program data and to subsequently mine, format, and QC the data, becomes redundant for all time-series datasets included in SPOTS. The required 815 input format of TOATS is also readily available by accessing the time-series product data through ERDDAP (Sect. 7.2). Further, detailed information on methods and their changes over time will become even more accessible once the ODIS user interface is online. This will enable a sophisticated information-driven data selection of (subsets of) time-series data to analyze the effects of method changes on detected trends without having to study multiple cruise reports. A similar advantage is provided through the possibility of selecting subsets of data based on the 820 assigned BP flags (Sect. 4.1). Lastly, the estimates of precision and accuracy included in the SPOTS pilot (Sect. 3.4) additionally enable confident uncertainty estimations of the trend analyses (uncertainties of the observations being a mandatory input in TOATS). Regarding the potential misuse of the SPOTS pilot, caution must be applied in interpreting the results, particularly because the use-case analysis includes values accompanied by BP flags 2 and 3. Simply assuming that the 825 determined trends (Fig. 7) are valid and interpreting differences across time-series programs could lead to false conclusions. Robust trend analysis also requires the user to acknowledge the impact of large data gaps in timeseries that inhibit the ability to constrain seasonal variability in many of the included datasets (e.g. CVOO), and make it impossible to remove periodic signals with confidence (second step of TOATS trend analysis). Following TOATS guidelines, we recommend applying TOATS to surface ocean biogeochemical data with at least regular 830 seasonal measurements or to restrict the trend analysis to specific seasons. Increasing the number of samples using additional interpolation and computational techniques could relax this restriction (e.g., multivariate linear regression (MLR); Vance et al., 2022), but computations accompanied by large uncertainties might also harm the robustness of the trend analyses. Note that in the case of interpolating concentrations of single variables vertically, we recommend using a quasi-Hermetian piecewise polynomial (Key et al., 2010). And if techniques to increase 835 the data coverage involve using CO2SYS (van Heuven et al., 2011), we recommend using the carbonate dissociation constants of Lueker et al. (2000), the bisulfate dissociation constant of Dickson (1990), and the borateto-salinity ratio of Uppström (1974). Another large pitfall is neglecting the provided metadata and assuming that restricting the analyses to time-series data with a BP flag 1 erases all artifacts in the trend analyses. Such a restriction would increase the robustness of 840 the analysis, but unaccounted differences within the BP flag 1 (Sect. 4.1) would still bias the results. For example, ALOHA particulate phosphorus samples analyzed before 2012 are biased low but still fulfill all assessed BP requirements (Sect. 4.1). Similarly, some standardizations of the product resulted in the neglect of valuable timeseries details (e.g., information on ventilation events provided through the unique QC flags of CARIACO (Sect. 3.2)). We included all information in the additional metadata, made it easily accessible, and encourage users 845 consult it, particularly to check for any correlations of the trend analyses to method changes (Table S6) and/or specific time-series events. Even though this example highlights a multiple time-series study use-case, it depicts the benefits and especially the potential misuses for other applications of the SPOTS pilot. If the limitations of the product (e.g., data gaps 850 and varying baselines) are acknowledged, quality descriptors are utilized, and the data are used in conjunction with the supporting metadata, multiple applications can benefit from this time-series product. 25 Figure 7: Trend analysis of total alkalinity in the mixed layer using TOATS. Data symbols show the original time-series 855 observations (blue circles), the time series of monthly means (black circles), the de-seasoned monthly means (red squares), and the trend of the de-seasoned monthly means (red line) (From Sutton et al., 2022). The monthly anomalies (red squares) that are used for the trend analyses are shown as de-seasoned monthly means. The grey boxes include the yearly trend, adjusted R2 and the minimum trend detection time (TDT). An Asterix next to the yearly trend number indicates that the result is significant (two-sided t-test p-value < 0.05). Note that xand y-axis are not in synch among the different time series subplots. 860 26 6.3. Recommended Standard Operating Procedures (SOPs) for Ship-Based Time-Series Programs The process of generating the SPOTS pilot resulted in the formulation and recommendation of SOPs regarding metadata documentation, internal QC, and uncertainty estimation. These are directed at the data providers, i.e. those who help run the ship-based time-series programs. The proposed SOPs are briefly presented here, the full guidelines can be accessed at https://www2.whoi.edu/site/mets-rcn/. 865 1. Metadata documentation: The first SOP is the recommended metadata template (Table S2), which provides a structure for time-series programs to uniformly document the applied methodologies, thereby ensuring that relevant information, including differences between individual cruises, is recorded. It should be filled out for each cruise individually. The metadata enables detailed method comparisons of shipbased BGC EOV data such as the BP assessment of the participating sites of the data product. We 870 recommend that the metadata template be updated as the community re-determines, expands, and specifies the BPs for BGC EOV ship-based time-series data. 2. Consistent QC routine: The second SOP recommendation involves the use of a consistent routine to QC time-series data. The main goal is that scientists follow consistent criteria to flag single samples. Different characteristics of time-series programs - e.g., location (depth and seasonal influence), funding 875 opportunities (duration and frequency of visits), and scientific goals (variables measured) - preclude a “one-size-fits-all” QC method. Thus, a decision tree approach guides the user in choosing the appropriate type of QC for their dataset. All suggested semi-automated checks make particular use of comparisons with historical time-series data. To evaluate the flagging results, the SOP is accompanied by a comparison to the well-established HOT QC results. 880 3. Calculating uncertainty: The third SOP has been developed by the Oslo and Paris Conventions Commission (OSPAR), Hazardous Substances & Eutrophication Committee (OSPAR, 2011) and was originally intended for assessments of contaminants in biota and sediment done in OSPAR areas. It can also be applied to BGC EOV ship-based data. It provides detailed recommendations for a consistent estimation of one total measure of uncertainty, including exact formulas that combine the information 885 obtained through duplicate measurements (precision) and comparisons to reference material (accuracy). 27 7. Data Access and Availability 7.1. METS-RCN Website All information regarding the SPOTS pilot and the collaborative NSF EarthCube funded Marine Ecological Time Series Research Coordination Network (METS-RCN) can be accessed at https://www2.whoi.edu/site/mets-rcn/. 890 The SPOTS web page (https://www2.whoi.edu/site/mets-rcn/ts-data-product/) includes detailed information on the participating time-series programs, including: • contact person(s) • time-series website URL • relevant data repositories 895 • cruise reports and papers • detailed metadata on the BGC EOVs measured • recommended SOPs (Sect. 6.3) and in-depth information on the assigned BP flags • links to AtlantOS QC software and crossover toolbox used 900 The website also provides several options for users to download the SPOTS pilot (DOI: 10.26008/1912/bcodmo.896862.1), including: • Comma-separated value (CVS) format (directly from the website) • Link to the BCO-DMO repository (https://www.bco-dmo.org/dataset/896862, Lange et al., 2023) • GOOS-relevant ERDDAP server 905 (https://data.pmel.noaa.gov/generic/erddap/tabledap/bgc_ts_product.html) 7.2. Environmental Research Division's Data Access Program (ERDDAP) Providing the data through ERDDAP enables FAIR-compliant data access services and gives users significantly enhanced capabilities rather than just downloading the dataset directly from the website. Optional constraints 910 within the ERDDAP dataset enable downloading subsets of the dataset. The constraint options include amongst others variable-, stationand time selections. ERDDAP also enables downloading the dataset in several formats, such as tab-separated or netCDF. The latter format also entails additional metadata attributes, including alternative variable names (NERC P01 or following the recommendations from Liqing et al. (2022)). On the ERDDAP server, users find a link “Make a graph” (https://data.pmel.noaa.gov/generic/erddap/tabledap/bgc_ts_product.graph), 915 which enables plotting the data using the web-based ERDDAP tool. In addition to giving the users more degrees of freedom, hosting the dataset on the ERDDAP server has two important benefits. First, the dataset is machinereadable, enabling an automated transfer to other repositories and higher-level infrastructures (e.g., SeaDataNet, Copernicus Marine Environment Monitoring Service). Second, ERDDAP data managers are working to provide direct access to metadata information stored in the ODIS catalog, which, once achieved, will significantly improve 920 metadata interoperability. 7.3. ODIS catalog Through collaborating with ODIS, we developed two json-ld templates to publish time-series program metadata in a schema.org-friendly way (inspired by Science on Schema; Shepherd et al., 2022) and to enable FAIR metadata. 925 The first template (EventSeries) is designed to capture the general information about the time-series programs (e.g., location, time, principal investigators, funding, and related datasets). A “sub-events” section is used for more details about the individual cruise’s location, time, personnel, and vessel. That section also includes details about the applied measurement methodologies for each cruise and provides links to cruise reports. The second json-ld template (Dataset) is designed to describe the metadata of the related BGC discrete bottle datasets. Here, the 930 included variables and in particular, the applied semantics of the dataset are described. By using and linking these templates for each of the participating sites, we could include the metadata of the time-series sites and related datasets in the ODIS catalog. Here, the time-series programs are exposed on the web and machine-readable (interoperable) access to the metadata is guaranteed. Presently, these json-ld files are hosted by the METS-RCN GitHub repository (https://github.com/earthcube/METS-RCN). Eventually, the individual time-series program’s 935 data centers can host (and update) these files and assign unique identifiers. The metadata of the SPOTS pilot itself (Dataset) are also stored in the ODIS catalog, clearly linking all related metadata to the data synthesis product. 28 7.4. Fair Data Usage Agreement While the SPOTS pilot is made available without any restrictions (Creative Commons Attribution 4.0.), users of 940 the data should adhere to fair data use principles: For investigations that rely on data from a particular timeseries program, principal investigators should be contacted to explore opportunities for collaboration and coauthorship and if there are any uncertainties regarding methodological details or interpretation of datasets. The original dataset DOI and any articles where the data are described should be cited. Contacting principal investigators comes with the additional benefit of expert insight into the specific site under investigation. This 945 paper should be cited in any scientific publications that result from the usage of the SPOTS pilot. 29 8. Conclusion The SPOTS pilot synthesized data from 12 ship-based ocean time-series programs, each representative of a unique marine environment. Time-series data and metadata were compiled and assessed to provide an internally consistent data product. As a pilot study, for feasibility, the focus of this initial ship-based time-series data product was BGC 950 EOV data, which served as a use-case for the METS RCN and provided a template for a sustained living data product for ocean time-series. Through an external qualitative assessment of the applied methodologies, flags were assigned that reflect the degree to which BPs were followed, which determines the comparability of the data. The most recently applied methods typically met the required BPs, but measurements of oxygen and pH still show room for improvement. 955 Though the methods are adequately documented by many time-series programs, several others need to document their methods more thoroughly. The assessment also revealed the need to determine the level of granularity of both required documentation and required BPs for fully comparable data. The importance of inter-laboratory studies (e.g., QUASIMEME) must be highlighted in this context. In addition to the included precision and accuracy estimates, quantitative assessments yielded additional indicators that describe the consistency withinand across 960 the time-series programs. For time-series stations dominated by water masses that contribute negligible natural variability, the calculated minimum variabilities demonstrate a high continuity in measurement quality. Reasonable fits between GLODAP and the majority of the time-series programs further increase the confidence in the data quality. By making BGC EOV datasets from multiple sources consistent and ready to use, the SPOTS pilot facilitates an 965 improved understanding of the variability and trends of ocean biogeochemistry. It represents an important and necessary step forward in broadening our view of a changing ocean and maximizing the return on our continued investment in ship-based ocean time-series programs. It also enhances data readiness (Lindstrom et al., 2012) by implementing FAIR data practices for all included data. In particular, the implementation of ERDDAP and ODIS (Sect. 7.2) enables easy data integration into e.g., OceanOPS and Copernicus Marine Environment Monitoring 970 Service. On a higher level, this effort facilitates the consolidation of the international ship-based time-series network by collaborating closely with the participating time-series programs, developing, and recommending SOPs, and supporting the network to become more fit-for-purpose. 36 Olafsson, J., Olafsdottir, S.R., Benoit-Cattin, A., Danielsen, M., Arnarson, T.S., and Takahashi, T. 2009. Rate of 1245 Iceland Sea acidification from time series measurements. Biogeosciences, 6, 2661–2668. DOI: 10.5194/bg6-2661-2009 Olafsson, J., Olafsdottir, S. R., Benoit-Cattin, A., and Takahashi, T. 2010. The Irminger Sea and the Iceland Sea time series measurements of sea water carbon and nutrient chemistry 1983–2008. Earth Syst. Sci. Data, 2, 99–104. DOI: 10.5194/essd-2-99-2010 1250 Olsen, A., Key, R.M., van Heuven, S., Lauvset, S.K., Velo, A., Lin, X., Schirnick, C., Kozyr, A., Tanhua, T., Hoppema, M., Jutterström, S., Steinfeldt, R., Jeansson, E., Ishii, M., Pérez, F.F. and Suzuki, T. 2016. The Global Ocean Data Analysis Project version 2 (GLODAPv2) – an internally consistent data product for the world ocean. Earth System Science Data, 8(2), 297-323. DOI:10.5194/essd-8-297-2016 Orr, J.C., Fabry, V., Aumont, O. et al. 2005. Anthropogenic ocean acidification over the twenty-first century and 1255 its impact on calcifying organisms. Nature 437, 681–686. DOI: 10.1038/nature04095 Orr, J.C., Najjar, R.G., Aumont, O., Bopp, L., Bullister, J.L., Danabasoglu, G., Doney, S.C., Dunne, J.P., Dutay, J., Graven, H., Griffies, S. M., John, J.G., Joos, F., Levin, I., Lindsay, K., Matear, R.J., McKinley, G.A., Mouchet, A., Oschlies, A., Romanou, A., Schlitzer, R., Tagliabue, A., Tanhua, T., and Yool, A. 2017. Biogeochemical protocols and diagnostics for the CMIP6 Ocean Model Intercomparison Project (OMIP). 1260 Geosci. Model Dev., 10, 2169–2199. DOI: 10.5194/gmd-10-2169-2017 O’Brien, T. D., Lorenzoni, L., Isensee, K., and Valdés, L. (eds). 2017. What are Marine Ecological Time Series telling us about the ocean? A status report. IOC-UNESCO, IOC Technical Series, No. 129: 297. Paris: IOCUNESCO OSPAR commission 2011. JAMP Guidelines for estimation of a measure for uncertainty in OSPAR monitoring. 1265 Agreement 2011-3. HASEC 11/12/1, Annex 5 Pastor, M., Pelegri, J., Hernandezguerra, A., Font, J., Salat, J., and Emelianov, M. 2008. Water and nutrient fluxes off Northwest Africa. Cont. Shelf Res., 28, 915–936. DOI: 10.1016/j.csr.2008.01.011 Peng, T., Takahashi, T., Broecker, W.S., and Olafsson, J. 1987. Seasonal variability of carbon dioxide, nutrients and oxygen in the northern North Atlantic surface water: Observations and a model. Tellus, 39B, 439–458. 1270 DOI: 10.3402/tellusb.v39i5.15361 Reygondeau, G., Longhurst, A., Martinez, E., Beaugrand, G., Antoine, D., and Maury, O. 2013. Dynamic biogeochemical provinces in the global ocean. Global Biogeochem. Cy., 27, 1046–1058. DOI: 10.1002/gbc.20089 Sarmiento, J.L. and Gruber, N. 2006. OceanBiogeochemical Dynamics. Princeton University Press. xiii 503 pp. 1275 Princeton, Woodstock: Princeton University Press. Geological Magazine, 144(6), 1034-1034. DOI :10.1017/S0016756807003755 Schütte, F., Brandt, P. and Karstensen, J. 2016. Occurrence and characteristics of mesoscale eddies in the tropical northeastern Atlantic Ocean. Ocean Sci., 12(3), 663–685. DOI:10.5194/os-12-663-2016 Shepherd, A., Jones, M.B., Richard, S., Jarboe, N., Vieglais, D., Fils, D., Duerr, R., Verhey, C., Minch, M., 1280 Mecum, B., Bentley, N. 2022. Science-on-Schema.org v1.3.1. Zenodo. DOI: 10.5281/zenodo.7872383 Skjelvan, I., Falck, E., Rey, F., and Kringstad, S.B. 2008. Inorganic carbon time series at Ocean Weather Station M in the Norwegian Sea. Biogeosciences, 5, 549–560. DOI: 10.5194/bg-5-549-2008 Skjelvan, I., Lauvset, S.K., Johannessen, T., Gundersen, K., Skagseth, Ø. 2022. Decadal trends in Ocean Acidification from the Ocean Weather Station M in the Norwegian Sea. Journal of Marine Systems, 1285 Volume 234,2022,103775. ISSN 0924-7963. DOI: 10.1016/j.jmarsys.2022.103775 Stefansson, U. 1962. North Icelandic Waters, Rit Fiskideildar, 3: 1–269 Stramma, L., Hüttl, S., and Schafstall, J. 2005. Water masses and currents in the upper tropical northeast Atlantic off northwest Africa. J. Geophys. Res., 110, C12006. DOI :10.1029/2005JC002939 Stramma, L., Johnson, G. C., Sprintall, J. and Mohrholz, V. 2008. Expanding Oxygen-Minimum Zones in the 1290 Tropical Oceans. Science, 320(5876), 655–658. DOI:10.1126/science.1153847 Stramma, L., and Schmidtko, S. 2019. Global Evidence of Ocean Deoxygenation. Ocean Deoxygenation: Everyone’s Problem: Causes, Impacts, Consequences and Solutions, edited by D. Laffoley and J. M. Baxter (Gland, Switzerland: IUCN, 2019), pp. 25–36 Stuiver, M., Quay, P.D., and Ostlund, H.G. 1983. Abyssal water carbon-14 distribution and the age of the world 1295 oceans. Science, 219, 849–851. DOI: 10.1126/science.219.4586.849 37 Sutton, A.J., Roman, B., Brendan, C., Wiley, E., Newton, J., Simone, A., Bates, N.R., Wei-Jun, C., Currie, K., Feely, R.A., Sabine, C., Tanhua, T., Tilbrook, B., and Wanninkhof, R. 2022. Advancing best practices for assessing trends of ocean acidification time series. 2022. JOURNAL=Frontiers in Marine Science. VOLUME=9. DOI: 10.3389/fmars.2022.1045667 1300 Takahashi, T., Olafsson, J., Broecker, W. S., Goddard, J., Chipman, D.W., and White, J. 1985. Seasonal variability of the carbon-nutrient chemistry in the ocean areas west and north of Iceland. Rit Fiskideildar, 9, 20–36 Takahashi, T., Sutherland, S.C., Feely, R. A., and Wanninkhof, R. 2006. Decadal change of the surface water pCO2 in the North Pacific: A synthesis of 35 years of observations. J. Geophys. Res., 111, C07S05. DOI:10.1029/2005JC003074 1305 Tanhua, T., van Heuven, S., Key, R.M., Velo, A., Olsen, A., and Schirnick, C. 2010. Quality control procedures and methods of the CARINA database. Earth Syst. Sci. Data 2: 205– 240. DOI:10.5194/essd-2-35-2010 Tanhua, T., Orr, J.C., Lorenzoni, L., and Hansson, L. 2015. Increasing ocean carbon and ocean acidification. WMO Bull. 64, 48–51 Tanhua, T., McCurdy, A., Fischer, A., Appeltans, W., Bax, N., Currie, K., DeYoung, B., Dunn, D., Heslop, E., 1310 Glover, L.K., Gunn, J., Hill, K., Ishii, M., Legler, D., Lindstrom, E., Miloslavich, P., Moltmann, T., Nolan, G., Palacz, A., Simmons, S., Sloyan, B., Smith, L.M., Smith, N., Telszewski, M., Visbeck, M., and Wilkin, J. 2019. What We Have Learned From the Framework for Ocean Observing: Evolution of the Global Ocean Observing System. Front. Mar. Sci. 6:471. DOI: 10.3389/fmars.2019.00471 Tanhua, T., Lauvset, S., Lange, N., Olsen, A., Álvarez, M., Diggs, S., Bittig, H., Brown, P., Carter, B., Cotrim da 1315 Cunha, L., Feely, R., Hoppema, M., Ishii, M., Jeansson, E., Kozyr, A., Murata, A., Pérez, F., Pfeil, B., Schirnick, C., and Key, R. 2021. A vision for FAIR ocean data products. Communications Earth & Environment. 2. 136. DOI: 10.1038/s43247-021-00209-4 Tomczak, M. 1981: An analysis of mixing in the frontal zone of South and North Atlantic Central Water off NorthWest Africa. Prog. Oceanogr., 10, 173–192. DOI:10.1016/0079-6611(81)90011-2 1320 Tsurushima, N., Nojiri, Y., Imai, K., and Watanabe, S. 2002. Seasonal variations of carbon dioxide system and nutrients in the surface mixed layer at Station KNOT (44°N, 155°E) in the subarctic North PacificDeep Sea Res., Part II, 49, 5377–5394. DOI: 10.1016/S0967-0645(02)00197-2 Uppstrom, L.R., 1974. The boron/chlorinity ratio of deep-sea water from the Pacific Ocean. Deep-Sea Res. Oceanogr. Abstr. 21, 161–162. DOI: 10.1016/00117471(74)90074-6 1325 Valdés, L., and Lomas, M.W. 2017. New light for ship-based time series. In: What are Marine Ecological Time Series telling us about the ocean? A status report. pp. 11–17. Ed. by T. D. O'Brien, L. Lorenzoni, K. Isensee, and L. Valdés. IOC UNESCO, IOC Technical Series, No. 129. 297 pp. Valdés, L., Bode, A., Latasa, M., Nogueira, E., Somavilla, R., Varela, M.M., González-Pola, C., and Casas, G. 2021. Three decades of continuous ocean observations in North Atlantic Spanish waters: The RADIALES 1330 time series project, context, achievements and challenges. Progress in Oceanography, Volume 198, 102671. ISSN 0079-6611. DOI: 10.1016/j.pocean.2021.102671 Vance, J.M., Currie, K., Zeldis, J., Dillingham, P., Law, C. 2022. An empirical MLR for estimating surface layer DIC and a comparative assessment to other gap-filling techniques for ocean carbon time series. Biogeosciences, 19(1). DOI: 10.5194/bg-19-241-2022 1335 Velo, A., Jesús, C., Fiz, F.P., Toste, T., and Lange, N. 2021. AtlantOS Ocean Data QC: Software packages and best practice manuals and knowledge transfer for sustained quality control of hydrographic sections. DOI: 10.5281/zenodo.4532402 Vescovali, I., Oubelkheir, K., Chiaverini, J., Pizay, M.D., Stock, A. and Marty, I.C. 1998. The Dyfamed timeseries station: a reference to coastal studies in the Mediterranean sea. IEEE Oceanic Engineering Society. 1340 OCEANS'98. Conference Proceedings (Cat. No.98CH36259), Nice, France, 1998, pp. 1785-1789 vol.3. DOI: 10.1109/OCEANS.1998.726393 Wakita, M., Watanabe, S., Honda, M., Nagano, A., Kimoto, K., Matsumoto, K., Kitamura, M., Sasaki, K., Kawakami, H., Fujiki, T., Sasaoka, K., Nakano, Y., and Murata, A. 2013. Ocean acidification from 1997 to 2011 in the subarctic western North Pacific Ocean. Biogeosciences, 10, 7817–7827. DOI: 10.5194/bg-1345 10-7817-2013 Wakita, M., Nagano, A., Fujiki, T., and Watanabe, S. 2017. Slow acidification of the winter mixed layer in the subarctic western North Pacific. Journal of Geophysical Research: Oceans, 122, 6923–6935. DOI: 10.1002/2017JC013002122 38 Weller, R.A., Gallage, C., Send, U., Lampitt, R.S., and Lukas, R. 2016. OceanSITES: Sustained Ocean Time 1350 Series Observations in the Global Ocean. vol. 2016 Wilkinson, M.D., Dumontier, M., Aalbersberg, I.J.J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J-W., Santos, L.B.D.S., Bourne, P.E., Bouwman, J., Brookes, A.J., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., Evelo, C.T., Finkers, R., Gonzalez-Beltran, A., Gray, A.J.G., Groth, P., Goble, C., Grethe, JS., Heringa, J., 't Hoen, P.A.C., Hooft, R., Kuhn, T., Kok, R., Kok, J., Lusher, S.J., Martone, M.E., Mons, 1355 A., Packer, A.L., Persson, B., Rocca-Serra, P., Roos, M., van Schaik, R., Sansone, S-A., Schultes, E., Sengstag, T., Slater, T., Strawn, G., Swertz, M.A., Thompson, M., van der Lei, J., van Mulligen, E., Velterop, J., Waagmeester, A., Wittenburg, P., Wolstencroft, K., Zhao, J., and Mons, B. 2016. The FAIR Guiding Principles for scientific data management and stewardship. Scientific data, vol. 3, 160018. DOI: 10.1038/sdata.2016.18 1360 133 7 Synthesis 134 Synthesis 135 7 Synthesis The overarching goal of this thesis was to improve the BGC data landscape through the manifestation of BGC synthesis products as an integral part in the BGC ocean observing system. This overarching goal was addressed by working towards four individual objectives. In this section, the main contributions from this thesis for each objective are presented. Eventually, the limitations of the presented work are discussed. 7.1 Individual Research Goals For the overarching goal, the following four objectives were identified at the onset of this thesis (Section 1.2): O1. Evaluating BGC EOV synthesis products in support of the elimination of any weaknesses O2. Continuously updating existing living BGC EOV synthesis products with new data O3. Developing and implementing improvements O4. Expanding the BGC EOV synthesis product landscape to previously overlooked observations As outlined in Section 1.2, each publication addressed, in particular, one objective. In the following, the main results of each publication are set into the context of the corresponding objective, and the contributions are briefly summarized and discussed. Additional results of this thesis that go beyond the individual publications, but are relevant for these objectives, are also included. However, parallel efforts that this thesis did not support, but which also relate to these objectives are not presented, e.g. the annual updates of SOCAT (Bakker et al., 2023). 7.1.1 Evaluating BGC EOV synthesis products in support of the elimination of any weaknesses In the larger context, the FOO readiness of an (BGC) ocean observing system depends on the maturity of all its components, input, processes, and output (Section 1.3.1, Lindstrom et al., 2012). However, for “Data and Information Products”, i.e. the output of the FOO engineering approach to ocean observations, and more particularly its sub-category “Data Products”, assessments of the maturity are missing (IOCCP 2017a-h). Aiming at closing this gap, and supporting the work towards more sophisticated methods to determine the maturity of ocean observing systems designed to capture BGC phenomena, an objective assessment of BGC data synthesis products with a clear scoring system was developed (Section 3). Accordingly, for the evaluation of BCG synthesis products, an objective scoring scheme was developed, utilizing a criteria catalog that has been created based on the FOO readiness level concept. The novel scoring scheme thus introduced and applied the FOO readiness concept to the evaluation and guidance of synthesis data products, and for the first time provides an objective measure of a synthesis products maturity. Four BGC EOV data synthesis products, SOCAT (Bakker et al., 2016), GLODAP (Lauvset et al., 2022), MEMENTO (Kock and Bange, 2015), and GO2DAT (Grégoire et al., 2021), were selected to which this new evaluation scheme was applied. The four synthesis products represent the entire spectrum of BGC EOV data synthesis products, spanning different stages of development (maturity), focusing on different EOVs (Inorganic Carbon, Nitrous Oxide, and Oxygen), and implementing an EOV-based, and a platform-based synthesis approach. The evaluation showed that of these four products, SOCAT is the most mature one reaching a “Mature” status, followed by GLODAP being in the “Pilot” phase, and MEMENTO and GO2DAT both being in the “Concept” phase. The reliability of the ranking and hence the developed evaluation system was further proven by the product’s identified impact, as approximated through the number of publications citing a product. Overall, the results underline the Synthesis 136 importance of further improving synthesis data products (and thereby the maturity of “Data and Information Product”), as for multiple BGC EOVs, e.g. oxygen, and associated BGC phenomena, the output of the ocean observing systems and in particular the sub-category “Data Products” appears to be a weak link in the ocean observing value chain (Section 1.3.1, Section 3, IOCCP 2017a-h). During the development and application of the evaluation, several critical features of BGC synthesis data products were identified that should guide new and existing BGC data synthesis products to realize their full potential, above all: • Theme-oriented concept • Achieving FAIR (Wilkinson et al., 2016) data (original data and synthesis product) with, in particular o the recognition and attribution of original (meta)data o the adaptation of common community standards o the implementation of Interoperable (meta)data • Incorporating an automated data flow as much as possible • Implementing a (customized) QC that is traceable • Feeding back to requirements (FOO) • Obtaining sustainable project-independent funding 7.1.2 Continuously updating existing living BGC EOV synthesis products with new data A successful synthesis product involves the continuous inclusion of newly obtained data, i.e. it is a “living synthesis”. As an integral part of this thesis annual updates of GLODAP could be realized, releasing v2.2019 (Olsen et al., 2019), v2.2020 (Olsen et al., 2020), v2.2021 (Lauvset et al., 2021), and v2.2022 (Lauvset et al., 2022), with v2.2023 on the horizon. The updates harmonized, QC’ed, archived, and added a total of 361 new cruises 21 to GLODAPv2 (Olsen et al., 2016), averaging about 90 new cruises per update. v2.2022 now includes inorganic carbon-relevant bottle data from a total of 1085 cruises. The total amount of samples has increased from 999,448 samples to 1,381,248 (Table 8 in Lauvset et al., 2022), and the coverage in time has been extended from 1972 until 2013 (v2) to 1972 until 2021 (v2.2022). Most of the newly added cruises cover repeat hydrographic sections in the Pacific Ocean (about 51%), particularly the North West Pacific. However, a large amount of newly added data can also be attributed to the inclusion of selected cruises from CODAP-NA (Jiang et al., 2022) and the GEOTRACES intermediate data product (GEOTRACES, 2021), the addition of Davies Strait cruise data, Line-P cruise data, as well as the extension of Weather Station M in the Norwegian Sea (Skjelvan et al., 2008; 2022) with an additional 10 years of data. Even though not many cruises from the Indian Ocean could be added, the few that were, have proven to be of significant importance, e.g. global anthropogenic carbon inventory estimates (Müller et al., 2023). Most of the data can be attributed to cruises being younger than 2014 (58%), but also newly (re-) discovered old datasets, e.g. from cruises taking place in 1982, were added during those updates. The upper 100 m remain the best-sampled part throughout all updates, and the number of observations steadily declines with depth until 1,000 m, caused by the reduction of ocean volume towards greater depth (Figure 11 in Lauvset et al., 2022). During the course of this thesis, a few changes have been applied to the product. Since v2.2020 GLODAP includes discrete fCO2 samples, which led to the adaption of the calculation scheme for “missing” inorganic carbon sub-variables using CO2SYS (Olsen et al., 2020). Also, since v2.2020, no internal consistency evaluation procedures of the inorganic carbon system were used to assess or correct sub-variables of the inorganic carbon system. This change followed studies demonstrating that such evaluations are prone to “[…] errors owing to incomplete understanding of the thermodynamic constants, major ion concentrations, measurement biases, and potential contribution of organic 21 v2.2019: 116 new cruises; v2.2020: 106 new cruises; v2.2021: 43 cruises; v2.2022: 96 new cruises Synthesis 137 compounds or other unknown protolytes to alkalinity (Takeshita et al., 2020) […]” (Olsen et al., 2020, p. 3661) leading to pH-dependent offsets in calculated pH (Álvarez et al., 2020; Carter et al., 2018). On the other hand, comparisons to calculated data from neural networks (CANYON-B and CONTENT), which are trained using GLODAPv2 (Bittig et al., 2018), were used more extensively in both (external) 1st QC and 2nd QC for all core variables. Since v2.2022, also SF6 undergoes the full 2nd QC process of GLODAP extending the number of “core variables” to 13 (Section 4). Hence, the 2nd QC methods applied have been adapted slightly, however, the crossover analysis remained the main method throughout all updates. Most importantly, adjustments were only applied when detected offsets to the earlier data product release were not attributed to natural variability or anthropogenic trends. Often, the thorough 2nd QC detected incomplete applied QA and 1st QC and instead of applying depthindependent bias corrections, the data generators could resolve the problem once supported by the QC results, e.g. by correcting wrongly applied unit conversions, or by correcting towards measured CRMs. Generally, the trend towards detecting fewer biases for more recent measurements remained (Figure 8 in Lauvset et al., 2022), reflecting the improvements in data quality due to the widespread adaption of standardized samplingand analyzing practices, and the implementation of thorough QA, in particular the usage of CRMs. Furthermore, by applying the formula for the propagation of uncertainty 22 to the global consistencies of all releases since v2, (e.g. Table 7 in Lauvset et al., 2022), the overall improvement in consistency due to the (2nd) QC of data (and the corresponding corrections of detected systematic offsets), is evident (Table 5). The improvements were strongest for v2, however, the corrections of each annual release improved the consistency of newly added cruises to the previous release for all variables. In particular, SiOH4 adjustments applied to analyses using different standards (North Pacific, Section 4) stand out when only considering the updates. Note that the shown consistencies and corresponding improvements can vary strongly regionally and that pH consistencies are not shown as the corresponding data were not estimated for v2. Table 5: Evolution of GLODAP consistency estimates and overall improvements. Overall improvements were calculated using the law of uncertainty propagation. v2 v2.2019 v2.2020 v2.2021 v2.2022 Overall unadj. adj. unadj. adj. unadj. adj. unadj. adj. unadj. adj. unadj. adj. Sal (x1000) 4.1 3.1 3.5 3.5 2.4 2.4 2.9 2.9 1.3 1.3 3.7 3.0 O2 (%) 1.7 0.9 1.0 0.8 0.5 0.5 1.0 1.0 0.5 0.4 1.5 0.8 NO3 (%) 1.7 1.2 0.8 0.8 0.5 0.5 1.5 1.1 0.4 0.4 1.5 1.1 SiOH4 (%) 2.8 1.7 1.3 1.1 1.0 0.8 1.7 1.2 1.4 0.6 2.4 1.5 PO4 (%) 2.2 1.3 1.0 0.9 0.8 0.8 2.2 1.8 0.7 0.7 1.9 1.2 DIC (μmol kg-1) 4.4 2.6 4.2 4.0 2.2 1.9 2.6 2.4 2.4 2.4 4.0 2.7 TA (μmol kg-1) 5.8 2.8 3.3 2.7 2.4 2.1 3.2 3.0 1.9 1.8 5.0 2.7 22 √(wv2*σ2v2 + wv2*σ2v2.2019 + wv2*σ2v2.2020 + wv2*σ2v2.2021 + wv2*σ2v2.2022), with w=number of cruises of version/ total number of cruises; σ=consistency estimate of version Synthesis 138 7.1.3 Developing and implementing improvements Over the last decade, GLODAP has matured with a set of well-documented protocols and development of dedicated software (Section 5). With the onset of this thesis and the annual GLODAP updates several further improvements to GLODAP itself (besides adding data) and to the underlying data flow were made. These improvements can mainly be linked to four software developments (Table 6). The corresponding advancements and implications for the data flow of GLODAP are discussed below. Table 6: Crucial software developments for GLODAP during the course of this thesis. Note that the consultation efforts focused on the aspects pertaining to the applicability to GLODAP, rather than the coding. New Software Application in GLODAP Status Thesis’ Contribution AtlantOS QC 1st QC Finished Consultation Python-based Crossover Tool 2nd QC Ongoing Consultation Python-based “Make Ocean” Routine Merging Finished Co-development Digital Earth Viewer Visualization Finished Consultation To begin with, the development of AtlantOS QC for the 1st QC of hydrographic data (Velo et al., 2021; Section 2.5.3) was embedded in the data flow of GLODAP since v2.2019. Its utilization improved the QC of the data itself, enabled a transparent and traceable flagging of data, and also directly addresses F2 of FAIR (Section 1.3.2). Regarding the 2nd QC of GLODAP, developments of a Python-based crossover tool are still ongoing. The existing beta version already shows significant improvements in computational speeds, and user-friendliness, and provides much more flexibility in the crossover analysis (e.g. manually excluding questionable crossover-pairs for calculating total mean offsets). Further progress is initiated by working towards a direct connection between the Python-based crossover tool and the GLODAP adjustment table 23 implementing a more automated, streamlined, and interoperable data flow. In this context, the currently applied merging Python routine (Section 2.6) that harmonizes, applies flag changes, applies adjustments, assigns QC flags, interpolates, estimates missing variables, and produces the regional and global GLODAP datasets, also represents a major improvement resulting from this thesis. The routine generally follows the “rules” set out in Key et al. (2004) and Olsen et al. (2016). However, apart from erasing detected “bugs” in the product (e.g. wrongly assigned flags) and/or code and adding extra columns to the data product (DOI, expocodes, SF6 2nd QC flags), several important improvements have been implemented in the merging routine since the release of GLODAPv2.2019: • Approximating bottom depth using ETOPO1 (Amante and Eakins, 2009) instead of the Terrain Base (National Geophysical Data Center/NESDIS/NOAA/U.S. Department of Commerce, 1995) • Adaption of calculation scheme for missing carbon system variables due to the inclusion of fCO2 • Inclusion of conversion routine to calculate fCO2 from pCO2 values • Calculating neutral density following Jackett and McDougall (1997) instead of using the polynomial approximation of Sérazin (2011) 23 https://glodapv2-2022.geomar.de/ Synthesis 139 The last major software development of GLODAP is the utilization of the DigitalEarthViewer 24 for visualizing GLODAP. The viewer enables a 4D presentation of all data included in the bias-corrected products, as well as of the mapped GLODAP climatology. Moreover, the development of the DigitalEarthViewer in combination with the individual files generated by the merging routine, makes an interactive and flexible extraction system that enables customized sub-setting, as well as the provision of synthesized unadjusted data, more feasible. All of these software developments further follow the guideline to use open-source software. Besides these software developments, it is also important to mention the development and dissemination of an official GLODAP cruise submission requirement document that includes clearly articulated mandatory and optional requirements for inclusion into GLODAP. For the future, GLODAP developed a clear vision that builds upon these recent developments: "[the] GLODAP team now strive for advancements on two fronts towards a semi-automated system that reduces the work intensity and associated errors. Firstly, implementing a uniform, semi-automatic, and standards-compliant data ingestion system that will facilitate the data submission and quality control (QC) procedures. […] Secondly, upgrading to a modern and versatile data extraction system that provide users more flexibility and options […] "(Tanhua et al., 2021). The overall vision is to reduce existent bottlenecks, manual work intensity, and associated errors, and implement fully FAIR data. 7.1.4 Expanding the BGC EOV synthesis product landscape to previously overlooked observations Even though the spatial footprint of fixed time-series stations is still limited (10% - 15% of the global ocean, Henson et al., 2016), time-series programs represent one of the most powerful vehicles for monitoring marine ecosystem changes. During the past decade, several studies underlined their collective value regarding our understanding of BGC phenomena, e.g. of ocean acidification (e.g. Bates et al., 2014, O'Brien et al., 2017). However, the BGC ship-based time-series community is as of now neither an official GOOS network nor had the community clearly articulated community standards for data management. Collaborating closely with the recently established METS-RCN and generating the pilot of SPOTS (Section 6), addressed these shortcomings and followed the mandate to work towards fit-for-purpose ocean BGC time-series data (Benway et al., 2020, Telszewski and Palacz, 2022). For the generated pilot the focus was set on BGC ship-based time-series programs that measure BGC EOVs, the latter focus links to the general concept of FOO for GOOS (Section 1.3.1). In total, 108,332 water samples of 12 ship-based time-series programs, representative for different marine environments, and different program structures, were included. Besides the collection, and harmonization of data and metadata, the synthesis included optional 1st QC, and archiving, as well as the implementation of a set of 2nd QC methods. Both 1st QC and 2nd QC were developed with BGC shipbased time-series in mind. The 2nd QC constituted of three complementing approaches, all aiming at comparable data (Section 6): 1. A qualitative assessment of the applied methodologies: This entailed the development of community-agreed method recommendations, and particularly resulted in significant improvements in metadata documentation of individual time-series programs. Overall, “the most recently applied methods typically met the required BPs, but measurements of oxygen and pH still show room for improvement”. 2. Comparisons to GLODAP: For the comparisons the crossover routine was adapted and employed. Generally, a good consistency between SPOTS and GLODAP could be identified even though robust comparisons were limited. 24 https://www.digitalearthviewer-glodap.geomar.de/ [Document text truncated for crawler view.]