CHARACTERISATION OF SEDIMENT PATTERNS AND BENTHIC MEGAFAUNA DISTRIBUTION USING AUTOMATED UNDERWATER IMAGE ANALYSIS DISSERTATION in fulllment of the requirements for the degree of Doctor rerum naturalium (Dr. rer. nat) of the Faculty of Mathematics and Natural Sciences at the Christian-Albrechts-Universität zu Kiel Submitted by BENSON ONYANGO MBANI GEOMAR Helmholtz Center for Ocean Research Kiel Kiel, October 2024 1
Examiners: 1. Prof. Dr. Jens Greinert 2. Prof. Dr.-Ing. Reinhard Koch Date of disputation: 25.10.2024 2
“How we spend our days is, of course, how we spend our lives” Annie Dillard 3
Abstract Seaoor habitat classication and marine biodiversity assessment studies are core fundamental activities in marine science research. These studies are crucial because they provide foundational data for the establishment of accurate and reliable baseline marine information. Established marine biodiversity and habitat baselines provide a sound scientic basis for tracking changes in ecosystem health, thereby supporting evidence-based decision making for the overall protection of marine environments. Mapping these vast remote marine ecosystems is typically achieved using (medium resolution) acoustics methods, whereas ground truthing and detailed investigations of specic target sites are usually performed using high resolution optical imaging. Recent technological advancements in imaging sensors and storage memory have enabled the acquisition of huge volumes of seaoor images even from a single scientic expedition. While marine scientists have used images to study marine ecosystems for decades, the manual annotation approaches that were traditionally used to interpret the images are no longer feasible in this marine big data regime. To complement these manual approaches, automated image analysis workows are required to not only improve the visibility of degraded raw images, but to also expedite the transformation of these terabyte-scale seaoor images (and videos) into semantic habitat classes or megafaunal taxa. The workows should also implement methodologies for investigating potential sampling and scaling biases e.g as a result of varying altitude (or speed) of the imaging platform, as well modules for correcting artifacts introduced by these biases. These corrections involve e.g the standardization of visual footprints that vary depending on respective image acquisition heights, or the establishment of xed-size sampling units within which to convert absolute megafaunal counts into abundances that are normalized relative to actual observed seaoor area. Finally, the workows must also allow for seamless integration of the generated annotations with spatio-ecological models. This integration is crucial because it provides geographic context to the image-derived annotations, which in turn enables comprehensive characterisations of multi-scale spatial distribution patterns of habitats and biodiversity. The integration also allows for an objective assessment of environmental drivers (e.g bathymetry, temperature, salinity, etcetera) that inuence the observed distribution patterns. Therefore, this thesis implements integrated workows centered around the above mentioned aspects, and reports detailed ndings based on specic case studies from scientic expeditions to the Clarion-Clipperton Zone in the Pacic, as well as the tropical North Atlantic. 5
Zusammenfassung Studien zur Klassizierung von Lebensräumen am Meeresboden und zur Bewertung der marinen Biodiversität sind grundlegende Kernaktivitäten der meereswissenschaftlichen Forschung. Diese Studien sind von entscheidender Bedeutung, da sie grundlegende Daten für die Erstellung genauer und zuverlässiger Basisinformationen über die Meere liefern. Festgelegte Basisdaten zur marinen Biodiversität und zu Lebensräumen bieten eine solide wissenschaftliche Grundlage für die Verfolgung von Veränderungen der Ökosystemgesundheit und unterstützen so eine evidenzbasierte Entscheidungsndung zum allgemeinen Schutz der Meeresumwelt. Die Kartierung dieser riesigen abgelegenen Meeresökosysteme erfolgt in der Regel mithilfe von (mittlerer Auösung) akustischen Methoden, während die Bodenwahrnehmung und detaillierte Untersuchungen bestimmter Zielstandorte normalerweise mithilfe hochauösender optischer Bildgebung durchgeführt werden. Jüngste technologische Fortschritte bei Bildsensoren und Speichermedien haben die Erfassung riesiger Mengen von Meeresbodenbildern sogar von einer einzigen wissenschaftlichen Expedition ermöglicht. Während Meereswissenschaftler seit Jahrzehnten Bilder zur Untersuchung mariner Ökosysteme verwenden, sind die manuellen Annotationsmethoden, die traditionell zur Interpretation der Bilder verwendet wurden, in diesem marinen Big-Data-Regime nicht mehr durchführbar. Als Ergänzung zu diesen manuellen Ansätzen sind automatisierte Bildanalyse-Workows erforderlich, die nicht nur die Sichtbarkeit von Rohbildern mit schlechter Qualität verbessern, sondern auch die Umwandlung dieser Meeresbodenbilder (und -videos) im Terabyte-Maßstab in semantische Habitatklassen oder Megafauna-Taxa beschleunigen. Die Workows sollten auch Methoden zur Untersuchung potenzieller Stichprobenund Skalierungsverzerrungen implementieren, die beispielsweise durch unterschiedliche Höhen (oder Geschwindigkeiten) der Bildgebungsplattform entstehen, sowie Module zur Korrektur von Artefakten, die durch diese Verzerrungen verursacht werden. Diese Korrekturen umfassen beispielsweise die Standardisierung visueller Fußabdrücke, die je nach den jeweiligen Bildaufnahmehöhen variieren, oder die Einrichtung von Stichprobeneinheiten mit fester Größe, innerhalb derer absolute Megafauna-Zählungen in Häugkeiten umgewandelt werden können, die relativ zur tatsächlich beobachteten Meeresbodenäche normalisiert sind. Schließlich müssen die Workows auch eine nahtlose Integration der generierten Anmerkungen in räumlich-ökologische Modelle ermöglichen. Diese Integration ist von entscheidender Bedeutung, da sie den aus den Bildern abgeleiteten Anmerkungen einen geograschen Kontext verleiht, was wiederum eine umfassende Charakterisierung mehrskaliger räumlicher Verteilungsmuster von Lebensräumen und Artenvielfalt ermöglicht. Die Integration ermöglicht auch eine objektive Bewertung von Umweltfaktoren (z. B. Bathymetrie, Temperatur, Salzgehalt usw.), die die beobachteten Verteilungsmuster beeinussen. Daher implementiert diese Arbeit integrierte Arbeitsabläufe, die sich auf die oben genannten Aspekte konzentrieren, und 6
berichtet über detaillierte Ergebnisse basierend auf spezischen Fallstudien von wissenschaftlichen Expeditionen in die Clarion-Clipperton-Zone im Pazik sowie in den tropischen Nordatlantik. 7
List of Abbreviations A.I Articial Intelligence ADCP Acoustic Doppler Current Proler ANOSIM Analysis of similarities AUV Autonomous Underwater Vehicle BIIGLE Bio-Image Indexing and Graphical Labeling Environment CCZ Clarion-Clipperton Zone CLAHE Contrast Limited Adaptive Histogram Equalization CTD Conductivity, Temperature, and Depth GPU Graphics Processing Unit LED Light Emitting Diode LISA Local Indicators of Spatial Association nm-MDS Non-metric Multidimensional Scaling OFOS Ocean Floor Observation System PCA Principal Components Analysis POC Particulate Organic Carbon ROV Remotely Operated Vehicle R-CNN Region-based Convolutional Neural Networks RV Research Vessel SIMPER Similarity Percentages USBL Ultra Short Baseline XOFOS Extended Ocean Floor Observation System 8
(Freestone et al., 2011). The decline in marine biodiversity could also be attributed to anthropogenic factors such as bottom trawling, pollution, introduction of invasive species, and human-induced climate change which raises ocean temperatures and causes acidication (Sala and Knowlton, 2006) To better quantify and address this biodiversity decline, globally coordinated eorts are required to support an increase in not only the number of scientic expeditions, but also spatial-temporal extents covered by respective sampling campaigns(Canonico et al., 2019). Specically for the deep sea, some of these coordinated initiatives should be directed towards the development of acoustic and optical imaging platforms that allow for detailed non-invasive investigations of remote benthic ecosystems (Levin et al., 2019). To maximize the resourcefulness of the acquired images, videos, and auxiliary sensor readings, corresponding investments are also required to incentivise innovation in marine data science workows (Guidi et al., 2020). These innovations would signicantly expedite the process of transforming the huge volumes of (potentially unstructured) raw marine datasets into actionable insights, which can be used to inform policy decisions e.g regarding delineation of marine protected areas. Additional technological investments should also be directed towards upgrading computational infrastructures such as federated data portals and GPU clusters (Guidi et al., 2020). Recent advances in data science methods have demonstrated remarkable capability for the ecient extraction of semantic annotations from terabyte-scale underwater images and videos (Mbani et al., 2023a, 2022a). State-of-the art computer vision models that have been pre trained on large benchmark datasets (Deng et al., 2009) can now be freely downloaded from repositories such as TensorFlow Hub, PyTorch Hub and HuggingFace Model Hub (Pang et al., 2020) (Paszke et al., 2019) (Jain, 2022). Downloaded models can be straightforwardly deployed as-is to perform common image analysis tasks such as visibility enhancement, classication, segmentation, object detection and tracking. Considering that pretrained models were originally trained using a large set of terrestrial images containing common objects (Shin et al., 2016), the models become very eective after ne-tuning against marine-specic images (Luo et al., 2019). This is because pretraining already exposes the models to diverse abstract features and patterns in terrestrial images, causing the ne-tuning to quickly converge after adapting the model to the unique properties of the marine environment (Pedersen et al., 2019). However, the ne-tuning process still requires a substantial amount of example images that are each labeled (e.g with a corresponding habitat class or megafaunal taxa). This requirement proves challenging in marine sciences, in part because annotating images after every expedition is costly and unscalable (Mbani et al., 2023a). Besides, marine environments (especially in the deep sea) naturally exhibit patterns of reducing megafaunal abundance with depth (Haedrich et al., 1980) (Saeedi et al., 2022). As a result, the majority of acquired images do not contain visible megafauna, and even those with megafauna are typically unknown to annotators ahead of time (Mbani et al., 2023a). As a result, inspecting the entire image dataset sequentially just to nd 15
the few with visible megafauna to be annotated is not only laborious, but also undesirable. This challenge calls for the development of automated annotation workows that oer high accuracy, eciency, repeatability, and that can scale to even bigger-sized marine datasets. The automated workows must also be capable of integrating with spatio-ecological models to geographically contextualize the image-derived annotations. Figure 2: Illustration of the clear upward trend in the performance of computer vision models relative to the human performance baseline. For clarity, the gure only shows baseline comparisons for image classication (imageNet Top-5) models. Data for performance benchmarks obtained from the AI Index Report 2024 (Maslej et al., 2024) The distribution of benthic megafauna in geographic space is neither random nor uniform (Legendre and Fortin, 1989). Instead, biotically-similar megafauna generally exhibit geographic clustering patterns that are in turn inuenced by variability in environmental factors such as bathymetry, geomorphology and temperature gradients, as well as ecological processes such as POC ux, predator-prey interactions, breeding patterns etcetera (Lacharité and Metaxas, 2018). Accurate characterization of these megafaunal distribution patterns is therefore required in 16
order to develop fundamental baseline data for monitoring the health of marine ecosystems (Harris and Baker, 2012), as well as for developing predictive models that simulate e.g potential environmental impacts of deep sea polymetallic Mn-nodule mining (Peukert et al., 2018). Therefore, this thesis builds upon the foregoing foundational conceptual frameworks. The overarching goal of the thesis is to develop a set of automated data science workows that are capable of quick but objective semantic annotation of benthic habitat types and megafaunal taxa from large sequences of high-resolution optical images. The generated annotations are then used downstream to characterize neand regional-scale spatio-ecological distribution patterns. Examples are provided based on case studies from deep seabed areas of the Pacic and tropical North Atlantic. 1.1 Motivation and Objectives 1.1.1 Marine Science Perspective Technological advances in marine sciences are driving up demand for marine data science expertise (Guidi et al., 2020). In particular, underwater imaging platforms are nowadays equipped with multiple sensors that enable detailed investigations of large extents of marine ecosystems (Purser et al., 2019). These surveys typically generate huge volumes of high-resolution seaoor images (and videos) that require both streamlined data management protocols (Schoening et al., 2018), as well as automated data analysis workows that are built upon emerging digital technologies such as data science (Guidi et al., 2020). Examples of these marine technological advances include tethered platforms such as the Ocean Floor Observation Systems (OFOS) (Purser et al., 2019) and Remotely Operated Vehicles (ROVs) (Huvenne et al., 2018). These tethered platforms are typically connected to the ship using an umbilical ber optic cable that delivers both power and propulsion to the platform, while simultaneously transmitting live video feed back to the computer on the ship for real time annotation using e.g the OFOP software. The OFOS is usually towed behind the ship along predened waypoints, which makes it suitable for visual investigations of the seabed along linear transects (Mbani et al., 2022a). In contrast, ROVs are typically piloted and therefore possess the maneuverability to investigate complex benthic terrains, including specic sites of interest as deemed t by the marine scientist (during the survey). Non-tethered imaging platforms such as Autonomous Underwater Vehicles (AUVs) operate independent of direct control by an operator, and can therefore be programmed to self-complete pre planned missions for surveying large areas of marine ecosystems using both acoustics and optical imaging modalities (Huvenne et al., 2018). This autonomy of AUVs allows other sampling operations to proceed in parallel, which maximizes ship time, widens the range of marine observations, and increases the productivity of 17
the expedition overall. In addition, the acquired marine datasets can also be georeferenced because the imaging platforms also record time-indexed navigation information based on either acoustic transponders (Hegrenæs et al., 2009) or ultrashort baselines (USBL) (Rigby et al., 2006). Depending on the specic marine science question under investigation, the platforms can either be deployed ad-hoc, or as networks of interconnected long-term observation stations (Wang et al., 2022). The above marine technological advances generate marine big datasets that are normally curated and archived on federated data portals such as PANGAEA(Felden et al., 2023) and BIIGLE (Langenkämper et al., 2017). Therefore, the need to eciently transform these marine big datasets into actionable insights is a strong motivation for investing in marine data science. Figure 3: Extended Ocean Floor Observation System (XOFOS) being lowered into the water column to survey deep seabed areas of the tropical North Atlantic during cruise M182. Marine scientists need to perform routine characterization of habitats and biodiversity because these are key indicators of a healthy and functioning ecosystem (Galparsoro et al., 2014). To 18
eectively perform the characterizations, methods are required to convert qualitative visual information (in seaoor images and videos) into objects that marine researchers can manipulate computationally e.g through statistical and process-based modeling (Durden et al., 2016). This is a strong motivation for adopting data science techniques such as multi-scale feature extraction (Lu et al., 2023), which allow for a straightforward encoding of abstract visual information into compact vector representations that can be manipulated computationally. For example, hand-crafted texture and entropy features are capable of representing the distribution of Mn-nodules that appear on seaoor images as dark near-circular patches (Mbani et al., 2022a). On the other hand, pre-trained convolutional neural networks such as Inception V3 have shown remarkable capabilities for encoding regions of the seabed exhibiting subtle variability in substrate composition (Szegedy et al., 2015). Marine scientists are also interested in identifying microhabitats that are nested within the larger dominant habitats e.g to understand processes that drive habitat selection decisions by species (Vaudo and Heithaus, 2013). These microhabitats can be easily detected through data science methodologies for anomaly and novelty detection e.g the isolation Forest (iForest) algorithm (Liu et al., 2008). Marine scientists also need high-resolution habitat maps to accurately parameterize ecological models, since these maps allow modelers to avoid making incorrect assumptions regarding random or uniform spatial distribution of substrate characteristics (Evans et al., 2015). These habitat maps can be accurately produced using supervised machine learning techniques like deep learning, which have so far demonstrated capacity to accurately determine both linear and nonlinear decision boundaries for classication (provided there are sucient labeled training examples) (Shin et al., 2016). Even in the absence of labeled annotations, unsupervised methods such as K-means have proven useful for quick preliminary sorting of images based on natural groupings, after which semantic habitat classes can be manually assigned later (Mbani et al., 2022a). Benthic biodiversity assessments are also required by marine scientists to understand the status of marine ecosystem functioning and services (Galparsoro et al., 2014) . To this end, data science models such as Faster R-CNN (Ren et al., 2015) and YOLO (Du, 2018) have demonstrated capacity for automated detection, localization and counting of organisms from images and videos with high accuracy. Finally, there is a need by marine scientists to formulate hypotheses explaining the inuence of environmental variables on the observed spatial distribution patterns of habitats and biodiversity (Legendre and Fortin, 1989). These needs can be addressed through dimensionality reduction techniques in data science (e.g principal components analysis and non-metric multidimensional scaling) (Bakker, 2024). The fact that classical marine science workows already employ multivariate statistics for modeling phenomena simplies integration with data science methods (also deeply rooted in statistical formulation) (“Community ecology in the age of multivariate multiscale spatial analysis - Dray - 2012 - Ecological Monographs - Wiley Online Library,” n.d.). For example, marine scientists typically map out statistically signicant hotspots of megafaunal abundances 19
in ne-scale using spatial autocorrelation techniques (local Morans’ I), which is a special case of autocorrelation in statistics. Also, marine scientists typically assess interrelationships between taxa and environmental factors by projecting the variables onto a two-dimensional ordination feature space e.g using non-metric multidimensional scaling (Bakker, 2024); these ordination techniques are a special case of dimensionality reduction in data science. This seamless integration allows marine researchers to appropriately leverage the strengths of either marine science and data science workows whenever appropriate. Specically, pattern recognition capabilities of data science models can be used to automate time-consuming error-prone tasks like cleaning, missing value imputation, and semantic annotation of potentially unstructured marine datasets. On the other hand, spatio-ecological models that are typically formulated around mechanistic simulation of ecological processes (and have been tested over several decades) can be used to provide nuanced interpretations to observed patterns (Cuddington et al., 2013). Beyond annotations, marine scientists can use (aspects of) data science formulations as an alternative for some spatio-ecological models that by default make unrealistic statistical assumptions of normality among input variables (Legendre and Fortin, 1989). In such situations, algorithms like random forests and deep learning can be used since they do not make assumptions about the underlying data distribution. The need to comply with binding regulatory requirements aimed at protecting marine environments (at both regional and global scale) is another motivation for data science adoption in marine sciences (Guidi et al., 2020). Examples include: The Marine Strategy Framework Directive that requires EU member states to achieve Good Environmental Status (GES) in their jurisdictional marine environments e.g by incentivizing routine surveys of marine ecosystems to monitor biodiversity (European Commission. Joint Research Centre and International Council for the Exploration of the Sea (ICES), 2010). Additionally, the European Green Deal initiative that aims to ensure climate neutrality in Europe by 2050 also requires mandatory periodic reporting of the impacts of pollution and climate change on marine biodiversity (Fetting, n.d.). Beyond the policy level, the European Marine Observation and Data Network (EMODnet) that curates authoritative marine datasets also motivates the need for enhanced data processing capabilities of data science algorithms (Martín Míguez et al., 2019). These data science methods would enable EMODnet (stakeholders) to eciently transform raw marine data into reliable actionable insights that support evidence-based decision making. Outside of the EU, the International Seabed Authority (ISA) that is responsible for regulating deep sea mining of polymetallic Mn-nodules also has a duty to ensure that mining activities are conducted in an environmentally sustainable manner (Bräger et al., 2020). In this regard, the ISA will inevitably motivate the adoption of data science technologies in marine sciences e.g by funding and facilitating independent scientic investigations by academia, requiring comprehensive impact assessment reports from mining companies, and detecting unauthorized mining activities and 20
potential violations of the (forthcoming) Mining Code (Lodge, 2011). These initiatives from the broader marine science community will undoubtedly motivate adoption of data sciences. 1.1.2 Data Science Perspective Supervised machine learning models are generally mature, well-established, and produce high accuracy whenever enough labeled data is available (Burkart and Huber, 2021). As a result, these models readily nd applicability in regression or classication analysis of (partially) annotated marine datasets. In the context of marine sciences, the annotations can be obtained either in real time during image acquisition (e.g using the OFOP software), or during dedicated annotation sessions through platforms such as BIIGLE (Langenkämper et al., 2017). The high accuracy capabilities of supervised models is desirable because it ensures the reliability of fundamental components of marine science research e.g seaoor classication (Mbani et al., 2023b) and biodiversity assessment (Mbani et al., 2023a). Accurate quantication of marine biodiversity and habitat heterogeneity is essential because these indices support evidence-based decision making in marine sciences regarding e.g estimation of spatial coverage and abundance of polymetallic Mn-nodules (Schoening et al., 2017), as well as protection of vulnerable marine ecosystems from anthropogenic activities such as unregulated exploitation of Mn-nodules (Glasby, 2002). In addition, accurate baseline data on habitats and megafaunal abundance provides a strong foundation upon which marine researchers can track changes to justify follow up investigations. Also, supervised machine learning techniques such as Generalized Linear Models (GLMs) and Bayesian models allow marine researchers to explicitly specify the statistical distribution of the data to be analyzed (e.g image-derived annotations) (van de Schoot et al., 2021). This ability to decide the data distribution is useful for modeling in marine sciences, where measured biotic and abiotic variables (megafaunal taxa counts, abundances, temperature, depth etc) do not always follow a gaussian distribution as assumed by most classical models (Legendre and Fortin, 1989). Another motivation is that traditional supervised machine learning models such as random forests and support vector machines are computationally cheap (compared to e.g deep learning) (Li et al., 2016), and also do not require a lot of labeled examples. Considering also that most of these traditional machine learning algorithms are designed to operate on data matrices (rows and columns), the models nd ready applicability in marine data analysis where available annotations are typically organized (or converted) in tabular format. Finally, training supervised models involves learning a function that maps from inputs to class labels. This training process is achieved by optimizing an objective function, which formalizes our (usually subjective) beliefs regarding the nature of the relationship between inputs and corresponding labels. Therefore, the fact that it is possible to (mathematically) customize this objective function to capture nuances and assumptions in the 21
marine domain is a strong motivation for incorporating supervised machine learning in marine sciences. Unsupervised machine learning, on the other hand, is useful for discovering underlying structure and hidden patterns from unlabeled dataset (Hachaj and Mazurek, 2020). This is already a strong motivation to incorporate unsupervised methods into marine science workows, where the majority of (image and video) datasets are unlabeled. The ability of unsupervised clustering algorithms to automatically group together visually similar images also nds ready applicability in marine sciences. This is because clustering reduces the complexity of marine datasets, which means that domain scientists only need to interpret and assign semantic labels to the few generated clusters rather than manually inspecting the entire dataset. In marine image analysis for example, these clusters would typically be obtained by rst encoding visual information from images onto a potential high dimensional feature vector (Bengio et al., 2014), and then using e.g euclidean or cosine metric to measure how similar the vector representations are to each other in feature space. The assumption here is that similar vectors (dependent on the chosen metric) should ideally map close to each other in feature space, and can therefore be easily detected as distinct groupings using a clustering algorithm like K-means. This capability to cluster images based on vector similarity makes unsupervised methods readily applicable to marine sciences, because it allows the the same dataset to be represented (and visualized) in multiple ways depending on the use case e.g based on biotic composition, substrate characteristics, water mass properties etcetera. The only requirement is that the features extracted from images be representative and capable of encoding the phenomena of interest. For example, hand-engineered texture, color, and entropy features may be suitable for representing heterogenous seabed covered with densely distributed Mn-nodules (Mbani et al., 2022b), whereas abstract features extracted from pre-trained deep learning models may be more suitable for capturing subtle variability in an otherwise homogeneous seabed of the abyssal plains (Bengio et al., 2014). Therefore, combining vector space representation with clustering techniques allows for quick, but objective sorting of potentially unstructured marine dataset into arbitrary categories. Furthermore, the ability to detect patterns based on vector representation nds direct applicability in the detection of anomalous patterns in marine sciences (e.g rare taxa or ne-scale microhabitats), since these unusual attributes would be grouped together into a cluster that deviates signicantly from the overall pattern (Mbani et al., 2023a). Finally, unsupervised dimensionality reduction techniques such as principal components analysis are able to eciently project the high-dimensional vector representations onto a 2D feature space (Mika et al., n.d.). This unlocks the capability to graphically visualize e.g the entire seaoor at-a-glance, which is very appealing in marine science domains because it can inform decision making e.g where to sample next during a cruise, or which regions of the working area have been overor under-sampled. 22
Weakly supervised machine learning methods comprise a relatively recent family of models with a promising ability to learn patterns from noisy or imprecisely labeled datasets (Yi et al., 2022). This ability nds direct applicability in marine science use cases, where annotations are available from previous unrelated studies, except not in a format that is machine learning-ready. An example would be megafaunal taxa that were comprehensively annotated on a photomosaic using mouse clicks (point annotations), yet training a standard object detection model requires annotated bounding box coordinates. The promising capability of modern weakly supervised models to learn a signal from such less-than-perfect annotations provides a strong motivation for considering their incorporation into marine science workows. This is because the weakly supervised models will unlock the potential to productively re-use marine datasets that were acquired, annotated, and archived in institutional data management portals over several decades of sea going activities. The wide availability of well documented open source data analysis frameworks have signicantly lowered the barrier for implementation of machine learning models (Pedregosa et al., 2011). Tasks such as data cleaning and exploratory data analysis that previously required writing several lines of code can now be achieved through straightforward function calls to stable and production-ready libraries such as pandas and matplotlib (Mckinney, 2011). Additionally, numerical computing libraries such as SciPy, scikit-learn and NumPy expose ultra-optimised routines for data matrix manipulation and statistical hypothesis testing (Pedregosa et al., 2011). Recently, deep learning libraries such as PyTorch and TensorFlow have abstracted away most of the low-level coding requirements that previously discouraged adoption by scientists from outside the computer science domains (Paszke et al., 2019) (Pang et al., 2020). In parallel, geospatial analytics has also witnessed remarkable progress in the open source space e.g through regularly updated desktop GIS softwares like QGIS, as well as through well maintained programmatic computing libraries like PyGMT and GeoPandas for vector processing, and Rasterio for reading and writing geographic raster les directly into NumPy arrays. Once rasters are loaded in NumPy format, image processing libraries such as OpenCV and torchvision and scikit-image greatly simplify tasks such as superpixel segmentation and ne-tuning of convolution neural networks. Furthermore, frameworks such as Dask and joblib have simplied the process of scaling parallelizable workows to multiple CPU cores or even to distributed computing clusters. The interoperability of these (and other) scientic computing libraries (e.g through intermediate data structures like NumPy) have also standardized data science workows, which is a strong motivation for integration with marine science workows. 23
1.2 State of research 1.2.1 Niche of the thesis This study belongs to the wider discipline of remote sensing of marine environments. In this context, remote sensing generally refers to the use of various imaging (or sensor) modalities such as optical images, videos and acoustics to document the ecology, biology and geology of marine ecosystems. Specically, this thesis conceptualizes and implements automated data science workows that eciently generate semantic annotations from large sequences of high resolution optical images, with specic case studies from the Pacic and tropical North Atlantic. The data science workows implemented in this thesis are also seamlessly integrated with multivariate spatio-ecological techniques in order to provide geographic context to the image-derived annotations. In this way, this thesis provides a holistic framework for characterizing habitats, megafaunal abundance, as well as environmental drivers that inuence these patterns in marine ecosystems. Therefore, the niche of this thesis is at the intersection of marine sciences and data sciences (specically computer vision). 1.2.2 Contribution to science This thesis makes the following contributions to the marine imaging community: a) Image processing: A novel workow was developed for automatically detecting laser points from a large sequence of deep sea images. The novelty of this implementation lies in the seamless integration of image processing (band arithmetics) and geometric set theory (intersection) methodologies into an end-to-end workow that is computationally fast, accurate and easily scalable through concurrent execution across multiple CPU cores. The utility of the implemented laser point detection workow is demonstrated here through its ability to guide the automatic selection of a reference image with the highest scale for use in color normalization through histogram matching (in the absence of logged camera altitude information). The detected laser points also demonstrate applicability in the determination of an optimal reference scale relative to which the altitude-dependent visual footprints of respective images can be standardized. 24
This involved the use of choropleth maps to qualitatively pick out obvious patterns of geographic clustering, as well as the use of spatial autocorrelation analysis to quantitatively map out statistically signicant hotspots of megafaunal abundance. Environmental drivers that inuenced these observed distribution patterns were also investigated using ordination techniques. The ordination involved the use of non-metric multidimensional scaling to project both biotic and abiotic variables onto a two dimensional feature space, allowing for a straightforward graphical assessment of inter-relationships among the variables. Finally, image datasets, annotations, and software code that were either acquired or generated form in this thesis have all been published to pangaea, gitlab, or as supplementary materials associated with respective peer reviewed publications. These research products are freely available for use and re-use by the marine scientic community, with the aim of promoting transparency, open science and collaboration to advance the discipline of marine data science. 1.4 Outline of scientic chapters and declaration of contribution 1.4.1 Scientic Paper 1: Implementation of an automated workow for image-based seaoor classication with examples from manganese-nodule covered seabed areas in the Central Pacic Ocean (Mbani et al. 2022) This chapter presents the Automated and Integrated Seaoor Classication Workow (AI-SCW) that was implemented to partition deep seabed areas in the Pacic into semantic habitat classes. The classication was based on visual information extracted from a sequence of optical images collected along linear transects, using the OFOS imaging platform that sampled the seabed at a constant frequency of 0.1 Hz. This chapter provides comprehensive descriptions for all constituent components of the AI-SCW workow. These include automatic laser point detection, determination of altitude-dependent image scales, visibility improvement transformations, and standardization of visual footprints. Training schedules and performance evaluation metrics (for both supervised and unsupervised seaoor image classiers) are also provided. Finally, spatial distribution patterns of the observed seaoor habitats are analyzed in the context of multibeam bathymetry (and its derivatives). I conceptualized the workow, wrote the software code for all the image processing, seaoor classication, geospatial analysis tasks. I also wrote the manuscript. 31
1.4.2 Scientic Paper 2: An automated image-based workow for detecting megabenthic fauna in optical images with examples from the Clarion–Clipperton Zone (Mbani et al. 2023) This chapter comprehensively describes the Megabenthic Fauna Detection with Faster R-CNN (FaunD-Fast) workow, which was implemented to automatically detect, localize and classify megafauna from seaoor images of the Pacic. First, an innovative methodology is presented that semi-automatically generates weak megafaunal bounding box annotations based on analysis of anomalous superpixels. Subsequent sections describe the steps for post processing the anomalies to remove false positives, allowing for manual assignment of semantic morphospecies labels to only the truly anomalous superpixels. Additional sections describe how the auto generated annotations were used to train a state-of-the-art benthic fauna detection model, as well as how the performance of the trained model was evaluated based on the standard COCO detection metrics. The developed model was also benchmarked against comparable state-of-the-art megafauna detection models. Further details are provided for the conversion of absolute megafauna counts into abundances, as well as characterization of the spatial distribution patterns of biota in the context of structuring environmental drivers. I conceptualized the workow and wrote the software code for all the image processing, benthic fauna detection, and geospatial analysis tasks. I also wrote the manuscript. 1.4.3 Scientic Paper 3: Automated image-based workows reveal the composition and spatial distribution patterns of megabenthic communities in the tropical North Atlantic Ocean (Mbani et al., In review SciRep). This chapter integrates the seaoor classication and megafauna detection workows from the rst two scientic chapters. Since both workows were originally developed and tested against seaoor images from the Pacic, one main objective of this study was to investigate the generalizability of the workows when applied to a new dataset from the tropical Atlantic. Rather than focusing on technical algorithmic development, this chapter focuses more on detailed investigations of spatio-ecological distribution patterns of annotated benthic habitats and megafaunal taxa. Descriptions are rst provided for the geologic features, geographic extents and water mass properties of the new working area in the Atlantic, followed by a brief description of the training, evaluation and inference setups for the two annotation workows. Next, the magnitude of sampling, scaling and (megafaunal) double counting biases along respective camera deployment tracks is investigated. Details are also provided for the solution to these biases through the establishment of xed-size sampling units within to normalize absolute 32
megafauna counts relative to actual observed seaoor area. A comprehensive description of the use of spatial autocorrelation analysis to reveal local hotspots of megafaunal abundances is also provided, along with descriptions for how ordination techniques based on non-metric multidimensional scaling were used to graphically assess the inter-relationships between biotic and abiotic variables in feature space. I conceptualized the workow and wrote the software code for all the image processing, seaoor classication, benthic fauna detection, and geospatial analysis tasks. I also wrote the manuscript. 33
2. Scientic Chapters 2.1 Implementation of an automated workow for image-based seaoor classication with examples from manganese-nodule covered seabed areas in the Central Pacic Ocean Metadata: Mbani, B., Schoening, T., Gazis, IZ. et al. Implementation of an automated workow for image-based seaoor classication with examples from manganese-nodule covered seabed areas in the Central Pacic Ocean. Sci Rep 12, 15338 (2022). https://doi.org/10.1038/s41598-022-19070-2 Motivation and Objectives: Classication of the seaoor into habitat categories is crucial towards the establishment of comprehensive marine baseline information. Accurate and regularly updated baseline data in turn supports informed decision making in use cases such as delineation of marine protected areas, assessment of environmental impacts of deep sea mining, as well as tracking changes in marine ecosystem health. Whereas large scale seaoor mapping is typically performed using medium resolution acoustic imagery e.g from multibeam echosounders, optical imaging is the preferred method for detailed close-range investigation of selected sites of interest. In this regard, optical images can serve as ground-truths for evaluating the accuracy and reliability of acoustics-based mapping methods, while also being used independently to reveal subtle variability in habitat characteristics and other localized sedimentological patterns (e.g along survey transects) that may not be detectable at the resolution of acoustics imagery. However, it is costly and infeasible to manually annotate the huge volumes of optical seaoor images that are nowadays acquired during scientic expeditions. Therefore, one primary objective of this study was to develop an automated image-based workow for eciently annotating seaoor images in order to reveal sediment patterns in the Mn-nodule covered deep seabed areas of the Clarion Clipperton Zone. Specic objectives included: 1. To automate the detection of laser points from every image. This was useful for determining altitude-dependent image scales to be used for converting measurements (e.g visual footprints) from pixels to real world metric units (e.g square meters). The distribution of image scales was also used to determine the reference median scaling factor relative to which all images would be resized, ensuring a consistent visual footprint representation on the seabed. 34
2. To color normalize and improve the overall visual quality of raw images. The focus here was to account for degradations resulting from wavelength-dependent attenuation and scattering eects of the propagating LED light (from the imaging platform) through the water column, as well as inconsistencies in scene brightness caused by the varying altitude of the imaging platform during image acquisition. 3. To (semi) automate seaoor classication by mapping each image into one of predened seaoor habitat classes. This involved deploying both supervised and unsupervised image analysis workows that were each trained based on training examples obtained semi-automatically e.g through nearest neighbor sampling in feature space (for supervised classication), as well as stratied cluster-based sampling (for unsupervised classication). 4. To characterize spatial distribution patterns of the observed seaoor sediment patterns. This involved integrating the annotations with other auxiliary georeferenced datasets (e.g depth, slope and terrain ruggedness), with the aim of providing geographic context to image-derived annotations e.g by visualizing clustering patterns on choropleth maps plotted along camera deployment tracks, as well as through statistical assessments of the inuence of environmental drivers on the observed patterns. Materials and Methods: This study used 40,678 high resolution seaoor images from the Clarion-Clipperton Zone (CCZ) in the Pacic. The images were acquired using a Canon EOS 5D Mark IV camera attached to the Ocean Floor Observation System (OFOS) during SONNE cruise SO268 to the German and Belgian contract areas for Mn-nodule exploration. Twelve video investigations conducted in the two areas covered a combined track length of 92.5 km at an average water depth of 4,280 meters. Within the German area, a chain dredge was used to disturb the sediment as part of a small-scale experiment to simulate the potential spatio-temporal extent of re-depositioned sediment plume. The seaoor in the German contract area was therefore photographed twice (before and after the disturbance experiment). Automatic laser point detection involved a combination of image processing and geometrical analysis. First, three visible laser points were annotated from a well illuminated reference image. A 250-pixel buer was then established around the triangular geometry dened by the three 35
laser points, within which laser points from all the images were expected to be located. Next, an image processing pipeline was implemented through a linear combination of (RGB) channels, which produced an intermediate signal (image) for which a local-peak nding algorithm identied potential candidate laser point locations. Finally, the actual laser points were obtained by intersecting the buer mask and the candidate laser point coordinates. The image scale was determined as the ratio between the average distance separating the detected laser points (in pixels) to their calibrated distance (of 40 centimeters). Illumination and color normalization involved four key steps. First, the light cone eect was corrected by pixelwise z-score normalization applied to batches of sequentially ordered images. Second, image contrast was maximized using adaptive histogram equalization that redistributed image intensities in local image tiles resulting in improved overall image contrast. Third, uneven brightness among images was corrected by adjusting the intensity distribution (histograms) of all images to match the histogram of a manually chosen reference image (with good overall scene brightness). Finally, the images were rescaled (relative to the median scale) and then center cropped to represent standardized visual footprints (of 1.6 square meters each) on the seabed. Supervised seaoor classication involved ne-tuning an instance of the Inception V3 convolutional neural network, and then applying the trained model (in inference mode) to predict habitat class labels of respective images. The habitat classes included: Class Seafloor A that comprised images for which the seabed was covered with no or only few Mn-nodules. Images from this class also exhibited visible dredge marks and turned-over sediment indicative of the impact from the chain dredge experiment; Class Seafloor B comprised patchy Mn-nodules that were qualitatively small sized and only partially covered the seabed; Class Seafloor C comprised Mn-nodules that were densely distributed per unit area; and class Seafloor Dcomprised qualitatively large sized Mn-nodules relative to those in classes Seafloor B and C. On the other hand, unsupervised classication was achieved by applying K-means clustering to the data matrix of six-dimensional texture and entropy features extracted from the entire images. The clustering grouped together visually similar images, after which a domain expert manually inspected the clustering (and corresponding image subsamples) in order to assign each cluster a semantic class label. Performance of the classication models was evaluated using standard precision, recall and confusion matrix (for the supervised Inception V3), while the silhouette score was used to assess the quality of clustering (for the unsupervised K-means). Choropleth maps (based on photo center coordinates) were used to visually assess the spatial distribution patterns of seaoor habitats along camera deployment tracks. In addition, boxplots 36
were also generated to show the variability of environmental drivers in respective dives, as well as the correlation between the drivers and dierent habitat classes. Key Findings: The color normalization workow improved the visual quality of images by (a) signicantly reducing the visual eect of gradual reduction of light towards the image edges, (b) maximizing the distribution of pixel intensities to span the entire 8-bit dynamic range, and (c) eliminating the greenish haze to improve the overall sharpness and clarity of images. However, it was observed that laser points were detected with high accuracy from raw images compared to the color normalized images that produced many false positives. This observation was the result of a side eect of the applied adaptive histogram equalization transformation, which desirably improved the overall image contrast but in the process reduced the distinctiveness of the red laser points relative to local background pixels. Another explanation was that the color normalization transformation could have introduced (or amplied) noise in respective images, which reduced the signal-to-noise ratio of the red laser points, thereby making the lasers indistinguishable from the background. Figure 9: Example images showing visibility improvement transformation. (A) shows a well illuminated reference image whose (B) intensity distribution was used to transform (C) a raw 37
image into (D) transformed image Additionally, the use of the same xed-size triangular buer mask to lter down to the three actual laser points from a pool of candidate laser locations allowed for a straightforward concurrent implementation of the laser detection workow. This is because laser point detections in respective images occurred independent of each other, which made it possible to computationally map the workow to multiple CPU cores simultaneously. It was observed that this approach increased processing speeds (by up to a factor of 3) compared to alternative methods involving sequential processing. Projecting seaoor images to feature space and then performing nearest neighbor sampling around a few manual annotations signicantly expedited the process of generating labeled training examples for supervised classication. This approach also allowed for a quick at-a-glance visualization of seabed characteristics, which enabled expert annotators to factor in rich and nuanced contextual information about e.g the properties of neighboring clusters, or whether inter-class transition patterns are subtle or abrupt. Such underlying contextual seaoor patterns and relationships may be missed when images are manually inspected sequentially (in isolation). In addition, the fact that the projection to feature space is deterministic implies that the annotation workow can easily be generalized to other seabed areas, provided that the visual features extracted from images accurately capture the sediment patterns. For example, domain knowledge on the Mn-nodule variability and appearance (as approximately round blobs of dark pixels) may lead to the choice of texture and entropy as appropriate features to encode visual information from seaoor images of the CCZ. This projection-based approach also provides a natural way to prevent class imbalance, since annotators can easily (visualize and) oversample regions of the seabed (or feature space) that are underrepresented, while uniformly sampling the rest of the feature space to ensure balanced representation across habitat classes. While image features extracted using pre-trained models are in general very expressive, from this thesis indicate that in situations where one has an idea of the seaoor characteristics (either after visualizing projections in feature space projection, or from prior domain knowledge), then manual feature selection may produce superior classication results especially in unsupervised setting. 38
Figure 10: Projection of the seaoor into feature space, color coded by respective seaoor classes. Notice how the unsupervised classication produces abrupt class boundaries. Randomly sampling images from a batch of seaoor images to train an unsupervised K-means classier is the best strategy only if the metric being optimized against is computational speed. The downside of random sampling is the class imbalance that may arise when dominant regions of the seabed are overrepresented in the sample, considering that substrate patterns in the deep sea vary relatively slowly (in kilometer scale). Training models with a class-imbalanced training set results in poor generalization performance, or misleading conclusions. Stratied cluster-based sampling was found to be the best strategy when the metric of comparison was the quality of cluster groupings (as measured by the silhouette score). This was because the initial overclustering step allows for a more uniform and representative sampling of the seabed, including anomalous regions e.g regions of the Clarion Clipperton Zone where there are no deposits of polymetallic Mn-nodules. Regardless of the chosen sampling strategy, ndings indicate that unsupervised classication generally produces abrupt class boundaries (in feature space), which is not typical of deep sea environments where inter-class transitions are smooth and gradual. Closer investigations revealed that these sharp class boundaries were caused by the 39
K-means objective function, whose optimization minimizes the within-cluster sum of squares resulting in distinct and well-dened clusters. In contrast, the non convex objective functions that are used in deep learning models (e.g cross-entropy loss) naturally produce fuzzy class boundaries. Moreover, deep learning models typically output probability distributions over predicted class labels rather than single deterministic class labels; these distributions can also be used to fuzzify class boundaries. Still, both results of supervised and unsupervised seaoor classications exhibited good agreement overall, as evidenced by a Cohen's Kappa coecient of 0.6 that indicates the classications are more consistent than would be expected from chance. 40
the parameter inuences how sensitive the model is to anomalies. Thus, although setting the contamination factor to a high value produced many false positives, this choice was preferable in this study since it allowed domain experts to be directly involved e.g to closely inspect false positives, or to determine whether a superpixel showing what appears to be a partially burrowed ophiuroid is indeed anomalous. The auto generated weak bounding box annotations were not sucient (in quality and quantity) for a comprehensive characterisation of benthic biodiversity. This is because the weak annotations were generated using an unsupervised (superpixel-based) approach for which false negatives are highly probable. Therefore, such weak annotations should be regarded as computationally cheap sources of training examples to be later rened, semantically labeled, and eventually used for ne-tuning a state-of-the-art object detection model. For example, this study trained a Faster R-CNN model using semantically labeled bounding box annotations and achieved an accuracy of 78.1% (at an intersection-over-union threshold of 0.5), which was on a par with other state-of-the-art benthic object detectors that were trained using manually annotated training examples. This good performance demonstrates that even though superpixel-based annotations are computationally cheap to generate, they are reliable and therefore easily scalable to even larger image datasets as long as there is sucient computing power. This superpixel-based approach to expedited annotation readily nds applicability in environments where computational capacity may be limited e.g on scientic expeditions on board a small research vessel, or in real-time analysis of live video feed. Benchmarking with other comparable state-of-the-art models showed that while two stage detectors like Faster R-CNN produce high accuracy, the models are slow (in comparison) and therefore not suitable in certain applications e.g on edge devices like NVIDIA jetson. Another key nding was that detection of small-sized megafauna was consistently challenging across all the benchmarked benthic object detectors. This diculty could be because small-sized megafauna occupy fewer pixels and are therefore easily occluded or camouaged relative to the background environment. Another reason could be the technical design of the network, in which the sequence of convolution and pooling operations gradually reduce the resolution of the feature maps. This downsampling ultimately leads to the disappearance of the small-sized megafauna deeper into the network. 47
Figure 14: Examples of annotated megafauna that were automatically detected using the trained Faster R-CNN model. Also shown (in red) is the set of false negatives comprising mostly small-sized objects and those that are visually similar to background seaoor. Megafaunal abundance was generally low (< 1 ind. per sq. m) in the surveyed area of the Clarion Clipperton Zone. In particular, regional comparison showed that the German contract area exhibited higher megafaunal abundances (0.247 ind. per sq. m) when compared to the Belgian contract area (0.200 ind. per sq. m). The German area also exhibited higher diversity of megafauna, with a Shannon diversity index of 2.4 compared to the Belgian area (1.7). Considering that the German area is on average shallower (-4121 m) compared to the Belgian area (-4510 m), this variability in both abundance and diversity could be explainable by the East-to-West reduction in POC ux that avails more food to the German seabed in the form of sinking organic material. 48
Figure 15: Map showing the spatial distribution of megafaunal abundances along respective survey transects in the Clarion-Clipperton Zone of the Pacic. Further investigations revealed that the majority (68%) of megafauna in the Belgian area occupied regions of the seaoor covered with large sized Mn-nodule seabed, despite the region being predominantly covered with both large sized and densely distributed Mn-nodules per unit area. The hypothesis here is that densely distributed Mn-nodules cover most of the seabed 49
surface area, which does not allow sucient space for soft sediment dwellers. This is in comparison to large sized Mn-nodules that leave more space among themselves, thereby creating ne-scale heterogeneity that accommodates both epifauna and soft-sediment dwellers such as crustaceans and echinoderms. Another nding was that ophiuroids and xenophyophores exhibited high abundance and diversity in both the German and Belgian contract areas. The co-occurrence of these morphospecies could be because xenophyophores create complex structures on the seabed that provides habitat and shelter for the ophiuroids. The ophiuroids may also be attracted to the organic materials accumulated in xenophyophore tests. 2.3 Automated image-based workows reveal the composition and spatial distribution patterns of megabenthic communities in the tropical North Atlantic Ocean. Metadata: Mbani, B., & Greinert, J. (In Review, SciRep) https://doi.org/10.31223/X5CQ6B Motivation and Objectives: Detailed documentation of megafaunal communities and their habitat characteristics is key towards the overall monitoring of marine ecosystem functioning. Although there have been past studies that successfully used data science techniques for substrate characterisation and marine biodiversity assessments, most of these models have not been extensively tested in seabed areas other than where they were originally trained. As a result, the robustness and generalization capabilities of these automated workows when applied to varying substrate types, water masses, and megafaunal communities remains unknown. In addition, many studies that develop data science-based workows for benthic characterisation typically focus more on technical aspects such as network architecture design, optimization strategies, and performance benchmarking to demonstrate (sometimes marginal) improvements against existing state-of-the-art models. Few studies go further to provide spatio-ecological interpretations of the image-derived annotations. As a result, there is a need for comprehensive research that integrates both aspects of technological advancements as well as spatio-ecological interpretations. Therefore, a major objective of this study was to investigate the generalization capabilities of two automatic image-based workows for habitat characterization and megafauna annotation. Both workows were originally developed and tested against images of the Pacic (see papers 1 and 2 above). For this study, however, the models were applied to a new dataset from the tropical North Atlantic. 50
Thus, specic objectives for this study included: 1. To automatically annotate habitats and megafaunal taxa in deep seabed areas of the tropical North Atlantic using ne-tuned data science-based workows from previous related studies (see papers 1 and 2). 2. To investigate potential manifestations of sampling and scaling biases along camera deployment tracks, as well as to account for these biases by pooling and normalizing annotations relative to xed-length sampling units. 3. To reveal statistically signicant hotspots, coldspots and other distribution patterns of megafaunal abundance based on analysis of spatial autocorrelation. 4. To investigate the relative contributions of the dierent taxa groups towards similarity and/or dissimilarity of megafaunal communities. 5. To investigate the inuence of environmental drivers (e.g depth, slope, temperature and terrain ruggedness) on the observed megafaunal distribution patterns. Materials and Methods: The dataset for this study comprised 8,838 high resolution seaoor images from an East-West section of the tropical Atlantic located oshore Mauritania and North of Cape Verdes. The images were collected using an XOFOS during cruise M182, whose overall aim was to investigate the inuence of mesoscale eddies on biogeochemical processes and modulation of organic carbon transport in the Eastern boundary upwelling systems. Previously developed Automated and Integrated Seaoor Classication Workow (AI-SCW) was trained using a subset of this new image dataset from the Atlantic, and then applied in inference mode to predict habitat class labels for the entire dataset. Specically, the unsupervised classier implemented in AI-SCW was used to cluster seaoor images into natural groupings based on similarity in visual features, afterwhich the clusters were manually inspected and assigned semantic labels. In addition to the AI-SCW, previously developed Fauna 51
Detection with Faster R-CNN (FaunD-Fast) workow was deployed as-is to automatically detect megafauna from the sequence of images. Since FaunD-Fast was also originally trained using images from the Pacic, the detected objects automatically became weak annotations that needed to be manually rened, semantically labeled, and eventually used as training examples for retraining FaunD-Fast. This retrained version of FaunD-Fast is the model that was used to detect, localize, classify and count instances of megafaunal taxa from the entire image dataset. Graphs showing variability in number of images, sampling speed and observed visual footprint were generated to assess potential sampling and scaling biases along respective survey transects. In addition, the average distance between successive images was compared against the average length of the along-track image axis to check for potential double counting of megafauna due to image overlap; there was overlap if the image length was on average shorter than the distance separating the images. To account for these biases, xed-length (100-meter-long) sampling units were dened within which absolute megafaunal counts were converted into abundances relative to actual observed visual footprint. Qualitatively, the spatial distribution of megafauna was visualized on a choropleth map obtained by color-coding centroid coordinates of respective sampling units based on binned megafaunal abundances. Quantitatively, measures of spatial autocorrelation were calculated to identify statistically signicant hotspots and coldspots of megafaunal abundance. Specically, Moran’s I scatter plot was used to classify Local Indicators of Spatial Association (LISA) statistics, thereby revealing regions of the seabed where megafaunal abundances were signicantly higher or lower compared to average local neighborhood abundances. Finally, non-metric multidimensional scaling (nm-MDS) ordination was used to visually assess inter-relationships among megafaunal taxa groupings and environmental variables. The nm-MDS ordination involved projecting all the 232 xed-length sampling units onto a two-dimensional feature space, and then superimposing environmental drivers (e.g depth, slope, topographic position index, terrain ruggedness, salinity, temperature and longitude) onto the same feature space. The goal of the ordination was to identify the (subset) of environmental factors that potentially explain the observed distribution patterns of megafaunal taxa. This ordination was also used to generate hypotheses about the underlying ecological processes that inuence megafaunal community structures. Key Findings: Automated seaoor classication in the tropical North Atlantic revealed seven clearly distinct habitat classes. On close inspection of subsamples, each cluster was found to contain images from the same dive where the physical characteristics of seaoor substrates was more-or-less similar. Furthermore, projecting the images in feature space revealed that dives that were in close geographical proximity to each other also mapped close to each other in feature space. 52
Color coding the feature space projection with clustering results revealed a left-to-right gradient in the intensity of biogenic activity within an otherwise homogeneous seabed. In particular, the intensity of (visible) sediment reworking was low for clusters on the left half of the feature space, while the degree of sediment disturbance increased signicantly towards the right half of the feature space. Therefore, the rst (horizontal) axis of the nm-MDS clearly partitioned the seaoor based on bioturbation intensity, but it was not immediately obvious what attribute the second axis was partitioning the seaoor against. Additionally, sub-partitions were also observed within most of the major clusters, which could be interpreted as subtle variations in sedimentological properties along the camera deployment tracks at short spatial scales. Figure 16: Feature space representation of seaoor classes in the tropical North Atlantic. Images from the Eastern region that showed visible signs of sediment reworking due to bioturbation are grouped together on the right half of the feature space. Also note the subpartitions within respective dives. A total of 10,189 organisms belonging to 13 taxa groups were detected across both the Eastern and Western regions. Although ndings showed that there was potential double counting of 53
megafauna in dive 19 due to overlapping images, this did not aect abundance estimates because absolute taxa counts in respective sampling units were normalized by dividing relative to actual observed visual footprints (which would also be doubled because of the overlap). In terms of proportions, Foraminifera, Echinodermata and Lebensspuren were the most dominant taxa that accounted for more than 76% of all detections in both Eastern and Western regions. The rest of the taxa groups accounted for less than 10% each in relative proportions, including: Porifera, Arthropoda, Cnidaria, Sponge-Skeleton, Mollusca, Chordata, Annelida, Ctenophora and Chaetognatha. Regional comparisons revealed statistically signicant dierences in both abundances and diversity between the shallower closer-to-shore Eastern region, and the deeper Western region. Specically, the Eastern region exhibited an average abundance of 0.44 (ind. per sq. m) compared to 0.03 in the Western region. Variability in megafaunal abundances (also an indicator of ecological heterogeneity) was higher in the shallower Eastern region (standard deviation = ± 0.35) compared to the Western region (standard deviation = ± 0.02). Similarity percentages (SIMPER) analysis further revealed that 50% of the dissimilarity in biotic composition between the Eastern and Western regions could be explained by only four taxa groups, including: Porifera (14.46%), Lebensspuren (13.27%), Cnidaria (11.97%) and Mollusca (9.65%). The high megafaunal abundances and diversity in the Eastern region might be explainable by the high food availability in the form of sinking organic matter, as well as by the relatively warmer temperatures that enhance metabolic rates of megafauna. 54
Figure 17: Distribution of megafaunal abundances for respective taxa in the tropical North Atlantic. The shallow Eastern region records higher abundances and diversity among all the taxa groups. Megafaunal abundances exhibited clear patterns of geographic clustering at both local and regional scale. The Eastern region exhibited relatively high abundances consistently throughout respective observational transects (except in the relatively at terrains of dive 131). Topographically complex areas in the Eastern region in particular exhibited signicantly higher abundances e.g the sides of the submarine canyon (in dive 144) as well as the top of the seamount in dive 145. The deeper Western region generally exhibited low abundances of megafauna, with topographically complex features here similarly exhibiting statistically signicant hotspots of megafaunal abundances e.g the top of the seamount (in dive 32), as well as the pair of abyssal hills in dive 28. 55
Figure 18: Prole view showing statistically signicant megafaunal hotspots and coldspots along respective survey transects in the tropical North Atlantic. Topographically complex regions of the seabed contained most megafaunal hotspots e.g the sides of the submarine canyon in dive 144. Projecting the sampling units onto a two-dimensional ordination feature space revealed two main clusters of megafaunal communities corresponding to the Eastern and Western regions. Megafaunal taxa that primarily inuenced the deeper Western region included Echinodermata, Foraminifera, and Arthropoda. The remaining ten taxa predominantly inuenced the shallower Eastern region, potentially explaining the high number of biogenic structures (Lebensspuren) observed in the Eastern region. Another observation was that bathymetric drivers such as slope, ruggedness, and roughness predominantly inuenced megafaunal communities in the deeper Western region. This observation could be because topographies in the deeper regions vary slowly in kilometer scale, such that even minor variability in bathymetric derivatives within these deeper seabed areas results in a more pronounced inuence on e.g hydrodynamic eects like current patterns, as well as nutrient distribution. In contrast, the already complex topographies in shallower parts naturally disrupt hydrodynamic ows, so that the eect of minor changes in bathymetric derivatives are not as pronounced. 56
3. Conclusion, challenges and future outlook This thesis conceptualized, implemented and presents a sound framework for integrating data science techniques with conventional marine science methods. Workows such as AI-SCW and FaunD-Fast that were developed in this thesis have so far demonstrated the capabilities of data sciences to automate repetitive and time-consuming aspects of image-based benthic characterisation studies (e.g through accurate annotation of megafauna and sediment patterns). This automation oers signicant speed gains to marine scientists by reducing the human eort required to execute routine tasks, allowing scientists to focus on core scientic research and innovation. Although fully automated end-to-end workows are ecient and even ideal for certain applications e.g periodic environmental monitoring to enforce regulatory compliance (or broadly estimate production capacity) for deep sea polymetallic Mn-nodule mining operations (Volkmann and Lehnen, 2018), one key conclusion from this thesis is that domain expertise must still be incorporated into the analytical workows for a number of reasons: First, the complexity of deep sea environments introduces biases that vary unpredictably from expedition to expedition (or even within a particular dive). As a result, direct application of o-the-shelf image processing and data science techniques (without expert input from marine scientists) will either introduce non-obvious artifacts, or lead to outrightly misleading conclusions. For example, the high degree of sediment disturbance due to biogenic activity in the tropical North Atlantic may alter the composition of suspended particles and organic matter in the water column, which in turn inuences the optical properties of the water (and therefore image quality) dierently compared to the relatively stable and less disturbed sediment in the Pacic; Second, an objective determination of the quality and completeness of autogenerated benthic annotations requires foundational understanding of geologic context, as well as spatio-ecological processes that inuence interactions between biotic and abiotic variables. For example, while the use of anomalous superpixels as weak bounding box annotations did expedite the overall process of generating megafaunal annotations, failure to incorporate domain expertise may easily have caused certain morphospecies to be misclassied as false negatives e.g ophiuroids that partially burrow in sediment to feed on organic detritus (Warner, 1982) until they visually almost resembled background seaoor; Third, the serious nancial and legal consequences that accompany non-compliance with regulatory requirements (e.g E.U Habitats Directive (Evans, 2006)) does necessitate the incorporation of marine domain expertise to provide independent oversight to automated benthic characterisation workows. Collectively, this thesis shows that integration of Marine and Data Sciences should be complementary in nature. Specically, data science techniques should be deployed to eciently transform potentially unstructured raw marine big datasets into information, whereas the task of providing rich, nuanced and contextual interpretations to the auto generated information must be reserved for marine domain experts. 63
However, the adoption of data science technologies in conventional image-based marine science workows is still relatively low. This is due to a number of open challenges: First, raw underwater images are typically not directly usable as inputs to data science models because the images suer from degradations and other biases e.g due to variability in sampling gears, acquisition altitude, lighting conditions, water mass properties etcetera. The choice between whether to account for these degradations separately, or to directly develop generalizable workows that are by-design agnostic to the above variabilities introduces extra layers of complexity that slows down adoption. Future work could explore potential data-centric approaches that organically incorporates marine science assumptions and nuances into the core of neural network architectures e.g by strategically formulating objective functions to be optimized; Second, navigating stringent regulatory landscape (e.g the EU Articial Intelligence Act (Veale and Borgesius, 2021)), as well as competing (inter)national interests may also slow down the mainstreaming of data sciences technologies into marine sciences, especially in sensitive (or contested) applications involving e.g granting or rejection of Mn-nodule mining licenses in the Clarion-Clipperton Zone. Future work could explore the feasibility of exible sandbox approaches that allow for controlled testing of the strengths, weaknesses and risks of emerging data science technologies relative to agreed-upon parameters (e.g data privacy), with the experimentation being conducted under a relaxed version of supervision by regulators (Truby et al., 2022). Lessons learned from these controlled experiments could then inform future regulatory decisions; Third, the lack of comprehensive standardized benchmark datasets for advancing marine data science research also poses adoption challenges. Specically, it is not yet obvious how to account for the concept drift (Langenkämper et al., 2020) that arises from the dierent sampling strategies employed by scientists who collect, annotate and eventually publish their labeled datasets. Future work could explore the development of open protocols that standardize sampling campaigns, as well as possibilities for eciently generating synthetic training datasets based on physics-based rendering (Nimier-David et al., 2019) that faithfully reproduce complexities in marine environments; Fourth, integration of data science techniques into established marine science workows may not be obvious in certain situations, either because of incompatible data formats or non-aligned approaches to annotation between the marine and data science communities. For example, annotations to be used for training data science models are required to be statistically balanced across the dierent taxa classes, which is not necessarily a consideration in marine sciences where the focus is more on revealing existing patterns and formulating hypotheses about underlying processes. Another example of non-alignment is that ‘annotations’ in data science typically means the chosen subset of the entire dataset to be used as training examples, whereas in marine sciences annotations are the actual semantic information that is derived from raw images. In addition, there is usually a default assumption in data science workows to resize (or crop) images to conform to the standard dimensions of the input layer of the convolutional neural network architecture. While 64
trivial in data science workows, this arbitrary resizing operation is consequential in marine sciences because it interferes with the semantic interpretability of image content e.g actual visual footprint on the seabed. Therefore, future work could explore innovative strategies for establishing a common frame of reference between marine and data science domains e.g by harmonizing terminologies, assumptions, and data formats in a way that facilitates seamless bidirectional data exchange between data science models and spatio-ecological workows. Future outlook suggests that marine sciences will continue to experience a surge in marine big datasets, along with a corresponding increased demand for marine data analytical expertise. This is evidenced by the ever growing global consensus on the central role of the oceans in addressing modern societal challenges such as climate change and green energy transition (Bigg et al., 2003). Recommendations from forward-looking policy documents such as the European Marine Board’s Future Science Briefs (“Future Science Briefs | European Marine Board,” n.d.) strongly advocate for increased funding for collaborative, transdisciplinary, and cross cutting partnerships to support marine science research activities e.g through programs such as Horizon Europe (Tenhunen-Lunkka and Honkanen, 2024). In addition, the possible adoption of the International Seabed Authority’s Mining Code to regulate deep-sea mining activities (Singh, 2021) will undoubtedly drive up demand for marine data science workows to support e.g evidence-based monitoring of regulatory compliance, as well as objective environmental impact assessments by stakeholders (Peukert et al., 2018). In general, such policy initiatives will incentivize future investments in integrated marine big data handling and analytical workows centered around overarching themes such as: (a) long-term monitoring of marine ecosystem functions, services and biodiversity; (b) responsible natural resource exploitation e.g polymetallic Mn-nodules to balance ecological, economic and social interests; (c) maximizing the value of federated data portals that adhere to Findable, Accessible, Interoperable and Reusable (FAIR) principles (Schoening et al., 2018) 65
Bibliography Albano, P.G., Azzarone, M., Amati, B., Bogi, C., Sabelli, B., Rilov, G., 2020. Low diversity or poorly explored? Mesophotic molluscs highlight undersampling in the Eastern Mediterranean. Biodivers. Conserv. 29, 4059–4072. https://doi.org/10.1007/s10531-020-02063-w Bakker, J.D., 2024. Types of Ordination Methods. Bengio, Y., Courville, A., Vincent, P., 2014. Representation Learning: A Review and New Perspectives. Bigg, G.R., Jickells, T.D., Liss, P.S., Osborn, T.J., 2003. The role of the oceans in climate. Int. J. Climatol. 23, 1127–1159. https://doi.org/10.1002/joc.926 Bräger, S., Romero Rodriguez, G.Q., Mulsow, S., 2020. The current status of environmental requirements for deep seabed mining issued by the International Seabed Authority. Mar. Policy, Environmental governance of deep seabed mining - scientic insights and food for thought 114, 103258. https://doi.org/10.1016/j.marpol.2018.09.003 Brown, A., Thatje, S., 2014. Explaining bathymetric diversity patterns in marine benthic invertebrates and demersal shes: physiological contributions to adaptation of life at depth. Biol. Rev. 89, 406–426. https://doi.org/10.1111/brv.12061 Burkart, N., Huber, M.F., 2021. A Survey on the Explainability of Supervised Machine Learning. J. Artif. Intell. Res. 70, 245–317. https://doi.org/10.1613/jair.1.12228 Canonico, G., Buttigieg, P.L., Montes, E., Muller-Karger, F.E., Stepien, C., Wright, D., Benson, A., Helmuth, B., Costello, M., Sousa-Pinto, I., Saeedi, H., Newton, J., Appeltans, W., Bednaršek, N., Bodrossy, L., Best, B.D., Brandt, A., Goodwin, K.D., Iken, K., Marques, A.C., Miloslavich, P., Ostrowski, M., Turner, W., Achterberg, E.P., Barry, T., Defeo, O., Bigatti, G., Henry, L.-A., Ramiro-Sánchez, B., Durán, P., Morato, T., Roberts, J.M., García-Alegre, A., Cuadrado, M.S., Murton, B., 2019. Global Observational Needs and Resources for Marine Biodiversity. Front. Mar. Sci. 6. https://doi.org/10.3389/fmars.2019.00367 Community ecology in the age of multivariate multiscale spatial analysis - Dray - 2012 - Ecological Monographs - Wiley Online Library [WWW Document], n.d. URL https://esajournals.onlinelibrary.wiley.com/doi/full/10.1890/11-1183.1?casa_token=R uXWXB_IzncAAAAA%3A8Auq90GXr5EnKgiilG-T_8uRaCeaX20TRgxKrdJzy5eSBB1_slhzCzwoMLbdKyC784XOzOBPUUxdFUL (accessed 9.21.24). Cuddington, K., Fortin, M.-J., Gerber, L.R., Hastings, A., Liebhold, A., O’Connor, M., Ray, C., 2013. Process-based models are required to manage ecological systems in a changing world. Ecosphere 4, art20. https://doi.org/10.1890/ES12-00178.1 66
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., Fei-Fei, L., 2009. ImageNet: A large-scale hierarchical image database, in: 2009 IEEE Conference on Computer Vision and Pattern Recognition. Presented at the 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 248–255. https://doi.org/10.1109/CVPR.2009.5206848 Du, J., 2018. Understanding of Object Detection Based on CNN Family and YOLO. J. Phys. Conf. Ser. 1004, 012029. https://doi.org/10.1088/1742-6596/1004/1/012029 Durden, J.M., Schoening, T., Althaus, F., Friedman, A., Garcia, R., Glover, A.G., Greinert, J., Stout, N.J., Jones, D.O.B., Jordt, A., Kaeli, J.W., Köser, K., Kuhnz, L.A., Lindsay, D., Morris, K.J., Nattkemper, T.W., Osterlo, J., Ruhl, H.A., Singh, H., Bett, M.T.& B.J., 2016. Perspectives in Visual Imaging for Marine Biology and Ecology: From Acquisition to Understanding, in: Oceanography and Marine Biology. CRC Press. European Commission. Joint Research Centre, International Council for the Exploration of the Sea (ICES), 2010. Marine Strategy Framework Directive : task group 2 report : non-indigenous species, April 2010. Publications Oce, LU. Evans, D., 2006. The Habitats of the European Union Habitats Directive. Biol. Environ. Proc. R. Ir. Acad. 106B, 167–173. Evans, J.L., Peckett, F., Howell, K.L., 2015. Combined application of biophysical habitat mapping and systematic conservation planning to assess eciency and representativeness of the existing High Seas MPA network in the Northeast Atlantic. ICES J. Mar. Sci. 72, 1483–1497. https://doi.org/10.1093/icesjms/fsv012 Felden, J., Möller, L., Schindler, U., Huber, R., Schumacher, S., Koppe, R., Diepenbroek, M., Glöckner, F.O., 2023. PANGAEA - Data Publisher for Earth & Environmental Science. Sci. Data 10, 347. https://doi.org/10.1038/s41597-023-02269-x Fetting, C., n.d. THE EUROPEAN GREEN DEAL. Flannery, E., Przeslawski, R., 2015. Comparison of sampling methods to assess benthic marine biodiversity. Are spatial and ecological relationships consistent among sampling gear? (Report). Geoscience Australia. https://doi.org/10.11636/Record.2015.007 Freestone, A.L., Osman, R.W., Ruiz, G.M., Torchin, M.E., 2011. Stronger predation in the tropics shapes species richness patterns in marine communities. Ecology 92, 983–993. https://doi.org/10.1890/09-2379.1 Frutos, I., Kaiser, S., Pułaski, Ł., Studzian, M., Błażewicz, M., 2022. Challenges and Advances in the Taxonomy of Deep-Sea Peracarida: From Traditional to Modern Methods. Front. Mar. Sci. 9. https://doi.org/10.3389/fmars.2022.799191 Future Science Briefs | European Marine Board [WWW Document], n.d. URL https://www.marineboard.eu/publication/future-science-brief (accessed 9.28.24). Galparsoro, I., Borja, A., Uyarra, M.C., 2014. Mapping ecosystem services provided by benthic 67
habitats in the European North Atlantic Ocean. Front. Mar. Sci. 1. https://doi.org/10.3389/fmars.2014.00023 Gamfeldt, L., Lefcheck, J.S., Byrnes, J.E.K., Cardinale, B.J., Duy, J.E., Grin, J.N., 2015. Marine biodiversity and ecosystem functioning: what’s known and what’s next? Oikos 124, 252–265. https://doi.org/10.1111/oik.01549 Glasby, G.P., 2002. Deep Seabed Mining: Past Failures and Future Prospects. Mar. Georesources Geotechnol. 20, 161–176. https://doi.org/10.1080/03608860290051859 Guidi, L., Guerra, A.F., Canchaya, C., Curry, E., Foglini, F., Irisson, J.O., Malde, K., Marshall, C.T., Obst, M., Ribeiro, R.P., Tjiputra, J., Heymans, S.J., Alexander, B., Piniella, Á.M., Kellett, P., Coopman, J., 2020. Big Data in Marine Science. European Marine Board. https://doi.org/10.5281/zenodo.3755793 Hachaj, T., Mazurek, P., 2020. Comparative Analysis of Supervised and Unsupervised Approaches Applied to Large-Scale “In The Wild” Face Verication. Symmetry 12, 1832. https://doi.org/10.3390/sym12111832 Haedrich, R.L., Rowe, G.T., Polloni, P.T., 1980. The megabenthic fauna in the deep sea south of New England, USA. Mar. Biol. 57, 165–179. https://doi.org/10.1007/BF00390735 Harris, P.T., Baker, E.K., 2012. 1 - Why Map Benthic Habitats?, in: Harris, P.T., Baker, E.K. (Eds.), Seaoor Geomorphology as Benthic Habitat. Elsevier, London, pp. 3–22. https://doi.org/10.1016/B978-0-12-385140-6.00001-3 Hegrenæs, Ø., Gade, K., Hagen, O.K., Hagen, P.E., 2009. Underwater transponder positioning and navigation of autonomous underwater vehicles, in: OCEANS 2009. Presented at the OCEANS 2009, pp. 1–7. https://doi.org/10.23919/OCEANS.2009.5422358 Huvenne, V.A.I., Robert, K., Marsh, L., Lo Iacono, C., Le Bas, T., Wynn, R.B., 2018. ROVs and AUVs, in: Micallef, A., Krastel, S., Savini, A. (Eds.), Submarine Geomorphology. Springer International Publishing, Cham, pp. 93–108. https://doi.org/10.1007/978-3-319-57852-1_7 Jain, S.M., 2022. Hugging Face, in: Jain, S.M. (Ed.), Introduction to Transformers for NLP: With the Hugging Face Library and Models to Solve Problems. Apress, Berkeley, CA, pp. 51–67. https://doi.org/10.1007/978-1-4842-8844-3_4 Lacharité, M., Metaxas, A., 2018. Environmental drivers of epibenthic megafauna on a deep temperate continental shelf: A multiscale approach. Prog. Oceanogr. 162, 171–186. https://doi.org/10.1016/j.pocean.2018.03.002 Langenkämper, D., van Kevelaer, R., Purser, A., Nattkemper, T.W., 2020. Gear-Induced Concept Drift in Marine Images and Its Eect on Deep Learning Classication. Front. Mar. Sci. 7. Langenkämper, D., Zurowietz, M., Schoening, T., Nattkemper, T.W., 2017. BIIGLE 2.0 - 68
Browsing and Annotating Large Marine Image Collections. Front. Mar. Sci. 4, 83. https://doi.org/10.3389/fmars.2017.00083 Legendre, P., Fortin, M.J., 1989. Spatial pattern and ecological analysis. Vegetatio 80, 107–138. https://doi.org/10.1007/BF00048036 Levin, L.A., Bett, B.J., Gates, A.R., Heimbach, P., Howe, B.M., Janssen, F., McCurdy, A., Ruhl, H.A., Snelgrove, P., Stocks, K.I., Bailey, D., Baumann-Pickering, S., Beaverson, C., Beneld, M.C., Booth, D.J., Carreiro-Silva, M., Colaço, A., Eblé, M.C., Fowler, A.M., Gjerde, K.M., Jones, D.O.B., Katsumata, K., Kelley, D., Le Bris, N., Leonardi, A.P., Lejzerowicz, F., Macreadie, P.I., McLean, D., Meitz, F., Morato, T., Netburn, A., Pawlowski, J., Smith, C.R., Sun, S., Uchida, H., Vardaro, M.F., Venkatesan, R., Weller, R.A., 2019. Global Observing Needs in the Deep Ocean. Front. Mar. Sci. 6. https://doi.org/10.3389/fmars.2019.00241 Li, J., Tran, M., Siwabessy, J., 2016. Selecting Optimal Random Forest Predictive Models: A Case Study on Predicting the Spatial Distribution of Seabed Hardness. PLOS ONE 11, e0149089. https://doi.org/10.1371/journal.pone.0149089 Liu, F.T., Ting, K.M., Zhou, Z.-H., 2008. Isolation Forest, in: 2008 Eighth IEEE International Conference on Data Mining. Presented at the 2008 Eighth IEEE International Conference on Data Mining (ICDM), IEEE, Pisa, Italy, pp. 413–422. https://doi.org/10.1109/ICDM.2008.17 Lodge, M.W., 2011. International Seabed Authority. Int. J. Mar. Coast. Law 26, 463. Lotze, H.K., 2021. Marine biodiversity conservation. Curr. Biol. 31, R1190–R1195. https://doi.org/10.1016/j.cub.2021.06.084 Lu, S., Ding, Y., Liu, M., Yin, Z., Yin, L., Zheng, W., 2023. Multiscale Feature Extraction and Fusion of Image and Text in VQA. Int. J. Comput. Intell. Syst. 16, 54. https://doi.org/10.1007/s44196-023-00233-6 Luo, X., Qin, X., Wu, Z., Yang, F., Wang, M., Shang, J., 2019. Sediment Classication of Small-Size Seabed Acoustic Images Using Convolutional Neural Networks. IEEE Access 7, 98331–98339. https://doi.org/10.1109/ACCESS.2019.2927366 Martín Míguez, B., Novellino, A., Vinci, M., Claus, S., Calewaert, J.-B., Vallius, H., Schmitt, T., Pititto, A., Giorgetti, A., Askew, N., Iona, S., Schaap, D., Pinardi, N., Harpham, Q., Kater, B.J., Populus, J., She, J., Palazov, A.V., McMeel, O., Oset, P., Lear, D., Manzella, G.M.R., Gorringe, P., Simoncelli, S., Larkin, K., Holdsworth, N., Arvanitidis, C.D., Molina Jack, M.E., Chaves Montero, M. del M., Herman, P.M.J., Hernandez, F., 2019. The European Marine Observation and Data Network (EMODnet): Visions and Roles of the Gateway to Marine Data in Europe. Front. Mar. Sci. 6. https://doi.org/10.3389/fmars.2019.00313 Maslej, N., Fattorini, L., Perrault, R., Parli, V., Reuel, A., Brynjolfsson, E., Etchemendy, J., Ligett, K., Lyons, T., Manyika, J., Niebles, J.C., Shoham, Y., Wald, R., Clark, J., 2024. 69
Articial Intelligence Index Report 2024. https://doi.org/10.48550/arXiv.2405.19522 Mbani, B., Buck, V., Greinert, J., 2023a. An automated image-based workow for detecting megabenthic fauna in optical images with examples from the Clarion–Clipperton Zone. Sci. Rep. 13, 8350. https://doi.org/10.1038/s41598-023-35518-5 Mbani, B., Schoening, T., Gazis, I.-Z., Koch, R., Greinert, J., 2022a. Implementation of an automated workow for image-based seaoor classication with examples from manganese-nodule covered seabed areas in the Central Pacic Ocean. Sci. Rep. 12, 15338. https://doi.org/10.1038/s41598-022-19070-2 Mbani, B., Schoening, T., Gazis, I.-Z., Koch, R., Greinert, J., 2022b. Implementation of an automated workow for image-based seaoor classication with examples from manganese-nodule covered seabed areas in the Central Pacic Ocean. Sci. Rep. 12, 15338. https://doi.org/10.1038/s41598-022-19070-2 Mbani, B., Schoening, T., Greinert, J., 2023b. Automated and Integrated Seaoor Classication Workow (AI-SCW) [WWW Document]. https://doi.org/10.3289/SW_2_2023 Mckinney, W., 2011. pandas: a Foundational Python Library for Data Analysis and Statistics. Python High Perform. Sci. Comput. Mika, S., Scholkopf, B., Smola, A., n.d. Kernel PCA and De-Noising in Feature Spaces 7. Montes, E., Lefcheck, J.S., Guerra-Castro, E., Klein, E., Kavanaugh, M.T., de Azevedo Mazzuco, A.C., Bigatti, G., Cordeiro, C.A.M.M., Simoes, N., Macaya, E.C., Moity, N., Londoño-Cruz, E., Helmuth, B., Choi, F., Soto, E.H., Miloslavich, P., Muller-Karger, F.E., 2021. Optimizing Large-Scale Biodiversity Sampling Eort: Toward an Unbalanced Survey Design. Oceanography 34, 80–91. Nimier-David, M., Vicini, D., Zeltner, T., Jakob, W., 2019. Mitsuba 2: a retargetable forward and inverse renderer. ACM Trans Graph 38, 203:1-203:17. https://doi.org/10.1145/3355089.3356498 Oliver, T.H., Heard, M.S., Isaac, N.J.B., Roy, D.B., Procter, D., Eigenbrod, F., Freckleton, R., Hector, A., Orme, C.D.L., Petchey, O.L., Proença, V., Raaelli, D., Suttle, K.B., Mace, G.M., Martín-López, B., Woodcock, B.A., Bullock, J.M., 2015. Biodiversity and Resilience of Ecosystem Functions. Trends Ecol. Evol. 30, 673–684. https://doi.org/10.1016/j.tree.2015.08.009 Pang, B., Nijkamp, E., Wu, Y.N., 2020. Deep Learning With TensorFlow: A Review. J. Educ. Behav. Stat. 45, 227–248. https://doi.org/10.3102/1076998619872761 Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., Chintala, S., 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library, in: Advances in Neural Information Processing Systems. Curran Associates, Inc. 70
Payne, R.J., Egan, J., 2019. Using palaeoecological techniques to understand the impacts of past volcanic eruptions. Quat. Int., Distal Eects of Volcanic Eruptions on Pre-Industrial Societies 499, 278–289. https://doi.org/10.1016/j.quaint.2017.12.019 Pedersen, M., Haurum, J.B., Gade, R., Moeslund, T.B., 2019. Detection of Marine Animals in a New Underwater Dataset with Varying Visibility 9. Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., 2011. Scikit-learn: Machine Learning in Python. Mach. Learn. PYTHON 6. Peukert, A., Schoening, T., Alevizos, E., Köser, K., Kwasnitschka, T., Greinert, J., 2018. Understanding Mn-nodule distribution and evaluation of related deep-sea mining impacts using AUV-based hydroacoustic and optical data. Biogeosciences 15, 2525–2549. https://doi.org/10.5194/bg-15-2525-2018 Purser, A., Marcon, Y., Dreutter, S., Hoge, U., Sablotny, B., Hehemann, L., Lemburg, J., Dorschel, B., Biebow, H., Boetius, A., 2019. Ocean Floor Observation and Bathymetry System (OFOBS): A New Towed Camera/Sonar System for Deep-Sea Habitat Surveys. IEEE J. Ocean. Eng. 44, 87–99. https://doi.org/10.1109/JOE.2018.2794095 Ren, S., He, K., Girshick, R., Sun, J., 2015. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks, in: Advances in Neural Information Processing Systems. Curran Associates, Inc. Rigby, P., Pizarro, O., Williams, S.B., 2006. Towards Geo-Referenced AUV Navigation Through Fusion of USBL and DVL Measurements, in: OCEANS 2006. Presented at the OCEANS 2006, pp. 1–6. https://doi.org/10.1109/OCEANS.2006.306898 Saeedi, H., Warren, D., Brandt, A., 2022. The Environmental Drivers of Benthic Fauna Diversity and Community Composition. Front. Mar. Sci. 9. https://doi.org/10.3389/fmars.2022.804019 Sala, E., Knowlton, N., 2006. Global Marine Biodiversity Trends. Annu. Rev. Environ. Resour. 31, 93–122. https://doi.org/10.1146/annurev.energy.31.020105.100235 Schoening, T., Jones, D.O.B., Greinert, J., 2017. Compact-Morphology-based poly-metallic Nodule Delineation. Sci. Rep. 7, 13338. https://doi.org/10.1038/s41598-017-13335-x Schoening, T., Köser, K., Greinert, J., 2018. An acquisition, curation and management workow for sustainable, terabyte-scale marine image analysis. Sci. Data 5, 180181. https://doi.org/10.1038/sdata.2018.181 Shin, H.-C., Roth, H.R., Gao, M., Lu, L., Xu, Z., Nogues, I., Yao, J., Mollura, D., Summers, R.M., 2016. Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer Learning. IEEE Trans. Med. Imaging 35, 1285–1298. https://doi.org/10.1109/TMI.2016.2528162 71
Singh, P.A., 2021. The two-year deadline to complete the International Seabed Authority’s Mining Code: Key outstanding matters that still need to be resolved. Mar. Policy 134, 104804. https://doi.org/10.1016/j.marpol.2021.104804 Stratmann, T., Soetaert, K., Wei, C.-L., Lin, Y.-S., van Oevelen, D., 2019. The SCOC database, a large, open, and global database with sediment community oxygen consumption rates. Sci. Data 6, 242. https://doi.org/10.1038/s41597-019-0259-3 Szegedy, C., Vanhoucke, V., Ioe, S., Shlens, J., Wojna, Z., 2015. Rethinking the Inception Architecture for Computer Vision. ArXiv151200567 Cs. Tenhunen-Lunkka, A., Honkanen, R., 2024. Project coordination success factors in European Union-funded research, development and innovation projects under the Horizon 2020 and Horizon Europe programmes. J. Innov. Entrep. 13, 7. https://doi.org/10.1186/s13731-024-00363-x Truby, J., Brown, R.D., Ibrahim, I.A., Parellada, O.C., 2022. A Sandbox Approach to Regulating High-Risk Articial Intelligence Applications. Eur. J. Risk Regul. 13, 270–294. https://doi.org/10.1017/err.2021.52 van de Schoot, R., Depaoli, S., King, R., Kramer, B., Märtens, K., Tadesse, M.G., Vannucci, M., Gelman, A., Veen, D., Willemsen, J., Yau, C., 2021. Bayesian statistics and modelling. Nat. Rev. Methods Primer 1, 1–26. https://doi.org/10.1038/s43586-020-00001-2 Vaudo, J.J., Heithaus, M.R., 2013. Microhabitat Selection by Marine Mesoconsumers in a Thermally Heterogeneous Habitat: Behavioral Thermoregulation or Avoiding Predation Risk? PLOS ONE 8, e61907. https://doi.org/10.1371/journal.pone.0061907 Veale, M., Borgesius, F.Z., 2021. Demystifying the Draft EU Articial Intelligence Act — Analysing the good, the bad, and the unclear elements of the proposed approach. Comput. Law Rev. Int. 22, 97–112. https://doi.org/10.9785/cri-2021-220402 Volkmann, S.E., Lehnen, F., 2018. Production key gures for planning the mining of manganese nodules. Mar. Georesources Geotechnol. 36, 360–375. https://doi.org/10.1080/1064119X.2017.1319448 Wang, C., Mei, D., Wang, Y., Yu, X., Sun, W., Wang, D., Chen, J., 2022. Task allocation for Multi-AUV system: A review. Ocean Eng. 266, 112911. https://doi.org/10.1016/j.oceaneng.2022.112911 Warner, G., 1982. Food and Feeding Mechanisms: Ophiuroidea, in: Echinoderm Nutrition. CRC Press. Webb, T.J., Berghe, E.V., O’Dor, R., 2010. Biodiversity’s Big Wet Secret: The Global Distribution of Marine Biological Records Reveals Chronic Under-Exploration of the Deep Pelagic Ocean. PLOS ONE 5, e10223. https://doi.org/10.1371/journal.pone.0010223 72
5 Vol.:(0123456789) Scientif㘶c Reports | (2022) 12:15338 | https://doi.org/10.1038/s41598-022-19070-2 www.nature.com/scientificreports/ shows the results of the same images after color normalization by histogram matching the reference ECDF; we point out that these color normalized images have already been transformed to a consistent spatial resolution (in pixel/cm), and then center cropped to represent a standard spatial footprint of 1.6 square meters on the seabed (see details in “Image enhancements” section). From this figure, the resulting ECDF of the color normalized images are identical to the reference distribution. Qualitatively, this transformation results in normalized photos which have a good overall scene brightness. Seaf㘶oor classif㘶cation assessment. This section presents results of the quantitative assessment aimed at evaluating the performance for both the supervised and unsupervised classifiers on the test set images. Also provided are the results of the PCA projection of all classified images onto a two-dimensional feature space, which can be used to visualize how semantically similar images group together in the feature space. For example, whereas images of large sized Mn-nodules group together on the upper region of this feature space, images of densely distributed Mn-nodules group together on the left. Performance evaluation of the fine-tuned Inception V3 classifier. The confusion matrix used to evaluate the performance of the supervised classifier is shown in Fig.5A. The off-diagonal elements of the matrix tabulate the number of misclassifications made by the supervised classifier when it made predictions on the test images. These test images comprised 12.2% of the labeled images (612 in total), since the other 7% was used as validation set during hyper-parameters tuning. The matrix shows that some of the test images labeled as Seafloor A were wrongly predicted to belong to Seafloor B. This could be because in some images of Seafloor A the sediment blanket cover over the Mn-nodules was not complete (not high enough), and some that were only partly Mn-nodules are still visible; this caused them to be misclassified as Seafloor B. The same reasoning also explains the confusion between Seafloor A and Seafloor D, depending on the general distribution of Mn-nodules (small patches as in Seafloor B or homogenous distribution of large Mn-nodules (Seafloor D). Figure3. Comparison between sets of (A) images corrected for illumination drop-off and, (B) adaptive histogram equalized images. Also shown as insets are the intensity histograms showing the distribution of the pixel values for each RGB channel.
6 Vol:.(1234567890) Scientif㘶c Reports | (2022) 12:15338 | https://doi.org/10.1038/s41598-022-19070-2 www.nature.com/scientificreports/ Figure5B shows the development of the learning curves used to track the progress of model training after every iteration step. The shape of the cross-entropy loss curve decreases with every iteration, and converges with a final loss value of 0.09. On the other hand, the shape of the accuracy curve rises steadily with each iteration until model convergence. F1 score was used as the accuracy metric to evaluate the performance of the fine-tuned Inception V3 classifier; it provides the harmonic mean of precision and recall, and ranges between a worst score of 0 to the best score value of 1. Using this approach, the F1 score of our fine-tuned classifier was found to be 0.93. Figure4. (A) Reference image used for color normalization and, (B) The Empirical Cumulative Distribution Function for the reference image. (C) Comparison between adaptive histogram equalized images and, (D) color normalized and rescaled center cropped images.
7 Vol.:(0123456789) Scientif㘶c Reports | (2022) 12:15338 | https://doi.org/10.1038/s41598-022-19070-2 www.nature.com/scientificreports/ Figure5C shows the distribution of prediction confidence scores for each of the seafloor classes. The prediction confidence was greater than 0.6 in more than 80% of the images, for all four classes. Figure5D shows three example images for each of the four seafloor classes as predicted by the fine-tuned Inception V3 classifier. Performance evaluation of sampling strategies used in fitting k-means classifiers. Figure6 shows the results of the comparison of the four training set sampling strategies. As mentioned, each strategy was evaluated based on the time it took from generating the training images to fitting a k-means classifier, and also on the quality of the resulting clusters. Random sampling was the quickest with a time-to-fit of 0.12s. Stratified cluster-based sampling resulted in the best quality of clusters, with an absolute silhouette score of 0.4. Spatially uniform sampling was the slowest with a time-to-fit of 7s, despite the quality of its clusters being equal to both random sampling and probabilistic weighted resampling. The optimal sampling strategy was chosen as stratified cluster-based sampling. This is because it achieved the highest silhouette score of 0.4, and was only 1.7s slower than the second fastest strategy of the probabilistic resampling. Treating the Inception V3 classification results as ground truths, the Fowlkes-Mallows Index (FMI) was used to quantify the success of our k-means classifier in successfully defining clusters that are similar to the ground truth set of classes. The FMI is the geometric mean of the precision and recall that makes no assumption about the cluster structure29,30. Using this approach, we obtained an FMI score of 0.5, which indicates good similarity in the classification accuracies. Figure7 shows the PCA projection of all the images color coded by the results of both unsupervised and supervised classification. The class boundaries of the unsupervised classifier are abrupt and well-defined. On the other hand, the supervised classifier results in fuzzy class boundaries, which is the case in reality since the transition between classes is subtle. Overall, the two classifiers generate results that agree with a Cohen’s kappa coefficient of 0.6. Figure8 shows the PCA projection of all the images without color coding, in which semantically similar images (e.g., those with large Mn-nodules) can be seen to group together in the feature space. Spatial distribution of seaf㘶oor classes. This section puts the image-based classification results into geographic perspective, by mapping out the distribution of the seafloor classes spatially within both the German and Belgian working areas. Figure5. (A) Confusion matrix for supervised classifier performance evaluation. (B) Loss and accuracy curves used to monitor progress of model training. (C) Distribution of prediction confidence scores for each seafloor class. (D) Examples of classifier predictions for each seafloor class.
8 Vol:.(1234567890) Scientif㘶c Reports | (2022) 12:15338 | https://doi.org/10.1038/s41598-022-19070-2 www.nature.com/scientificreports/ Figure6. Performance evaluation of unsupervised classifiers fit using training data generated using each sampling strategies. Figure7. PCA projection of all images onto feature space color coded by results of (A) unsupervised classification and, (B) supervised classification.
9 Vol.:(0123456789) Scientif㘶c Reports | (2022) 12:15338 | https://doi.org/10.1038/s41598-022-19070-2 www.nature.com/scientificreports/ The spatial distribution of the seafloor classes for both German and Belgian contract areas is shown in Fig.9, by color-coding the image locations along the deployment tracks based on the supervised Inception V3 classification results. The northern half of the Belgian area comprises dense nodule covered seafloor (Seafloor C), intermingled with patches of larger nodules of Seafloor D at local elevations. Seafloor D becomes the dominant class in the depressions in the northern part of the area, and along the two tracks in the southern part. Very few instances of Seafloor class A and B exist in the Belgian area. The German contract area typically comprises Seafloor B in the northern part, and a mix of Seafloor B and D in the southern part. The German area does not contain a significant amount of Seafloor C, highlighting that the Mn-nodules are generally smaller compared to the Belgian area. The occurrences of instances of Seafloor A are clearly correlated to the locations of the dredge experiment conducted in the German area. In Fig.10, preand post-dredge deployment tracks are shown, along with example images. After the dredge experiment, the proportion of Seafloor A increased by over 30%, caused by the sediment turnover/ploughing and subsequent settling of the suspended sediment onto Mn-nodules after a certain time (see also31). In order to allow easy visualization of the impact of the dredge experiment, only the post-dredge track that has a corresponding pre-dredge track is shown in Fig.10. The other two post-dredge tracks also fit with the occurrences from Seafloor A. Discussion For structuring the discussion for clarity, we briefly summarize what has been presented above, and make relevant connections to various aspects. In general, it can be said that recent developments in both underwater imaging technologies and hydro-acoustics have allowed marine researchers to characterize and ground-truth seabed substrate classes32–34. In this paper, we present an Automated and Integrated Seafloor Classification Workflow (AI-SCW), which includes a module that automatically detects laser points from images, and uses them for scale determination. This is then used to correct for illumination artefacts, and to color normalize images recorded at different times during different camera deployments. This generates a visually homogenous image data set that can be used as training datasets for the automated classification. To this end, the workflow also includes a module for semi-automatic labeling, which reduces the manual effort required to annotate images needed for training classification models. Both supervised Inception V3 and unsupervised k-means classifiers have been trained, and their classification results compared. The Inception V3 classifier is trained using semi-automatically generated labels, while the k-means classifier is trained using a subset of unlabeled images (those have been generated using a stratified cluster-based sampling strategy). When the results of both the Inception V3 and k-means classifiers were compared, they showed a good agreement, with a Cohen’s kappa coefficient of 0.6. Below, the various aspects of AI-SCW workflow are first discussed in detail. Finally, we discuss the spatial distribution of the derived seafloor classes in the context of terrain and backscatter properties of the seafloor. As part of the laser point detection workflow, the subtraction of the red channel from the linear combination of blue and green channels worked very well, producing an intermediate image with very high values around the laser points and very low values elsewhere; the intensity maxima of this intermediate image contained the three lasers. We observed that the linear combination had to be scaled by a coefficient to reduce the effect of the correlation among the RGB channels. After some iterations, we found that coefficient values between 0 and 1 produced good results (true positives), and in particular, a coefficient value of 0.2 produced the best results for this dataset. This value may differ when the workflow is applied in other datasets e.g., depending on the lighting condition, depth, type of device recording the images etc. We further observed that detections based on the contrast enhanced images resulted in many false laser spot detections. This is because the contrast enhancement transformation reduced the intensity of the red laser points, by adaptively equalizing the distribution of intensity values in local image tiles. The laser point detection accuracy increased significantly when the detection was done on the original photos. This is because the lasers are targeted towards the center of the field of view of the Figure8. PCA projection of all images onto feature space without color coding. Semantically similar images (e.g., those with large Mn-nodules) can be seen to group together.
10 Vol:.(1234567890) Scientif㘶c Reports | (2022) 12:15338 | https://doi.org/10.1038/s41598-022-19070-2 www.nature.com/scientificreports/ camera, a region that was already well illuminated by the artificial light source; which made the red laser points prominent and easy to detect. In addition, since the mask used to filter the intensity maxima coordinates was the same for all images, it allowed us to implement the laser point detection workflow concurrently across all CPU cores. This increased the processing speed by up to 3 times, compared to a sequential implementation. When compared to the DELPHI system implemented in a previous study5, our implementation relies on manually hand labeled information from just one single example image, compared to the DELPHI system that iteratively uses 70 manually annotated images. As our dataset specifically comprised three laser points projected from a camera altitude of between 1 to 4m, future research could improve this laser point detection workflow to accommodate other possible scenarios e.g., where there are fewer than three laser points visible in the image. Our semi-automated approach for generating the labeled training set to be used for fine tuning the supervised Inception V3 classifier was both straight forward and quick to execute. The semi-automation involved some manual annotation of example images by a human analyst, followed by an automated nearest neighbor sampling of additional training images in the feature space. This set-up was used to explicitly include domain knowledge during the labeled training data generation process. Incorporation of domain knowledge was important because nearest neighbor sampling only makes sense when semantically similar images are mapped close together in feature space. In our case, domain knowledge was incorporated by encoding texture and entropy into the feature Figure9. Maps of deployment tracks in both the German (BGR) and Belgian (GSR) contract areas. The deployment tracks are color coded by supervised classification results, and overlaid on our gridded multibeam bathymetric dataset. The boundaries of polymetallic nodule exploration areas were sourced from International Seabed Authority (https:// www. isa. org. jm), while the base map in the overview map was sourced from GEBCO (https:// www. gebco. net/).
11 Vol.:(0123456789) Scientif㘶c Reports | (2022) 12:15338 | https://doi.org/10.1038/s41598-022-19070-2 www.nature.com/scientificreports/ vectors that represent the images in feature space, because the seabed in our working area was mostly characterized by Mn-nodules that appear as irregular blobs of black pixels in the image. The use of texture features to encode seabed characteristics has been reported in previous studies, which found that they adequately reveal subtle variation in seabed characteristics19–21. In addition, domain knowledge was incorporated by interactively choosing example images that are uniformly distributed in the two-dimensional PCA-projected feature space; this reduced the class imbalance in the sample training set. We emphasize that the above-described feature space was only used for generating labeled training data, after which deep learning was employed for supervised classification. Focusing on supervised classification, we observed that the Inception V3 model, which was finetuned using the semi-automatically generated labeled dataset achieved a good classification performance. This can be attributed to the ability of deep learning architectures to extract useful abstract features during training, as pointed out e.g. by35. Achieving such a good performance using a handful of human-labeled examples partly addresses the bottleneck of manual annotation of large data sets. Therefore, our approach contributes towards making the adoption of state-of-the art computer vision models into image based marine studies much easier. When exploring the results of the unsupervised classification in the PCA-projected feature space, we observed that the inter-cluster separation distance was small. This observation is typical for images of the deep-sea, where the variation in seabed characteristics is subtle over kilometer scale, as was also pointed out in previous studies such as36. Our comparison of the four strategies for selecting the appropriate training data for fitting an unsupervised classifier showed that with such images, random sampling may be the ideal approach for generating this training data set, if quick computation is most important. Despite the speed, however, random sampling may cause imbalance in the training data, since samples will be disproportionately drawn from regions of the feature space that are over-represented. This is likely to occur e.g., when images are collected in the largely homogenous abyssal Mn-nodule plains with only infrequent regions of no (sediment covered) nodules or other rare seabed characteristics. In a previous study, Prati etal.37 also corroborate this relationship between random sampling and class imbalance. Furthermore, training a classifier using imbalanced data may result in poor generalization performance, as was also pointed out by38. When the metric of interest is not speed but quality of the resulting clusters, our experiment showed that stratified cluster-based sampling is a better approach. This becomes evident Figure10. Deployment tracks in the German contract area (BGR), which show the seafloor classification results before and after the dredge experiments. The dredge experiment increased the proportion of sedimentcovered seabed (Seafloor A) by 30%. For context, the ‘before’ track is shown in gray overlaid as background in the bottom panel. For clarity, only one observation track is shown from after the dredge experiment.
12 Vol:.(1234567890) Scientif㘶c Reports | (2022) 12:15338 | https://doi.org/10.1038/s41598-022-19070-2 www.nature.com/scientificreports/ as it resulted in the highest absolute silhouette score, while being only 2s slower than random sampling. The high silhouette score implies that of all the other methods we compared, the stratified cluster-based sampling approach resulted in clusters that are clearly distinguished from each other. This is because the approach incorporates an over-clustering step that first partitions the entire entropy-defined feature space, which allows samples to be drawn uniformly across all regions of the feature space, naturally reducing class imbalance. Data from a more heterogenous seafloor like in shallow water environments may not exhibit a similar feature space distribution, and therefore a different sampling strategy may be required that needs to be tested in future studies. Visualizing the unsupervised seafloor classification results of the k-means clustering in feature space showed abrupt inter-class boundaries. This is an artifact of the k-means objective function that results in convex clusters with well-defined boundaries, as was also pointed out in a previous study by39. In the literature, there have been attempts to explicitly introduce fuzzy class boundaries in seafloor classification e.g.40. Others such as41 attempted to customize k-means by modifying its objective function using least-squares criterion to implicitly introduce fuzziness. However, the authors clearly point out that these fuzzy clustering approaches are very sensitive to cluster size imbalance, and can result in clusters whose semantic meaning is hard to interpret. In our supervised case of the Inception V3, the boundaries were logically fuzzy in the PCA-projected space because the convolutional neural network outputs a set of probabilities. This allows the model to make soft classification by assigning each image a likelihood of belonging to each of the classes using a confidence score. Despite these general differences, our quantitative assessment of both the supervised Inception V3 and unsupervised k-means classification results indicate a good agreement, with a Cohen’s kappa coefficient of 0.6. McHugh42 also quantified the extent to which two different classifiers assign the same score to a variable using the Cohen’s kappa statistic. From these assessments, it can be concluded that a fast, preliminary seafloor classification while at sea may be performed with unsupervised methods. The main advantage of using k-means is the low computing power needed, which nevertheless produces reasonable classification results. Supervised methods which require GPU computing resources may be better executed for seafloor classifications as part of a detailed habitat characterization in a second step. They are better suited to pick out subtle transitions between classes, which is usually the case in reality. Regarding the geospatial distribution of seafloor classes and their correlation with bathymetry, the visualization of our seafloor classes in map view makes it obvious that their spatial distribution is not random/ homogenous but clustered. In the Belgian area, the dense Mn-nodules (Seafloor C) and large sized Mn-nodules (Seafloor D) occupy 94% of the OFOS-inspected area at about 4500m water depth. The German contract area pre-dominantly contains sparsely distributed Mn-nodules (Seafloor B) in the North, while the south comprises a mixture of both Seafloor B and Seafloor D. Overall, 96% of sparsely distributed Mn-nodules make up the seafloor in the German area at water depths of approximately 4100m. The seafloor class D with big Mn-nodules lies in deeper regions of the seafloor with rugged terrain, whereas sparsely distributed nodules occupy shallower regions of the seafloor with flat terrain (please see further details in the supplementary information). This is consistent with the findings of prior studies, which showed that a higher number of Mn-nodules occur in areas with a rugged seafloor than flat plains e.g.,41,43,44. At the site of the dredge experiment, stretches of sediment covered, or dredged/ploughed seabed (Seafloor A) appears after the experiment; the proportion of sediment covered seabed increased by 30% (see Fig.10). Comparing our approaches to related works, our color normalization approach based on automatically detected laser points introduced a simple yet novel workflow, which significantly reduced the underwater illumination artifacts on our images. Similar to previous works e.g. by24, our approach rescaled and normalized the color of each image depending on its altitude above the seafloor: images recorded farther from the seafloor were transformed more than those close to the seafloor. In contrast, however, our approach was different since we did not have access to the altitude of each image, and therefore we implemented a novel approach that infers them from automatically detected laser points. Furthermore, our approach uses histogram matching to correct for uneven scene brightness among the photos, which is simple, parallelizable, and does not even require knowledge of parameters required for reconstruction of the path of light rays through the water column e.g. as demanded by physics-based approaches24,45,46, or photogrammetric structure from-motion47 and simultaneous localization and mapping (SLAM)48,49. Moreover, performing color normalization on the raw images improved the accuracy of seafloor classification; this is similar to observations by previous works such as by24,45. Particularly in our case, the raw images acquired at varying altitude were not directly comparable; these images represent regions ofthe seafloor with varying spatial foot print and scene brightness. With respect to generating annotations for training classifier, our semi-automated labeling approach greatly reduces human effort similar to previous works by11,50,51. However, our approach is novel because it allows for the incorporation of domain knowledge through feature space engineering, rather than only relying on similarity in geographic space and spatial auto correlation e.g. as proposed by52. Regarding optical image-based seafloor classification, our results revealed seabed substrate classes that had semantic meaning, similar to previous works by20,21,24,25,53,54. However, our workflow is novel since we implement both supervised (Inception V3) and unsupervised (k-means) classifiers, and we further demonstrate that both of them show a good overall agreement in seafloor classification accuracy. Thus, our study provides an unsupervised seafloor classification workflow that can be used at sea where computers are not very powerful, as well a supervised workflow that is suitable for office settings where there is access to computers with increased memory and GPU hardware. Conclusion and recommendations for future research This study contributes to the current understanding of image-based deep seabed classification, and its novelty can be summarized as follows: First, we implement a new approach that automatically detects laser points from a sequence of underwater images. This is useful for calculating scale, which allows marine researchers who make measurements on the images to convert from pixel units to real world metric units e.g., meters. The scale also
13 Vol.:(0123456789) Scientif㘶c Reports | (2022) 12:15338 | https://doi.org/10.1038/s41598-022-19070-2 www.nature.com/scientificreports/ allows the scientists to infer the height above the seafloor from which the image was acquired, which is useful for determining the spatial footprint of the image; Second, we implement a semi-automated approach that drastically reduces the required human effort during image labeling. This is useful for marine scientists who routinely work with large volumes of underwater images, since it reduces the burden and fatigue of generating manual annotations, while still allowing them to train high-accuracy classification models (e.g., convolutional neural networks, random forests etc.) that automatically analyze the images and subsequently classify the seafloor into habitat types; Third, we propose stratified cluster-based sampling as good strategy for generating a subset of images to be used for training an unsupervised classifier. When faced with huge volumes of images following an expedition, this proposed sampling strategy is useful to help marine scientists when deciding how to generate a training set that fits in standard computer memory, while simultaneously being class balanced and not requiring specialized hardware such as GPU. Although we propose stratified cluster-based sampling as the optimal unsupervised classifier, this conclusion is specific to underwater photos recorded in the deep-sea abyssal plains, for which the seafloor is mostly homogenous. This may not hold in shallower depths with heterogenous seafloor characteristics, and future research could address this by extending our AI-SCW feature extraction module to investigate e.g., if color differences could be useful in classifying seafloor covered by other substrate classes such as seagrass, rocks and sand. In addition, future research could explore the possibility of embedding automated image-based workflows such as AI-SCW onto the data acquisition software of ROV/UAV. This would enable real time on-the-fly image analysis, which provides significant savings in time and post processing effort for applications such as seabed classification, megabenthic fauna detection as well as the quantification of crusts and nodules; these are ongoing developments at GEOMAR and other centers. As a concluding remark, this study comprehensively presents a set of methods that together form an automated workflow for image-based seafloor classification and mapping. These include: automated laser point detection; semi-automated image annotation; as well as supervised and unsupervised seafloor classification. The applicability of these methods is demonstrated using a case study involving seafloor images from Mnnodule covered deep seabed areas of the Clarion Clipperton Zone in the Central Pacific Ocean. In so doing, we clearly demonstrate potential ways of incorporating recent advances in machine learning and computer vision into marine research, especially for purposes of generating actionable insights from the huge volumes of optical underwater imagery, which are nowadays recorded during scientific expeditions by imaging systems such as AUV, ROV and OFOS. We believe that these insights significantly contribute towards the broader aim of understanding our marine ecosystems, which in turn enables appropriate measures to be established for their management and sustainable use. Materials and methods Working area. As a case study, AI-SCW was applied to an underwater image dataset recorded during an expedition to the German and Belgian contract areas for manganese nodule (Mn-nodule) mining, at the Clarion Clipperton Zone (CCZ) in the central Pacific Ocean. The expedition was part of the second phase of the JPIOceans project MiningImpact, and was executed on board the German research vessel SONNE during cruise SO268. The project aimed at assessing how potential mining of polymetallic nodules on the seafloor would impact the deep-sea environment. A total of 12 video investigations were performed in two different contract areas (Table1). Within the German contract area, a small-scale sediment plume experiment was conducted, using a chain dredge to observe the re-depositioning of plume sediments before and after the disturbance. Three camera deployments were conducted to photograph the seafloor after the dredge experiment, while one deployment was conducted before the experiment for comparison55. Setup of the image acquisition system. The underwater images were acquired using an Ocean Floor Observation System (OFOS), which was towed at a speed of 0.5 knots at 1 to 4m above the seafloor. It comprised a steel frame equipped with both still and video cameras. The still images were recorded using a Canon EOS 5D Mark IV camera with a 24mm lens, whereas video was recorded using the HD-SDI camera with 64° × 40° view. These two cameras were spaced 50cm apart and directed vertically towards the seafloor alongside two strobe lights (Sea&Sea YS-250), four LED lights (SeaLite Sphere), three lasers spaced 40cm apart, one altimeter and one USBL system for tracking the position of the OFOS. Whereas the video camera recorded continuously, the still camera took an image once every 10s. Both cameras had a dome port that did not alter the field of view as long as the lens was centered properly within the dome. Camera calibration was done by photographing a camera calibration target on deck before the actual deployment. Further details regarding the image acquisition set up can be found on page 65 of the SO268 cruise report55. Image dataset. The image dataset comprised 40,678 underwater still images, which were recorded during the 12 deployments of the towed Ocean Floor Observation System (OFOS). The respective data are published on PANGAEA56, and can be accessed online (https:// doi. panga ea. de/ 10. 1594/ PANGA EA. 935856). In addition to the PANGAEA dataset, the images can also be accessed upon request as services through the BIIGLE portal (https:// annot ate. geomar. de/ proje cts/ 44). BIIGLE is an online image annotation platform, specifically developed to facilitate the annotation of benthic fauna from underwater images57. The OFOS deployments were conducted in an average water depth of 4,280m, covering a track length of 92.5km in total. After acquisition, the images were georeferenced by matching each image’s acquisition time in UTC to the USBL navigation data. Three laser pointers in a triangular configuration are used for scaling. Light is provided through several lights focusing on the central area below the OFOS frame illuminating the field of view of the vertically downward looking camera.
14 Vol:.(1234567890) Scientif㘶c Reports | (2022) 12:15338 | https://doi.org/10.1038/s41598-022-19070-2 www.nature.com/scientificreports/ Detailed information about these images and their acquisition can be found in the SO268 cruise report55. Table1 provides an overview of the OFOS deployments during the expedition. Software and APIs. AI-SCW has been implemented using the Python programming language. Some of the major libraries used include scikit-image, scikit-learn, TensorFlow and pandas. Supplementary TableS1 provides a complete list of all the specific python libraries used in AI-SCW as well as a brief description of their use. In addition, the specific python scripts used in implementing each component of AI-SCW is shown in supplementary TableS2. All of these scripts can be found online in this public Gitlab repository (https:// git. geomar. de/ opensource/ AISCW), where the complete source code files for the AI-SCW project are open source. Alongside the source code files, a detailed guide for setting up the programming environment and executing the respective scripts for each component of AI-SCW is provided. Image enhancements. Scale determination by laser point detection. Three laser points projected onto the seafloor and photographed in each image are used to determine the image scale. This scale is useful for marine scientists who rely on measurements made on the images, since it forms the basis for the conversion from pixel units to real world units. The scale is calculated as a ratio between the distance separating the three laser points (in pixel units) and their calibrated distance measured in real world units (centimeters). Manually annotating the laser points from thousands of images is laborious and non-scalable. This provides the primary motivation for automating the laser point detection. Below, we describe an approach for automatically detecting lasers from each image, and using these detections to calculate the scale for each corresponding image. Each image I of the data set consists of three-color channels (I(R), I(G), I(B)) and each of these channels has a pixel width w = 4480 and pixel height h = 6720. One training image with well visible laser points is manually selected and annotated. The annotations provide an estimate for the pixel coordinates of laser points in all images. This is done by creating a mask Mlp as triangle which connects the three annotated laser points. To allow for variability in laser point coordinates caused by varying OFOS altitude, a buffer is added around Mlp. This buffer is chosen as 250 pixels in our implementation. It was determined by randomly checking different images of varying altitudes to evaluate if laser points indeed fall within the buffered mask. The three annotated laser points provide the average pair-wise distance dlp between laser points in pixel units. To detect the red laser points in an image I, first a linear combination of its color channels is used to generate a laser signal image I(LS) : Equation1 was derived from the observation that since the pixels around the laser points were bright red in color (high values in the red color channel), it follows logically that the blue and green color channels had low values in the same region of the image. Therefore, subtracting a linear combination of blue and green color channels from the red color channel would result in an intermediate image, which has very high values around the laser points and very low values elsewhere. (1) I(LS)=I(R)− 0.2 (I(B)+I(G)) Table 1. Overview of the OFOS deployments in the Clarion Clipperton Zone during cruise SO268. Station Contract area Depth (m) Approx. track length (km) Approx. bottom time (h) Number of pictures taken at the seafloor SO268-1_21-1_ OFOS02 German 4538 9.2 11.0 3921 SO268-1_30-1_ OFOS03 German 4070 10.7 12.0 4981 SO268-2_100-1_ OFOS05 German Dredge (before) 4247 5.0 8.6 2749 SO268-2_117-1_ OFOS06 German 4109 7.7 8.5 2956 SO268-2_126-1_ OFOS07 German Dredge (after) 4117 8.8 10.0 3492 SO268-2_160-1_ OFOS11 German Dredge (after) 4115 11.0 10.0 3526 SO268-2_164-1_ OFOS12 German Dredge (after) 4118 5.0 9.0 3075 SO268-2_177-1_ OFOS13 German 4127 7.3 9.5 3414 SO268-1_63-1_ OFOS04 Belgian 4478 10.5 12.5 4532 SO268-2_128-1_ OFOS08 Belgian 4486 5.5 8.0 2749 SO268-2_147-1_ OFOS09 Belgian 4519 5.8 8.0 2981 SO268-2_153-1_ OFOS10 Belgian 4522 6.0 8.0 2302
1 Vol.:(0123456789) Scientif㘶c Reports | (2023) 13:8350 | https://doi.org/10.1038/s41598-023-35518-5 www.nature.com/scientificreports An automated image‑based workf㘶ow for detecting megabenthic fauna in optical images with examples from the Clarion–Clipperton Zone Benson Mbani 1*, Valentin Buck 1 & Jens Greinert 1,2 Recent advances in optical underwater imaging technologies enable the acquisition of huge numbers of high‑resolution seaf㘶oor images during scientif㘶c expeditions. While these images contain valuable information for non‑invasive monitoring of megabenthic fauna, f㘶ora and the marine ecosystem, traditional labor‑intensive manual approaches for analyzing them are neither feasible nor scalable. Therefore, machine learning has been proposed as a solution, but training the respective models still requires substantial manual annotation. Here, we present an automated image‑based workf㘶ow for Megabenthic Fauna Detection with Faster R‑CNN (FaunD‑Fast). The workf㘶ow signif㘶cantly reduces the required annotation ef㘶ort by automating the detection of anomalous superpixels, which are regions in underwater images that have unusual properties relative to the background seaf㘶oor. The bounding box coordinates of the detected anomalous superpixels are proposed as a set of weak annotations, which are then assigned semantic morphotype labels and used to train a Faster R‑CNN object detection model. We applied this workf㘶ow to example underwater images recorded during cruise SO268 to the German and Belgian contract areas for Manganese‑nodule exploration, within the Clarion–Clipperton Zone (CCZ). A performance assessment of our FaunD‑Fast model showed a mean average precision of 78.1% at an intersection‑over‑union threshold of 0.5, which is on a par with competing models that use costly‑to‑acquire annotations. In more detail, the analysis of the megafauna detection results revealed that ophiuroids and xenophyophores were among the most abundant morphotypes, accounting for 62% of all the detections within the surveyed area. Investigating the regional dif㘶erences between the two contract areas further revealed that both megafaunal abundance and diversity was higher in the shallower German area, which might be explainable by the higher food availability in form of sinking organic material that decreases from east‑to‑west across the CCZ. Since these f㘶ndings are consistent with studies based on conventional image‑based methods, we conclude that our automated workf㘶ow signif㘶cantly reduces the required human ef㘶ort, while still providing accurate estimates of megafaunal abundance and their spatial distribution. The workf㘶ow is thus useful for a quick but objective generation of baseline information to enable monitoring of remote benthic ecosystems. Modern digital underwater imaging platforms such as the Ocean Floor Observation Systems (OFOS)1, or Automated Underwater Vehicles (AUVs)2 are increasingly used for the exploration and monitoring of marine seabed ecosystems by researchers, the military, as well as other stakeholders in the private sector2. This is because these platforms offer affordability, ease of deployment, and the ability of repeatable seafloor sampling across varying scales with high temporal and spatial resolution3. As a result of the recent technological developments in both hardware and software, these imaging platforms are nowadays fitted with large memory storage capabilities, as well as high-resolution photo and video camera sensors4,5. Consequently, camera deployments during scientific expeditions now generate huge volumes of high-resolution images of the seafloor6. These images carry a lot of OPEN 1DeepSea Monitoring Group, GEOMAR Helmholtz Center for Ocean Research Kiel, Wischhofstraße 1-3, 24148 Kiel, Germany. 2Institute of Geosciences, Kiel University, Ludewig-Meyn-Str. 10-12, 24118 Kiel, Germany. *email:
[email protected]
2 Vol:.(1234567890) Scientif㘶c Reports | (2023) 13:8350 | https://doi.org/10.1038/s41598-023-35518-5 www.nature.com/scientificreports/ valuable information and insights into deep sea ecosystem, such as the characteristics of seafloor substrate7, as well as the megabenthic fauna that inhabits these ecosystems8. However, the lack of automated techniques for analyzing and interpreting these huge volumes of image datasets limits both the quality and quantity of information that can be derived from them e.g., by marine scientists focusing on deep sea geological and ecosystem monitoring9. Furthermore, the current manual approaches that involve the inspection and interpretation of each image by a human analyst are no longer feasible in this huge data regime, because manual annotation is expensive, subjective and thus prone to human bias10. Despite these challenges, underwater imaging has shown remarkable capability for documenting new discoveries in the deep ocean, using both color images and videos. In particular, the use of underwater imaging in scientific publications from domains such as marine ecological monitoring, animal behavior observation, and time-lapse imaging for temporal studies, is estimated to have increased two-fold5. This preference for marine imaging over traditional sampling is as a result of the ability of photographs to represent more taxa, and also because the spatial extent of the surveyed area can be determined accurately11. Therefore, automated workflows are needed to analyze the acquired underwater images to support these domain-specific applications. Depending on the application, these automated workflows can involve tasks such as semantic/instance segmentation, image classification, as well as object detection. Machine learning techniques have demonstrated the potential to automate both underwater image classification and object detection tasks12. While image classification involves assigning a single class label to describe the content of an entire image scene (e.g. a habitat class), object detection goes further to include the identification and localization of individual instances of objects visible in the image, typically by drawing bounding boxes around them13. This makes object detection models particularly useful for marine scientists who aim at identifying, measuring and counting underwater objects e.g. to estimate their density and abundance14. While modern object detection models such as Faster R-CNN15 can be trained to detect objects in images with relatively high accuracy16, they require a lot of manually annotated bounding box coordinates along with their corresponding class labels, which is very expensive and tedious to obtain17. Even when expert annotators are available, the selection of example images containing megafauna to be presented to the annotators can be very challenging; this is more pronounced in underwater image datasets of deep seabed areas because the frequency and diversity of megafauna is very low at greater depths, which implies that only a small proportion of the underwater image dataset contain visible megafauna18. In OFOS/AUV deployments where tens to hundreds of thousands of images have been recorded, the task of selecting example images with visible megafauna does pose a serious challenge. A proposed workflow for automated detection of megabenthic fauna should therefore incorporate (semi) automated ways of reducing and/or complementing the effort of human annotators e.g. by efficiently expediting the generation of annotations from the optical underwater images19,20. This automation should facilitate both the selection of example images with visible megafauna to be presented to the annotators, as well as the generation of a set of weak annotations to be refined later. In this context, weak annotations are imprecise or noisy annotations that can be obtained cheaply using unsupervised approaches21. An example of a set of weak annotations would be bounding box coordinates that only partly cover the body of an ophiuroid (e.g., its central disk) while leaving out its arms. When available, these weak annotations greatly reduce the effort required from expert annotators, since their tasks are essentially reduced to: (a) refining the provided bounding box coordinates to precisely cover the entire megabenthic fauna; (b) annotating additional megabenthic fauna that are not part of the provided weak annotations; and (c) assigning the correct morphotype class labels17. One computationally cheap way of generating these weak annotations is through the analysis of image superpixels22. Superpixels are partitions of an image where each partition comprises a group of pixels with similar perceptual characteristics23. In underwater images recorded from a relatively homogenous seabed substrate e.g. sandy or muddy bottoms in the deep sea, these superpixels generally correspond to the objects occurring on the seafloor, such as megabenthic fauna, rocks or marine litter24. Since the frequency of megabenthic fauna on the deep seafloor is very low compared to background objects such as the soft sediment or rock debris18, those superpixels that correspond to megabenthic fauna can be considered anomalous. This is because their visual properties are clearly different to the background seafloor25. However, in order to automatically distinguish between normal and anomalous superpixels, their visual properties must first be extracted and encoded into feature vectors. Although this can be achieved by manually identifying the distinguishing characteristics of the superpixels (e.g. color and texture), this process requires significant amount of domain expertise and experience to be done correctly6. An alternative approach is to automatically learn these properties directly from the superpixels e.g., by using convolutional variational autoencoders for feature extraction26. Anomaly detection algorithms such as iForest27 can then be applied to these features, so that anomalous superpixels can be detected and presented to expert annotators as weak annotations for refinement, labeling, and subsequent training of an object detection model e.g., Faster R-CNN15. Past studies have proposed various approaches for seafloor substrate classification28,29, and in particular for underwater object detection. Traditional image processing techniques have been used to estimate the coverage of seagrass meadows in Croatia through classification of irregular image segments30, as well as in the Palma bay using regular square image tiles31. In the same direction, a saliency-based workflow was implemented to approximate background regions of the image to detect underwater ‘foreground’ objects32, whereas contrast stretching and adaptive thresholding has been used to segment and subsequently detect underwater objects33. Further, a combination of Laplacian filtering, histogram equalization and blob detection has also been used to detect underwater objects34. Regarding the detection of objects and human artifacts on the seafloor, a region-based approach was used to detect marine litter in Greek waters24, whereas geometric reasoning was employed for the detection of pipelines on the seabed35. In another study36, underwater robots were used to perform color restoration in real time in order to improve accuracy when detecting and tracking mobile objects, whereas template matching was used to detect and track objects from images recorded using an underwater robot platform37. By
3 Vol.:(0123456789) Scientif㘶c Reports | (2023) 13:8350 | https://doi.org/10.1038/s41598-023-35518-5 www.nature.com/scientificreports/ modeling the propagation of light through the water column, a workflow was implemented to detect underwater objects using monocular vision38, while another one was proposed to detect underwater objects by leveraging collimated regions of artificial lighting39. Most recent studies employ deep learning approaches: an architecture was proposed for detecting objects in complex underwater imaging environments based on feature enhancements and anchor refinement40, whereas an augmentation strategy was used to simulate e.g. overlaps and occlusions to improve underwater object detection accuracy41. Similarly, an architecture was proposed to detect underwater objects by accounting for underwater image degradation through the joint learning of color conversion and object detection42. Finally, a variational autoencoder architecture was used to distinguish salient regions from the background based on reconstruction residuals25. In this study, we propose a three-stage workflow for automatically detecting megabenthic fauna from optical underwater images; examples of target megabenthic fauna classes (morphotypes) for this study are shown in Fig.1A, whereas the proposed workflow is conceptualized schematically in Fig.1B. The first stage involves generating superpixels from a small subset comprising e.g., 500 images per dive/camera tow, which are randomly sampled to reduce computational cost in this stage. A variational autoencoder is Figure1. Overview of our optical image-based megabenthic fauna detection framework. (A) Examples of target morphotypes, including litter, that were detected on the seafloor. (B) Schematic diagram of our threestep workflow: The first step (automatically) generates superpixels from a small subset of sampled images, and (automatically) extracts their features for training an anomaly detection model. The second step detects anomalous superpixels (automatically) from a larger subset of images, and (semi-automatically) proposes them as weak annotations ready to be post-processed and assigned semantic morphotype labels (manually). The final step uses the semantic annotations to (automatically) train a Faster R-CNN object detection model, which then detects instances of benthic megafauna visible in the entire underwater image dataset (automatically), allowing for the estimation of megafaunal abundance, diversity and spatial distribution (manually).
4 Vol:.(1234567890) Scientif㘶c Reports | (2023) 13:8350 | https://doi.org/10.1038/s41598-023-35518-5 www.nature.com/scientificreports/ then applied to these superpixels to extract feature vectors, which are used to train an iForest anomaly detection model. The second stage applies the trained iForest model dive-by-dive to detect anomalous superpixels from a much larger subset of underwater images e.g., comprising six out of twelve dives. A binary classifier is used for post-processing the anomalous detections to remove false positives. The bounding boxes of the truly anomalous superpixels are then presented as a set of weak annotations to an expert annotator, who assigns semantic morphotype labels to them. The final stage uses the semantic annotations to train and evaluate a Faster R-CNN object detection model, which is subsequently used to detect and classify megafauna visible in all images from all dives. These georeferenced detections are finally used to estimate abundance, diversity and spatial distribution of megabenthic fauna within the working area. Our approach significantly reduces the required human annotation effort, since the user input is only required to post-process the automatically generated weak annotations, and assign them semantic labels. Furthermore, we have also open sourced the python scripts implementing each component of the above-described workflow, along with detailed documentation to guide users to get started using and/or extending our workflow. Thus, our approach offers a convenient underwater image annotation solution for the marine imaging community, allowing them to quickly generate accurate baseline information that allows for efficient and repeatable characterization of ecological and spatial distribution of remote marine benthic communities, including their habitats, at varying spatio-temporal scales. Results Visualization of superpixel separation. This section provides projections of both normal and anomalous superpixels onto a two-dimensional feature space for visualization purposes. These projections are obtained by applying Principal Components Analysis (PCA) onto the data matrix of feature vectors extracted from the superpixels. A grid view of truly anomalous superpixels is also provided. Superpixels for training the iForest anomaly detector. Figure2 shows the feature space representation of superpixels used to train the iForest model. The figure clearly shows that superpixels representing the background seafloor are densely distributed around the center of the feature space since they are visually similar, while those with unusual visual characteristics are distributed farther away towards the periphery of the feature space. Therefore, the background seafloor superpixels are obviously the majority, and were considered the ‘normal’ in this study. Detected anomalous superpixels. Figure3 shows the feature space representation of the anomalous superpixels that were detected from dive 126. While some of the detected anomalies are false positives e.g., the red laser Figure2. Feature space projection of superpixels whose features were used to train the anomaly detection model. Those representing the background seafloor are densely distributed around the origin of the feature space, whereas few anomalous superpixels are sparsely distributed further away towards the periphery of the feature space.
5 Vol.:(0123456789) Scientif㘶c Reports | (2023) 13:8350 | https://doi.org/10.1038/s41598-023-35518-5 www.nature.com/scientificreports/ points and unusually dark objects on the seabed, the rest of the detected anomalies indeed represent interesting objects e.g., megabenthic fauna, or other unusual objects worth investigating. Weak annotations (truly anomalous superpixels). As mentioned above, some of the detected anomalous superpixels are false positives that do not represent megabenthic fauna. Thus, it was necessary to remove these false positives, and retain only the truly anomalous superpixels during further processing. Below, we provide the results of two post-processing strategies that we attempted: setting a threshold on the anomaly score; and training a supervised binary classifier. Figure4A shows the results of post processing obtained by setting a 75th percentile threshold on the anomaly scores assigned to the anomalous superpixels; superpixels with anomaly scores greater than the set threshold were marked as truly anomalous. While the superpixels are visually anomalous in some way, some of them still represent objects that are not of interest in this study e.g., the red laser points, and white spots surrounded by black pixels. Because of this, we concluded that thresholding based on anomalous scores alone was not sufficient to distinguish truly anomalous superpixels from false positives. There was also no obvious way of determining the suitable anomaly score threshold. Figure4B shows the truly anomalous superpixels that were obtained by using our supervised binary classifier. Unlike the thresholding approach, the binary classifier correctly identified the set of truly anomalous superpixels. Bounding box coordinates of these truly anomalous superpixels were then proposed as a set of weak annotations. Training and evaluating FaunD‑Fast model. The weak annotations still lack semantic morphotype labels, and are therefore not directly usable. In this section, we provide the results of the semantic labeling exercise involving an expert annotator, as well as the results of the performance evaluation of the Faster R-CNN object detection model that was trained using these annotations. Semantic labeling of the weak annotations. A human expert manually inspected all the weak annotations and assigned them semantic morphotype labels. The expert also annotated instances of megafauna that were visible in the images, but missing from the weak annotations. This semantic labeling exercise was repeated twice (after shuffling the weak annotations) to reduce biases e.g., due to human fatigue. Supplementary FigureS1 shows a screenshot of our superpixel annotation software during an active semantic labeling session. The left panel of the software shows all truly anomalous superpixels, whereas the right panel displays their bounding box extents overlaid on the respective parent images. The bottom panel shows the morphotypes that were considered in this study. These include: anemone, coral, fish, gastropod, holothurian, ophiuroid, sea urchin, shrimp, sponge and xenophyophore. Figure3. Feature space projection of the anomalous superpixels detected from images in dive 126. While some false positives such as red lasers and dark pixels of the water column were also detected, the rest of the anomalous detections represent potential instances of megafauna whose bounding boxes can be proposed as a set of weak annotations.
6 Vol:.(1234567890) Scientif㘶c Reports | (2023) 13:8350 | https://doi.org/10.1038/s41598-023-35518-5 www.nature.com/scientificreports/ Figure5 shows the distribution of the annotated morphotypes. In terms of proportions, the dominant morphotypes were ophiuroids (31%), sponges (18%), xenophyophores (17%) and anemones (11%). The other morphotypes had occurrences of less than 10%. These are the annotations we used to train and evaluate our FaunDFast model; we have provided these annotations as a csv file in supplementary TableS1. Performance evaluation of FaunD-Fast model. Our FaunD-Fast model achieved an average precision (AP.50) score of 78.1% at an IoU threshold of 0.5. The model performance was higher when detecting large-sized objects/ megafauna, as can be shown by the values of (APlarge) and (ARlarge) metric categories that are both greater or equal to 70% (see Table1). On the other hand, the model’s performance was lower when detecting small-sized objects, since both their average precision (APsmall) and recall (ARsmall) values were less than 20%. When compared to competing state-of-the art models from the empirical evaluation inLütjenset al43, their best model (CM-X-101/Synth-Blcd) performed better than ours with regards to the (AP.50:0.95) metric category, which is obtained by averaging the precision values over multiple IoU thresholds. In contrast, our model performed better than all the compared models with regards to the (AP.50), which is the precision at a single (absolute) IoU threshold of 0.5. In addition, our model also performed better than the others with regards to the Figure4. Grids of image patches showing truly anomalous superpixels obtained by (A) Thresholding the anomaly scores, and (B) Binary classifier trained with examples of both true and false positives. Thresholding produces undesired results e.g., the red laser points and the dark patches from the water column. On the other hand, the binary classifier results in a set of truly anomalous superpixels that are clearly instances of megabenthic fauna. These were proposed as weak annotations. Figure5. Distribution of the annotated morphotypes after exporting from the annotation software. Ophiuroids, sponges and xenophyophores were among the dominant morphotypes in the annotated dataset.
7 Vol.:(0123456789) Scientif㘶c Reports | (2023) 13:8350 | https://doi.org/10.1038/s41598-023-35518-5 www.nature.com/scientificreports/ (APlarge) and (AR1) metric categories; this implies that our model was good at detecting large-sized megafauna. However, all the compared models reported very low precision and recall scores when detecting small-sized objects. To show our model’s performance on the semantic morphotype classes, we present the confusion matrix in Supplementary FigureS2. The confusion matrix shows that majority of the morphotypes were correctly localized and identified. In particular, xenophyophores and ophiuroids contributed towards the largest proportion of false negatives. This could be because the visual characteristics of some xenophyophores and partially burrowed ophiuroids are similar to the seabed substrate, which makes them difficult to detect. In addition to this, Fig.6 shows that in instances characterized by associations among morphotypes e.g., between ophiuroids and sponges/ corals, the model made incorrect or low-confidence predictions. Although none of the compared models (in Table1) was able to achieve the highest score across all the metric categories, these quantitative evaluation results show that overall, the performance of our model was on a par with the best performing state-of-the art alternative(s), yet our approach required less manual annotation effort. Abundance, diversity and spatial distribution of the detected megabenthic fauna. Figure7A shows qualitative examples of correctly identified and localized megafauna as detected by our FaunD-Fast model. In total, 27,954 individual instances of megabenthic fauna were detected from the entire image dataset. Furthermore, we estimated the megafaunal abundance within German area to be approximately 0.247 ind. m−2 while in the Belgian area it was approximately 0.200 ind. m−2. Figure7B shows the distribution of the detected morphotypes. Ophiuroids and xenophyophores were the most dominant morphotypes accounting for 62% of all the detections. Other species are sponges (9.6%), sea urchins (7.8%), gastropods (6.1%), anemones (5.7%), corals (4.0%) and holothurians (3.1%). The rest such as fish and shrimp have occurrences of less than 1%. Apart from ophiuroids which are abundant in both contract areas, the German seabed is predominantly occupied by xenophyophores (22.8%) and sponges (10%); the Belgian contract area was predominantly occupied by sea urchins (34.9%), anemones (14.5%) and sponges (10%). Figure7C shows few examples of detected morphotypes, while a table summarizing all the detections is provided in the Supplementary TableS2. Table 1. Performance comparison relative to other state-of-the art benthic fauna detection models43. The highest scores per metric category are indicated in bold. Model AP.50:.95 AP.50 APsmall APmedium APlarge AR1AR10 AR100 ARsmall ARmedium ARlarge FaunD-Fast (Ours) 46.5 78.1 12.7 42.0 69.7 50.0 52.0 52.4 16.2 50.0 73.2 CM-X-101/Baseline 41.7 68.2 25.3 29.3 54.7 21.6 51.6 55.2 25.4 45.1 70.8 CM-X-101/Synth 48.8 71.0 27.4 39.1 62.8 24.7 58.8 64.2 27.9 57.3 77.1 CM-X-101/Synth-Blcd 51.8 76.7 27.5 40.2 66.1 25.7 59.0 63.9 27.9 55.7 77.9 CM-X-101/Trad. Augm 48.8 75.0 26.9 38.6 58.5 23.0 55.3 58.9 27.2 50.1 72.6 CM-X-101/Fusion 51.7 74.1 27.1 42.1 65.1 24.9 57.6 61.6 27.5 52.2 77.6 CM-V-99/Synth 47.9 72.0 27.9 37.0 62.8 23.6 56.6 61.9 28.3 52.6 77.1 CM-L-M/Synth 27.3 48.6 19.1 19.0 40.0 18.3 39.1 43.7 20.0 34.4 59.5 M-X-101/Synth 33.3 53.2 13.2 22.7 53.0 20.7 39.2 40.0 13.2 30.6 60.7 R-X-101/Synth 47.8 70.7 27.9 37.1 62.2 24.2 56.6 61.9 28.4 53.8 76.7 Figure6. Examples images showing correctly detected instances of megabenthic fauna, as well as instances of both false positives (FP) and false negatives (FN). Morphotypes whose visual characteristics is similar to the seafloor substrate (e.g. xenophyophores and partially burrowed ophioroids) resulted in a higher proportion of false negatives. Also, incorrect detection/localization was observed in instances where morphotypes formed associations with each other e.g. between ophiuroids and sponges.
8 Vol:.(1234567890) Scientif㘶c Reports | (2023) 13:8350 | https://doi.org/10.1038/s41598-023-35518-5 www.nature.com/scientificreports/ The spatial distribution of the detected megabenthic fauna is shown in Fig.8. The map shows that the German contract area contains a higher abundance of megafauna compared to the Belgian area (see details in the “Discussion” section below). For ease of visualization, the detections (points) were spatially clustered by first gridding them into square blocks of size 200m, and then normalizing the absolute count of megafauna based on the visual footprint of each respective block (in square meters). Thus, the symbology size is proportional to the abundance of megafauna within each spatial cluster/block. Discussion The proposed megabenthic fauna detection workflow comprised the generation of weak annotations from superpixels, semantic morphotype labeling of the proposed weak annotations, and finally the usage of these annotations to train our FaunD-Fast model. Below, we discuss key aspects of these proposed workflow steps, and provide a more detailed discussion of the spatial distribution, density and diversity of the detected megabenthic fauna. We also suggest a few recommendations for further research. The hyperparameter settings of the segmentation algorithm control the geometrical properties of the generated superpixels e.g., shape (regular or irregular), and size (large or small). Given that the used seafloor images comprised background seabed substrate (Mn-nodules) and other objects of varying shapes and sizes, we had to manually determine the optimal values of hyperparameters such as scale (pixel size) and width of the gaussian filter that smooths the image prior to segmentation. These parameters must be properly tuned if the workflow is applied to other underwater image datasets. If this is not done thoroughly, the generated superpixels may be of low quality hence negatively affecting the accuracy of downstream analysis. In our case, we observed that a poor choice of these hyperparameters led to inaccurate segmentation of certain morphotypes of interest, especially those with extended arms and spikes e.g., ophiuroids and sea urchins. In a related previous study using a fish dataset44, the authors also emphasize that segmentation hyperparameters must be properly optimized before being applied to underwater images recorded from challenging environments e.g. where both illumination conditions and background seafloor properties vary within and between datasets. In addition to the hyperparameter settings, we also had to select a subset of images whose superpixels would be used to train the iForest anomaly detection model; our subset comprised 500 randomly sampled images that generated 125,000 superpixels. We Figure7. (A) Qualitative examples of detected instances of megabenthic fauna (B) Distribution of morphotypes that were detected by our FaunD-Fast model. This distribution is similar in shape to that of annotations (see Fig.5), except the FaunD-Fast detected a lot more instances of megafauna. (C) Grid view showing megafauna examples grouped by morphotypes in every row of the grid; the morphotype label for each row follows the same order as in panel (B).
9 Vol.:(0123456789) Scientif㘶c Reports | (2023) 13:8350 | https://doi.org/10.1038/s41598-023-35518-5 www.nature.com/scientificreports/ chose this sample size because it fit into our CPU memory at train time (a larger subset should be used if more memory and compute is available e.g., in HPCs). In any case, we observed that this random sampling approach potentially resulted in a more representative subset compared to e.g., a manual sampling approach. This is because the huge volume of images could easily cause the analyst to mistakenly choose a subset of images that represent more or less the same region of the seafloor (consecutive images were recorded every 10s, and are stored on disk in order of their acquisition time), or those which look ‘interesting’. The majority of the superpixels generated for training the anomaly detection model represented background seafloor, compared to the relatively few megafauna superpixels (see Fig.2). This observation was expected since the abundance of megabenthic fauna in the deep ocean is typically very low, due to the low organic carbon flux/little food availability in great water depth18. Similar findings have been reported in studies that examined the relationship between megabenthic fauna communities and bathymetric gradients e.g.45–48. As a result, we Figure8. Map view showing spatial distribution of detected megabenthic fauna along camera deployment tracks in both the German and Belgian contract areas. The German seabed contained higher abundance of megafauna, probably because of availability of food in form of sinking organic material since it is on average shallower than the Belgian seabed. The map was generated using the open source QGIS software v3.2 (https:// www. qgis. org/).
10 Vol:.(1234567890) Scientif㘶c Reports | (2023) 13:8350 | https://doi.org/10.1038/s41598-023-35518-5 www.nature.com/scientificreports/ observed that the detection of anomalous superpixels was relatively straightforward: they are visually rather different from the background, and are thus distributed further from the center of the feature space where the majority of the background superpixels clustered (see Fig.3). The approach of analyzing the visual properties of superpixels has been employed in previous studies aimed at identifying the boundaries of interesting objects on seafloor images32,44, as well as on terrestrial images25. In contrast to these three publications, our approach is different because we do not make assumptions regarding the region of the image in which the foreground objects are expected e.g., in the central portion of the image. Instead, we assume that our objects of interest will be located anywhere on the image, and thus our trained iForest model detects anomalous superpixels based purely on features extracted from the superpixels. The detected anomalies still had to be post-processed to remove false positives, which occurred because we intentionally set iForest’s ‘contamination factor’ setting to a high value (0.4); this caused it to flag as many anomalous superpixels as possible (both obvious and subtle). We did this because the visual properties of some megafauna of interest such as xenophyophores are very similar to background seafloor substrate, yet we wanted the iForest model to detect them as well. In a previous study on image-based megafauna community assessment in the DISCOL area of the south Pacific Ocean49, the authors also point out the difficulty that even expert human annotators face when it comes to distinguishing certain morphotypes from background seafloor. The bounding box coordinates of the detected anomalous superpixels were proposed as a set of weak annotations to be inspected and labeled by an expert annotator. This semi-automated approach significantly reduced the human effort required in generating training annotations, because it was no longer necessary for the expert annotator to manually inspect a large number of images with the aim of identifying the few that contain visible fauna, and then mark these fauna manually. The available bounding box further reduced the work of the expert annotators to just verifying and assigning semantic morphotype labels. Since these annotations were generated from unsupervised segmentation and anomaly detection methods, they are not sufficient on their own (in quantity and quality) to estimate the abundance of megafauna on the seabed for the entire dataset. They just represent training examples for a state-of-the art object detection model, which can then be applied to the entire image dataset. After using the generated annotations to train our FaunD-Fast model, we achieved a good performance (78.1%) that is on a par with other state-of-the art object detection models, which were trained in previous studies43 using underwater image dataset comparable to ours (see Table1). Given that our FaunD-Fast model achieves this good performance for a fraction of the annotation effort implies that it is scalable to other applications involving huge volumes of underwater imagery. Another observation is that since our model uses the twostage Faster R-CNN architecture that prioritizes prediction accuracy over speed15, the results of our comparison with other state-of-the art models (in Table1) shows that our approach is suitable for deployment on workstations with good processing capability e.g. GPU and memory resources. For faster detection on edge devices and smaller computers, the comparison implies that a single stage detector is probably more suitable; future research could explore this further. Also, none of the compared state-of-the-art models achieved the highest score across all the metrics, which could be because each model are trained to optimize a different loss function50. Finally, the comparison revealed that the prediction accuracy for small-sized megafauna was consistently lower than for large-sized megafauna across all the compared models. This could be caused by the convolutions and pooling layers in the object detection architecture, which gradually reduce the size (and resolution) of the image deeper into the network, making small sized objects harder to detect51. Future research could explore backbone network architectures that improve the model’s detection of small-sized objects. Moreover, future research could explore how to extend the FaunD-Fast model to re-use image features extracted from earlier stages of the workflow so as to reduce computation cost, especially for real time megafauna detection while at sea. Based on the detection results of our trained FaunD-Fast model, we found that the megabenthic fauna abundance in our working area was relatively low. This finding is consistent with previous studies from the CCZ that also reported megafaunal abundances of less than one individual per square meter52–55. In terms of regional differences between the two contract areas, a higher abundance was observed in the German area (0.247 ind. m−2) compared to the Belgian area (0.200 ind. m−2). Similarly, the megafaunal diversity was higher in the German area, with a Shannon diversity index of 2.4 compared to the 1.7 in the Belgian area. Because the German area is located approximately 1050km east of the Belgian area, the observed high diversity and abundance could be as a result of the east-to-west reduction in the particulate organic carbon flux (POC), as has also been reported in previous studies56,57. Also, the difference in abundance between the two contract areas could be explained by availability of food source in the form of sinking organic material through the water column; availability of food is higher in German area because it is on average shallower (−4121m) than the Belgian area (−4510m). This relationship between food availability and abundance of megafauna in the Pacific has also been reported in a previous study58. Considering the relationship between megabenthic fauna abundance and the seafloor substrate classes from28, the Belgian area contained more than 68% of detected megabenthic fauna occupying the large-sized nodules, even though the seabed substrate comprised both largeand densely-distributed nodules. A lower proportion of megabenthic fauna was observed in densely distributed manganese nodules, probably because this substrate class does not allow enough space for soft-sediment dwellers, as was also pointed out in a previous study59. On the other hand, analysis in the German area revealed that 57% of megabenthic fauna occurred in seafloor substrates comprising patchy nodules. In addition to being the dominant seafloor class in this area, patchy nodules also provide a natural balance between soft and hard substrates, which would accommodate both hard and soft sediment dwellers, as was also reported in a previous study59. In both contract areas, we found that ophiuroids and xenophyophores were the most abundant and diverse morphotypes, and that they occur in association with each other, while occupying both hard and soft bottom substrates. Similar conclusions were also drawn from previous studies59–61.
! Peer review status:! This is a non-peer-reviewed preprint submitted to EarthArXiv.
Automated underwater image analysis reveals sediment patterns and megafauna distribution in the tropical Atlantic Benson Mbani1* and Jens Greinert1,2 1DeepSea Monitoring Group, GEOMAR Helmholtz Center for Ocean Research Kiel, Wischhofstraße 1-3, 24148 Kiel, Germany 2Institute of Geosciences, Kiel University, Ludewig-Meyn-Str. 10-12, 24118 Kiel, Germany. * Correspondence: Corresponding Author
[email protected] Abstract The deep sea environment comprises diverse ora, fauna and habitats, whose characterisation is key towards our collective understanding of ocean health and resilience. Whereas direct sampling allows for detailed investigation of the vertical variability of seabed characteristics at small spatial scales, optical imaging is suitable for high-resolution assessment of the spatial distribution of habitats and their benthic megafauna across multiple scales. These assessments are typically facilitated by scientic expeditions that survey extensive seabed areas using e.g. continuous imaging techniques, generating huge volumes of high-resolution images for which manual inspection and annotation is costly, non-scalable and therefore infeasible. Transforming these terabyte-scale images (and videos) into actionable insights requires automated workows that expedite both the generation of baseline information, as well as downstream spatial-ecological analysis. Here, we deployed two A.I workows to automate the annotation of seabed substrates and megafaunal taxa from still images, which we acquired during seven camera deployments along an 18° N East-West section in the tropical Atlantic north of Cabo Verdes. We manually inspected the auto generated annotations for quality, and subsequently assigned them semantic labels. Thereafter, we used clustering, feature space visualisation and multivariate statistical analysis techniques to classify the seaoor into habitats, estimate megafaunal abundance and spatial distribution patterns, as well as environmental drivers that inuence the identied patterns. Our results show that the seabed can
be partitioned into seven clearly distinct clusters, with each of these clusters showing visible sub-partitions. Investigations revealed a clear gradient in terms of sediment disturbance due to biogenic activity, with images showing little-to-no sediment disturbance grouping together on one half of the feature space, whereas those images with visibly vigorous signs of sediment reworking clustered on the other half. Our results also show that megafaunal abundance was on average 14 times higher in the Eastern region of our study area, which was approximately 700 metres shallower and closer to shore than the Western region. This observed high abundance could be attributed to higher POC ux that transports more organic matter to the shallower seabed, as well as due to relatively warmer temperatures that enhance metabolic rates of benthic fauna. Our results further reveal geographic hotspots of megafauna in topographically complex features such as the sides of a submarine canyon and the top of seamounts. The complex topography of these features introduces heterogeneity that creates diverse microhabitats and unique niches that megafauna exploit. Finally, we observed that while co-varying depth and longitude variables generally explained the separation between the two main megafaunal communities in our East-West oriented working area, bathymetric drivers like slope and ruggedness had a more pronounced inuence in the deeper Western region (-3698m) compared to the shallower Eastern region (-2477m deep). Collectively, these ndings demonstrate that the integration of A.I workows into classical spatio-ecological methods does expedite the transformation of large volumes of marine image datasets into actionable insights, thereby signicantly contributing to our understanding, monitoring and sustainable use of ocean resources. Introduction The deep sea comprises a wide range of benthic habitats, is home to diverse sets of oral and faunal communities, and is the largest biome on earth1. Despite this, the biodiversity within these remote ecosystems is still largely under-sampled2and/or patchily documented3, even after accounting for the increased frequency of scientic expeditions over the past decades4. This is because of logistical, technological and nancial challenges that constrain the overall spatio-temporal extents that can be reasonably investigated 5. Besides, in-situ and/or visual characterization of organisms in the deep sea can sometimes be non-trivial, either because organisms in these environments are new to science, or because their distribution patterns (and ecosystem processes) are not yet properly understood 6.
Recent scientic studies have also provided conclusive evidence showing a global decline in marine biodiversity as a result of both natural and anthropogenic factors e.g. overshing, pollution, coastal development, natural climate variability, and long-term geological processes like sedimentation and tectonic/hydrothermal activities 7. To better quantify and address this biodiversity decline, globally coordinated eorts are required to not only increase the frequency and spatial extent of marine ecosystem surveys, but also to expedite the analysis of the acquired datasets. These datasets include high resolution images and videos collected using platforms such as Autonomous Underwater Vehicles (AUVs), Remotely Operated Vehicles (ROVs), and towed Ocean Floor Observation Systems (OFOS) 8. While imaging sensors attached onto these platforms conveniently allow for non-invasive surveying of deep-sea environments in high resolution, they generate huge volumes of imagery for which manual interpretation is unfeasible 9. As a result, automated workows based on emerging digital technologies are required to expedite the processing and annotation of images, thereby providing comprehensive baseline information on geological, sedimentological and biological properties of marine ecosystems10. Modern machine learning techniques have demonstrated the capacity for rather quick yet accurate extraction of semantic information from large sequences of image and video datasets 11. In particular, pre-trained computer vision models based on convolutional neural networks are nowadays readily available for download from open-source repositories (e.g TensorFlow Hub and PyTorch Hub), and can be directly deployed as-is to accomplish common tasks such as image enhancement, classication, object detection and dense pixel segmentation 12. Given that most of these pre-trained models were originally trained to identify common objects on terrestrial images using benchmark datasets like ImageNet 13, the models require ne-tuning using annotated underwater images before they can be useful for applications such as marine habitat mapping and biodiversity assessment 14. This requirement poses signicant bottlenecks in at least two dimensions: First, annotating images after every scientic expedition is costly, unscalable and therefore undesirable; Second, marine environments naturally exhibit low density of megafauna with increasing depth, which implies that organisms will be visible on only a handful (out of possibly tens of thousands) of acquired images that are typically unknown apriori 15. Addressing these challenges requires automated A.I-based seaoor classication and megafaunal detection workows that not only work well in a specic working area, but that are easily generalizable to
other marine ecosystems 16. Such a generalised approach saves human analysts the trouble of annotating datasets from scratch, allowing them to concentrate on rening and assigning semantic morphospecies labels only to auto-generated annotations 9. The semantic annotations can then form the basis of downstream assessment of spatio-ecological distribution patterns of habitats, megafauna and environmental drivers. Megafaunal species are not distributed randomly in space 17. Instead, they cluster together into biotically-similar communities that are in turn structured by processes and variables such as bathymetric gradients, geomorphology, food availability, chemical/physical bottom water conditions, as well as sediment or hardground properties related to settling, hiding or breeding 18. Characterization of these megafaunal patterns is typically performed using multivariate statistical analysis techniques 19, which are also applicable to this present study given that our image-derived annotations comprise abundances for multiple taxa. Before using these statistical techniques to assess the distribution of megafauna, however, it is necessary to rst account for the inconsistent visual footprints among respective images due to their variable acquisition heights 16. This inconsistency can be resolved by systematically dening standardised sampling units (e.g equal-area quadrats or xed-length linear transects), within which megafauna counts are pooled and normalised relative to the actual observed area 20. Collectively, these sampling units encode biotic information as abundances that can simply be binned and plotted on a choropleth map to visualise spatial distribution of megafauna. Alternatively, ordination techniques such as non-metric multidimensional scaling (nm-MDS) can be used to graphically display inter-relationships among the dierent taxa in feature space 19. Furthermore, an arbitrary number of relevant environmental variables can also be superimposed on the ordination plot, allowing for a more nuanced visual assessment of the (subset of) abiotic factors that inuence the dierent clusters of megafauna 21. Finally, spatial autocorrelation analysis may also be used to reveal megafaunal hotspots, coldspots and outliers 22. Past studies have proposed various workows and approaches for semi-automating the annotation of underwater images. Supervised approaches have been used extensively for tasks such as image-based seaoor classication because they are capable of generating accurate annotations (in inference mode) whenever sucient number of labelled examples are available for training 23. To facilitate rapid innovation, experimentation, reproducibility and evaluation of supervised models,
there have been studies aimed at curating standardised (labelled) benchmark datasets from both real 24 and simulated marine environments 25. Whereas classical machine learning techniques such as random forests 26 and support vector machines 27 were predominantly incorporated in marine image analysis workows in the past decade, recent studies almost exclusively use convolution neural networks 28. In particular, models such as YOLO 29, RetinaNet 30 and Faster R-CNN 31 are now widely used for detecting, localising and classifying ora and fauna from images and videos after training with just hundreds of training examples per class 11. There is also evidence that these deep learning models are computationally resource-intensive only during model training, otherwise the models are remarkably ecient when making predictions in inference mode 32. Unsupervised approaches such as template matching 33 and superpixel-based segmentation have also been used in previous studies 34, typically as an initial preliminary step e.g. to cheaply generate weak annotations 35, or to quickly sort images based on natural groupings 36. Some studies still rely exclusively on human workforce to exhaustively annotate their datasets, which is accurate (and arguably the gold standard) but also very costly and non-scalable 37. Regardless of the chosen annotation strategy, the generated annotations are normally used as inputs to downstream spatio-ecological workows that rely on e.g multivariate statistics and measures of spatial autocorrelation to characterise abundances, diversity, and spatial distribution patterns of megafauna 19 17 21 Here, we investigated seaoor habitats and benthic megafaunal distribution patterns in the tropical North Atlantic using the conceptual workow in Figure 1. Specically, we ne-tuned A.I workows that we previously developed for classifying seaoor habitats 9and detecting benthic megafauna 23 in the Clarion-Clipperton Zone. We used the ne-tuned models to expedite the annotation of a new dataset comprising seaoor images from the tropical North Atlantic. Given that the A.I workows were originally used for benthic assessments in the Pacic, one broad objective for this study was to investigate the generalizability of the two workows when presented with dataset from a completely dierent area and geological setting. Specic objectives were: (1) to reveal subtle variability in seaoor habitat classes using unsupervised machine learning techniques; (2) to semi-automate the detection, localisation and classication of megabenthic taxa from sequences of high-resolution images; (3) to characterise the spatial distribution patterns of the annotated megafauna; (4) to estimate megafaunal abundance, diversity and community
composition; and nally, (5) to assess the inuence of environmental drivers on the observed distribution patterns.
Figure 1: Flow diagram showing the interconnected components of our proposed workow that comprises three key steps: First, we enhance the visibility of images before deploying two A.I workows to classify seaoor images into habitat classes, and also to detect megafaunal taxa; Second, we inspect and assign semantic taxa labels to the auto-generated weak annotations, before converting the absolute taxa counts into abundances relative to actual observed area within our predened xed-size sampling units; Finally, we use the abundances to characterise spatial distribution patterns of megafauna, and also to graphically display interrelationships among biotic and abiotic variables in ordination feature space. Study Area Our working area oshore Mauritania and North of Cape Verdes (Figure 2) followed an East-West orientation, with a total of seven camera deployment stations distributed between the Eastern region (comprising dives 131, 144, 145) and Western region (dives 19, 32, 28, 78). The Eastern region was shallower with depths ranging between -2470 m up to -2970 m. This region was characterised by topographically complex features such as a seamount (in dive 145), as well as the submarine canyon at -2920 metres water depth (in dive 144). The canyon exhibits steep near-vertical 20-metre-high walls with a cross section that is approximately 500 metres wide, marking a visibly distinct narrow passage on the seabed. CTD proles further show that the water masses in the Eastern region are relatively warmer, with average temperatures of (2.85°C ± 0.12). In contrast, the Western region was deeper and relatively colder, with depth ranges of between -3128m and -3693 m, and average temperatures of (2.50°C ± 0.08). There was also a seamount in this region that rose approximately 200 m high from the seabed (in dive 32), as well as a pair of adjacent 40-metre high locally elevated abyssal hills (in dive 28). For further details on the physical and water mass properties for respective dives, please refer to the CTD proles in Supplementary Figure S1.
Figure 2: Map showing the OFOS (camera) deployment tracks during cruise M182 to the tropical North Atlantic, which we conducted on board RV Meteor between May - July 2022. Note that in this map we only show camera deployments from deep sea environments (> 2,000 metres water depth). Also notice that dive 145 is shorter than the other dives because the camera malfunctioned after just 1 hour 13 minutes of bottom time.
Image dataset We surveyed the working area between May 31st and July 10th 2022 on board RV METEOR during expedition M182 (Greinert et al., 2024; link to cruise report for later). The aim of the cruise was to study the inuence of mesoscale eddies on (a) biogeochemical processes in the Eastern boundary upwelling systems, and (b) modulation of organic matter transport from the surface waters down to the seaoor 6. The sampling campaign involved the deployment of several gears and systems such as the extended Ocean Floor Observation Systems (XOFOS), CTDs, MultiNets, biogeochemical landers, AUVs, and ship-based multibeam bathymetry. For this particular study, we used still images collected by the XOFOS, which is an imaging platform comprising a topside unit on the ship (for power, data connection and live video feed), as well as a subsea unit that is lowered into the water column by a winch system to survey the seaoor (up to 6000 metres deep). The subsea unit comprises a heavy metal frame that houses forwardand downward-looking 24-megapixel digital Ocean Imaging Systems camera (DSC 24000). The XOFOS records Images automatically at a constant frequency that is set before deployment, as well as through hotkey functionality for recording adhoc images or random events of interest . In addition to the camera, the XOFOS is also tted with downward facing LED lights/ashers and a USBL positioning system that tracks the platform position during image acquisition. Additional sensors such as ADCPs, CTDs and other loggers may also be attached to the XOFOS, allowing for a straightforward integration of auxiliary datasets and images based on synchronised timestamps. Based on the above set up, we obtained 8838 still images by photographing the seabed at constant frequencies of 0.07 Hz (in dives 28, 32), 0.10 Hz (in dives 78, 131, 144, 145) and 0.2 Hz (in dive 19). The seven XOFOS dives covered a total track length of 22.7 kilometres, which represents a visual footprint of approximately 73,616 m2on the seaoor. We estimated this visual footprint based on the xed opening angles of the camera (48° horizontal, 33° vertical) and the acquisition heights of respective images above the seaoor. Table 1 below provides an overview of our camera deployments, while the cruise report contains further technical details regarding the image acquisition setup (Greinert et al., xxxx).
Figure 4: Variability in number of images, XOFOS speed, and actual observed area within xed-size (100-metre-long) sampling units along respective survey transects. (A) shows an obvious inverse relationship between the speed of the XOFOS and the number of acquired images, whereas in (B) the relationship between number of images and observed visual footprint is not obvious, especially in topographically complex terrains like seamounts. Note that the relatively high number of images at the start of some transects is caused by the initial stabilisation phase, where the deployed XOFOS rst experiences twists and turns in more or less the same location before eventually maintaining a linear transect. Assessing the spatial distribution of megafauna Characterising the spatial distribution of megabenthic fauna allows us to provide geographic context to the image-derived abundances. Here, we used the centroid coordinates of each sampling unit to plot their locations in map view. We applied quantile classication to bin abundances into eight distinct classes that we used to colour-code the choropleth maps (Figure 13). This visual representation allowed for a straightforward interpretation of the variability in megafaunal abundances relative to the background bathymetry that we plotted as a basemap. In addition to the planar map view, we also plotted the abundances along elevation proles of each transect, to investigate variability at local heights. To complement the qualitative choropleth mapping, we used quantitative measures of spatial autocorrelation to reveal regions of the seaoor where geographic clustering of megafauna was statistically signicant (beyond what would be expected from random chance). In this context, spatial autocorrelation quanties the degree to which the abundance of megafauna in a given sampling unit is similar to the average abundances of neighbouring sampling units. Thus, the choice of the optimal neighbourhood size is key because it directly inuences the outcome and subsequent interpretation of hotspot analysis results: overly large neighbourhood sizes may smooth away local spatial patterns, whereas overly small neighbourhoods may be very sensitive to noise and other spurious artefacts in the abundance data matrix. Here, we dened our optimal neighbourhood size to comprise six nearest neighbours, after empirically observing that for most dives, the rate change in spatial autocorrelation (Moran’s I) does not change signicantly from around the sixth-order neighbourhood. (Figure 5).
Figure 5: Spatial correlation values plotted against dierent sizes of k-th order neighbourhoods. A rule of thumb for choosing the optimal neighbourhood size for hotspot analysis is to look for the inection point of the curve, which represents a balance between too few and too many neighbours. To detect megafaunal hotspots and coldspots, we rst used k-nearest neighbour algorithm41 to construct a graph that connects each sampling unit to its six nearest neighbours (based on geographic proximity). Based on this neighbourhood graph, we calculated Local Indicators of Spatial Association (LISA) statistics for each sampling unit, which identied localised regions where megafaunal abundances were signicantly higher or lower than would be expected from spatial randomness. To classify these geographic clusters (as either hotspots or coldspots), we projected the LISA statistics onto a Moran’s scatterplot (Supplementary Figure S2), which shows the relationship between the abundance of each sampling unit versus the average abundances of its neighbours (spatial lag). Depending on where a given sampling unit was located on this scatterplot, we classied it as either a hotspot (high-high abundances), coldspot (low-low abundances) or an outlier (low-high or high-low abundances). Finally, we assessed the statistical signicance of the observed spatial patterns (or LISA statistics) by conducting a randomised hypothesis test under the null hypothesis of complete spatial randomness.
Assessing megafaunal biodiversity and community composition Benthic biodiversity assessments are key towards understanding the overall ecosystem health and functionality. Here, we used standard deviation of megafaunal abundances and Shannon diversity index to measure diversity in both the Eastern and Western region. This regional comparison of diversity allowed us to simultaneously assess both spatial variability and depth-wise zonation patterns of megafauna, since the Eastern and Western regions vary by water depth and distance to shore (with distinct dierences in carbon export to the seaoor, upwelling processes, and input of terrigenous material). We identied clusters of megafaunal communities using non-parametric multivariate statistics. First, we applied double-root transformation to the abundance data matrix to stabilise the variance and moderate the inuence of dominant taxa (abundance data typically contains many low values and few high values). Next, we used the transformed abundances to generate a Bray-Curtis similarity matrix that captures the degree of biotic (dis)similarity among the sampling units. We then applied hierarchical agglomerative clustering (with group-average linking) to this similarity matrix, thereby revealing clusters of sampling units with similar biotic composition. To formally test whether the dierences among the major clusters of megafaunal communities was statistically signicant, we performed an analysis of similarity (ANOSIM). ANOSIM calculates a test statistic Rthat captures the average dierence between interand intra-community similarities, with the null hypothesis H0 dened as: There is no significant difference in biotic composition among the main megafaunal communities. In addition, we used a similarity percentages analysis (SIMPER) to identify taxa that contributed the most towards the separation among respective clusters of megafaunal communities. Finally, we visualised the inter-relationships among megafaunal communities by projecting the sampling units onto a two-dimensional ordination (feature) space using non-metric multidimensional scaling (nm-MDS). We also superimposed onto the ordination plot taxa and environmental variables that we sampled at the centroids of respective sampling units. These abiotic variables included: depth, slope, topographic position index, terrain ruggedness, salinity, temperature and longitude. This graphical representation allows for a convenient visual interpretation of the association between biotic and abiotic variables, together with their inuence
on the identied megafaunal communities. (Note that we omitted latitude since all our deployments were along an East-West transect. Also, longitude here is proportional to distance from shore but not to depth, even though the two variables are correlated to some extent). Results This section presents ndings from our unsupervised seaoor classication, along with a description of megafaunal taxa that we detected in the area. We further describe the spatial distribution patterns of these megafauna, as well as environmental drivers that inuence their distribution in both Eastern and Western regions. Visibility improvement Figure 6 shows qualitative results of our colour correction workow. The reduction in image intensity towards the edges is now accounted for, and the overall contrast is enhanced in the transformed image. The correction also removes the greenish haze that was prevalent in the raw images, resulting in good distribution of colours over the entire enhanced image. Collectively, these transformations produce well-illuminated scenes that reveal biota and substrate characteristics with clear contrast e.g. highlighting animal tracks and sediment disturbance due to biogenic activities.
Figure 6: Examples showing visibility improvement transformation from (A) original images, to (B) colour normalised images. The megafaunal taxa, animal trails and bioturbation-driven sediment disturbance are now clearly visible in the transformed images. Seaoor substrate classication Figure 7 shows results of our unsupervised classication of the seaoor. The classication is based on the extracted visual information for the entire image, which encodes both biogenic and
abiogenic properties of the photographed seabed (e.g. bioturbation, lebensspuren, burrows, seaoor morphology, sediment colour, etcetera). Each point corresponds to an individual image mapped in feature space, while colour coding is based on the twelve seaoor classes assigned to the respective images using unsupervised K-means classication. Figure 7: Projection of images (as points) in feature space, colour coded by one of 12 seaoor substrate classes. Images from respective dives cluster together because they represent the same geographic region on the seabed and thus have most similar sedimentological and benthic properties. The images also show sub-partitions within dives, which is an indication of subtle dierences in seabed substrates at small spatial scales. Overall, there is a clear variability for PC1 that links to increasing intensity of sediment disturbance by bioturbating organisms (higher PC1 values). Sediment sampling during cruise M182 showed that the seaoor is composed of soft sediment of dierent grain size and composition. Towards the East and in closer proximity to land, the amount
of ne grained (silt) terrigenous sediment increases, while towards the West sediments are strongly dominated by foraminifera shells. The clustering shows that the seabed exhibits subtle dierences at small range along a survey transect, while there were clear dierences at regional scale separating the dierent surveys from each other. By manually inspecting subsample images from each survey-cluster (Figure 8), we observed that these dierences reected the extent of sediment disturbance by biogenic activities (on feeding and moving tracks), and in the sediment (burrow holes, sediment mounds). The disturbance was mostly pronounced in the texture of images from the shallower Eastern region (surveys 131, 144, 145), characterised by burrows, pits and rosette-like structures resembling a sweeping polychaete arm. The feature space projection captures this dierence, by showing a left-to-right (West to East) gradient in terms of bioturbation intensity.
Figure 8: Image subsamples showing that variability in substrate characteristics between the Eastern and Western regions was inuenced by the degree of sediment disturbance from biogenic
activity. Megafaunal abundance and diversity Our FaunD-Fast model detected 10189 megafaunal organisms belonging to 13 taxa groups. To check for potential double counting of megafauna due to overlapping images, we compared the average distance between successive images against the average length of the along-track image axis (oriented in the direction of image acquisition). The results of this comparison are shown in Figure 9, where we only found overlap in dive 19 out of the seven dives. Note that dive 19 was also where the sampling frequency was highest (0.2 Hz). Figure 9: Relationship between average distance between successive images and the average image length in the direction of image acquisition. For a given dive, there was overlap if the image length was shorter than the distance between images.
Figure 13: Choropleth maps showing the distribution of megafaunal abundances along respective survey transects. Overall, the dives in the shallow eastern region exhibit high abundance consistently along transects whereas the abundances are low in the deeper Western region, except in topographically complex habitats like on top of seamounts. Figure 14 shows a representation of hotspots, coldspots and outliers that we plotted in prole view. Compared to the choropleth map above, only locations with statistically signicant clusters of high/low megafaunal abundances (relative to local neighbourhoods) are colour coded. Figure 14: Prole view of geographic clustering showing the distribution of statistically signicant hotspot, coldspots and outliers along respective transects. The size of the symbol is proportional to the megafaunal abundance in the corresponding sampling unit at that location. At our chosen scale of analysis (100 metres) and small neighbourhood size (of 6), the gure shows that hotspots of megafauna are predominantly found in complex topographic features e.g the top of seamount (dive 32), abyssal hills (dive 28) and on the sides of a submarine canyon (dive 144). The Eastern region was characterised by a statistically signicant hotspot of high megafaunal abundances at the start of the steep side of the submarine canyon in dive 144, with the other half of the cross section exhibiting coldspots of relatively lower abundances. Contrary to expectations, we
did not observe statistically signicant geographic clusters on the seamount in dive 145, potentially because the seabed here was undersampled due to camera malfunction (Note the shorter length of dive 145). In the Western region, statistically signicant hotspots of megafauna were observed in complex physical landscapes e.g on top of the seamount (in dive 32) as well as on top of local abyssal hills (in dive 28). Short stretches of coldspots were located in relatively at abyssal plains (at the start of dive 78 and at the end of dive 28). We did not observe signicant outliers. Inuence of biotic and abiotic drivers Figure 15 shows the ordination of all the 232 sampling units projected onto a two-dimensional nm-MDS feature space. The 64 sampling units from the Eastern region mostly group together into a small tight cluster/community on the extreme left of the ordination space, whereas the remaining 168 sampling units from the Western region are scattered throughout (although they span mostly the right half of the feature space). A few smaller sub-clusters are also visible in the Western region. In terms of biotic/community composition, the taxa that predominantly inuenced the deeper Western region include Echinodermata, Foraminifera, and Arthropoda. The remaining majority of taxa predominantly inuenced the shallower Eastern region, which may explain the high number of biogenic structures (Lebensspuren) that we observed in the Eastern region. Superimposing environmental variables onto the ordination plot shows that the horizontal axis distinguishes megafaunal communities based on temperature and depth variables. This is obvious considering that the two variables map close to the boundary separating communities in the Eastern and Western regions. On the other hand, bathymetric derivatives such as slope, ruggedness, roughness and positioning index predominantly inuenced megafaunal communities at great depths in the Western region.
Figure 15: Projection of sampling units onto nm-MDS ordination space. Colour coding is based on the dives whereas the symbols distinguish between the two regions. Also superimposed in the ordination plot are taxa and environmental drivers that potentially structure megafaunal communities in the two regions. It is clear that there is a distinction between the two main megafaunal communities in the Eastern and Western region, and also that bathymetric drivers predominantly structure communities in the deeper Western region. Discussion So far, we have demonstrated that incorporating A.I into marine science workows accelerates the characterisation of habitats and megafauna distribution in deep sea environments. Here, we provide further interpretations for the observed patterns, and contextualise the ndings relative to other studies. Our semi-supervised seaoor sediment classication workow involved clustering based on encoded visual features, followed by manual interpretation of the clusters to assign semantic
meaning. We chose this approach because modern implementations of the K-means clustering algorithm are fast, accurate and straightforward to use 42, thereby enabling automated sorting of huge volumes of unlabelled images into manageable representative groupings. In conducting benthic sediment mapping for the Australian National Marine Bioregionalisation project, Lucieer et al. 43 also point out that statistical clustering allows for more objectivity and repeatability in image-based seaoor classication when compared to manual interpretation. In another study, Diesing et al. 44 emphasise the key role of feature space projection methods such as principal components analysis (PCA) towards enabling visual interpretability of seaoor clusters. Clustering also ensures consistency in seaoor annotation since visually similar images are almost guaranteed to be assigned the same labels, as was also pointed out by Lathrop et al. in previous benthic habitat characterisation study in New York Bight 45. However, the clustering performance will depend on the method used to encode visual information from images: hand-engineered features like texture are best for representing obviously heterogeneous seabed e.g as was previously used to characterise the distribution of Mn-nodules in the Clarion Clipperton Zone 9 46. For this study, we extracted high-level features using a pre-trained convolutional neural network that have been shown in past studies e.g by Yamada et al. 47 to be capable of capturing subtle variability in seabed substrate composition. Colour-coding the feature space using clustering labels produces a graphical display that allows for a quick (qualitative) visual rst impression of sediment characteristics. This kind of display proves useful for decision making by marine scientists during expeditions e.g to determine where to sample next, or to help in the choice of an appropriate image dataset for studying a given phenomenon of interest (e.g from repositories like BIIGLE 48 or PANGAEA 49). We used this graphical display in feature space as our main interface for semi-automated annotation because (a) it was more convenient to assign semantic labels to clusters compared to labelling individual images, (b) it was easy to leverage contextual information e.g characteristics of neighbouring clusters to adapt the annotation accordingly based on the underlying structure of the data, and (c) it was straightforward to detect any anomalous patterns and/or artefacts as these clusters would be unusually isolated in the feature space. In approaching seaoor classication this way, we consider automated algorithms to be useful agents for preliminary sorting, while reserving the nal call to annotators with the domain expertise to resolve nuanced, granular and subtle variability in substrate composition that may be missed by algorithms.
Deploying our FaunD-Fast model to detect megafaunal taxa produced weak annotations that still needed to be manually inspected, rened and re-labeled, as was also done previously by Tang et al. 50 and Zhang et al 51. In our case, the model was able to correctly detect instances of megafauna in images (with an accuracy of 78.1%), even though taxa labels assigned to the detections were sometimes incorrect even for organisms that were present in both the Atlantic and the Pacic. While overtting and poor generalisation may explain the misclassication of previously unseen taxa 24, it is not obvious why some organisms (e.g Holothurians and Ophiuroids) that were present in both the Pacic and Atlantic were correctly detected yet misclassied, despite colour normalising the two datasets in the same way. A possible explanation is oered by Zurowietz et al. 52 who previously pointed out that concept drift poses a big challenge for knowledge transfer across marine environments, especially in studies involving non-endemic taxa. In this context, concept drift is the phenomenon where the statistical distribution of visual properties of marine organisms shifts across datasets either gradually or suddenly, as was also highlighted by Langenkämper et al. 53. Therefore, exactly how to develop a species detection model that generalises across oceans remains a challenging open problem that needs further investigation. In principle, such a generalizable model must be altogether agnostic to the distinct dierences in terms of ecological habitats, intraand inter-species appearance, water column properties etcetera. Recently, there have been eorts aimed at addressing these generalisation challenges by developing well-curated standardised benchmark image datasets. For example, the openly available global image database FathomNet by Katija et al. 54 provides annotations that cover a wide range of taxa categories from dierent ocean environments. The goal of these benchmark datasets is to enable training and evaluation of deep learning models that would be more robust and generalizable, since the models would have been exposed to diverse visual features of marine organisms 53. Another potential solution for poor generalisation is data augmentation, which involves the application of random geometric and photometric transformations e.g random scaling, rotations and ipping in order to articially increase the volume and variety of training examples, as has previously been demonstrated by Tan et al. 55. How well these (and other) solutions work is a promising direction for further research. Despite the aforementioned challenges, our pre-trained FaunD-Fast model is still directly useful in situations where one only cares about binary fauna/non-fauna detections e.g. to distinguish marine organisms from other background objects in a live OFOS/ROV video feed 56 57.
Absolute megafaunal taxa counts that we obtained from our fauna annotation workow required standardisation due to potential sampling bias and lack of a consistent (spatial) scale of reference. However, the choice of an optimal length (or resolution) of the sampling unit within which to standardise the annotations is not obvious but depends on the problem and ecological context, as was also previously pointed out by Enrichetti et al. 20 and Dominguez-Carri et al. 58 In our case, we chose a xed-size length of 100 metres because we were interested in capturing localised megafaunal distribution patterns in high resolution, along linear transects whose variability in substrate characteristics was very subtle. According to guidelines from a previous study on transects and quadrats in ecology by Murray et al. 59, we consider our 100-metre-long sampling units to be high resolution considering that the average length of our transects was 3200 metres. Murray et al 59 recommended the use of high resolution sampling units whenever possible (and resource permitting), since the ne resolution allows for the capturing of granular localised distribution patterns e.g megafauna adapted to microhabitats, specic depth gradients or substrate type over short distances. Choosing larger-sized sampling units (e.g with resolutions of 500 metres or greater) may average out small scale spatio-ecological patterns, and are best suited for providing generalised information regarding overall trends in community structure e.g as was previously argued by Montaña et al 60. In any case, we surveyed the seaoor at suciently high frequency (maximum 0.2 Hz), resulting in an average of 38 images within each of our 232 sampling units of each 100m length. This sample size is sucient for unbiased and robust biodiversity assessments using multivariate statistics, as has previously been demonstrated by Forcino et al. 61. Our chosen resolution was also convenient purely from a computational perspective 62, because visualising the 232 sampling units in the ordination feature space did not require too much memory or computing resources. There were major dierences in the distribution and abundance of megafauna between the Eastern and Western regions. As Ramos et al. also point out in their previous study of marine biodiversity o Mauritanian deep waters 63, the regional dierences in abundance may be explainable by variability in depth and geomorphological complexity of the seabed. Considering that the average depth dierence between the Eastern and Western regions of our working area was approximately 700 metres, the high megafaunal abundances and diversity in the shallower Eastern region might be due to food availability in the form of sinking organic matter 64. Since the Eastern region is also
relatively closer to Mauritanian shore, the region benets more from both land-based nutrient sources as well as from nutrient enrichment from upwelling currents, as has also been previously reported by Scepanki et al. 65. Moreover, our CTD proles in Supplementary Figure S1 show that the shallower waters in the Eastern region also exhibit relatively warmer temperatures (2.85°C ± 0.12) compared to the Western region (2.50°C ± 0.08). These warmer temperatures may also have contributed to the observed high megafaunal abundance in the Eastern region, since elevated temperatures have been shown to enhance metabolic rates while also supporting a wider range of functional traits, e.g as shown by Puerta et al. 66 and Sweetman et al. 67. Both the transect depth proles and the hill-shaded bathymetric grid (Figure 2) show the presence of structurally complex features like seamounts and local elevations in the Eastern and Western regions. In general, we observed high megafaunal abundances in these complex habitats compared to at terrains because the complex topographies create microhabitats e.g rocky outcrops and sediment pockets, which provide shelter and protection for megafauna while also inuencing hydrodynamic conditions to create stronger currents that promote nutrient distribution and richer food webs 68. Still, we observed higher abundances on seamounts in the Eastern region compared to those in the deeper Western region, which could be because POC ux is higher in shallow seamounts due to stronger upwelling eects, as was also previously shown by Victorero et al. 69. Regarding the inuence of environmental drivers on the observed spatial patterns, our ordination plot showed that bathymetric drivers such as slope, ruggedness, and roughness predominantly inuenced megafaunal communities in the deeper Western region. This could be because habitats in the deeper seabed areas are in general more stable, with topographies that vary slowly in kilometre scale 70. As a result, even minor variability in the bathymetric derivatives in the deeper seabed results in a more pronounced inuence on hydrodynamic eects like current patterns and nutrient distribution. This is in contrast to the already complex topographies in shallower parts that naturally disrupt hydrodynamic ows, so that the eect of minor changes in bathymetric derivatives are not as pronounced e.g as was also pointed out by Kaiser et al. 71. Collectively, the above ndings demonstrate that incorporation of A.I into conventional image-based marine science workows does contribute towards expedited characterisation of substrate types and megafaunal distribution as follows: First there are obvious speed gains since human eort is required to semantically label only the A.I generated annotations instead of the
entire dataset. Second, our workow practically demonstrates how projecting seaoor images onto feature space does allow for a quick at-a-glance visualisation (and interpretation) of natural benthic habitat groupings, including any anomalies that require further investigations. Third, outputs from our A.I workows (e.g data matrices) seamlessly integrate with existing classical spatio-ecological workows such as ordination and spatial autocorrelation analysis. This integration allows us to automate only the necessary repetitive time-consuming tasks (like annotation), while avoiding unnecessary re-invention of the wheel in downstream analysis. Fourth, despite the occasional misclassications, our model does generalise across oceans to the extent that it correctly detects and localises organisms in images. This is directly useful in applications such as rapid underwater video analysis to e.g extract relevant frames for subsequent semantic labelling. Therefore, we conclude that automated image analysis workows have the capacity to eciently extract actionable insights from terabyte-scale seaoor imagery, which is necessary to complement both ongoing and planned development of timely marine baseline information for monitoring remote benthic ecosystems at regional and global scale. Acknowledgements We acknowledge that the seaoor images used in this research were acquired during RV METEOR cruise M182, under the framework of the REEBUS project that is funded through BMBF grant 03F0815. We appreciate the operators of the XOFOS system on board the vessel, and to all the crew and scientists who made the data acquisition campaign successful. We also acknowledge Daphne Cuvelier for manually inspecting and assigning semantic taxa labels to the weak megafauna annotations. The rst author wants to specically thank the Helmholtz School for Marine Data Science (MarDATA), Grant No. HIDSS-0005, for direct nancial support. All authors declare that this research was conducted in the absence of any commercial or nancial relationships that could be construed as a potential conict of interest. This is publication 65 of the DeepSea Monitoring Group at GEOMAR Helmholtz Centre for Ocean Research Kiel.
Author contribution statement B.M conceived, implemented and programmed the computer vision and spatio-ecological analysis workows, and also drafted the manuscript. J.G was the chief scientist in cruise M182, and also participated in the conception, design, and overall coordination of the research, as well as providing geospatial interpretations and drafting the manuscript. All authors read and approved the nal manuscript. Data availability statement The datasets presented in this study can be found online in BIIGLE here https://annotate.geomar.de/projects/65. Intermediate data generated during the analysis is also provided in the supplementary materials as an excel le. To request the data used in this study, please contact Prof. Dr. Jens Greinert (
[email protected]) References 1. Danovaro, R., Snelgrove, P. V. R. & Tyler, P. Challenging the paradigms of deep-sea ecology. Trends Ecol. Evol. 29, 465–475 (2014). 2. Riehl, T., Wöl, A.-C., Augustin, N., Devey, C. W. & Brandt, A. Discovery of widely available abyssal rock patches reveals overlooked habitat type and prompts rethinking deep-sea biodiversity. Proc. Natl. Acad. Sci. 117, 15450–15459 (2020). 3. Webb, T. J., Berghe, E. V. & O’Dor, R. Biodiversity’s Big Wet Secret: The Global Distribution of Marine Biological Records Reveals Chronic Under-Exploration of the Deep Pelagic Ocean. PLOS ONE 5, e10223 (2010). 4. Montes, E. et al. Optimizing Large-Scale Biodiversity Sampling Eort: Toward an Unbalanced Survey Design. Oceanography 34, 80–91 (2021).
5. Howell, K. L. et al. A Blueprint for an Inclusive, Global Deep-Sea Ocean Decade Field Program. Front. Mar. Sci. 7, (2020). 6. Dumke, I. et al. Underwater hyperspectral imaging as an in situ taxonomic tool for deep-sea megafauna. Sci. Rep. 8, 12860 (2018). 7. Sala, E. & Knowlton, N. Global Marine Biodiversity Trends. Annu. Rev. Environ. Resour. 31, 93–122 (2006). 8. Huvenne, V. A. I. et al. ROVs and AUVs. in Submarine Geomorphology (eds. Micallef, A., Krastel, S. & Savini, A.) 93–108 (Springer International Publishing, Cham, 2018). doi:10.1007/978-3-319-57852-1_7. 9. Mbani, B., Schoening, T., Gazis, I.-Z., Koch, R. & Greinert, J. Implementation of an automated workow for image-based seaoor classication with examples from manganese-nodule covered seabed areas in the Central Pacic Ocean. Sci. Rep. 12, 15338 (2022). 10. Martin Ludvigsen, Johnsen, G., Sørensen, A. J., Lågstad, P. A. & Ødegård, Ø. Scientic Operations Combining ROV and AUV in the Trondheim Fjord. Mar. Technol. Soc. J. 48, 59–71 (2014). 11. Moniruzzaman, Md., Islam, S. M. S., Bennamoun, M. & Lavery, P. Deep Learning on Underwater Marine Object Detection: A Survey. in Advanced Concepts for Intelligent Vision Systems (eds. Blanc-Talon, J., Penne, R., Philips, W., Popescu, D. & Scheunders, P.) 150–160 (Springer International Publishing, Cham, 2017). doi:10.1007/978-3-319-70353-4_13. 12. Du, Y., Liu, Z., Li, J. & Zhao, W. X. A Survey of Vision-Language Pre-Trained Models. arXiv.org https://arxiv.org/abs/2202.10936v2 (2022). 13. Deng, J. et al. ImageNet: A large-scale hierarchical image database. in 2009 IEEE Conference on