International Journal of Data Science and Analytics (2025) 20:3965–3979 https://doi.org/10.1007/s41060-024-00704-9 REGULAR PAPER Automated cetacean detection in UAV imagery using AI models: a case study on Delphinid species João Canelas1,3 ·Luana Clementino2·André Cid3·Joana Castro3,4 ·Inês Machado2,5 ·Susana Vieira1 Received: 30 June 2024 / Accepted: 13 December 2024 / Published online: 10 January 2025 © The Author(s) 2024 Abstract The identification and quantification of marine mammals is crucial for understanding their abundance, ecology and supporting their conservation efforts. Traditional methods for detecting cetaceans, however, are often labor-intensive and limited in their accuracy. To overcome these challenges, this work explores the use of convolutional neural networks (CNNs) as a tool for automating the detection of cetaceans through aerial images from unmanned aerial vehicles (UAVs). Additionally, the study proposes the use of Long-Short-Term-Memory (LSTM)-based models for video detection using a CNN-LSTM architecture. Models were trained on a selected dataset of dolphin examples acquired from 138 online videos with the aim of testing methods that hold potential for practical field monitoring. The approach was effectively validated on field data, suggesting that the method shows potential for further applications for operational settings. The results show that imagebased detection methods are effective in the detection of dolphins from aerial UAV images, with the best-performing model, based on a ConvNext architecture, achieving high accuracy and f1-score values of 83.9% and 82.0%, respectively, within field observations conducted. However, video-based methods showed more difficulties in the detection task, as LSTM-based models struggled with generalization beyond their training environments, achieving a top accuracy of 68%. By reducing the labor required for cetacean detection, thus improving monitoring efficiency, this research provides a scalable approach that can support ongoing conservation efforts by enabling more robust data collection on cetacean populations. Keywords Unmanned aerial vehicles ·Convolutional neural networks ·Long-short-term-memory ·Machine learning · Marine mammals detection ·Photo identification Luana Clementino, André Cid, Joana Castro, Inês Machado and Susana Vieira have authors contributed equally to this work. BJoão Canelas [email protected] Luana Clementino
[email protected] André Cid [email protected]g Joana Castro [email protected]g Inês Machado
[email protected] Susana Vieira [email protected] 1IDMEC, Instituto Superior Técnico, Av. Rovisco Pais, 1049-001 Lisbon, Portugal 2WavEC Offshore Renewables, Edifício Diogo Cão, Doca de Alcântara Norte, 1350-352 Lisbon, Portugal 1 Introduction Cetaceans play a key role in maintaining ecosystem stability, acting as sentinel or indicator species that reflect the overall state of the ocean’s health [1,2]. Monitoring and safeguarding the diversity and abundance of cetaceans is imperative to support conservation efforts (e.g., through conventions and agreements) and achieve Good Environmental Status (GES) in European waters [3]. Achieving Good Environmental Sta3AIMM - Associação para a Investigação do Meio Marinho, Rua Maestro Fred. Freitas N15-1, 1500-399 Lisbon, Portugal 4MARE - Marine and Environmental Sciences Centre/ARNET - Aquatic Research Network, Laboratório Marítimo da Guia, Faculdade de Ciências da Universidade de Lisboa, Av. Nossa Senhora do Cabo, 939, 2750-374 Cascais, Portugal 5MARE - Marine and Environmental Sciences Centre/ARNET - Aquatic Research Network, Faculdade de Ciências da Universidade de Lisboa, Campo Grande, 1749-016 Lisbon, Portugal 123
3966 International Journal of Data Science and Analytics (2025) 20:3965–3979 tus (GES) in European waters is a key objective under the Marine Strategy Framework Directive (MSFD), which was adopted by the European Union to evaluate and maintain the health of the marine environment. GES is defined by eleven descriptors that assess various aspects of marine ecosystems, enabling a comprehensive evaluation of marine conditions and the pressures from human activities. This approach aligns with similar international conventions, such as the United Nations Sustainable Development Goal 14, which targets the conservation and sustainable use of oceans, seas, and marine resources, as well as the OSPAR Convention, which focuses on the protection of the North-East Atlantic marine environment. These frameworks collectively contribute to a more resilient and sustainably-managed global marine ecosystem. Monitoring and assessing the achievement of GES is particularly challenging given that cetaceans are highly mobile species, distributed over large areas, and moving across various marine habitats subject to diverse anthropogenic pressures. These pressures include incidental by-catch in fishing gear, bioaccumulation of pathogens and toxins, harmful algal blooms, collisions with ships, underwater noise and climate change [4–7]. More recently, the advancement of offshore renewable energy further intensified these challenges. Such projects often target large marine areas,commonlyoverlapping with cetacean habitats, thereby escalating the pressure between conservation needs and energy exploitation [8]. The vulnerability of these species and exploitation of their habitats underscores the importance of their conservation, thus, it is imperative to improve our current understanding of cetacean distribution patterns. However, such studies are excessively costly, posing a significant barrier to advancing conservation efforts. Traditional methods to study and monitor marine mammal populations involve visual surveys from a defined platform (e.g., aerial, ship-based, or land-based), acoustic surveys [10,11], observation of Very High Resolution (VHR) satellite images [12, 13], and observation methods that allow a more thorough understanding (e.g., capture-recapture) [14]. Furthermore, emerging methodologies, such as remote sensing through photo detection and identification, present a promising tool for complementing such methods while reducing associated costs and risks [15–18]. Unmanned Aerial Vehicles (UAVs) are equipped with imaging sensors that can collect extremely high-resolution data, thus becoming an increasingly used tool for researchers to observe marine wildlife and study cetaceans. UAVs are a non-invasive method [19] that allows the detectability of animals in subsurface waters, thus increasing the time available for detection [20]. UAVs have an increasing number of applications, such as monitoring abundance and distribution [21], photo identification [22], behavioral studies [23], among others [20]. Nonetheless, this new technology still presents limitations and challenges, particularly associated with data management. The high volume of data generated requires efficient processing solutions, as manual inspection is often impractical and prone to human error [24,25]. As machine learning and computer vision advance rapidly, automated computer vision models present a promising solution for automating the inspection process. Since the scientific developments by Krizhevsky et al. [26] proving the efficacy of deep learning algorithms in image recognition, Convolutional Neural Networks (CNNs) have become the model of choice for image detection and identification, achieving results on par with human performance in detection and identification tasks [27].Thesemodelshavebeensuccessfullyappliedtoindividual identification of whales [28–30] and dolphins, [31] with methods that can be adapted to other cetacean species [29]. Therehavealsobeenapplications forwhalecountingthrough satelliteVHRimages[32],wherecombiningtheusualcounting procedure with an initial detection of whale presence has improved model accuracy and computational efficiency [32]. However, these are limited to species of larger size due to the spatial resolution of images and are difficult to develop due to the lack of open VHR image datasets. Models developed for image detection, however, are limited to the events occurring in one single frame, potentially missing important information and context from the video sequences obtained with UAVs. That said, algorithms capable of handling video frames, such as Recurrent Neural Networks (RNNs), leverage the temporal continuity and contextual information provided by video sequences while reducing information missed, thus improving detection capability. Still, the development of such models represents a higher degree of complexity, and studies exploring their efficacy on marine mammal detection are relatively limited [33–35]. The main objective of this study is to develop machine learning models capable of automating the detection of cetaceans, specifically delphinids, using UAV data. The presentworkexplorestheimplementationofwell-documented CNN architectures directed to image detection, while also proposing the use of a Deep Fake detection algorithm based on Guera and Delp’s “Deepfake video detection using recurrent neural networks” [36], applied to the detection of marine mammalsinvideosequences.Thisapproachbuildsuponcurrent methodologies, but also explores new avenues through the use of a Long-Short-Term-Memory (LSTM) network, a specific RNN model, seeking to harness the additional information provided through video analysis. This study explores the synergies between deep learning and marine science, focusing on the potential to enhance environmental monitoring and impact assessment strategies in offshore environments, particularly for the conservation of dolphin populations. These findings provide a method123
International Journal of Data Science and Analytics (2025) 20:3965–3979 3967 ological basis for improving data quality, which can support future efforts aimed at advancing sustainable management and conservation efforts in marine ecosystems. 2 Data acquisition In-situ collection of a sufficient volume of remote sensing datasuitableforthedevelopmentofanefficientidentification model is challenging due to the high cost of equipment and logistical constraints associated with ocean surveys. Furthermore,datasetsonspecieswithbroaddistributionranges,such as cetaceans, are scarce, and publicly available datasets tailored for aerial detection are largely non-existent. As a result, the data used to build the models were obtained by collecting scraped video files sourced from various online sources (e.g., YouTube, Pexels, Dailymotion, etc). Given the challenging nature of developing such datasets, data gathered for this study were limited to species of the family Delphinidae, as they are among the most accessible cetaceans in publicly available footage, and with the intent of gathering as much data as possible, no pre-selection criteria such as location or time period were applied during collection. 2.1 Training dataset The videos obtained were recorded in diverse locations and under varying conditions, leading to significant variability in water characteristics such as hue, brightness, and foam, as well as differences among delphinid species. This diversity enhances the model’s ability to generalize across different environmental settings. Furthermore, to achieve representative samples where no dolphins are present, the videos also include objects or subjects such as boats, boards, and swimmers which served as potential confounding elements for the model. Theresultingdataconsistedof138aerialvideosofvarying durations and settings. Some videos exclusively containing cetacean footage, others solely featuring water scenes, and some combining both elements. These 138 videos were then processed to create two distinct collections of data: one tailored for image classification and the other for video classification. For image classification, individual frames were extractedfromthevideos,providingstaticsamples.Forvideo classification, the original video segments were retained to capture dynamic features. While data for both methodologies were derived from the same set of videos, these were processed to suit the specific requirements of each classification type. 2.1.1 Image data Image data were generated by deconstructing the original 138rawvideosintoimagesbyextractingframesataspecified rate of one frame every three seconds using the open software FFmpeg. This rate can be adjusted based on user needs and the source of extraction, as some videos may include more or less irrelevant data. In this study, images were categorized into two distinct classes, based on the presence or absence of cetaceans: “Cetacean” and “No Cetacean”. Images were initially filtered to exclude the frames that were poor representatives of their class, such as cases where subjects were obstructed or not in the frame. Additional manual selection was also conducted on frames that were good representatives. The resulting set of data consisted of 2451 images, divided into its respective classes. The “No Cetacean” class included images where no cetaceans were present, as well as images with other surface or subsurface artifacts that could lead the model to incorrectly label them as containing a cetacean. Including these artifacts within the “No Cetacean” class helps to correct for potential false positives by exposing the model to non-cetacean images that may resemble cetaceans. The “Cetacean” class included images where at least one cetacean was present. The classification process resulted in 776 images (approximately 31.1%) representing the “No Cetacean” class, and 1720 images (about 68.9%) representing the “Cetacean” class. The notable imbalance in the number of images per class is due to the limited variation in water surface patterns over time. Frames extracted within a few seconds of each other are often nearly identical, providing minimal additional value. On the other hand, a dataset heavily composed of cetacean images could bias model predictions, increasing the rate of false positives. To mitigate this, the number of images in the “No Cetacean” class was increased by artificially generating new sea images from existing ones. This was achieved by introducing random variations in brightness, hue, and saturation to all newly generated samples. Additionally, further transformations were applied with varying probabilities: sharpness enhancement (25%), random mirroring (25%), blurring (25%), random rotations (15%), and random cropping (30%). The described set of transformations was applied a total of 944 times on randomly selected samples from the “No Cetacean” class, generating an additional 944 images. This augmentation was performed to equalize the number of samples with that of the “Cetacean” class. The resulting balanced data were composed of a total of 3420 images, with an equal distribution of 1720 (50%) images per class. 123
3968 International Journal of Data Science and Analytics (2025) 20:3965–3979 2.1.2 Video data Video data were generated by deconstructing the same 138 raw videos into several smaller videos (clips) of eight seconds each and subsequently extracting a total of 64 frames from each of these smaller videos. The initial fragmentation process of the original videos was performed using the software Adobe Premiere Pro 2020, version 14.0 (Adobe Inc., San Jose, California). Firstly, intervals of eight seconds were manually selected to accurately represent each class. Simple transformations, such as mirroring, cropping, varying brightness, and hue, were applied to some of the samples to introduce variation. Each segment was then exported to create new video samples. This video length was selected based on careful analysis of initial data acquired, balancing the goal of capturing comprehensive information on dolphin behavior and movement within a concise timeframe. This interval proved effective for segmenting original videos with frequent transitions and various added content such as overlays, logos, or artifacts that could otherwise cause unwanted model responses. A longer interval would have significantly reduced the number of usable samples, while a shorter window risked losing contextual details, as many segments showed minimal movementoverbriefdurations.Theeight-secondlength,therefore, provided an optimal compromise, enabling ample sample quantity while retaining sufficient information for model training. After this segmentation, each clip is processed using Python’s OpenCV library to extract frames at a rate of eight frames per second, resulting in a batch of images containing a total of 64 frames per clip. The choice of frame extraction rate allowed for capturing as much information on dolphin behavior variations over time, while minimizing the number of images. The resulting data post-processing operations consisted of 1216 videos, of which 622 belong to the “Cetacean” class(approximately 51.2%),whiletheremaining594videos belong to the “No Cetacean” class (approximately 48.8%). This equates to 1216 batches of 64 images each, totaling 77824 images spanning both classes. 2.2 Test dataset To monitor and understand model performance over the course of training, models are tested on data not involved in their learning process. This practice provides insights into expected performance and generalization by evaluating samples the model’s parameters were not directly adjusted to, providing a general understanding of model progress and anticipated behavior within similar data samples. In this study, the test data was derived from the original dataset outlined in Sect.2.1, from which a small portion was retracted. This division creates two distinct subsets: training data,comprising80%of theoriginal dataset,used toteach the model to recognize class patterns, and test data, making up the remaining 20%, to verify the state of models. Empirical studies suggest optimal results when reserving 20–30% of data for testing while using 70–80% for training [37]. This separation was done randomly from all available samples while keeping a proportional number of samples from each class, resulting in 688 (20%) test and 2752 (80%) training samples for image classification, and 244 (20%) test and 972 (80%) training samples for video classification. 2.3 Validation dataset The validation data for this study were provided by Associação para a Investigação do Meio Marinho (AIMM), which supported the research by supplying UAV data from previous expeditions. Data were acquired on the coastal region in south Portugal within the Faro district. Specifically, the study area is located approximately 12 km offshore from the coastline of Albufeira, extending into the Atlantic Ocean. This region is a significant habitat for various cetacean species, especially delphinids such as common dolphins (Delphinus delphis) and bottlenose dolphins (Tursiops truncatus)[38– 40]. A total of seven campaigns conducted between 2022 and 2023wereanalyzed.One campaign wasincludedinthetraining data to better adapt to local environmental conditions and UAV settings, while the remaining six campaigns were used for evaluation. These surveys were conducted in the morning, between 10:30 and 12:00, under favorable sea conditions defined by a sea state of ⩽3 according to the Beaufort scale, swells <1.5 m, good visibility (>5km), and no precipitation. Figure1offers a comprehensive view of the region under study, providing information on various expeditions, including dates, times, and the precise locations of dolphin sightings. The UAV-based remote sensing data used in this study were collected using a Mavic 2 Pro multi-rotor UAV (DJI, Shenzhen, China). The UAV captured videos at a resolution of 3840 ×2160 pixels using a 1-inch CMOS RGB imaging sensor with a maximum resolution of 20-megapixel, coupled with a 3-axis gimbal and a 28mm equivalent, f/2.8-f/11 lens, providing a field of view of approximately 77◦. The drone missions were conducted at different flight altitudes, depending on several factors. These factors included whether there were any dolphin sightings at the time and the size of the group of dolphins, with higher flights preferably usedforgreaterseacoveragewhen nosightingswere present, and lower flights for a more detailed view when a group was located. Figure2presents a box plot of the flight altitudes recorded by the UAV during the different expeditions. Of the six flights, three were conducted at a maximum altitude of 123
International Journal of Data Science and Analytics (2025) 20:3965–3979 3969 Fig. 1 Overview of the study area for acquiring data Fig. 2 Box plot: flight altitude distribution for each expedition nearly 80 m, while the remaining three were flown below 50 m. In general, the UAV was observed to operate at an altitude of around 20 m, except for the second campaign, where flights predominantly occurred at higher altitudes. The UAV-based imagery collected resulted in 35min and 40s of footage. Similar to previously acquired data, this footage was processed to create two datasets from the same source, this time for validation. The first, tailored for the validation of image-based models, was obtained by extracting and labeling frames from the original raw video data at a rate of one frame every five seconds, effectively processing the entire footage. Each sample was labeled as belonging to either the “Cetacean” or “No Cetacean” class based on inference from the original video imagery captured, allowing to discern the presence of dolphins on samples that would otherwise be challenging to identify correctly. The resulting data comprises a total of 428 image samples, with 247 (approximately 57.7%) classified as “Cetacean” and 181 (approximately 42.3%) as “No Cetacean”. These samples capturea varietyofdolphinpositions,cameraangles, anddistances, providing sufficient variation to support a robust and comprehensive assessment. The second dataset, designed for the validation of video-based models, was obtained by manuallydividing theoriginalvideo dataintosmallereight-second clips, extracting them and subsequently converting them into 64 images. The resulting video data consists of 232 videos, 120 of which were classified as belonging to the “Cetacean” class (approximately 51.7%), and 112 classified as belonging to the “No Cetacean” class (approximately 48.3%). Table 1provides a summary of the sample distribution across the training, testing, and evaluation datasets for both image and video data. Each dataset is divided into “Dolphin”and“Ocean”samples,correspondingtothe“Cetacean” and “No Cetacean” classes, with the training and testing sets holding 80% and 20% of the original data, respectively. The evaluation dataset includes an additional set of samples that covers 100% of its allocated data, ensuring comprehensive assessment of the models. This division maintains a balanced classrepresentationwithin eachsubset,withanear-equaldistributionbetween“Dolphin”and“Ocean”samplesacrossthe datasets, facilitating robust training and performance evaluation. While the evaluation dataset was enriched with cetacean images to ensure sufficient data for testing the model, it is acknowledged that in real-world applications, ocean-only imagesarelikelytobefarmoreprevalentthancetaceansight123
3970 International Journal of Data Science and Analytics (2025) 20:3965–3979 ings. Consequently, this enrichment may slightly underestimate the rate of false positives under operational conditions. However, by maintaining a balanced dataset for evaluation, the accuracy metric becomes more representative of the model’s true performance. 3 Implementation The models and pipelines employed in this study were developed within the research platform Google Colab Pro, taking advantage of its cloud computing capabilities. The primary programming language used was Python version 3.9 with PyTorch’s library as the foundation of this project’s machine learning framework, which allowed for an easy implementation of state-of-the-art deep learning techniques. 3.1 Image-based models In order to analyze distinct image identification models, the following CNN architectures were individually employed, as outlinedinTable2.Theselectedmodelshaveatrackrecordin the field and known performance in image classification [41– 45]. It is worth noting that while certain models may display superior performance on average, real-world outcomes can varysignificantlybasedonthespecificnatureof theproblems being addressed. Alongside the model names, specifications such as the number of parameters and the top accuracies achieved when these architectures were trained on ImageNet are also provided. Initially, these models were set up according to their predefinedarchitecturesandinitializedwith randomparameters, making them essentially empty frameworks incapable of making meaningful predictions. However, through transfer learning, parameters from models with identical architectures that have been trained on extensive datasets, such as ImageNet, can be transferred to these models. ImageNet, for example, comprises over a million samples and covers a wide range of classes, including animals like gray whales, dugongs, orcas, and sea lions. While it does not encompass thespecific“dolphin”class,thefeaturesthatdistinguishthese related classes can be invaluable for the identification of dolphins. The models presented are structured in two sections, the first being the feature extractor, mainly consisting of the input layer and a series of hidden layers. It constitutes the majority of the network and is designed to identify specific features within the input data through its convolutional layers. These layers have been previously trained on the ImageNet dataset, therefore, they already possess the ability to effectively recognize a wide range of characteristics from a thousand different classes. Consequently, their parameFig. 3 Combined CNN model pipeline ters are “frozen” during further training to ensure that the extracted features remain consistent. The second section of the model comprises the classifier, typically a fully connected linear layer with a softmax activation function at the end of the network. This classifier interprets the features obtained from the hidden layers and assigns a class label to each data sample. The classifier, originally designed to classify 1000 classes, was adapted to identify only two classes. This was achieved by replacing the final linear layer, which initially had 1000 output nodes, with a new linear layer containing only two output nodes, corresponding to the two target classes. In addition to the individual architectures, a combined model was developed to leverage multiple feature extractors concurrently. In this approach, the feature extraction layers from VGG16, ConvNext, and a straightforward set of convolutional layers were merged into a unified feature extractor. The simple convolutional model implemented along VGG16 andConvNextconsistsofthreeconvolutionallayersandthree max-pooling layers, complemented by certain Rectified Linear Unit (ReLU) layers in between as activation functions. The intent behind this design was to tailor the model to identify dolphin-specific features from the training dataset, instead of those that represent other subjects as in the case of transfer learning. Ultimately, a shared classifier with two layers and 1024 nodes in its middle layer was employed to receive and effectively process the features extracted from all architectures, resulting in a collective prediction. Notably, this implementation was carried out in two instances: one that omitted the Convolutional Layers, CombinedModel (1), while the other usedallthreearchitectures asexplainedCombinedModel (2). Figure3provides a general overview of this combined model implementation and how data were shaped through it. 3.2 Video-based models A video-based identification approach incorporating modern deepfake detection techniques was also adopted to leverage 123
International Journal of Data Science and Analytics (2025) 20:3965–3979 3971 Table 1 Dataset train-test split Image dataset Video dataset Train samples Test samples Eval samples Train samples Test samples Eval samples 2752 (80%) 688 (20%) 428 (100%) 972 (80%) 244 (20%) 232 (100%) Dolphin Ocean Dolphin Ocean Dolphin Ocean Dolphin Ocean Dolphin Ocean Dolphin Ocean 1376 1367 194 194 247 181 497 475 125 119 120 112 (50%) (50%) (50%) (50%) (57.7%) (42.3%) (51.1%) (48.9%) (51.2%) (48.8%) (51.7%) (48.3%) Table 2 Common CNN architectures and respective performances on ImageNet dataset Model Top-1 Acc (%) Top-5 Acc (%) Parameters (M) ResNet50 80.858 95.434 25.6 InceptionV3 77.294 93.450 27.2 VGG16_bn 73.360 91.516 138.4 ConvNext_Base 84.062 96.870 88.6 EfficientNet_V0 77.692 93.532 5.3 the unique temporal continuity feature of videos. Unlike images, videos are composed of a sequence containing numerous frames, where adjacent frames display a substantial correlation and temporal continuity. The method implemented, based on the work by Guera and Delp [36], involves using a CNN without a classifier to extract features from individual video frames and feed the resulting sequence offeaturesinto anLSTMto analyzepatternsintheirtemporal evolution. Two CNN architectures were used for this purpose: InceptionV4 and ConvNext. InceptionV3 architecture was replicated from the original work, while the ConvNext model architecture was selected based on results from the imagebasedclassificationmethodologycounterpart.Inlinewiththe previous approach, models were established based on their respective architectures, initialized with random parameters, and subsequently refined by transferring parameters from pre-trained models on ImageNet. Since the CNNs within this CNN-LSTM pipeline are used exclusively for feature extraction, their parameters were “frozen” to maintain the consistency of the extracted features. Simultaneously, the classifiers were removed, enabling direct passage of the featuresidentifiedfromthehiddenlayerstotheLSTM.Different CNN architectures have specific input size requirements: InceptionV3 and ConvNext have input sizes of 299 ×299 and 224 ×224 pixels, respectively, and extract 2048 and 1024 features, respectively. To accommodate the distinct feature structures obtained from each CNN architecture, two distinct LSTM architectures were developed. Each was designed to handle a specific input size for the transferred features, aligned with its corresponding CNN. Both models were created with two recurrent layers of 256 nodes each. This means that for each model, two LSTMs were stacked together to form a stacked LSTM, with the second taking in the outputs from the first to compute a new output at each time step. This setup enabled the LSTMs to iteratively produce 256 values, representing their hidden states, for every frame in the video sequence. To conclude the CNN-LSTM pipeline, a classifier was introduced to process the output produced by the LSTM cell and make predictions. The classifier implemented features a linear layer with 16384 nodes on its input side. At each time step, the LSTM processes a frame from the input video, thus generating 256 values, which correspond to the 256 nodes in its hidden layers, representing the hidden state at that specific time step. To maximize the amount of information used within the classifier, all the hidden states produced by the LSTM cell were aggregated. This aggregation results in a total of 16,384 nodes on the classifier’s input side since all input videos are pre-processedtoconsist of64frames.Furthermore,itsoutput layer consists of two nodes representing the two available classes and utilizes a softmax activation to convert the raw output values into probabilities. Figure 4provides a general overview of the pipeline created, its inner workings, and how the data were shaped through this system. 3.3 Pre-processing data Effective pre-processing is essential for preparing the training and testing datasets. Key steps include organizing data into manageable batches and applying transformations compatible with pre-trained models, which optimize learning and enhance model performance. To maximize computational efficiency and improve learning precision, all data samples within the training and testing datasets were grouped into batches of 64 samples each. This 123
3972 International Journal of Data Science and Analytics (2025) 20:3965–3979 Fig. 4 CNN-LSTM pipeline structure aggregation allows the simultaneous processing of multiple samples, providing a more stable gradient during backpropagation by combining losses from diverse samples within each batch. A batch size of 64, in particular, is a commonly employedchoicethatoftenworkswellforvariousdeeplearning tasks. Further transformations were applied to leverage the patterns learned from pre-trained models. Given that weights from these were trained on data with a specific distribution, it is essential to normalize the input data accordingly. This normalization involves subtracting the mean and dividing by the standard deviation values of the dataset used to pretrain the models. For ImageNet-trained models, these values are (0.485, 0.456, 0.406) for the mean and (0.229, 0.224, 0.225) for the standard deviation across the three color channels. Following normalization, the input data are resized and cropped to match the dimensions required by each CNN architecture, with most models accepting (224 ×224) input, except InceptionV3, which requires (299 ×299). 3.4 Model training The process of adapting a neural network to fit a specific problem involves iteratively assessing its performance on the training dataset, and adjusting its parameters each time to achieve predictions as close as possible to the desired values. To accomplish this, a Cross-Entropy Loss function, commonly utilized in multiclass problems such as this one was defined. This function serves as a guide to determine how the model’s parameters should be updated. Subsequently, the loss function is utilized to compute a loss value, produced for each batch, which in turn is used to optimize the parameters of the model. This optimization is conducted through an Adam optimizer with an initial learning rate of 0.001, due to its adaptive learning rate and ease of use with fewer hyperparameters. Additionally, a dropout layer with a dropout probability of 20% was added to the classifier at the end of each model during training, immediately before the final linear layer. This allows to randomly delete activations from the nodes carrying features before entering the classifier with a probability of 20% for each feature. This step proved to help the learning process of all nodes in the classifier and reduce data overfitting significantly by allowing nodes of undervalued importance to suffer larger adjustments. Finally, models underwent training by iteratively processing the training dataset, one batch at a time, over several iterations, continuously assessing predictions made and adjusting their parameters accordingly. During this process, the dataset is completely processed multiple times, and models are stored for future use with their most up-to-date parameters and key performance metrics after each iteration. 4 Results and discussion The predictive performance of the trained models was initially evaluated using the test dataset, offering an early indication of their effectiveness before validation on the evaluation dataset. This assessment includes metrics such as accuracy (Acc), precision (Prec), recall (Rec), f1-score (F1), and loss, providing a preliminary baseline of each model’s generalization capacity. Table 3summarizes the best-performing models, selected based on their f1-score, reflecting the balance between precision and recall. Figure 5shows the training curves for the top two models from each classification approach, highlighting the trends in lossandaccuracyoverepochs.Thesevisualizationshelpclarify model stability and learning dynamics, setting the stage for a more detailed validation using the evaluation dataset in the subsequent section. 4.1 Model validation To validate the effectiveness of the models studied and confirm the quality of their applicability, field observations were simulatedusingdatacollectedduringfieldworkconductedby AIMM.Subsequently,theperformanceofthepre-established models in training was assessed within the evaluation dataset detailedinSect.2.3.Thequantitativeresultsfromthis assessment are presented in Table 4. The results show a clear fall in the overall performance of all models. This is to be expected since both training and testing data share the same origin source, and therefore, bear far more similarities. From the presented values it is possible to infer that the image-based identification achieved better performance than its video-based counterpart, establishing it as the superior methodology for this task within applied models. Models based on the ConvNext architecture experienced a smaller drop in performance. In particular, the ConvNext 123
International Journal of Data Science and Analytics (2025) 20:3965–3979 3973 Table 3 Performance on test dataset Model Acc Rec Prec F1 Loss EfficientNet 0.834 0.875 0.809 0.841 0.006 ConvNext 0.969 0.965 0.974 0.969 0.001 ResNet50 0.898 0.916 0.885 0.900 0.004 InceptionV3 0.923 0.930 0.917 0.924 0.003 VGG16_bn 0.924 0.945 0.908 0.926 0.005 ConvolutionslLayers 0.850 0.863 0.841 0.852 0.018 CombinedModel (1) 0.956 0.965 0.949 0.957 0.004 CombinedModel (2) 0.955 0.945 0.964 0.954 0.006 CNN-LSTM (Incep.V3) 0.898 0.950 0.856 0.900 0.453 CNN-LSTM (ConvNext) 0.939 0.924 0.948 0.936 0.628 Fig. 5 Training curves from the most relevant models implemented model stands out as the model of choice, achieving the highest accuracy and f1-score with values of 83.9% and 82.0%, respectively. This model is expected to produce overall good predictions, having achieved a better balance between recall andprecision.Eventhougha decrease inrecallwasobserved, the value of 86.7% still provides the model with a reduced number of false negatives, while the precision value of 77.7% leads to a reduced number of false positives compared to the remaining image-based models, meaning that predictions where the model finds the presence of dolphins are more trustworthy. The notable performance of this model can be traced back to its training and testing curves, represented in Fig.5, which, when compared to the curves of other models, display a higher degree of similarity, to the extent that they overlap. This suggests a great generalization ability by showing the model’scapacitytoobtainhighaccuracyvalueswithoutoverfitting training data. RemainingConvNext-basedarchitectureshavealsodemonstratedgoodperformance.Notably, theperformanceofCombinedModel (2) improved compared to CombinedModel (1), achieving the highest recall value of 90.6%, which 123