On the use of deep learning for fish species recognition and quantification on board fishing vessels
Abstract
17 pages, 8 tables, 9 figures.-- Under a Creative Commons license
Full text
Marine Policy 139 (2022) 105015 Available online 9 March 2022 0308-597X/© 2022 The Author(s). Published by Elsevier Ltd. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). On the use of deep learning for fish species recognition and quantification on board fishing vessels Juan Carlos Ovalle a , Carlos Vilas a , Luís T. Antelo a , * a Bioprocess Engineering Group, IIM-CSIC, c/ Eduardo Cabello,6, 36208 Vigo, Spain ARTICLE INFO Keywords: Fishing catch characterization Fisheries management and compliance with regulations Remote Electronic Monitoring Species identification and length estimation Deep learning ABSTRACT The development and effective compliance of efficient fishing policies that guarantee both the sustainability of marine resources and fishing activity is one of the main challenges that policymakers nowadays face. At EU level, successful implementation of the Common Fisheries Policy (CFP) depends, at a large extent, on the capacity to quantify catches on board commercial vessels. Because of the large number of fishing vessels and the high number of trips to be monitored classic, monitoring methods, mainly based on inspections, are not effective. Therefore, the use of electronic devices to quantify fishing catches is gaining relevance. The data provided by such devices, in combination with mathematical models, may be used to assess the state of the different fishing stocks and to optimize the fishing activity. In this work, we consider different algorithms based on Deep Learning (DL) for species identification and length estimation. On the one hand, for the instance segmentation task, we have adapted the Mask R-CNN algorithm to the problem of fish species identification. On the other hand, the MobileNet-V1 convolutional neural network is used for the estimation of the length of each individual. The results show that, when overlapping among individuals is moderate to low, both the identification and length estimation models are able to satisfactorily quantify the catch. In situations where overlapping among individuals is large, results need further improvements. 1. Introduction The Common Fisheries Policy (CFP) of the European Union (EU) [8] aims to ensure that fisheries are environmentally, economically, and socially sustainable. To achieve this, it proposes: increasing the selectivity of the fishing fleets; a gradual elimination of the practice of discarding unwanted catches; the mitigation of accidental catches of protected species; and the persecution of unreported and unregulated (IUU) fishing. However, its success depends, at a large extent, on the implementation of effective, modern, and technologically advanced monitoring, control, and enforcement systems that allow to precisely determine the degree of compliance with the proposed measures. To this aim, Regulation (EC) No. 1224/2009, often called the Fisheries Control Regulation [7], was created and came into force in 2010 to provide a system of monitoring, inspection and enforcement for fishing operations in EU waters and activities of the EU fleet globally. Despite the significant advances attained, important loopholes and weaknesses have been identified in its implementation and need to be solved to ensure that the CFP aims are fully met [10,11,13,36]. For instance, a key requirement to achieve a successful implementation of the CFP is the need to monitor unwanted catches at sea. This is a particularly ambitious problem due to the large number of fishing vessels and the high number of trips to be monitored. The European Fisheries Control Agency (EFCA) emphasized that the detection of non-compliance with the discard veto continues to be difficult when it only depends on classic surveillance, based on inspections, due to the fleet-based nature of discards at sea, which can occur at any time during the fishing campaign [10,11,13]. Therefore, and to solve the referred loopholes of the CFP, on 2018, the European Commission published a proposal for the revision of the fisheries control system [9] based on the incorporation of technological solutions that implement digital transformation and ecologic transition in the fishing industry by: (i) using modern, innovative technologies to monitor fishing activities; (ii) improving enforcement and control activities; and (iii) updating of certain standards. More in detail, the Commission’s proposal introduces requirements for more complete fisheries data collection and publication, including the installation of an electronic tracking system for all EU fishing vessels; mandatory cameras * Corresponding author. E-mail address: [email protected] (L.T. Antelo). Contents lists available at ScienceDirect Marine Policy journal homepage: www.elsevier.com/locate/marpol https://doi.org/10.1016/j.marpol.2022.105015 Received 26 August 2021; Received in revised form 21 February 2022; Accepted 22 February 2022
Marine Policy 139 (2022) 105015 2 for boats over 12 m in length; fully digitized reporting of caches with electronic logbooks and landing declarations; vessel power control; and rules for recreational fishermen to declare all catches. So, the availability and quality of fisheries data should be improved. Besides, there is a need towards systematic ways for sharing data among all relevant entities, including fisheries scientists. This revision also proposes that an accurate recording and accountability of by-catches of sensitive species, such as birds and mammals, and of marine biological resources are essential for an ecosystem approach to fisheries and for a sound stock assessment, which are, in turn, the foundation of responsible and sustainable fisheries management. As a conclusion, and from a legal point of view, this document, that it is nowadays under EU trilogue discussion, proposes a revision of the enforcement rules aiming to provide the necessary regulatory framework so that the new fisheries control system is not perceived as invasive, but as fair. In this new legal framework, the digital revolution has to contribute to ensure accurate catch registration data not only for control purposes by European or national administrations but for scientific evaluation of stocks/populations and self-monitoring of fleets and fishermen associations. In addition, digitalization will improve the verification of measures on fishing capacity applicable to vessels engine power, better traceability of fisheries products and improved catch certification schemes. So, digitalization and advanced tools applied to fisheries, (such as Remote Electronic Monitoring (REM or EM) Systems, artificial intelligence (AI), machine learning tools, sensor data and high-resolution satellite imagery) have an enormous potential to enhance our ability to collect and analyze data towards the optimization of fishing operations and the improvement of the monitoring and control capabilities of policymakers and regulatory administrations. Nowadays, in addition to on board inspections at sea, there are other options to monitor fishing activity [23,37]. These include vessel monitoring systems (VMS); electronic logbooks; on board observers; in-land observers for acquired video analysis (the so-called dry observers); or self-sampling by fishermen. Increasingly though, different technology has quickly developed during the last years to provide vision-based, REM systems, at lower costs, and with more potential to cover large areas than traditional monitoring strategies. As fully described in [23], the main drawbacks detected in the EM systems that currently exist include: (i) distrust and interference of the crew in these camera-based systems; (ii) video-based off-line/ground evaluation of catches, with great effort and cost in terms of equipment and specialized ground personnel; (iii) they do not allow biological sampling to provide some data of interest (therefore carrying out scientific campaigns with observers on board is still necessary) and; (iv) impossibility of monitoring the entire fishing campaign due to the lack of video storage and/or transmission/remote accessing capacity due to the huge amount of data that composes it. As a consequence, REM systems providing reliable and precise data of the whole catch of fishing fleets are required [12]. Such data can be combined with mathematical models to assess and predict the state of different stocks; to optimize the fishing activity [32,38]; and to define proper traceability tools or methods of the fish catches at all levels of the food supply chain. These mechanisms will allow to satisfy the increasing demand of sustainable marine food products made by consumers [17,31] and will generate a great opportunity of adding value to the catches. In recent years, the scientific community is exploding the enormous potential that artificial intelligence, and more precisely deep learning (DL) offers to obtain reliable fishing catch data [5]. This is an emerging research field that should be cross-disciplinary, bringing together marine scientists, maritime (including fisheries) authorities, IT specialists, and governance experts. Deep learning has been applied in recent years to provide automatic fish identification, counting, and sizing. For the case of unconstrained underwater, various automatic computer-based fish sampling solutions have been presented [28,39,40]. However, an optimal solution for automatic fish detection and species classification does not exist. This is mainly because of the challenges present in underwater videos due to environmental variations in luminosity, fish camouflage, dynamic backgrounds, water murkiness, low resolution, shape deformations of swimming fish, and subtle variations between some fish species [22]. Several approaches have been followed for the development of DL tools to obtain fully documented catches during the fishing operations on board. For the case of long-liners, [29] proposed a method that identifies different species of tuna and swordfish from images obtained with cameras installed on the decks of the fishing vessels. [35] proposed an approach to identify and quantify harvested fish (mainly tuna and swordfish) in videos from REM systems. However, in terms of complexity, this scenario can be considered the easiest one since the specimens are caught individually, and the number of different species captured is reduced (highly specific fisheries). Moreover, different issues such as occlusions by fishermen; image blur caused by rain; or other external, uncontrollable factors may affect the fish detection capacity. For the case of gillnets, [4] analyzed the capabilities of dedicated fisheries remote electronic monitoring (REM) camera to identify and quantify captures on board 5 small-scale fisheries Peruvian vessels. The proposed approach was effective in detecting and quantifying elasmobranch target catch and pinniped by-catch, but its performance is also affected by the problems detected for long-liners. Regarding trawlers, [14] developed a computer vision system that analyzes video from CCTV systems installed on fishing trawlers for the purpose of identifying discarded fish. This approach does not estimate the length and exhibits some problems when tracking fish between frames and when fish individuals go out of view temporarily due to occlusions. The work by [15] presents a different approach based on the processing of stereo images acquired by the so-called Deep Vision imaging system, directly placed in the trawl net. A Mask R-CNN architecture is used to locate and segment each individual fish in the images. The main drawbacks of this approach are the reluctance of the fleets to install this type of device in their fishing gears; and the risk of breakage or loss of the Deep Vision during the fishing operation. Moreover, a high percentage of the images (mostly acquired during the fishing steps before and after the trawling operation) contain no fish on them, which leads to extra storage and processing efforts. In [3], the authors developed a DL-based method to estimate European hake (Merluccius merluccius) catches weight in boxes at auction places. The proper collocation of individuals on the box maximizes the capabilities of the length estimation model. However, the main drawback of the approach is that it is only trained for identifying one species. The authors in [37] developed an EM system (the iObserver) for automatic, real-time quantification of the whole catch on board fishing vessels. This device is located above the conveyor belt where the fishermen separate the catch. During the separation process, this device takes images, analyzes them, and provides an estimation of the catch. The identification model used color, texture, and shape features to classify the fish. Although it provided good results when the fish individuals were separated, its precision is significantly reduced when the individuals are minimally overlapped or belong to species that are similar in color, texture, and shape (for instance specimen of Triglidae spp.). In this work, we improve the identification capabilities of the iObserver, as an automatic, low cost (without using expensive proprietary software) EM device to provide big volumes of accurate catch data whose proper analysis will help both the trawling fleets as well as to policymakers. One the one hand, trawling fleets will significantly increase their level of compliance with the legal framework set by the CFP by performing a more sustainable activity. On the other hand, policymakers will have a decision-making tool based on real-time fishing data to define more effective and selective fishing strategies and/or regulations. In this aim, we propose new algorithms for length estimation and species identification by exploiting the potential of the above mentioned novel AI tools. The weight of each individual can be computed from its length using a simple algebraic equation, whose parameters depended on the species, [18,34]. Instead of using color, texture, and shape J.C. Ovalle et al.
Marine Policy 139 (2022) 105015 3 parameters for species identification purposes, the new models were developed in the framework of deep learning. In this regard, we have considered different detection, instance segmentation, and regression Artificial Neural Networks (ANNs). In particular, for the instance segmentation task, we have adapted the Mask R-CNN algorithm [1,19] to the problem of fish species identification. The model for length estimation uses the convolutional neural network MobileNet-V1 [20]. The ANNs were trained to identify and quantify fifteen different classes, fourteen of them correspond with objective species (species of interest to the fishing fleet operation in ICES regions 8c and 9a) whereas the remaining one contains the non-objective species. The training and validation of the models were carried out using images taken by the iObserver. These images contained both overlapped and separated individuals. Different case studies, with different degrees of overlapping fishes, were considered to test the predictive capabilities of the ANNs. 2. Materials and methods 2.1. iObserver description A complete description of the iObserver hardware has been provided in [37]. For the sake of completeness, we will summarize in this section the main features and the modifications carried out in the last year. The hardware components of the iObserver are: (i) an industrial vision camera, (ii) a lightning system, and (iii) an industrial computer equipped with open-source image acquisition software. The camera and the computer, on the one hand, and the different lights, on the other hand, are protected by metallic cases that allow IP66 certification, as shown in Fig. 1(a)-(b). The dimensions of the main box containing the camera and the computer are 40 ×23 ×26 cm 3 and its weight is around 18 kg. A Peltier cell system was installed in the main box to avoid water condensation that might damage the computer or blurry the pictures. The main hardware improvement with respect to the previous version [37] is the lighting system. In this case, it consists of 4 LED lights whose position and angle with respect to the horizontal plane can be changed. Polarizing filters have also been installed to minimize bright reflections, improving the sharpness and definition of the pictures. The iObserver was installed at an onshore location (Fig. 1(c)), on board an oceanographic vessel (Fig. 1(d)) and, on board a commercial vessel(Fig. 1(e)). As shown in these pictures, the iObserver should be, ideally, installed on the fishing park, over the conveyor belt. The iObserver automatically takes pictures during the fish separation process to record the whole catch. To avoid overlapping/repetition among pictures as well as missing parts of the haul, a system of sensors and a magnet is used. When a magnet pass close to the sensor, it sends a signal that triggers the shutter release of the camera (see [37] for details). In case this system fails, an optical flow software algorithm is used. 2.2. Image recognition algorithms The mathematical algorithms developed in the framework of machine learning use input data to learn from. In this regard, the learning algorithm, which in this case is an Artificial Neural Network (ANN), contains a number of parameters that are tuned by minimizing the error between the input data and the model results. Once such parameters are estimated, the model can be used to make predictions. For details about machine learning algorithms, the reader is referred to the literature, for instance [30,33]. Fig. 1. (a) iObserver box. (b) Internal view of the iObserver. (c)-(e) The iObserver installed at different locations: onshore; oceanographic vessel and commercial vessel. J.C. Ovalle et al.
Marine Policy 139 (2022) 105015 4 In this work, the input data of the algorithm are images taken by the iObserver, usually with annotations made by human observers specifying the species and size of the specimens present in the image. Such images are divided into three sets: •Training images. The images in this set are used to estimate, in an iterative procedure, the parameters of the ANN that minimize the desired cost function which measures the distance between the outputs of the model and the desired output. •Validation images. This set is used to check whether the model is approaching or moving away from the expected value and to adjust the algorithm hyperparameters (number of layers, type, and number of neurons, etc.). These images are also used to avoid the specialization of the model on the images of the training set (i.e., it minimizes model overfitting issues). In this way, the predictive capabilities of the model are improved. •Test images. These images are used to assess the prediction capabilities (performance) of the ANN. Therefore, it is important that the ANN has never “seen” these images before. In other words, the Fig. 2. Illustrative examples of the images used for training, validation and test of the iObserver. (a)-(b) Individual samples. (c) Several separated individuals. (d) Several overlapped individuals. (e)-(f) Non-objective species. (g) Objective and non-objective species. (h)-(i) Mosaics for data augmentation. J.C. Ovalle et al.
Marine Policy 139 (2022) 105015 5 images contained in this set must be different from those used for training and validation purposes. In the following sections, we will describe the type of images used for model training and validation as well as the different algorithms used for detection, instance segmentation, and fish length estimation. 2.3. Images used to develop and test the models As mentioned above, the artificial neural network (algorithm) must be trained and validated before being able to detect and classify the different objects in an image. The iObserver was used to obtain the pictures used for such purposes. Two iObserver installations, an onshore location (for controlled and fine: (i) hardware testing and tuning; (ii) image acquisition and; (iii) algorithm testing, training, and validation) and on board an oceanographic vessel during the campaign DESCARSEL0819 of the Spanish Institute of Oceanography (IEO) (which aims to study strategies to reduce discards and unwanted species, including selectivity and survival in trawling in Cantabrian-Northwest zone) were considered. The library of images used to train and validate the model can be classified into different groups (see Fig. 2): •Images containing one fish belonging to a species of interest to the fishing fleet operation in ICES regions 8c and 9a (objective species). Two examples of these pictures are shown in Fig. 2(a)-(b). Sample location within the image as well as its orientation were randomly selected. An effort was made to cover a large number of location and orientation combinations. The fish length was measured and recorded by experienced observer biologists from IEO for the individuals in these images. •Images containing several individuals of different objective species (Fig. 2(c)-(d)). Two groups of images were generated, one group with separated individuals and the other with overlapped individuals. The specimens were arranged in such a way that none presented more than 50% occlusion by other specimens. Again, different locations and orientations of the individuals were considered. As in the previous case, fish length was recorded for the individuals in these images. •Images containing samples of the non-objective species (Fig. 2(e)- (f)). These images were considered so that the iObserver is trained to reject the samples of non-objective species. No distinction is made among the species in this group and they are recorded with the code 000. Besides, fish length was not recorded. Two groups of images were generated, one group containing only one individual and another one containing several individuals with a distribution similar to the ones found on real hauls. •Mixture of individuals of both objective and non-objective species (Fig. 2(g)). The size of individuals is only recorded for those belonging to objective species. •Data augmented images with mosaics of overlapping individuals created using random rotations and translations of masked individuals from the previous images (Fig. 2(h)-(i)). Data augmentation techniques increase the performance of the model, making it more robust and accurate for a larger variety of situations. It multiplies the value of the manually annotated data set, thereby reducing the laborious data collection and annotation effort. It also helps with classes with few sample data and imbalance in the number of annotations. Besides, it alleviates the overfitting problem by increasing the volume of the training data set and thus improving the generalization of the model. Table 1 shows the number of individuals of each species (N. Samples) used in the development of the iObserver recognition software. It also summarizes the number of pictures that contain one individual of the considered species (N. Pictures) and the number of objects that were assigned by a human observer to each species using the LabelMe tool (N. Annotations). It must be noted that the same individual might appear and be annotated in different pictures. Row Combined refers to pictures with more than one individual. It is important to mention that, in the case of RJM and RJN, the number of individuals is very low. Unfortunately, no more specimens could be obtained during this work because these species are not highly abundant in the fishing areas of the present case study. Therefore, the results obtained for these species cannot be considered significant. 2.4. Bounding box detection algorithm Initially, a bounding box detection algorithm, that used the Tensorflow [21,24] object detection framework, was developed. Besides, a convolutional neural network optimized to be used on hardware with limited resources, Mobilenet-V1 [20], was used. This network was pre-trained with the set MS COCO [25], a large-scale object detection, segmentation, and captioning data set, commonly used for algorithm benchmarking. Among the different COCO-trained models provided in https://gith ub.com/tensorflow/models/blob/master/research/object_detection/ g3doc/tf1_detection_zoo.md, the ssd_mobilenet_v1_fpn_coco was chosen because of its good compromise between speed and precision. This implementation includes the Single Shot Multibox Detector (SSD) [27], and the Feature Pyramid Networks (FPN) for object detection [26], which enables the possibility of making the models invariant to scale. It is also faster than previous pyramid models. In these algorithms, image labeling is performed by using the smallest rectangle containing each sample in the image, as shown in Fig. 3. The size of each individual as well as its species is recorded as information in the image labeling. Note that, in addition to the individual being considered, most rectangles contain a part of the conveyor belt and a part of other individuals. This prevents the use of some very useful data augmentation transformations like rotations (Fig. 2(h)-(i)) and somehow hinders the training of the algorithms. In fact, preliminary results showed that the quality of species identification is relatively poor when individuals are overlapped. In order to overcome these limitations, we propose the use of instance segmentation algorithms, in which the area occupied by each object instance is identified. Two different models were developed: an instance segmentation algorithm for individuals detection and species identification; and a regression algorithm for length estimation of detected individuals. Table 1 Number of samples, pictures and annotations per fish species used to train, validate and test the model of the iObserver. Objective species are labeled using the FAO 3A code whereas non-objective species are labeled as 000. Scientific name 3A code N. samples N. pictures N. annotations Trisopterus luscus BIB 62 532 476 Trigla gurnardus GUG 47 265 433 Trigla lyra GUN 37 621 546 Aspitrigla cuculus GUR 44 292 382 Chelidonichthys lucerna GUU 53 1173 1134 Merluccius merluccius HKE 153 1621 1362 Trachurus trachurus HOM 269 1522 1177 Lepidorhombus boscii LDB 116 898 587 Lepidorhombus whiffiagonis MEG 107 1228 651 Scomber scombrus MAC 75 604 788 Raja clavata RJC 77 579 460 Raja montagui RJM 2 28 26 Leucoraja naevus RJN 4 50 52 Micromesistius poutassou WHB 161 1684 1220 Other (000) – – 783 616 Combined – – 1833 – J.C. Ovalle et al.
Marine Policy 139 (2022) 105015 6 2.5. Instance segmentation algorithm In the instance segmentation algorithms, the labeling task is manually performed using software tools to accurately define the contour of each individual in the images. Fig. 4 shows a picture where the contours of the different fishes, green lines, have been manually drawn. A class, in this case the fish species, is assigned to each object. The open-source LabelMe (https://github.com/wkentaro/labelme) software was used for this purpose. This software was modified to adapt it to this case study. In particular, metadata file generation codes were modified so that the sample size, as well as the ratio pixel-meter in the image, could be indicated. The modifications also allowed to assign different unconnected areas to the same individual. This is specially useful when individuals in the images are overlapped (see, Fig 2(d),(f)- (j)). The user interface and image processing were also optimized in order to speed up the labeling process as much as possible. The advantages of the instance segmentation algorithms with respect to the bounding box algorithms are that: (i) they provide more accurate identification results; (ii) they enable the possibility of using data augmentation geometric techniques, through the segmentation masks; and (iii) they provide an output that contains the contour of the object instead of a box that contains the object and the background or other objects. Note that, to enable the second advantage, contours must be defined with precision. The main disadvantages of the instance segmentation algorithms are: (i) the contour definition, including the labeling process, is a slow and tedious task since, as mentioned above, it involves manually defining the contour of each sample in the images; (ii) the species identification task is slower as compared with the bounding box algorithms; and (iii) algorithms are more complex and require more computational resources. For the instance segmentation task, an implementation of the Mask R-CNN algorithm [1,19] in Keras and Tensorflow, was selected. TensorFlow is an open-sourced end-to-end platform, which contains a library for multiple machine learning tasks. Keras is a high-level deep learning framework library that runs on top of TensorFlow, i.e., Keras functions are a wrapper to the TensorFlow framework. In this regard, the user can define an algorithm with the Keras interface, which is easier to use, then go into TensorFlow when a specific functionality, not included in Keras, is required. Both Keras and Tensorflow, provide high-level Application Programming Interfaces (APIs) used for building and training models. As in the bounding box algorithm, the Feature Pyramid Network (FPN) for object detection was used. Besides, the 101 layer version of ResNet convolutional neural network [6] for classification (ResNet101), optimized to deal with the vanishing gradient problem of deep neural networks and pre-trained with the data set COCO, was selected. This implementation of the Mask R-CNN only admits images with a maximum of 1024 px in height. Therefore, images obtained by the iObserver have been resized to comply with this requirement. A total number of 5.975 images, which included 9.910 annotations, were manually labeled using the LabelMe software. This process, although tedious, can be performed by anyone, after basic training. Of course, the assignment of each instance to a given species should be supervised with the appropriate knowledge. The training set for the final species identification and length estimation model consisted of 34.346 images obtained in two different ways. In this regard: •A set of 4.782 images was randomly selected for training purposes from the 5.975 available labeled images (see the beginning of Section 2.3). These images were subject to random transformations (rotation Fig. 3. Example of an image used in the bounding box algorithms. The original image is shown at the left. In the image at the right, each sample is delimited by the smallest rectangle containing it. J.C. Ovalle et al.
Marine Policy 139 (2022) 105015 7 and shifting) so that we could obtain three “different” images from each of the 4.782. This resulted in 14.346 images. •Data augmentation techniques were used to generate 20.000 synthetic images (see Figure (Fig. 2(h)-(i))). A maximum number per image of 50 individuals, randomly located and oriented, was considered. The maximum overlapping area allowed among individuals as well as the maximum area of the individual outside the frame was 15%. Three different background colors were considered. Backgrounds were generated from images of typical conveyor belts. The validation set for the final model consisted of: •612 images randomly selected from the 1.193 labeled images that were not used in the training set. •500 synthetic images obtained following the same steps as in the case of the training set. The training of the algorithm was performed both at the installations of the IMEDEA–CSIC (using the GPU NVIDIA Quadro GV100 32 GB graphic card) and CESGA (using GPU NVIDIA Tesla V100 16 GB). The computations required to perform the training of the final model lasted for seven days. 2.6. Length estimation algorithm One objective of the iObserver is to quantify the total mass of each objective species that has been caught during the fishing activity. To that purpose, the weight of each individual (W, in Kg) is estimated from its length (L, in cm) through the following equation [34]: W=aLb(1) where parameters a and b differ among species [37]. This equation is, however, an approximation. Accuracy may be improved if, for the same species, different values of a and b are considered depending on fisheries, season, and region [18,34]. The length estimation algorithm consisted of a convolutional neural network (MobileNet-V1), which was trained from scratch. The results provided by the instance segmentation algorithm, described in the previous section, are used as input arguments in the length estimation algorithm. The original MobileNet-V1 only admits inputs in the form of images. Therefore, it has been modified in this work so that other useful information can be used as input. In particular, the following additional information is used in our modified MobileNet-V1 algorithm: •The name of the fish species of the individual being analyzed. A oneshot format was used to encode the fish species. This format consists of a vector with 15 elements, one per species. A number 1 is assigned to the element corresponding to the species being analyzed, the remaining elements are set to 0. •Coordinates of the corresponding bounding box for the individual. •The relative importance of the different elements/objects of the image. Three types of objects are considered to that purpose: individual being analyzed; other individuals; and background. The idea is to give more importance to the individual being analyzed as compared with the other objects in the picture. In this regard, the Fig. 4. Contour definition and sample labeling in the instance segmentation algorithms. J.C. Ovalle et al.
Marine Policy 139 (2022) 105015 8 information in the input images (values of each pixel) was scaled according to the following criteria: – If the pixel corresponds to the individual being analyzed, it is scaled within the interval [128,255]. – If the pixel corresponds to an individual different from the one being analyzed, it is scaled within the interval [64,127]. – If the pixel corresponds to the background, it is scaled within the interval [0,63]. The original MobileNet-V1 is a convolutional network for image classification, therefore, its output consists of a confidence value for each of the classes under consideration. In our implementation, we have modified the output layer to obtain a value representing the length of the individual to be measured. Images used to train and validate this algorithm are the ones described in Section 2.3 and Section 2.5. The distance from the camera to the conveyor belt may vary depending on the installation of the iObserver. Therefore, a calibration was required so that the ratio pixel/m is the same for all images. To that purpose, the dimensions of the region of interest 1 (ROI) were measured. These measurements as well as the number of pixels in the images were used to resize the images so that 1.024 pixels corresponded to 630 mm, which is the largest width of the ROI for the vessels considered in this project. Black color (pixel value 0) was assigned to the empty pixels resulting from the scaling procedure. The parameters of the length algorithm are computed, as in the case of the instance segmentation algorithm, by minimizing the cost function. In this case, such function computes the mean squared error between the predictions provided by the algorithm and the human observations. Given that the size and complexity of the MobileNet-V1 network for length estimation is considerably smaller than the size of the instance segmentation algorithm, the set of available training images was enough to train it and only 5 epochs were necessary before overfitting problems began to arise. Therefore, convergence to the optimal values was around one order of magnitude faster with the fish length model. It should be mentioned that both the Mask R-CNN, for species identification, and the MobiliNet-V1 network, for length estimation, are equipped with techniques to handle overfitting problems, such as normalization, regularization or dropout. 2.7. Evaluation of the iObserver performance The iObserver has been tested in different situations, namely: (i) using images of the same type as those used for training and validation from both an onshore installation and an oceanographic vessel; (ii) using images of actual hauls on an oceanographic vessel; and (iii) using images of actual hauls on a commercial vessel. For the first case, images of the same type as those used for training, all individuals in the images were labeled by a human observer. Therefore, a detailed analysis is performed using the confusion matrix for several classes [2]. Besides, the length of every individual belonging to the objective species was recorded by a human observer. Therefore, the mean absolute error (MAE) and the mean absolute percentage error (MAPE), per each of the objective species, are used to evaluate the performance of the length estimation algorithm. The MAE and the MAPE are computed, respectively, as: MAEx=1 Nx∑ Nx i=1 ∣Lm,x,i−Le,x,i∣,(2) MAPEx=100 Nx∑ Nx i=1 ∣Lm,x,i−Le,x,i Lm,x,i ∣,(3) where N x is the number of individuals of species x that were correctly identified. L m,x,i is the length of individual i (species x) measured by a human observer. L e,x,i is the length of individual i (species x) estimated by the model. In the tests carried out using images of actual hauls on board oceanographic and commercial vessels, the images were not labeled due to the large number of pictures/samples obtained in each haul. Therefore, the confusion matrix cannot be computed in these tests. Besides, the length of the individuals was not measured by a human observer. In these cases, the only information registered by the human observer is the weight of the group of individuals belonging to a given species. Note that a small number of individuals of a given species might be captured in a haul. The MAE or the MAPE, using the weight of each species instead of the length of each individual, would give the same importance to species with large and small number of individuals in the catch. Therefore, these indicators might be misleading in this case. To avoid this problem, we use the Weighted Average Percentage Error (WAPE) to quantify the performance of the iObserver: WAPE =∑Ns x=1∣Wm,x−We,x∣ ∑Ns x=1Wm,x ,(4) where N s is the number of species, both detected by the human observer and the iObserver. W m,x and W e,x are, respectively, the weight of the group of individuals of species x measured by the human observer and estimated by the model. 3. Results and discussion This section presents the analysis of the performance of the instance segmentation and length estimation models using images of the test set, i.e., those that were not used for training and validation. We consider different case studies in which the test sets were obtained in different situations, with increasing degree of difficulty: •Case study 1. The test set consisted of those images that belong to the original set of 5.975 pictures obtained both at the onshore installation and on board the oceanographic vessel, and that were not used for training and validation of the models. •Case study 2. The images in this test set were taken, as in the previous case, on the onshore installation and on board the oceanographic vessel. However, these images are relatively different from the original 5.975 labeled images. In this regard, individuals of different objective and non-objective species were randomly distributed within the ROI with a moderate degree of overlapping among them. •Case study 3. The test set consisted of images taken during a haul performed on an oceanographic vessel. The flux of fishes entering the conveyor belt was reduced as compared with the usual working conditions to reduce the overlapping among individuals. •Case study 4. As in the previous case, the test set consisted of images taken during a haul performed on an oceanographic vessel. However, in this case, the flux of fishes entering the conveyor belt was not modified, resulting in pictures with highly overlapped individuals. •Case study 5. The images of the test set were obtained during a haul performed on a commercial vessel with high overlapping among individuals. In the following, we will discuss the results obtained in each of these cases. 3.1. Case study 1 In this case, we use the 581 images that were not used in the training and validation sets. Note that the images in the different sets were randomly selected. Therefore, it is expected that the images in this test 1 The region of interest corresponds with the camera framing. J.C. Ovalle et al.
Marine Policy 139 (2022) 105015 9 set are “similar” to those in the training and validation sets. 3.1.1. Instance segmentation model results The confusion matrix in Table 2 summarizes the species identification results obtained with the iObserver for the test set, i.e. the set of images that have not been used during the training of the model. FAO 3A code is used to identify the species. The following issues are highlighted from Table 2: •Rows show the prediction results, i.e. the precision. The numbers between parentheses, beside the species 3A code, indicate the number of times that the model has identified the different objects in the image as the species considered in this row. For example, BIB (54) means that the model has identified 54 objects as BIB. The values indicated in the matrix elements are normalized with respect to the rows. For instance, from the 54 objects that the model has identified as BIB, 53 of them (98.1%) coincided with the label (BIB), provided by the human observer. The remaining object that the model has wrongly identified as BIB (1.9%) was identified by the human observer as WHB. For each of the objective species, the precision of the model was above 98%, except for GUN, GUR, and LDB with 95.0%, 96.6%, and 93.2%, respectively. •Columns represent the results of the analysis of the images performed by a human observer (ground truth). In this case, the number between parentheses, beside the species 3A code, indicates the number of times that the human observer has identified the different objects in the image as this species. The next number represents the recall, i. e. the percentage of times that the model has correctly identified the individuals of this species. For instance, the images belonging to the test set contained 55 individuals of BIB and the model has correctly identified 53 of them, i.e., 96.4% For each of the objective species, the recall was above 90%, except for RJN. In this case, 1 of the 5 (20%) individuals of RJN in the pictures was identified as BG by the model. Note, however, that the number of individuals of RJN in the pictures (5) is not large enough to conclude that the recall of the model for this species is 80%. •Abbreviation BG in the last row and column stands for background. In this regard: – Column BG indicates the percentage of times that the model has wrongly assigned a portion of background to a given species. In this case, 1.7% of the objects identified by the model as LDB and 0.9% identified as WHB were actually portions of the conveyor belt. – Row BG accounts for labeled individuals that were undetected by the model and, therefore, classified as background. The number between parentheses (27) indicates the total number of undetected objects. The number in each cell denotes the percentage of individuals of each species that has not been detected. For instance, from the 27 undetected objects, a 3.7% (1 individual) corresponded to BIB. The total precision and recall obtained from images in the test set were 98% and 95%, respectively. This shows the capacity of the iObserver to identify and quantify the catch under the conditions considered in this experiment. Fig. 5 illustrates the instance segmentation results for three different images, with overlapped individuals, of the test set. Different conveyor belts were considered. As shown in the figure, the model is able to obtain the contour of each individual with a satisfactory degree of accuracy. The smallest rectangle containing each individual is also shown. The name of the species, as well as the length of the individual obtained with the model, are indicated in the top-left corner of such rectangle. As mentioned earlier in Section 2.5, the maximum overlapping area allowed among individuals for data augmented synthetic training Table 2 Confusion matrix for the images in the test set. J.C. Ovalle et al.
Marine Policy 139 (2022) 105015 16 developed the deep learning-based approach, setting up the computing architecture; JCO and LTA performed the in-land image acquisition work and supervised the on board installations and work; JCO and CVF carried out the analysis of results. CVF led the manuscript writing. All authors actively collaborated in the writing of the manuscript. Moreover, they critically contributed to the drafts and gave final approval for publication. Acknowledegments This work was developed with the collaboration of Fundaci´ on Biodiversidad of the Ministry for the Ecological Transition and the Demographic Challenge of the Spanish Government, through Programa Pleamar, co-funded by the EMFF of the European Union. The authors also want to acknowledge the invaluable collaboration of the SICAPTOR project partners (Foundation CESGA, Spanish Institute of Oceanography - IEO and, OPROMAR) on the development and on board testing of the iObserver. References [1] Waleed Abdulla.Mask R-CNN for object detection and instance segmentation on Keras and Tensorflow. https://github.com/matterport/Mask_RCNN, 2017. [2] Khalid K. Al-jabery, Tayo Obafemi-Ajayi, Gayla R. Olbricht, Donald C. Wunsch, Data analysis and machine learning tools in MATLAB and Python. Computational Learning Approaches to Data Analytics in Biomedical Applications, Springer, 2020, pp. 231–290, https://doi.org/10.1016/B978-0-12-814482-4.00009-7. [3] Amaya Alvarez-Ellacuria, Miquel Palmer, Ignacio A. Catalan, Jose-Luis Lisani, Image-based, unsupervised estimation of fish size from commercial landings using deep learning, ICES J. Mar. Sci. 77 (4) (2020) 1330–1339, https://doi.org/ 10.1093/icesjms/fsz216. [4] David C. Bartholomew, Jeffrey C. Mangel, Joanna Alfaro-Shigueto, Sergio Pingo, Astrid Jimenez, Brendan J. Godley, Remote electronic monitoring as a potential Fig. 9. Identification results for two images of the test set corresponding to the fifth case study. Figures at the left correspond with the original images taken by the iObserver. Detection and identification results are shown in the figures at the right. Table 8 Comparison in weight between the measured and the predicted catch, case 5. N. individuals Weight (g) Species Predictions Ground truth Predictions Error (%) 000 672 – – – BIB 26 17,000 2255 87 GUX 78 13,000 10,133 22 HKE 670 147,000 75,251 49 HOM 7 – 163 – LEZ 374 24,000 33,768 41 RJX 425 – 475,521 – WHB 147 – 6092 – Total 2399 201,000 603,183 – J.C. Ovalle et al.
Marine Policy 139 (2022) 105015 17 alternative to on-board observers in small-scale fisheries, Biol. Conserv. 219 (2018) 35–45, https://doi.org/10.1016/j.biocon.2018.01.003. [5] Cigdem Beyan, Howard I. Browman, Setting the stage for the machine intelligence era in marine science, ICES J. Mar. Sci. 77 (4) (2020) 1267–1273, https://doi.org/ 10.1093/icesjms/fsaa084 (ISSN 1054-3139). [6] A. Dhar, J.J. Kinnunen, P. T¨ orm¨ a, Population imbalance in the extended fermihubbard model, Phys. Rev. B 94 (7) (2016), https://doi.org/10.1103/ physrevb.94.075116 (ISSN 2469-9969). [7] EC. Council Regulation (EC) No 1224/2009 of 20 november 2009 establishing a community control system for ensuring compliance with the rules of the common fisheries policy, amending regulations (EC) No 847/96, (EC) No 2371/2002, (EC) No 811/2004, (EC) No 768/2005, (EC) No 2115/2005, (EC) No 2166/2005, (EC) No 388/2006, (EC) No 509/2007, (EC) No 676/2007, (EC) No 1098/2007, (EC) No 1300/2008, (EC) No 1342/2008 and repealing regulations (EEC) No 2847/93, (EC) No 1627/94 and (EC) No 1966/2006. OJ L 343 Technical report, European Commission, 2009. [8] EC. Regulation (EU) No 1380/2013 of the European parliament and the council of 11 December 2013 on the common fisheries policy, amending council regulations (EC) No 1954/2003 and (EC) No 1224/2009 and repealing Council Regulations (EC) No 2371/2002 and (EC) No 639/2004 and Council Decision 2004/585/EC Technical report, European Commission, 2013. [9] EC. Proposal for a Regulation of the European parliament and of the council amending council regulation (EC) No 1224/2009, and amending council regulations (EC) No 768/2005, (EC) No 1967/2006, (EC) No 1005/2008, and Regulation (EU) No 2016/1139 of the European Parliament and of the council as regards fisheries control. COM/2018/368 final. Technical report, European Commission, 2018. [10] EFCA. Evaluation of compliance with the landing obligation in north sea demersal species 2016 - 2017 - Executive summary. Technical report, European Fisheries Control Agency (EFCA), 2019a.Available at 〈https://www.efca.europa.eu/en/cont ent/compliance-evaluation〉. [11] EFCA. Evaluation of compliance with the landing obligation, north western waters 2016 - 2017 Executive summary. Technical report, European Fisheries Control Agency (EFCA), 2019b.Available at 〈https://www.efca.europa.eu/en/content/co mpliance-evaluation〉. [12] EFCA. Technical guidelines and specifications for the implementation of Remote Electronic Monitoring (REM) in EU fisheries. Vigo (Spain).Technical report, European Fisheries Control Agency (EFCA), 2019c.Available at 〈https://www.efca. europa.eu/en/content/technical-guidelines-and-specifications-implementationremote-electronic-monitoring-rem-eun〉. [13] EFCA. Evaluation of compliance with the landing obligation, Baltic Sea 2017 - 2018 - Executive summary. Technical report, European Fisheries Control Agency (EFCA), 2021.Available at 〈https://www.efca.europa.eu/en/content/complian ce-evaluation〉. [14] Geoff French, Michal Mackiewicz, Mark Fisher, Helen Holah, Rachel Kilburn, Neil Campbell, Coby Needle, Deep neural networks for analysis of fisheries surveillance video and automated monitoring of fish discards, ICES J. Mar. Sci. 77 (4) (2020) 1340–1353, https://doi.org/10.1093/icesjms/fsz149. [15] Rafael Garcia, Ricard Prados, Josep Quintana, Alexander Tempelaar, Nuno Gracias, Shale Rosen, Havard Vagstol, Kristoffer Lovall, Automatic segmentation of fish using deep learning with application to fish size measurement, ICES J. Mar. Sci. 77 (4) (2020) 1354–1366, https://doi.org/10.1093/icesjms/fsz186. [16] Richard C. Gerum, Sebastian Richter, Alexander Winterl, Christoph Mark, Ben Fabry, C´ eline Le Bohec, Daniel P. Zitterbart, Cameratransform: a python package for perspective corrections and image mapping, SoftwareX 10 (2019), 100333, https://doi.org/10.1016/j.softx.2019.100333. [17] Marcella Giacomarra, Maria Crescimanno, Demetris Vrontis, Lluís Miret Pastor, Antonino Galati, The ability of fish ecolabels to promote a change in the sustainability awareness, Mar. Policy 123 (2021), 104292, https://doi.org/ 10.1016/j.marpol.2020.104292. [18] J.M.S. Gonçalves, L. Bentes, P.G. Lino, J. Ribeiro, A.V.M. Can´ ario, K. Erzini, Weight-length relationships for selected fish species of the small-scale demersal fisheries of the south and south-west coast of Portugal, Fish. Res. 30 (1997) 253–256. [19] Kaiming He, Georgia Gkioxari, Piotr Dollar, Ross Girshick. Mask R-CNN. In: Proceedings of the 16th IEEE International Conference on Computer Vision (ICCV), 2980–2988, 2017.10.1109/ICCV.2017.322. [20] Andrew G. Howard, Menglong Zhu, Bo Chen, D. Kalenichenko, W. Wang, Tobias Weyand, M. Andreetto, Hartwig Adam, Mobilenets: Efficient convolutional neural networks for mobile vision applications, ArXiv (2017) (abs/1704. 04861), https://arxiv.org/abs/1704.04861. [21] Kai Hwang, Tensor flow, keras, deepmind, and graph analytics. Cloud Computing for Machine Learning and Cognitive Applications, The MIT Press,, 2017, pp. 463–520. ISBN 978-0-262036-41-2. [22] Ahsan Jalal, Ahmad Salman, Ajmal Mian, Mark Shortis, Faisal Shafait, Fish detection and species classification in underwater environments using deep learning with temporal information, Ecol. Inform. 57 (2020), 101088, https://doi. org/10.1016/j.ecoinf.2020.101088 (ISSN 1574-9541). [23] K.M. James, N. Campbell, J.R. Vidarsson, C. Vilas, K. Plet-Hansen, J. P´ erezBouzada, E. van Helmond, O. Abad, E. Gonz´ alez, R.I. P´ erez-Martín, M. Quinz´ an, L. T. Antelo, J. Valeiras, C. Ulrich, Tools and technologies for the monitoring, control and surveillance of unwanted catches, in: S. Uhlmann, C. Ulrich, S. Kennelly (Eds.), The European Landing Obligation: Reducing Discards in Complex, Multi-Species and Multi-Jurisdictional Fisheries, Springer, 2019, pp. 363–382, https://doi.org/ 10.1007/978-3-030-03308-8. [24] Ang Li, Yi-xiang Li, Xue-hui Li. Tensor flow and keras-based convolutional neural network in CAT image recognition. In: Proceedings of the 2nd international conference on computational modeling, simulation and applied mathematics (CMSAM), DEStech transactions on computer science and engineering, pp. 529–533, 2017. (ISBN 978-1-60595-499-8). [25] Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ ar, C. Lawrence Zitnick, Microsoft COCO: common objects in context, in: David Fleet, Tomas Pajdla, Bernt Schiele, Tinne Tuytelaars (Eds.), Computer Vision - European Conference on Computer Vision ECCV, Springer International Publishing, 2014, pp. 740–755, https://doi.org/10.1007/ 978-3-319-10602-1_48. [26] Tsung-Yi Lin, Piotr Doll´ ar, Ross Girshick, Kaiming He, Bharath Hariharan, Serge Belongie. Feature pyramid networks for object detection. In: Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 936–944, 2017.10.1109/CVPR.2017.106. [27] Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, Alexander C. Berg, SSD: Single shot multibox detector, in: Bastian Leibe, Jiri Matas, Nicu Sebe, Max Welling (Eds.), Computer Vision - ECCV 2016, Springer International Publishing, 2016, pp. 21–37, https://doi.org/ 10.1007/978-3-319-46448-0_2. [28] Vanesa Lopez-Vazquez, Jose Manuel Lopez-Guede, Simone Marini, Emanuela Fanelli, Espen Johnsen, Jacopo Aguzzi, Video image enhancement and machine learning pipeline for underwater animal detection and classification at cabled observatories, Sensors 20 (3) (2020), https://doi.org/10.3390/s20030726. [29] Yi-Chin Lu, Chen Tung, Yan-Fu Kuo, Identifying the species of harvested tuna and billfish using deep convolutional neural networks, ICES J. Mar. Sci. 77 (4) (2020) 1318–1329, https://doi.org/10.1093/icesjms/fsz089. [30] M. Mohri, A. Rostamizadeh, A. Talwalkar, Foundations of Machine Learning, The MIT Press, 2018 (ISBN 9780262039406). [31] Jos´ e Oliveira, Jos´ eEvaristo Lima, Dimitrida Silva, Volodymyr Kuprych, Pedro Miguel Faria, Cl´ audio Teixeira, Estrela Ferreira Cruz, Ant´ onio Miguel Rosado da Cruz, Traceability system for quality monitoring in the fishery and aquaculture value chain, J. Agric. Food Res. 5 (2021), 100169, https://doi.org/10.1016/j. jafr.2021.100169. [32] Maria Grazia Pennino, Ana Helena Bevilacqua, M. Angeles Torres, Jose M. Bellido, Jordi Sole, Jeroen Steenbeek, Marta Coll, Discard ban: a simulation-based approach combining hierarchical Bayesian and food web spatial models, Mar. Policy 116 (2020), 103703, https://doi.org/10.1016/j.marpol.2019.103703. [33] S. Shalev-Shwartz, S. Ben-David, Understanding Machine Learning: From Theory to Algorithms, Cambridge University Press, 2014, https://doi.org/10.1017/ CBO9781107298019. [34] M.A. Torres, F. Ramos, I. Sobrino, Length-weight relationships of 76 fish species from the Gulf of Cadiz (SW Spain), Fish. Res. 127–128 (2012) 171–175, https:// doi.org/10.1016/j.fishres.2012.02.001. [35] Chi-Hsuan Tseng, Yan-Fu Kuo, Detecting and counting harvested fish and identifying fish types in electronic monitoring system videos using deep convolutional neural networks, ICES J. Mar. Sci. 77 (4) (2020) 1367–1378, https:// doi.org/10.1093/icesjms/fsaa076. [36] S. Uhlmann, C. Ulrich, S. Kennelly, The European Landing Obligation: Reducing Discards in Complex, Multi-Species and Multi-Jurisdictional Fisheries, Springer, 2019, https://doi.org/10.1007/978-3-030-03308-8. [37] C. Vilas, L.T. Antelo, F. Martin-Rodriguez, X. Morales, R. Perez-Martin, A. A. Alonso, J. Valeiras, M. Abad, E. Quinzan, M. Barral-Martinez, Use of computer vision onboard fishing vessels to quantify catches: the iObserver, Mar. Policy 116 (103714) (2020) 436–449, https://doi.org/10.1016/j.marpol.2019.103714. [38] Raul Vilela, Maria Grazia Pennino, Gonzalo Rodriguez-Rodriguez, Hugo M. Ballesteros, Jose Maria Bellido, The use of a spatial model of economic efficiency to predict the most likely outcomes under different fishing strategy scenarios, Mar. Policy 129 (2021), 104499, https://doi.org/10.1016/j. marpol.2021.104499. [39] Sebastien Villon, Corina Iovan, Morgan Mangeas, Thomas Claverie, David Mouillot, Sebastien Villeger, Laurent Vigliola, Automatic underwater fish species classification with limited data using few-shot learning, Ecol. Inform. 63 (101320) (2021), https://doi.org/10.1016/j.ecoinf.2021.101320. [40] Sebastien Villon, David Mouillot, Marc Chaumont, Emily S. Darling, Gerard Subsol, Thomas Claverie, Sebastien Villeger, A deep learning method for accurate and fast identification of coral reef fishes in underwater images, Ecol. Inform. 48 (2018) 238–244, 110.1016/j.ecoinf.2018.09.007. J.C. Ovalle et al.