scieee AI-readable full text Open interactive document viewer

Self-supervised learning techniques for monitoring industrial spaces

Magalhães, Viviana Figueira

Abstract

Este documento é uma Dissertação de Mestrado com o título ”Self-Supervised Learning Techniques for Monitoring Industrial Spaces”e foi realizada e ambiente empresarial na empresa Neadvance - Machine Vision S.A. em conjunto com a Universidade do Minho. Esta dissertação surge de um grande projeto que consiste no desenvolvimento de uma plataforma de monitorização de operações específicas num espaço industrial, denominada SMARTICS (Plataforma tecnoló gica para monitorização inteligente de espaços industriais abertos). Este projeto continha uma componente de investigação para explorar um paradigma de aprendizagem diferente e os seus métodos - self-supervised learning, que foi o foco e principal contributo deste trabalho. O supervised learning atingiu um limite, pois exige anotações caras e dispendiosas. Em problemas reais, como em espaços industriais nem sempre é possível adquirir um grande número de imagens. O self-supervised learning ajuda nesses problemas, ex traindo informações dos próprios dados e alcançando bom desempenho em conjuntos de dados de grande escala. Este trabalho fornece uma revisão geral da literatura sobre a estrutura de self-supervised learning e alguns métodos. Também aplica um método para resolver uma tarefa de classificação para se assemelhar a um problema em um espaço industrial.

Full text

Viviana Figueira Magalhães Self-Supervised Learning Techniques for Monitoring Industrial Spaces janeiro de 2023 UMinho | 2023 Viviana Magalhães Self-Supervised Learning Techniques for Monitoring Industrial Spaces Universidade do Minho Escola de Ciências Viviana Figueira Magalhães Self-Supervised Learning Techniques for Monitoring Industrial Spaces Dissertação de Mestrado em Matemática e Computação Trabalho efetuado sob a orientação dos Professora Doutora Maria Fernanda Pires Costa Professor Doutor Manuel João Oliveira Ferreira Universidade do Minho Escola de Ciências janeiro de 2023 Direitos de Autor e Condições de Utilização de Trabalho por Terceiros Este é um trabalho académico que pode ser utilizado por terceiros desde que respeitadas as regras e boas práticas internacionalmente aceites, no que concerne aos direitos de autor e direitos conexos. Assim, o presente trabalho pode ser utilizado nos termos previstos na licença abaixo indicada. Caso o utilizador necessite de permissão para poder fazer um uso do trabalho em condições não previstas no licenciamento indicado, deverá contactar o autor, através do RepositóriUM da Universidade do Minho. Atribuição-NãoComercial-CompartilhaIgual CC BY-NC-SA https://creativecommons.org/licenses/by-nc-sa/4.0/ i Acknowledgements First of all, I would like to thank my supervisors, Professor Doctor Fernanda Costa and Professor Doctor Luís Ferrás for their availability, kindness, and support during the realization of this dissertation. I also want to express my gratitude to Neadvance and my supervisor Professor Doctor Manuel João Ferreira for the great experience and for suggesting the theme of this thesis. Furthermore, I would like to thank my colleagues at Neadvance for the way they integrated me into the team, for all the kindness, and availability to help, and for constantly asking if it was already done, in the final tough stretch. I especially want to thank Tiago Pinto for his patience with me, for always finding time to help me even on very busy days, for checking up on me constantly, and for all the knowledge and motivation given. I would also like to thank my desk buddy Vitor Figueiredo for all the help, for answering my 1-minute and 1-hour questions, for looking at detail, and for all the kindness, support, and good advice. I was happy to be part of your team! To all my colleagues who have accompanied me on my academic journey. A special thanks to Vitor Hugo, who accompanied me from day one in this experience at Neadvance. Thank you for all the support and availability, for helping me without me having to ask, for the laughter and, above all, for the positivity that I so much needed. I want to thank my aunts: Maria, Isabel and Catarina for all the love, support and kind words. To my godfather Alberto, thank you for watching over me up in heaven. A special thank you to my favorite children Valentina, Eva and Tomás for their genuine love and laughter and for reminding me to appreciate the little things. Thank you to my friends, especially: Adriana, for remembering and being a sun on dark days; Isabel, for being a good friend and making me feel stronger; Helena, for the laughter and reminding me to enjoy the journey; Patrícia, who has been by my side since the 5th grade. Thank you for making all our steps less scary; Ponto, for being a good friend and always finding time; Nacional, for believing in me when I wouldn’t believe in myself. I want to thank my boyfriend Cristóvão for all the support, his daily patience, putting up with all my dramas and nonsense, and for always listening and having something peaceful and nice to say. To my brothers Philipe and Nando, who make life better, for all the love, laughter and companionship. For cheering me up on weaker days and helping me become the person I am today. A thank you will never be enough. Last, but most importantly, a big thank you to my mother Martinha, to whom I dedicate all my work. Thank you for always believing in me, for applauding all my victories no matter how small they were, thank you for all the effort that allowed me to be where I am today, thank you for passing on so much joy, positivity and light, and thank you for always being there. For you mother! ii Statement of Integrity I hereby declare having conducted this academic work with integrity. I confirm that I have not used plagiarism or any form of undue use of information or falsification of results along the process leading to its elaboration. I further declare that I have fully acknowledged the Code of Ethical Conduct of the University of Minho. iii Resumo Este documento é uma Dissertação de Mestrado com o título ” Self-Supervised Learning Techniques for Monitoring Industrial Spaces ”e foi realizada e ambiente empresarial na empresa Neadvance - Machine Vision S.A. em conjunto com a Universidade do Minho. Esta dissertação surge de um grande projeto que consiste no desenvolvimento de uma plataforma de monitorização de operações específicas num espaço industrial, denominada SMARTICS ( Plataforma tecnológica para monitorização inteligente de espaços industriais abertos ). Este projeto continha uma componente de investigação para explorar um paradigma de aprendizagem diferente e os seus métodos - self-supervised learning , que foi o foco e principal contributo deste trabalho. O supervised learning atingiu um limite, pois exige anotações caras e dispendiosas. Em problemas reais, como em espaços industriais nem sempre é possível adquirir um grande número de imagens. O self-supervised learning ajuda nesses problemas, extraindo informações dos próprios dados e alcançando bom desempenho em conjuntos de dados de grande escala. Este trabalho fornece uma revisão geral da literatura sobre a estrutura de self-supervised learning e alguns métodos. Também aplica um método para resolver uma tarefa de classificação para se assemelhar a um problema em um espaço industrial. Palavras-chave: Visão por computador, Deep Learning , Self-Supervised Learning , Pretext tasks , Contrastive Learning , Espaços Industriais iv Abstract This document is a Master’s Thesis with the title ” Self-Supervised Learning Techniques for Monitoring Industrial Spaces ” and was carried out in a business environment at Neadvance - Machine Vision S.A. together with the University of Minho. This dissertation arises from a major project that consists of developing a platform to monitor specific operations in an industrial space, named SMARTICS ( Plataforma tecnológica para monitorização inteligente de espaços industriais abertos ). This project contained a research component to explore a different learning paradigm and its methods - self-supervised learning, which was the focus and main contribution of this work. Supervised learning has reached a bottleneck as they require expensive and time-consuming annotations. In real problems, such as in industrial spaces it is not always possible to require a large number of images. Self-supervised learning helps these issues by extracting information from the data itself and has achieved good performance in large-scale datasets. This work provides a comprehensive literature review of the selfsupervised learning framework and some methods. It also applies a method to solve a classification task to resemble a problem in an industrial space and evaluate its performance. Keywords: Computer Vision, Deep Learning, Self-Supervised Learning, Pretext tasks, Contrastive Learning, Industrial Spaces v Contents 1 Introduction 1 1.1 Objectives of the Dissertation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 1.2 Dissertation’sStructure................................... 2 2 Basic concepts 4 2.1 MachineLearning...................................... 4 2.2 MachineLearningPipeline ................................. 4 2.2.1 ProblemDefinition................................. 5 2.2.2 Data Understanding and Processing . . . . . . . . . . . . . . . . . . . . . . . . 9 2.2.3 Modelling ..................................... 9 2.2.4 Evaluation..................................... 11 2.2.5 IterativeProcess.................................. 14 2.3 DeepLearning ....................................... 14 2.4 ArtificialNeuralNetworks .................................. 14 2.4.1 TheSimplePerceptron............................... 16 2.4.2 The Multilayer Perceptron . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 2.4.3 Convolutional Neural Networks . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 2.4.3.1 Convolutional Layer . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 2.4.3.2 Non-linearity .............................. 22 2.4.3.3 PoolingLayer.............................. 25 2.4.3.4 Fully Connected Layer . . . . . . . . . . . . . . . . . . . . . . . . . . 26 2.4.3.5 Softmax function . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27 2.4.4 Training Neural Networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27 2.4.4.1 Gradient-Based Learning . . . . . . . . . . . . . . . . . . . . . . . . 27 2.4.4.2 CNN Loss Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . 30 2.4.4.3 Weight Initialization . . . . . . . . . . . . . . . . . . . . . . . . . . . 31 2.4.4.4 Regularizations ............................. 31 2.4.5 Convolutional Neural Architectures . . . . . . . . . . . . . . . . . . . . . . . . . 35 2.4.5.1 AlexNet................................. 35 2.4.5.2 VGGnet................................. 36 vi MCC Matthews Correlation Coefficient. mIU mean Intersection over Union. ML Machine Learning. MLP Multilayer Perceptron. MSE Mean Squared Error. Noisy ReLU Noisy Rectifier Linear Unit. OCR Optical Character Recognition. PReLU Parametric Linear Unit. ReLU Rectifier Linear Unit. ResNet Residual Network. RL Reinforcement Learning. RMSE Root Mean Squared Error. RMSProp Root Mean Square Propagation. RReLU Randomized Leaky Rectifier Linear Unit. Semi-SL Semi-Supervised Learning. SGD Stochastic Gradient Descent. SL Supervised Learning. SSL Self-Supervised Learning. SwAV Swapping Assignments between multiple Views. Tanh Hyperbolic Tangent. TL Transfer Learning. xiii TN True Negatives. TP True Positives. UL Unsupervised Learning. W&B Weights & Biases. xiv 1 Introduction This dissertation is a research component of a major project that consists of developing a platform to monitor specific operations in an industrial space, named SMARTICS ( Plataforma tecnológica para monitorização inteligente de espaços industriais abertos ). It was developed in a business environment at Neadvance - Machine Vision S.A. company. This component was mainly to research and explore the new paradigm of self-supervised learning and use it to solve a real problem in an industrial space context. The rise of algorithms capable of detecting, recognizing, and tracking objects in uncontrolled environments, such as indoor and outdoor surveillance systems, traffic control of people and vehicles, intelligent parking, etc., opens a temporal window for the application of these technologies in the logistics areas in industrial environments, allowing a more efficient and ”intelligent” management of the factory space. All of these surveillance and monitoring activities are currently performed manually or with CCTV systems that monitor critical areas. CCTV systems are connected to monitors at a control post, where one or more operators manually perceive the monitored areas and detect abnormal situations. In this context, human errors are very common, increasing depending on the quantity of video images and the size of the industrial environment to be monitored. In addition to the need for constant and more efficient monitoring, maximizing the use of factory space has also become a growing concern, not only because of the high cost of construction and maintenance, but also to optimize the cadence of the production process and act quickly in case of malfunctions. In this way, the use of ”intelligent” logistics solutions, capable of increasing the ambient density, allows reducing the operating and maintenance costs of the production facilities. This type of solution also leads to a reduction in waiting times for production and machine downtime, reducing the interference of people in logistical tasks. Deep Learning has brought significant development in automated computer vision systems such as object detection [1], image classification [2], and image segmentation [3], which is useful for monitoring these industrial spaces. However, the success of these systems relies on supervised learning that requires a large amount of labeled data. In many situations, it is not possible to acquire a big amount of images, as is the case of industrial spaces. As a result, a large research effort is currently focused on systems that can adapt to new conditions without leveraging a large amount of expensive and time-consuming supervision. One alternative to overcome this is to use the self-supervised learning paradigm [4, 5, 6]. Self-supervised learning constructs feature representations without manual annotations using pretext tasks, which allows models trained in these tasks to extract useful information that can later improve downstream tasks. Further self-supervised learning methods use contrastive learning to push positive instances closer together, and negative ones further apart, in 1 the embedding space [7, 8]. These methods have achieved great performance closing the gap with supervised learning. Researchers proclaim that the next AI revolution will not be supervised but self-supervised [8]. 1.1 Objectives of the Dissertation The main objective of this dissertation was the implementation of a vision system that could classify if a shelf was empty or not, to assimilate to an industrial space context, in a self-supervised learning setting. To achieve such a goal, the following tasks were performed: • Study of the self-supervised learning paradigm as well as the state-of-the-art methods; • Creation of datasets to resemble industrial space problems; • Training of models, varying parameters; • Evaluation of the trained models and selection of the best parameters. 1.2 Dissertation’s Structure This dissertation is organized into 6 chapters: Introduction; Basic Concepts; Self-Supervised Learning; Development Tools and Datasets; Experiments and Results; and Conclusions and Future Work. The contents of each chapter are shortly described in the following topics: • Chapter 1: Introduction In this chapter, the problem addressed is presented and the objective of this dissertation. The structure of the dissertation is also highlighted; • Chapter 2: Basic Concepts This chapter covers the theoretical and basic concepts that lead to the focus of the thesis. Starting with the main definitions of Machine Learning, Deep Learning, and Artificial Neural Networks, going further to how Convolutional Neural Networks work. Followed by their hyperparameters and some architectures and finishing with a description of Computer Vision history; • Chapter 3: Self-Supervised Learning This chapter is a central part of this thesis. It describes how Self-Supervised Learning works and explains some state-of-the-art models of this learning paradigm; 2 • Chapter 4: Development Tools and Datasets This chapter presents the tools and datasets used for the development of the self-supervised learning method; • Chapter 5: Experiments and Results In this chapter the different experiments are reported and the results displayed and discussed. • Chapter 6: Conclusions and Future Work In this chapter are presented the conclusions and future work of this thesis. 3 2 Basic concepts This chapter serves as a theoretical introduction and covers basic concepts to help understand the theory and the methods applied in this thesis. We start with a broad overview on Machine Learning and it’s types. Followed by Deep Learning and Computer Vision. We then enter the scientific background of these concepts and focus on artificial neural networks, especially convolutional neural networks. 2.1 Machine Learning At the beginning of the 20th century, science fiction introduced the world to the concept of artificially intelligent robots. By the 1950s, we had a generation of mathematicians, scientists and philosophers who had culturally internalized the concept of artificial intelligence in their minds. One of them was Alan Turing, a British mathematician and computer scientist who wondered that if humans use available information and reason in order to solve problems and make decisions, then machines could do the same. Much has evolved since then, and artificial intelligence has already had a profound impact on the world. Weather forecasting, email spam filtering, Google search predictions and image classification are just a few examples. What these technologies have in common are Machine Learning algorithms that allow them to learn and consequently respond in real-time. But what is Machine Learning ? First of all Artificial Intelligence (AI) is a field of computer science concerned with not just understanding but also building intelligent entities — machines that can compute how to act effectively and safely in a wide variety of novel situations [9]. Machine Learning (ML) is a subset of AI which allows computers to learn from data without being explicitly programmed. The goal of ML is to design methods that learn using observations of the real world, without explicit definition of rules by humans and improve automatically through experience [10]. A vast set of ML algorithms has been proposed to cover the wide variety of data and problem types. 2.2 Machine Learning Pipeline A ML pipeline can be divided into three main steps: Data Collection, Data Modelling, and Deployment. The first step is gathering data. Data is information collected to be examined and used to help make decisions [11]. This can be a spreadsheet with multiple rows and columns of information, text, images, or even audio files. The machines learn from the given data, so it is very important to collect good and reliable data so that the ML model can find the right patterns. The next stage is modelling, where a ML algorithm is taught to gain insights from the collected data. And finally, deployment, where the models are deployed into production. 4 In data modelling, there are 4 phases as shown in the Figure 1: Problem Definition, Data Understanding and Processing, Modelling, and Evaluation. Figure 1: Machine Learning Pipeline 2.2.1 Problem Definition The first step is aligning the problem you’re trying to solve to a machine learning paradigm. These methods can be divided into three primary approaches: supervised learning, unsupervised learning and reinforcement learning [12] (Figure 2). Figure 2: Types of Machine Learning 1. Supervised Learning Supervised learning (SL) is the most common type of machine learning. In this type of learning, each example in the dataset has a corresponding label. The dataset is the collection of labelled examples {(x(i), y(i))}N i=1. Suppose that each element among Nis a feature vector x(i), i.e. a vector in which each dimension j= 1, ..., d contains a value describing the instance. This value designated as x(i) j is called a feature. So if each x(i)in our dataset represents a fruit, then the first feature, x(i) 1, might contain the colour of the fruit, the second feature, x(i) 2, the weight in grams, x(i) 3the width and so on. 5 The feature at position jin a feature vector x(i)contains the same type of information in all the examples of the dataset. That is, if x(i) 2contains the weight in grams in an example x(i), then x(k) 2also contains the weight in grams in each example x(k),k= 1, ..., N [13]. The input can be a feature vector, but can also be more complex data such as images, that we will see further on. Each pair (x(i), y(i))was generated by an unknown function y=f(x). The goal is to find a function hthat approximates to the true function f. The issue is not how well the function performs in the training set, but how well it handles inputs xit has not seen before. We evaluate this with a second sample of (x(i), y(i))pairs, which we call the test set. hgeneralizes well if it accurately predicts the outputs y(i)of the test set (Figure 3). In SL we can have two types of problems [12]: (a) Classification, which is assigning a label to an unlabelled example where the label belongs to a finite set of categories. It’s a binary classification problem if it only has two categories and a multi-class classification problem if it has more than two. (b) Regression, which is predicting a real-valued label given an unlabelled example, where the label is a real or a continuous value. Figure 3: Illustrated Example of a Supervised Learning Workflow - Classification problem 2. Unsupervised Learning In Unsupervised Learning (UL), the dataset is a collection of unlabelled examples {x(i)}N i=1. Assume again that each x(i)is a feature vector and the goal is to build a model that takes a feature vector as input and transforms it either to another vector or to a value that can be used to solve a practical problem [13]. The machine uses unlabelled data and acts on the information from the data without guidance, grouping the unsorted information according to patterns, similarities and differences without any prior training. It is 6 defined to find the hidden structure in the data itself. For example, suppose the model receives images of different fruits like bananas, pears, cherries and blueberries that it has never seen before. So the machine has no idea of the characteristics of these fruits, so we cannot categorise them as ′banana′,′cherry′, etc. However, it can categorise them based on their resemblance and put together images of fruits that have the same shape and colour, for example (Figure 4). Unsupervised learning is classified into two categories [14]: (a) Clustering is applied to group data based on different patterns our machine model finds, such as grouping different fruits. (b) Association is a rule-based ML technique that finds out useful relations between parameters of a large data set, such as: people that buy Xalso tend to buy Y. Figure 4: Illustrated Example of an Unsupervised Learning Workflow - Clustering 3. Reinforcement Learning Reinforcement Learning (RL) is based on developing a system that improves its performance by receiving feedback from the environment at each iteration. There is an agent that uses insights from the environment to perform actions with the goal of maximising the cumulative reward. It learns from the environment by interacting with it and receives rewards or punishments based on its actions [15]. Let us take a simple example using Robert Tryon’s experiment testing the ability of successive generations of rats in completing a maze [16]. So suppose we want a mouse to complete a maze. In this case, our agent is the mouse, the environment is the maze, the action is the movement of the mouse through the maze, moving left, right, etc., and there is a state that represents the position of the mouse in the maze in each iteration. Eating the pieces of food it finds when it goes the right way is its reward and motivates it to explore further. Its punishment is, for example, getting stuck in a narrow hole. The mouse’s goal is to maximise the rewards, 7 i.e. to eat the largest amount of food (Figure 5). Reinforcement Learning is a powerful tool that can help increase automation and optimize sophisticated systems such as robotics, autonomous driving tasks and manufacturing. Figure 5: Illustrated Example of an Reinforcement Learning Workflow It’s important to mention that modern research is not limited to these Machine Learning approaches. There are many other types of learning. The ones that are used during this work are: •Semi-Supervised Learning •Transfer Learning •Self-Supervised Learning Labelled data is often difficult, expensive, or/and time-consuming to obtain, as they require the effort of human annotators. In industrial scenarios it is not always possible to have a large amount of data and when possible they are normally unbalanced. On the other hand, unlabelled data is relatively easy to collect. SemiSupervised Learning (Semi-SL) addresses this problem by using both labelled and unlabelled data to learn from. The portion of labeled examples is usually quite small compared to the unlabelled example. The goal of Semi-SL is to understand how combining labeled and unlabelled data may change the learning behaviour, and design algorithms that take advantage of this combination [17] (Figure 6 (a)). Humans recognise and apply relevant knowledge from previous experiences when confronted with new tasks. The more a new task is related to a previous experience, the easier it is to solve. In contrast, common ML 8 long fiber called the axon (see Figure 11). Figure 11: Illustrative biological neuron The cell body of the neuron, which includes the neuron nucleus, is where most of the neural computation takes place. Neural activity is passed from one neuron to another in the form of electrical impulses that travel along the axon of the neuron by an electrochemical process. The axon can be thought of as a connecting wire. This transport process moves along the cell of the neuron, down the axon then through synaptic junctions across a synaptic space to the dendrites and/or soma of the next neuron at an average speed of 3 m/sec. Since a given neuron may have multiple synapses, a neuron can connect to many other neurons. Similarly, since there are many dendrites, a single neuron can also receive messages from many other neurons. In this way, the biological neural network is interconnected. Not all connections are equally weighted, some have a higher priority than others. Also, some are excitatory and others are inhibitory (to block the transmission of a message). These differences are caused by differences in chemistry and by the presence of chemical transmitters and modulatory substances within and near neurons, axons, and in the synaptic junction [26, 27]. Neuroscientists have discovered that the human brain learns by changing the strength of the synaptic connection between neurons when simulated repeatedly by the same impulse. The human brain is made up of about 100 billion neurons that are connected in complex ways that allow us to learn new tasks and perform regular activities. A single neuron performs only one simple modular function, which is to respond to the nerve activations coming from the transmitter neurons connected to its dendrite and to transmit its activation to the receiver neurons via axons. However, it is the composition of these simple functions that together can express complex functions. The structure of the biological neural system and the way it performs its functions inspired the idea of ANNs [28]. Analogous to the structure of the human brain, an ANN consists of a series of processing units - neurons. 15 Each neuron is connected to another neuron by means of directed communication links, each with an associated weight, just like biological neurons. ANNs can be described as a directed graph whose nodes correspond to neurons that perform the basic units of computation and whose edges correspond to the connection between the neurons [29, 30]. The basic motivation behind using an ANN model is to extract the most relevant features from the original attributes. By using a complex combination of inter-connected nodes, ANN models are able to extract much richer sets of features [28]. 2.4.1 The Simple Perceptron In 1958, Frank Rosenblatt, an American psychologist notable in the field of artificial intelligence, invented an artificial neuron - the perceptron [31]. A perceptron is a feed-forward neural network consisting of a single neuron that can receive multiple inputs and produce a single output. Feed-forward because the information flows only forward through the network from the input to the output. Perceptrons are used to classify linearly separable classes by finding an arbitrary m-dimensional hyperplane in the feature space that separates instances of two classes [27]. Figure 12 illustrates the basic architecture of a perceptron that takes ninput attributes: x1, x2, ..., xn, with xi∈R, and produces a binary output ˆy∈R. Each attribute xiis multiplied by a specific weight wi. The weighted link is used to emulate the strength of a synaptic connection between neurons. These products are summed and fed to a nonlinear function, an activation function Φ. This function determines if the neuron is activated or not, if its value is above a certain threshold. The perceptron has an additional input called the bias. The job of the bias Θis to shift the activation function to positive or negative values, making adjustments within neurons. Changing the bias value does not change the shape of the activation function, but together with the other weights determines when the perceptron fires. Training the perceptron aims at determining the optimal weights and bias values at which it fires [28, 26]. The general model takes the form: ˆy= Φ n X i=1 wixi+ Θ! 16 Figure 12: Illustrative simple perceptron 2.4.2 The Multilayer Perceptron A single perceptron can solve any classification problem for linearly separable classes. If given two nonlinearly separable classes, a single perceptron will fail to solve the problem of classifying them. To solve this type of problem a multilayer perceptron (MLP) network is needed. The decision boundaries in a multilayer perceptron network have a more complex geometric shape in the feature space than in a hyperplane [27]. A multilayer perceptron is a feedforward artificial neural network [32]. It generalizes the basic concept of a perceptron to more complex architectures of nodes capable of learning nonlinear decision boundaries. In this architecture, nodes are arranged in groups called layers. These layers are usually organized in the form of a chain so that each layer acts on the outputs of the previous layer. In this way, the layers represent different levels of abstraction that are sequentially applied to the input features [28]. Figure 13 shows a multilayer neural network architecture with a single hidden layer. Figure 13: Illustrative example of an artifical neural network with a single hidden layers 17 There are three types of layers: Input layer, hidden layer, and output layer. The input layer, the first layer of the network, is used to represent attributes from the data. These inputs are fed into the intermediate layers - the hidden layers - which consist of processing units known as hidden nodes. Each hidden node processes the signals it receives from the input nodes or hidden nodes of the previous layer and generates an activation value that is passed on to the next layer. A unit in one layer is connected to all units in the previous and subsequent layers, so they are called fully connected layers. Intuitively, we can think of each hidden node as a perceptron trying to construct a hyperplane, while the output node simply combines the results of the perceptrons to obtain the decision boundary. While the first hidden layer works directly with the input attributes and thus captures simpler features, subsequent hidden layers can combine them and construct more complex features. The use of hidden layers in ANN is based on the general assumption that complex high-level features can be constructed by combining simpler low-level features. The larger the number of hidden layers, the deeper the hierarchy of features learned by the network tends to be. This motivates learning ANN models with long chains of hidden layers known as deep neural networks. Unlike shallow neural networks, which have only a small number of hidden layers, deep neural networks can represent features at multiple levels of abstraction and often require many fewer nodes per layer to achieve similar generalization performance as shallow networks. From this perspective, an MLP learns a hierarchy of features at different levels of abstraction that are eventually combined in the output layer, which processes the activation values from the previous layer to make predictions for the output variables [10, 33, 29]. 2.4.3 Convolutional Neural Networks Convolutional neural networks (CNNs), also called convnets, are a specific type of artificial neural network. They are one of the best learning algorithms for understanding image content and have shown exemplary performance in image segmentation, classification, detection, and retrieval-related tasks [34]. The groundwork for convolutional neural networks goes back to the 1980’s and was laid by Fukushima and LeCun et al., which will be detailed in chapter 2.5. However, it only became popular in 2012, when a CNN model called AlexNet [2] won the annual ImageNet Large-Scale Visual Recognition Challenge (ILSVRC) [35]. The attractive features of CNN is its ability to exploit spatial or temporal correlations in data and significantly reduce the number of learn-able variables, reducing time and computational cost. For example, in image classification, if one initial layer recognizes edges, a follow up layer can recognize simpler shapes, and the next recognize higher-level features such as faces. The network extracts different abstract features as the input spreads into deeper layers [36]. A typical CNN architecture is generally divided into several learning stages consisting of a combination of 18 convolutional layers, non-linear processing units, and pooling layers, followed by one or more fully connected layers at the end (Figure 14). In the following, we will explain each stage. Figure 14: Illustrative example of a Convolutional Neural Network 2.4.3.1 Convolutional Layer The input image is a matrix of numbers, where each number corresponds to the intensity of a single pixel, ranging from 0 to 255. In the RGB model, the color image consists of three matrices corresponding to the three color channels: red, green, and blue. These matrices pass through the convolutional layer, which is the most important building block in convolutional neural networks. This layer performs an operation called convolution , which is a mathematical linear operation between matrices. Since the technique was designed for two-dimensional inputs, multiplication is performed between an array of input data and a two-dimensional array of weights called a filter or kernel . These kernels are a grid of discrete numbers and typically have a small spatial dimensionality, but are distributed over the entire depth of the input data [37, 10]. Such reduction of dimensions provides moderate invariance in the scale and position of objects. When using images, the inputs have very high dimensions and must be efficiently processed by large CNN models. Therefore, instead of defining convolutional filters that match the spatial size of the inputs, we typically define them to be significantly smaller compared to the input images. This design offers two key advantages: The number of parameters that can be learned is greatly reduced when smaller kernels are used, and small filters ensure that different patterns are learned from local regions corresponding, for example, to different object parts in an image. The size of the filter, height and width, which defines the spatial extent of a region that a filter can change at each convolution step, is called the filter’s receptive field [10, 38]. A dot product is applied between the filter and the filtered region of the input. A dot product is the sum of the element-wise products of these two matrices, resulting in a single value. Specifically, the filter is applied systematically to each overlapping filtered region of the input, left to right, top to bottom [38]. Repeated application of the same filter to the input image results in a map of activations called a feature 19 map , which indicates the position and strength of a recognised feature in the input. The full feature maps are obtained by using several different kernels. Such a weight-sharing mechanism has several advantages such as it can reduce the model complexity and make the network easier to train. Each kernel has a corresponding activation map that is stacked along the depth dimension to form the entire output volume of the convolutional layer. This systematic application of the same filter across an entire image is a powerful idea. If the filter is designed to detect a particular type of feature in the input, then the systematic application of that filter to the entire input image provides the filter with the ability to detect that feature anywhere in the image. The innovation of using the convolution operation in a neural network is that the values of the filter are weights that are learned during the training of the network. The network learns which types of features to extract from the input [38, 39]. Mathematically, the convolution operation is denoted with an asterisk and consequently, the feature map values are calculated with the following formula [39]: S(i, j) = (K∗I)(i, j) = X mX n I(i−m, j −n)K(m, n) where Idenotes the two-dimensional input image, Kdenotes the kernel, and iand jare respectively the indexes of rows and columns of the resulted matrix. Figure 15 illustrates an example of a convolution operation, where the input image and kernel only have one channel. Figure 15: Example of a 2D convolution operation There have been many advancements in the convolution operation originating variants such as Tiled convo20 lution , Transposed convolution , Dilated convolution , among others which are described in [40]. In the example (Figure 15), the filter moves across the image with a change of one pixel in the column for the horizontal movements and a change of one pixel in the row for the vertical movements. The amount of movement between applications of the filter to the input image is called the stride and is almost always symmetrical in height and width dimensions. This value is an integer, usually one, and can be changed, affecting how the filter is applied to the image and the resulting feature map [38]. Figure 16 illustrates this when applying a convolution with a 5×5×1input image and a 3×3×1kernel, but with different strides. In (a)with a stride of 1 the result is a3×3feature map and in (b)with a stride of 2 the result is a 2×2feature map. Figure 16: Illustrative examples when using different stride values One of the disadvantages of the convolution operation is the loss of information that might be present at the edges of the image. Since they are only captured when the filter slides, they have less chance of being seen. A very simple and effective method to solve this problem is the use of zero-padding , which is the process of applying a combination of pixels of value zero around the input. It also helps prevent output size from shrinking with depth [36]. For example, in Figure 17 we are using the same input and filter as the example in Figure 15 but with zero-padding. However, the output will be a different feature map with the same size as the original input in Figure 15. 21 Figure 17: Illustrative example when using zero-padding The output size of the Convolutional layer can be calculated through the following equation [36]: O= 1 + I+ 2P−K S where Ois the output size, Iand Kare the input and filter size, respectively; Sis the stride value and Pis the number of zero-padding layers (for example P= 1 in Figure 17). 2.4.3.2 Non-linearity Once the feature map is created, each value in the feature map is passed through a non-linear activation function. The non-linearity can be used to adjust or truncate the output generated [36]. The activation function takes a real-valued input and squashes it within a small range. Applying a nonlinear function after the weighting layers is very important because it allows the neural network to learn nonlinear features. A nonlinear function can also be understood as a selection mechanism that decides whether a neuron fires or not given all its inputs. The most common activation functions used in DNNs are Sigmoid , Hyperbolic Tangent ( Tanh ), and Rectifier Linear Unit ( ReLU ) functions [10]. 22 Figure 18: (a) Sigmoid (b) Tanh (c) ReLU activation functions •Sigmoid: The definition of the sigmoid function is as follows: f(x) = 1 1 + e−x Here, Df =Rand D′f=]0,1[. Figure 18 (a) shows the graphic of the Sigmoid function. •Tanh: The hyperbolic tangent function is defined as follows: Tanh(x) = sinh(x) cosh(x)=ex−e−x ex+e−x Here, sinh(x)is the hyperbolic sine, cosh(x)is the hyperbolic cosine, D Tanh =R, and D′Tanh(x) = ]−1,1[. Figure 18 (b) shows the graphic of the Tanh function. Neural networks using Tanh as the activation functions converge faster than those using Sigmoid . In addition, the networks using Tanh have lower classification errors comparatively to those using Sigmoid activation function.[41]. •ReLU: The definition of the rectifier linear unit function is: f(x) = max(0, x) =      x, se x⩾0, 0,se x < 0, Here, Df =Rand D′f= [0,+∞[. Figure 18 (c) shows the graphic of ReLU function. ReLU is computationally cheaper than Sigmoid and Tanh because it does not need to compute exponential functions, making it much quicker. It converges much faster than the previous activation functions and also allows the network to easily obtain sparse representation, making the network learn more efficiently [41]. 23 The effectiveness of ReLU has led to many variants, such as Noisy ReLU , Leaky ReLU , Randomized Leaky ReLU and Exponential Linear Unit [10]. Figure 19: (a) Noisy ReLU (b) Leaky ReLU (c) Randomized Leaky ReLU (d) Exponential Linear Unit activation functions •Noisy ReLU: The noisy ReLU adds a sample drawn from a Gaussian distribution with mean zero and a variance that depends on the input value (σ(x)) in the positive input. The noisy ReLU is defined as follows: f(x) = max(0, x +ϵ),with ϵ∼ N(0, σ(x)) Figure 19 (a) shows the graphic of Noisy ReLU function. •Leaky ReLU: The LReLU function is defined by: f(x) =      x, se x > 0, cx, se x⩽0, where cis typically a positive small value, such as 0,01. Figure 19 (b) shows the graphic of this function. Instead of reducing the output to zero when the input is negative, LReLU outputs a down-scaled version of the negative input. Consequently, there are no zero gradients and no neuron units can be “off” always. Experimental results have shown that the learning capabilities of the neural networks become more robust when using LReLU [10, 41]. The Parametric Linear Units ( PReLU ) function behaves a lot like the LReLU . Its definition is analogous to the previous one with the difference of replacing the constant cwith a learnable parameter. Experiments 24 be defined as: L=1 2nX n yd2+ (1 −y)max(0, m −d)2 where mis the margin and y∈[0,1] indicate whether the pairs are dissimilar or similar respectively. d is a distance measure, for example, the Euclidean distance. This loss function will be explained in more detail in chapter 3. The loss functions used for multi-class classification tasks are also applicable to binary classification tasks. However, the case reverse is not true unless a multi-class problem is divided into various one-vs-rest binary classification problems where independent classifiers are trained for each case using a binary classification loss [10]. There are many other loss functions such as SVM Hinge Loss ,l1 error and Triplet loss , which are explained in detail in [10] and [40]. 2.4.4.3 Weight Initialization Initializing neural networks plays an important role to stably train deep neural networks. Proper initialization of the weights is critical to its convergence. The weights are updated during the training of the new model, making the pre-trained model act as a weight initialization scheme when training the new model. This helps the new model converge faster since this initialization gives it a starting point at a suitable region that would otherwise be inaccessible by random initialization. The models learned by pre-training are more consistent and provide better generalization performance. [10, 38, 28]. 2.4.4.4 Regularizations As explained previously we want our neural network to perform well on the training data, but also on new unseen data. DNNs have a large number of parameters and tend to over-fit on the training set while learning. There are many strategies used that are explicitly designed to reduce the test error and prevent over fitting. These strategies are designated regularization approaches. There are many forms of regularization and developing new effective strategies has been one of the major research efforts in the field [39, 10]. Next, we will describe some regularization techniques. Dropout Dropout is a regularization technique to avoid overfitting and thus increase generalization performance. The main goal of dropout is to avoid learning spurious features at hidden nodes. In NNs, multiple connections that learn a nonlinear relation are sometimes co-adapted, so they show good training performance only when used 31 in highly selective combinations. Dropout randomly drops input and hidden nodes in the network during training to disrupt complex ”co-adaptations” in the learned features (Figure 22). Dropout can prevent the network from becoming overly dependent on any one neuron and can force the network to be accurate even in the absence of certain information [28, 40]. Dropout can be added to the model by adding new dropout layers, specifying the amount of nodes removed as a parameter. For example, if the dropout rate is 0.2, 20% of the neurons are randomly selected and ignored at each iteration of forward propagation. We generally use a small dropout value between 20% - 50% of neurons. A probability too low has little effect and a value too high makes the network not learn well enough [32]. Several methods have been proposed to improve dropout. Some are mentioned in [10] and [40]. Figure 22: Illustrative example of dropout applied in the hidden layers where the darker nodes represent the ”dropped” nodes in an iteration lpregularization This regularization changes the loss function by adding additional terms that penalize the model complexity. Suppose the loss function is L, then the regularized loss will be: E=L+αR(θ) where R(θ)is the regularization term and αis the regularization strength [40], being θthe vector of model parameters. Normally αis a very small value. When p= 1,R(θ) = ∥θ∥1and the l1regularization is equal to the sum of absolute value of the magnitude of the coefficients. When p= 2,R(θ) = 1 2∥θ∥2 2and the l2regularization is commonly referred to as weight decay and is the sum of the square of all weights [10, 40]. 32 Data Augmentation Data Augmentation is an easy and effective way of enhancing the generalization power of CNN models and also enlarging datasets that have a small number of training examples. Data augmentation consists in transforming the available data into new data without altering their natures. Some methods are rotation, cropping, flipping, scaling, and colour distortion. These operations can be performed separately or combined [10, 38]. Figure 24 illustrates some random augmentation techniques applied to Figure 23, such as zoom, rotation, shift, colour distortion and flip. Figure 23: Original image to apply augmentation Figure 24: Random augmented images of 23:zoom, rotation, shift, colour jittering, and flip Batch Normalization Batch Normalization (BN) is used to solve the problems related to internal covariance shifting within feature maps. Internal covariance shift refers to the change in the distribution of activations of each layer as parameters 33 are updated during training. BN solves this problem by using a normalization step that fixes the means and variances of the layer inputs, computing the estimates of mean and variance after each mini-batch rather than after the entire training set. This improves convergence and avoids network instability issues such as vanishing/- exploding gradients and activation saturation [34, 10, 40]. During training, in order to zero-center and normalise the inputs, the BN layer calculates the mean µand variance σ2over the current mini-bacth, with mthe number of instances in the mini-batch: µ=1 m m X i=1 xiσ2=1 m m X i=1 (xi−µ)2 Then normalizes the inputs ˆxi=xi−µ √σ2+ϵ and scales and shifts the normalized values in order to obtain zi zi=γˆxi+β where γand βare parameters learned during training, ϵis a constant added to the mini-batch variance for numerical stability [54]. BN is usually applied after the CNNs convolution layers, before applying the nonlinear activation function, and is used in state-of-the-art CNN architectures [10]. BN has many advantages. It stabilizes the training of deep networks and provides robustness to bad weight initializations. In addition, when used the training of the network is less sensitive to the choice of hyperparameters, such as the learning rate, and it greatly improves the convergence rate of the network. Finally, BN regularizes the model, and thus reduces the need for Dropout [10, 40]. Early Stopping Early stopping is a strategy applied to avoid overfitting. This is achieved by returning to the parameter setting at the point in time where the metric being monitored is the lowest, which is normally the validation set loss (Figure 25). Every time the error on the validation set improves, a copy of the model parameters is stored. When the training is done, we return those parameters, instead of the latest parameters. The algorithm terminates when no parameters have improved over the best-recorded validation loss for a pre-specified number of epochs. This strategy is simple and effective and therefore one of the most used forms of regularization [39]. 34 Figure 25: An illustration of a network overfitting during training and applying early stopping 2.4.5 Convolutional Neural Architectures A CNN is a combination of the layers described previously, arranged in a specific way. The number and the sequence of these layers depend on the way it is dimensioned. Designing a CNN consists on defining the sequence and the structure of the layers: defining the number of filters that each convolutional layer will process, as well as the size of the filters and stride; choosing the activation function to use; defining the pooling operation, along with the number of filters, size, and stride. Figure 14 schematises a CNN architecture. Following, we will introduce some successful CNN designs which are constructed using the basic building blocks that were explained so far. 2.4.5.1 AlexNet AlexNet was proposed by Krizhevesky et al. [2] and won the ImageNet Large Scale Visual Recognition Challenge (ILSVRC) in 2012. It achieved a large improvement in image classification performance compared to previous CNN architectures such as LeNet, which were smaller and not tested on large datasets such as the ImageNet dataset [35]. The architecture of the network is summarised in Figure 26. It contains five convolutional layers and three fully-connected layers (FCLs). ReLU is applied after each convolutional layer. The output of the last FCL layer is fed into a 1000-way softmax , which produces a distribution over the 1000 ImageNet class labels. Data augmentation, dropout, and normalization were also applied to avoid overfitting. AlexNet has 62.4 million parameters trained on 35 ImageNet with 1.2 million images. The deep CNN’s learning capability was limited at this time due to hardware limitations. To overcome this, two GPUs were used in parallel to train AlexNet [10, 55, 2]. Figure 26: An illustration of the AlexNet Architecture 2.4.5.2 VGGnet The VGGnet [56] was introduced in 2014 by Simonyan and Zisserman from the Visual Geometry Group (VGG) research lab at Oxford University. Although it did not win the ILSVRC’14, it became one of the most popular CNN models, due to its simplicity and the use of small-sized convolutional filters. There are many different configurations of this network, but the most successful are VGGnet-16 and VGGnet-19 [10]. The VGGnet architecture uses only 3x3 convolutional kernels with max-pooling layers and three full-connected layers. Each convolutional layer is followed by a ReLU layer. Padding was performed to maintain spatial resolution, and dropouts are used in the first two FCLs to avoid overfitting. The use of smaller filters leads to a relatively small number of parameters and thus efficient training and testing. In addition, smaller filters allow more layers to be stacked, leading to deeper networks and thus better performance in vision tasks. VGGnet set a new trend to use small filters in CNNs [10, 34, 56]. Figure 27 illustrates VGGnet-19 which has 144 million parameters. 36 Figure 27: An illustration of the VGGnet-19 Architecture 2.4.5.3 ResNet He et al. developed ResNet (Residual Network) [57], which won the 2015 ILSVRC competition. Their goal was to design an ultra-deep network that was free from the vanishing gradient problem, as compared to previous networks [55]. The concept behind ResNet was that despite its depth, the network is trained similarly to a shallow network by skipping after every 2 layers [58]. To perform a computation, both the input and the output were copied to the next layer, basically learning the residual of the previous computation. The details of the skip connection are shown in Figure 28. Given an input x, the convolutional layers implement a transformation function on that input, denoted F(x). In a residual block, the original input is added to this transformation, using a direct connection from the input that bypasses the transformation layers. The original mapping thus becomes x+F(x)(Figure 28). This connection is called a skip connection . In this way, the transformation function in a residual block is split into an identity term (representing the input) and a residual term, which helps to focus on the transformation of the residual feature maps. ResNet consists of several residual blocks stacked on top of each other. The residual connections are the key to better classification accuracy of deep networks [57, 10, 55]. 37 Figure 28: Residual Block in a ResNet The convolutional layers in the residual block are followed by a BN and a ReLU activation layer. Different types of ResNet have been developed based on the number of layers (starting from 34 layers up to 1202 layers). The most widely used type was ResNet50 (Figure 29), which has 49 convolutional layers plus a single FC layer [57]. Figure 29: An illustration of the ResNet50 2.5 Computer Vision Vision is the sense on which humans rely on most, and undoubtedly the one that provides most of the data he receives. The amount of information the higher centers of the brain receive from the eye must be at least two orders of magnitude greater than all the information they receive from the other senses. Of course, humans do 38 not store all this information, some is forgotten over time and some goes unnoticed. It is impossible to retain all the data received when data rates for continuous viewing are likely to exceed 10 Mbps [59]. In a world where we want to have machines do tasks just like humans, machines need a sense of vision. We are inundated with images. It has never been easier to take and share a photo or video. Youtube is one of the largests search engines and hundreds of hours of videos are uploaded every minute and billions of videos are watched every day.The internet is comprised of text and images. It is relatively easy to index and search text, but to index and search images, algorithms need to know what the images contain. For a long time, the content of images and videos remained opaque and was best described by the meta descriptions of the person who uploaded them. To get the most out of image data, we need computers that can ”see” an image and understand the content [38]. Figure 30: Human vision system and computer vision system Computer vision (CV) is a field concerned with developing techniques that help computers and systems see and understand the content of digital images, derive meaningful information from them - and take action or make recommendations based on that information. If AI enables computers to think, then computer vision enables them to see, perceive and understand their environment [60]. Typically, this involves developing methods that attempt to replicate the capabilities of human vision.This seems like a simple problem because it is so effortless for humans, even at a very young age. Yet it remains a largely unsolved problem based on the limited perception of human vision and the complexity of the visual world [38]. CV works much like human vision, except that humans have a head start. Human vision has the advantage of a lifetime of training to distinguish objects, how far away they are, whether they are moving, and whether there is something wrong in an image [60]. A true vision system must be able to see in any scene and still extract 39 something meaningful. Computers work well on tightly constrained problems, not on open-ended, unbounded problems like visual perception. Still, there has been progress in this area, especially in recent years with available optical character recognition and face recognition systems in cameras and smartphones. Computer vision is at an extraordinary point in its evolution. The subject itself has been around since the 1960s, but only recently has it been possible to build useful computer systems using ideas from CV [38]. In 1959 David Hubel and Torsten Wiesel, two neurophysiologists, placed electrodes into the primary visual cortex area of an anesthetized cat’s brain and tried to observe the neuronal activity in that region while showing the cat various images (Figure 32). Initially, they couldn’t get the nerve cells to respond to anything. However, after a few months, they accidentally saw that a neuron fired as they were slipping a new slide into the projector. They after realised that what got the neuron excited was the movement of the line created by the shadow of the sharp edge of the glass slide. The researchers concluded that there are simple and complex neurons and that visual processing starts with simple structures such as straight edges [61]. Figure 31: David Hubel and Torsten Wiesel’s experiment Also in 1959 Russell Kirsch and his colleagues developed an apparatus that allowed transforming images into grids of numbers, a binary language machines could understand. It is because of their work that we now can process digital images in diverse ways [62]. In 1963 Lawrence Roberts published his Ph.D. thesis named ”Machine perception of three-dimensional solids” [63], where he described the process of deriving 3D information about solid objects from 2D photographs. The objective was to process 2D photographs into a line drawing, transform the line drawing into a 3D representation, and finally, display the 3D structure with all the hidden lines removed, from any point of view. This was considered to be one of the precursors of modern CV. AI became an academic discipline in the 1960s. In 1966, Seymour Papert a MIT professor launched the Summer Vision Project [64], proposing a group of MIT students to construct a significant part of a visual system 40 To avoid shortcuts due to edge continuity and pixel intensity distribution, the authors leave a random gap between tiles. To solve the jigsaw puzzle, the model must learn to recognise how the parts in an object are assembled, their shapes and the relative position of the different parts of objects. Thus, the representations are useful for classification and recognition of difficult tasks. On the PASCAL VOC 2007 classification and detection tasks this method obtains 67,6% and 53,2% mAP, respectively. On the PASCAL VOC 2012 segmentation task it achieves a 37,6% mIU, outperforming the colorization pretext task. Figure 36: Illustration of how the jigsaw puzzle pretext task is generated and solved with the permutation in Figure 35 (b) (9,8,4,2,1,5,6,3,7) and index 27 3.1.3 Rotation Gidaris et al. [89] proposed a pretext task to learn image representations by training CNNs to recognize the geometric transformation that is applied to an input image. Each input image is rotated with 4 different 2D rotations, which are all fed to a CNN model. This model is trained with cross-entropy on a 4-class classification task to predict the transformation applied. Formally, Xis an input image, and G={g(X|y)}4 y=1 the set of geometric transformations as all the image rotations by multiples of 90◦, this is, 2D image rotations by 0◦,90◦,180◦and 270◦.g(X|y)is the operator that rotates Xby ydegrees, which yields the transformed image Xy=g(X|y). The model F(.) receives as input the image Xy∗(where y∗is unknown to the model) and outputs a probability distribution over all possible geometric transformations: F(Xy∗|θ) = {Fy(Xy∗|θ)}4 y=1. Where Fy(Xy∗|θ)is the predicted 47 probability for the geometric transformation with label yand θare the learnable parameters of the model (Figure 37). Figure 37: Illustration of the rotation pretext task where gis the geometric transformation (0º, 90º, 180º or 270º rotations). Fy(Xy∗)is the probability of rotation transformation ypredicted by F(.)when it receives as input an image that has been transformed y∗ The authors believe that it is essentially impossible for a CNN model to effectively perform the above rotation recognition task unless it has first learned to recognize and detect objects as well as their semantic parts in images. Despite the simplicity of the task, it provides a very powerful supervisory signal for semantic feature learning. Fine-tuning the learned features on the PASCAL VOC 2007 classification and detection tasks this method reaches 72,97% and 54,4% mAP, respectively. On the PASCAL VOC 2012 segmentation task it reaches a 39.1% mIU, outperforming the previous pretext tasks. 3.2 Contrastive Learning As mentioned previously, in the early stages of SSL, representation learning focused on exploiting pretext tasks. Though these approaches succeed in computer vision tasks, there is still a large gap between these methods and 48 supervised learning [93, 99, 89, 79, 8]. Recently, there has been a significant advancement in using Contrastive Learning (CL) , which significantly closes the gap between SSL methods and supervised learning [107]. CL is a discriminative model that uses positive and negative pairs to learn representations by distinguishing between views of the same image [79]. It aims at pushing different views of the same instance close together and different instances further apart in the representation space [87, 5, 7, 8]. Figure 38 (a) shows a batch of two images, where each image forms its own class. To create a positive pair, the image is augmented in two different ways. One of the augmented views is called the anchor of the image and the other is called a positive . Any image that differs from the anchor image is negative to the anchor [8]. In Figure 38 (b) the image denoted as xais the anchor image of x,x+is a positive pair to the anchor image, and x−a negative, as it is from a different class. The CL objective is to learn a representation space that pulls representations that come from the same images (xaand x+) and repel representations that come from different images (xaand x−). Figure 38: (a) For each image in the batch random augmentation is applied to get a pair of two images that represent different instances of the same image. (b) CL pulls the anchor and positive images close together and the negative image away. The CL framework can be divided into 5 parts [8, 108, 79]: 1. Data Augmentation Pipeline: The purpose of data augmentation in contrastive learning is to generate anchor, positive and negative images. Augmentation is applied to the input images to generate new 49 samples which preserve the same underlying semantics as the original input images [108, 109]. A good augmentation strategy is an important factor for CL, as it forces the network to learn rich and generalizable features in an SSL environment [8]. Research has shown that combining multiple data augmentation techniques boosts representations [110, 83]. Figure 24, in section 2.4.4.4, illustrates examples of augmentation. 2. Encoder: The encoder part of the network extracts the feature representations of the images. Given two augmented images xiand xjit extracts embedding vectors hiand hj(Figure 39), which are feature representations of x+ 1and x+ 2respectively [8, 108]. These representations affect how well a classification model learns to distinguish between different classes. It has been shown that features extracted from the later stages of the encoder are a better representation of the input than features extracted from the earlier stages [7]. ResNet and its variants are the most commonly used CNNs in CL [8, 7, 79]. 3. Projection Head: After extraction, the embedding vectors hiand hjpass through an MLP to produce embeddings ziand zjon which the contrastive loss is computed (Figure 39). It has been proven that adding the MLP achieves better results [8, 79]. Figure 39: Encoder extracting embeddings hiand hjthat pass trough a projection head producing ziand zjon which the contrastive loss is computed 4. Similarity Measure: The central idea in contrastive learning is to bring similar instances closer together and dissimilar instances far apart. A metric is needed to measure the closeness between representations [7]. Cosine similarity is one of the most commonly used similarity metrics, which measures the cosine of 50 the angle between two non-zero vectors. The cosine similarity is calculated for the ziand zjembeddings and is defined as [7, 79]: sim(zi, zj) = zi·zj ∥zi∥∥zj∥ where ∥·∥is the Euclidean norm of the vector and “.” is the dot product. This similarity ranges from 1 to -1 [79]. 5. Contrastive Loss: A contrastive loss function is defined to penalize the network for getting different representations of different versions of the same image. The original image and the transformed image should provide similar predictions and produce similar features in the intermediate representations. Therefore, the loss function is minimized when the similarity between the query image and the positive embedding is greater and maximized when the dissimilarity between the two images is greater. Based on the loss, the representations of the encoder and the projection head improve over time and the obtained representations place similar image instances closer in the space and negatives far away [8]. Widely used loss functions include the Noise-Contrastive Estimation Loss (NCE) [111], the Triplet Loss [112], and InfoNCE [113]. Current contrastive learning methods compare embeddings with a contrastive loss called NT-Xent loss (Normalized Temperature-Scaled Cross-Entropy Loss) [114, 110]. This loss function for a positive pair of examples (i, j)is defined as: ℓi,j =−log exp(sim(zi, zj)/τ) P2N k=1 1 [k=i]exp(sim(zi, zk)/τ) where Nis the number of samples, 2Nare the transformed (augmented) pairs and 2(N−1) negative pairs from other examples in the dataset. sim(zi, zj)represents the cosine similarity defined previously. The term in the numerator is the positive pairs and the terms in the denominator are the negative pairs. 1 [k=i]∈0,1is an indicator function evaluating to 1if and only if k=iand τdenotes a temperature parameter. The final loss is computed across all positive pairs, both (i, j)and (j, i). The goal is to identify positive pairs of each ziand repel others [8, 110]. Similar to other deep learning methods, CL uses a variety of optimization algorithms for effective training. The training involves learning the parameters of the network by minimizing the contrastive loss function. SGD and its variants are the most popular optimization algorithms used with CL methods [7]. Recently, CL methods for CV tasks have increased and some have begun to outperform supervised learning 51 methods. DeepCluster [84], ClusterFit [115], SimSiam [116], PIRL [117], MoCo [118], BYOL [119], SimCLR [110], and SwAV [83] are some CL methods. SwAV is described in the following section. 3.2.1 Swapping Assignments between multiple Views (SwAV) Caron et al. [83] proposed an online clustering-based self-supervised method for learning visual features named SwAV. The authors purposed a multi-crop augmentation strategy that generates multiple views of the same image instead of just one pair without quadratically increasing memory and computational requirements. They use two standard resolution crops and take Vadditional low resolution crops that cover only small portions of the image (Figure 40). Using low resolution images provides only a small increase in computational cost and the model becomes scale invariant. This strategy showed an improvement in performance not only for SwAV but also for other contrastive learning methods [83]. Next, random horizontal flips, color distortions, and Gaussian blur are applied to each resulting crop. Figure 40: Multi-crop: image xnis transformed into V+ 2 views: two global views and Vsmall resolution zoomed views Each image xis transformed into augmented views x1and x2by applying a transform tselected randomly from a set of image transformations T. For simplicity, only 2augmented views are listed here, but there can be many more by using multi-crop. The augmented views are passed through a CNN fθ(ResNet50) that follows a projection head and outputs vectors z1and z2(Figure 41). 52 Figure 41: SwAV Architecture The feature vectors z1and z2are then mapped to a set of Ktrainable prototype vectors c1, c2, ..., cK.C is the matrix whose columns are c1, c2, ..., ck. This mapping/code is denoted by Q= [q1,…qB](Figure 42). Figure 42: Assigning Bsamples to Ktrainable prototype vectors Qis optimized to maximize the similarity between the features and the prototypes, i.e, maxQ=Tr(QTCTZ) + ϵH(Q) where H(Q) = −Pij QijlogQij is the entropy function and ϵis a parameter that controls the smoothness of the mapping. In practice, ϵis kept low because using a high value generally leads to a trivial solution where all samples fall into a unique representation and all are assigned to all prototypes uniformly.An equal partition is enforced by constraining the matrix Qso that each prototype is selected the same amount of times. Once a continuous solution Q∗is found it takes the form of a normalized exponential matrix: Q∗=Diag(u)exp CTZ ϵDiag(v) where uand vare renormalization vectors in RKand RBrespectively. These vectors are computed using the iterative Sinkhorn-Knopp algorithm [120]. 53 A ”swapped” prediction problem is set up consisting of predicting the code q1from the feature z2and q2 from z1with the following loss function: L(z1, z2) = ℓ(z1, q2) + ℓ(z2, q1) where ℓ(z1, q2)is the cross-entropy loss between the code and the probability obtained by taking a softmax of the dot products of ziand all prototypes in C: ℓ(zt, qs) = −X k q(k) slogp(k) t, where p(k) t=exp((zT tck)/τ) Pk′exp((zT tck′)/τ) where τdenotes a temperature parameter. The loss function is minimized with respect to the prototypes Cand the parameters θof the encoder fθ. This method compares features z1and z2using intermediate codes q1and q2. If these two features capture similar information, it should be possible to predict the code from the other feature. The features are learned by Swapping Assignments between multiple Views (SwAV) of the same image. Figure 43 illustrates this idea. Figure 43: Swapped prediction problem between two views of the same image 3.3 Evaluation on Self-Supervised Learning After learning the representations, either by pretext task or by contrastive learning, these representations must be evaluated to ensure quality. There are three ways to do so: •Linear Classification: a linear classifier is trained on top of the CNN trained on the unlabeled dataset. The last fully connected layer is removed and the rest of the CNN is frozen, on which the classifier is trained. Evaluation is often performed on the same dataset that was used to train the network [8]. 54 •Fine-tuning with a % of labels: the CNN trained with an SSL method/model is fine-tuned with a certain percentage of the labeled images (usually 1% or 10% from the same dataset). This is considered semi-supervised learning since the model is still trained with a few labels [83]. •Transfer learning to downstream tasks: Transfer learned representations to downstream tasks such as image classification, semantic segmentation, object recognition, and action recognition, etc., on a different data set. The performance of transfer learning on these high-level tasks shows the generalization ability of the learned features. If the CNN can learn general features, then the pre-trained models can be used as a good starting point for other vision tasks that require the acquisition of similar features from images or even small datasets that require this additional information.[7] 55 4 Development Tools and Datasets The model was implemented on Kaggle [121], a cloud computational environment for data scientists and machine learning engineers. Kaggle allows users to find and publish datasets, build AI models, work with other enthusiasts, and even enter competitions to solve data science challenges. It provides a 4-core CPU with 30GB of RAM and an Nvidia Telsa P100 GPU with 13GB of RAM. GPUs are very helpful when using code that takes advantage of GPU-accelerated libraries. Each notebook editing session is provided with 12 hours of execution time for CPU and GPU and 20GB of auto-saved disk space. As for the development environment, Python was the programming language used and Tensorflow the main library. Tensorflow [122] is an end-to-end open-source ML library developed by Google for the implementation and deployment of models. It contains many built-in functions that allow the creation of deep neural networks. In addition, the following libraries were also used: Keras , Numpy , Scikit-Learn , and Matplotlip . Weights & Biases (W&B) [123] was used for experiment tracking and visualizations to develop insights for this work. It is a machine learning platform for developers to build better models faster. W&B has interoperable tools to track experiments, evaluate performance, reproduce models and visualize results. 4.1 Flowers Dataset One of the datasets used in this work was Tensorflow’s flowers dataset [124], which contains 3670 images and respective labels for flower classification. This dataset contains 5 different types of flowers: Dandelion (denoted by 0), Daisy (denoted by 1), Tulip (denoted by 2), Sunflower (denoted by 3), and Rose (denoted by 4). In Figure 44 are represented some image examples from the dataset. Figure 44: Examples of images from the flowers dataset 56 Figure 53: Validation loss during linear classifier training in all experiments Figure 54: Validation accuracy during linear classifier training in all experiments Given the experiments made, experiment Cwas the one that achieved better performance. In more detail Figure 55 shows the confusion matrix and metric values obtained. The model predicted incorrectly 71 images out of a total of 550 images. Dandelion is the class with the highest recall as it is the class that fewer images were misclassified, only 8.9%. Rose has the highest percentage of misclassified images - 15.9%. Daisy is the class with less false positives, having the highest precision. We can also observe that the average metrics precision, recall, and F1Score are all equal to 88% due to the fact that the F N = 65 = F P . These experiments could have been optimized to achieve better results but because this dataset was not the focus of the work no more experiments were made. 63 Figure 55: Confusion matrix and metric values from experiment C 5.2 Bookcase Dataset As described previously the bookcase dataset was created to assimilate to a real problem in an industrial space context. As experiment Cachieved the best performance in the flowers dataset, the SwAV training and linear classifier on the bookcase dataset were done in the exact same settings. The only difference made was an increase in the early stopping of SwAV training. The training was defined to stop if the loss did not improve in 30 epochs, instead of 15. The dataset was also divided into 85% training and 15% validation, corresponding to 3907 and 690 images, respectively. This experiment was defined as Fand the results are shown in Table 3 as are the plots in Figures 56 and 57. Exp Init. W O.Schedules I. lr DS E. lr P Stop Loss ImageNet Poly. Decay 0.1 5 0.0001 0.5 308*/500 2.301 Stop Best Epoch Loss Accuracy Precision Recall F1-Score MCCF 100 100 0.032 0.988 0.99 0.99 0.99 0.98 Table 3: Results of SwAV training on bookcase dataset and results of the linear classifier on SwAVs frozen features with the same dataset 64 Figure 56: Loss during SwAV training on bookcase dataset Figure 57: Validation accuracy (Top) and validation loss (Bottom) during linear classifier training on bookcase dataset During SwAV training F achieved a loss of 2.301 in 308 epochs. The model stopped training due to Kaggle’s 12 hours execution time limitation. However, the linear classifier on these frozen features still achieved 98.8% 65 accuracy, 98% MCC, and 99% on the remaining metrics. The score obtained shows that the binary classifier was able to predict the majority of positive and negative instances. Figure 58 displays the confusion matrix and the metric values from experiment F. The model only misclassified 6 images from a total of 690. Figure 58: Confusion matrix and metric values from experiment F The bookcase dataset was also trained in a fully supervised manner with hyperparameter optimization to indicate the best model possible. Training took several days and achieved a 99.6% accuracy. Despite the selfsupervised method obtaining a lower accuracy, it took around 12 hours with is much less compared to the supervised setting. There was an attempt to do more experiments. One was to increase the batch size from 32 to 64 as SwAV works well on small and large batches [83]. This was not possible to test due to an out of GPU memory problem. Another experiment made was the use of a different backbone model. The settings were equal to experiment F except for the backbone model, where ResNet50 was replaced by VGGnet-16 and VGGnet-19. The linear classifier on the frozen features achieved 81,3% and 77,4% accuracy, respectively. 5.3 Transferring to Downstream tasks Another experiment made was transfer learning to a downstream classification task. The learned representations from SwAV training on the bookcase dataset were used to solve the classification problem with the boxes dataset. A linear classifier is trained on the frozen features (learned with the bookcase dataset) with a single dense layer with 2 neurons, with softmax function and a l2regularizer. The dataset was divided into 70% for training and 30% for validation. The dataset was divided this way because it only contains 116 images. Dividing it for example 70% 66 for training, 20% for validation, and 10% for testing makes the test set contain only 12 images which would not be enough to evaluate the model. To increase the test set, the training set would have to be decreased which would give fewer images for the model to learn from. The model was trained during 500 epochs and completed the total number of epochs. Despite the 30 epochs defined for early stopping the model kept on improving (Figure 59). Figure 59: Validation loss of the linear classifier on the boxes dataset Figure 60 displays the results achieved. The model correctly predicted all of the 35 images, achieving 100% in all metrics. Figure 60: Confusion matrix of the linear classifier on the boxes dataset This dataset was also trained in a fully supervised manner where 10 models were trained with different hyperparameters to obtain the best possible model. For a fair comparison, the dataset was also divided into 70% for training and 30% for validation. Figure 61 shows the results obtained. The supervised model misclassified 8 images, achieving an accuracy of 81.2% and an MCC of 54.5%. The reason for these results is that the dataset 67 contains very few images for the model to learn. In this specific case and with these settings, the information gained on the bookcase dataset was useful to improve performance compared to the supervised model. Figure 61: Confusion matrix of the model in a supervised seeting However, when given new images the model did not classify them all correctly. Since the new images are a bit different from the images the model learned from (Example in Figure 62), this shows that the model did not generalize well. This also happens due to the dataset being small and very specific. During SwAV training on the bookcase dataset, certain features were learned. The linear classifier grabbed those features and tried to fit them into the two classes of the boxes dataset. However, the meaningful features of the bookcase problem can bring the model to associate the new images to incorrect classes when regarding the new problem. This happens because the learned features are not adequate for the new problem. One way to achieve better results would be to acquire more images with different noise and also train SwAV with this dataset. SwAV would extract specific features from this dataset which would improve performance. Figure 62: New image with different angle and background noise. The model predicted it incorrectly as 0 68 6 Conclusions and future work 6.1 Conclusion The main goal of this work was to explore and research this new type of learning: self-supervised learning and to develop a classification model in a self-supervised manner that would classify if a shelf is empty or not in an industrial space context. After an investigation of the available methods, SwAV was chosen. In five experiments, SwAV was trained with different parameters using the flower dataset from Tensorflow. After training, the models were trained with a linear classifier based on the frozen final representations of the ResNet-50 trained with SwAV. The results were analyzed using various metrics. Based on these results, SwAV was then trained on the custom bookcase dataset with the same settings and achieved an accuracy of 98.8% and an MCC of 98% for the linear classifier. The SSL setting achieved 98.8% in much less time compared to the SL setting which achieved 99,6%. However, SwAV and other SSL methods require a lot of computational power which can be a limitation. Depending on the problem and resources, SSL can be an option for being faster and getting close to fully supervised learning performance. Regarding transfer learning to downstream tasks the results were not as expected. Despite the self-supervised method improving performance compared to the supervised model, the model did not generalize well. Given new images the model failed to predict them correctly, showing that the features learned from the bookcase model were not adequate to solve the new problem with the boxes dataset. This is not surprising because SwAV was trained on a very specific dataset with only two classes. It does not extract the same features as if it was trained on, for example, ImageNet that has over a million images and 1000 classes. In general, SwAV extracts useful features that can add information to solve other tasks. However, this depends on the datasets, the task, and the features learned during SwAV training. 6.2 Future Work In terms of classifying whether a shelf is empty or not in an industrial setting, the aim of future work is to capture real images in an industrial environment in order to create a new dataset and test the developed model. Further, train SwAV with the boxes dataset and even with both bookcase and boxes datasets and evaluate. As for SwAV in particular, the experiments conducted with the other backbones should be optimized and the results compared. Also, with more resources, experiment F should be trained with a higher number of epochs and evaluated. Another goal would be to implement other recent methods of self-supervised learning methods such as BYOL , ReLICv2 , and MoCov3 , which should also be tested in an industrial context. 69 References [1] T. Darrell R. Girshick, J. Donahue and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. arXiv , 2014. [2] I. Sutskever A. Krizhevsky and G.E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems , volume 25, page 1097–1105, 2012. [3] I. Kokkinos K. Murphy L.C. Chen, G. Papandreou and A.L. Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. arXiv , 2017. [4] X. Zhai A. Kolesnikov and L. Beyer. Revisiting self-supervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 1920–1929, 2019. [5] A. Gupta P. Goyal, D. Mahajan and I. Misra. Scaling and benchmarking self-supervised visual representation learning. In Proceedings of the IEEE/CVF International Conference on computer vision , pages 6391–6400, 2019. [6] Z. Hou L. Mian Z. Wang J. Zhang X. Liu, F. Zhang and J. Tang. Self-supervised learning: Generative or contrastive. IEEE Transactions on Knowledge and Data Engineering , 2021. [7] M.Z. Zadeh D. Banerjee A. Jaiswal, A.R. Babu and F. Makedon. A survey on contrastive self-supervised learning. Technologies , 9(1):2, 2020. [8] K. Ohri and M. Kumar. Review on self-supervised image recognition using deep neural networks. Knowledge-Based Systems , 224:107090, 2021. [9] S. Russell and P. Norvig. Artificial Intelligence: A Modern Approach , chapter 1. Pearson, fourth edition, 2021. [10] S. Shah S. Khan, H. Rahmani and M. Bennamoun. A Guide to Convolutional Neural Networks for Computer Vision , chapter 1,3,4,5,6. Morgan and Claypool publishers, 2018. [11] Data definition. https://dictionary.cambridge.org/pt/dicionario/ingles/data. Last accessed: 05-09-22. [12] S. Sah. Machine learning: A review of learning types. preprints , 2020. [13] A. Burkov. The Hundred Page Machine Learning Book , chapter 1,5. 2019. 70 [14] Y. Zhang. New Advances in Machine Learning , chapter 3. In-Tech, 2010. [15] I. Sarker. Machine learning: Algorithms, real world applications and research directions. Springer Nature , 2021. [16] R. C. Tryon. Genetic differences in maze-learning ability in rats. Yearbook of the National Society for Studies in Education , pages 111–119, 1940. [17] X. Zhu and A. Goldberg. Introduction to Semi-Supervised Learning , page Abstract. Morgan and Claypool publishers, 2009. [18] M. Sober J. Benedito E. Olivas, J. Guerrero and A. López. Handbook of Research on Machine Learning Applications and Trends: Algorithms, Methods, and Techniques , chapter 11. Information Science Reference, 2009. [19] M.Kamber J.Han and J.Pei. Data Mining: Concepts and Techniques , chapter 3, 8.5. Elsevier, third edition, 2012. [20] D. Bourke A. Neagoie and Zero to Mastery. Tensorflow developer certificate in 2022: Zero to mastery. https://www.udemy.com/course/ tensorflow-developer-certificate-machine-learning-zero-to-mastery/, 2022. Appendix: Machine Learning and Data Science Framework (https://www.mrdbourke.com/ a-6-step-field-guide-for-building-machine-learning-projects/) Last accessed 05-09-22. [21] J. Brownlee. Master Machine Learning Algorithms - Discover how they work and Implement Them From Scratch , chapter 7. Machine Learning Mastery, 2016. [22] J. Brownlee. Machine Learning Mastery with Python: Understand Your Data, Create Accurate Models and Work Projects End-To-End , chapter 10.3. Machine Learning Mastery, 2016. [23] Rok Blagus and Lara Lusa. Class prediction for high-dimensional class-imbalanced data. BMC Bioinformatics , 11(1):523, 2010. [24] Davide Chicco and Giuseppe Jurman. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genomics , 21(1):6, 2020. [25] L. Deng and D. Yu. Deep Learning: Methods and Applications , chapter 1. Now Publishers Inc, 2014. 71 [26] D. Graupe. Principles Of Artificial Neural Networks: Basic Designs To Deep Learning , chapter 2,4. World Scientific Publishing Company, 3 edition, 2013. [27] M.B. Khan M. Mohammed and E.B.M Bashier. Machine Learning: Algorithms and Applications , chapter 6. CRC Press, 2017. [28] A. Karpatne P.N. Tan, M. Steinbach and V. Kumar. Introduction to Data Mining , chapter 6. Pearson Education, 2 edition, 2019. [29] J.D. Kelleher. Deep Learning , chapter 3, 5. MIT Press, 2019. [30] S. Shalev-Shwartz and S. Ben-David. Understanding Machine Learning: From Theory to Algorithms , chapter 20. Cambridge University Press, 2014. [31] F. Rosenblatt. The perceptron: a probabilistic model for information storage and organization in the brain. Psychological review , 65(6):386, 1958. [32] O. Campesato. Artificial Intelligence, Machine Learning and Deep Learning , chapter 4. Mercury Learning & Information, 2020. [33] L. Fausett. Fundamentals of neural networks: architectures, algorithms, and applications , chapter 1. Prentice-Hall, 1994. [34] U. Zahoora A. Khan, A. Sohail and A.S. Qureshi. A survey of the recent architectures of deep convolutional neural networks. volume 53, page 5455–5516. Kluwer Academic Publishers, 2020. [35] ImageNet Large Scale Visual Recognition Competition (ILSVRC). https://www.image-net.org/ challenges/LSVRC/index.php. Last accessed: 31-10-22. [36] S. Albawi. T.A. Mohammed and S. Al-Zawi. Understanding of a convolutional neural network. In 2017 International Conference on Engineering and Technology (ICET) , pages 1–6, 2017. [37] K. O’Shea and R. Nash. An introduction to convolutional neural networks. 2015. [38] J. Brownlee. Deep Learning for Computer Vision - Image Classification, Object Detection and Face Recognition in Python , chapter 1, 11, 12, 18. Machine Learning Mastery, 2019. [39] Y. Bengio I. Goodfellow and A. Courville. Deep Learning , chapter 6,7,8,9. MIT Press, 2016. 72 [108] W. Falcon and K. Cho. A framework for contrastive self-supervised learning and designing a new approach. arXiv preprint arXiv:2009.00104 , 2020. [109] X. Wang and G.-J. Qi. Contrastive learning with stronger augmentations. arXiv preprint arXiv:2104.07713 , 2021. [110] M. Norouzi T. Chen, S. Kornblith and G. Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning , pages 1597–1607. PMLR, 2020. [111] M. Gutmann and A. Hyvärinen. Noise-contrastive estimation: A new estimation principle for unnormalized statistical models. In Proceedings of the thirteenth international conference on artificial intelligence and statistics , pages 297–304. JMLR Workshop and Conference Proceedings, 2010. [112] M. Schultz and T. Joachims. Learning a distance metric from relative comparisons. Advances in neural information processing systems , 16, 2003. [113] Li Y. Oord, A. v.d and O. Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748 , 2018. [114] K. Sohn. Improved deep metric learning with multi-class n-pair loss objective. Advances in neural information processing systems , 29, 2016. [115] A. Gupta D. Ghadiyaram X. Yan, I. Misra and D. Mahajan. Clusterfit: Improving generalization of visual representations, 2019. [116] X. Chen and K. He. Exploring simple siamese representation learning, 2020. [117] I. Misra and L. van der Maaten. Self-supervised learning of pretext-invariant representations, 2019. [118] Y. Wu S. Xie K. He, H. Fan and R. Girshick. Momentum contrast for unsupervised visual representation learning, 2019. [119] F. Altché C. Tallec P.H. Richemond E. Buchatskaya C. Doersch B.A. Pires Z.D. Guo M.G. Azar B. Piot K. Kavukcuoglu R. Munos J-B. Grill, F. Strub and M. Valko. Bootstrap your own latent: A new approach to self-supervised learning, 2020. [120] M. Cuturi. Sinkhorn distances: Lightspeed computation of optimal transport. Advances in Neural Information Processing Systems , 26, 2013. 79 [121] How to use kaggle. https://www.kaggle.com/docs/notebooks, 2022. [122] Tensorflow. https://www.tensorflow.org/. Last accessed: 22-01-23. [123] L. Biewald. Experiment tracking with weights and biases. https://www.wandb.com/, 2020. Software available from wandb.com. [124] The TensorFlow Team. Flowers. http://download.tensorflow.org/example_images/ flower_photos.tgz, 2019. [125] J. Mairal P. Goyal P. Bojanowski M. Caron, I. Misra and A. Joulin. Unsupervised learning of visual features by contrasting cluster assignments. https://github.com/facebookresearch/swav, 2021. [126] SwAV-TF. A. thakur and s. paul. https://github.com/ayulockin/SwAV-TF, 2022. 80