scieee AI-readable full text Open interactive document viewer

Analysis of traffic signs and traffic lights for autonomously driven vehicles

Marques, Rafael Meleiro

Abstract

Os Advanced Driver Assistance Systems(ADAS) estão relacionados a vários sistemas nos veículos que se destinam a melhorar a segurança do tráfego rodoviário, ajudando os condutores a terem melhor consciência da estrada e dos seus perigos inerentes, bem como de outros motoristas nas proximidades. A deteção e o reconhecimento de sinais de trânsito são parte integrante do ADAS. Os sinais de trânsito fornecem informações sobre as regras de trânsito, condições das estradas, direções de rotas e auxiliam os motoristas para uma condução segura. Isto garante que o limite de velocidade atual e outros sinais de trânsito sejam exibidos para o motorista continuamente. O projeto de reconhecimento de sinais de trânsito tem sido um problema desafiador por muitos anos e, portanto, tornou-se um tópico de pesquisa importante e ativo na área de sistemas de transporte inteligentes. Esta tecnologia está a ser desenvolvida por uma variedade de fornecedores automóveis. Esta pode usar técnicas de processamento de imagem para detetar sinais de trânsito. Uma abordagem do problema de deteção e reconhecimento de sinais/semáforos usando Su pervised Learning é apresentada nesta dissertação para dois cenários diferentes. Nesta disser tação, para cada objetivo, são apresentadas duas abordagens diferentes de Supervised Learning, bem como um estudo estendido dos hiperparâmetros. Para os dois primeiros objetivos, foram de senvolvidas as abordagens para um robô produzido pela equipa de Condução Autónoma do Labo ratório de Automação e Robótica. O robô deve detetar e classificar corretamente o sinal/semáforo de trânsito apresentado. O terceiro objetivo é para a via pública e a abordagem desenvolvida deve detetar e classificar corretamente o sinal/semáforo de trânsito apresentado no conjunto de dados restritos.

Full text

Universidade do Minho Escola de Engenharia Rafael Meleiro Marques Analysis of Traffic Signs and Traffic Lights for Autonomously Driven Vehicles November 2021 UMinho | 2021 Rafael Marques Analysis of Traffic Signs and Traffic Lights for Autonomously Driven Vehicles Rafael Meleiro Marques Analysis of Traffic Signs and Traffic Lights for Autonomously Driven Vehicles Guimarães, November 2021 Rafael Meleiro Marques Analysis of Traffic Signs and Traffic Lights for Autonomously Driven Vehicles Dissertação de Mestrado Mestrado Integrado em Engenharia Eletrónica Industrial e Computadores Trabalho efetuado sob a orientação do Professor Doutor Fernando Ribeiro Guimarães, November 2021 DIREITOS DE AUTOR E CONDIÇÕES DE UTILIZAÇÃO DO TRABALHO POR TERCEIROS Este é um trabalho académico que pode ser utilizado por terceiros desde que respeitadas as regras e boas práticas internacionalmente aceites, no que concerne aos direitos de autor e direitos conexos. Assim, o presente trabalho pode ser utilizado nos termos previstos na licença abaixo indicada. Caso o utilizador necessite de permissão para poder fazer um uso do trabalho em condições não previstas no licenciamento indicado, deverá contactar o autor, através do RepositóriUM da Universidade do Minho. Licença concedida aos utilizadores deste trabalho Atribuição-CompartilhaIgual CC BY-SA https://creativecommons.org/licenses/by-sa/4.0/ Acknowledgements Throughout my academic journey, several people contributed with essential support that culminated in the completion of this dissertation. A very special thanks to my parents Manuel Marques and Maria Marques for all the constant support, love and belief in my capabilities. Without their constant support, this journey wouldn’t be achievable. They encouraged me throughout my worst moments and were always there when I needed them. A big thanks to my supervisor, Professor Fernando Ribeiro, for the excellent opportunity, the constant guidance and the availability demonstrated through this dissertation. To Tiago Ribeiro who guided me in all the stages of this dissertation in the correct direction. To the Laboratório de Automação e Robótica members for the great work environment and the shared knowledge. To my close friends André Gonçalves, Bruno Sousa and Rafael Araújo, whom I was able to share unforgettable memories throughout this journey and in their own way supported and encouraged me to reach this goal. Also to my classmates, who helped me all these years. A special thanks to Mariana Sarmento for the constant support and continuous encouragement that were essential for the development of this dissertation. I am thankful for those who directly or indirectly had contributed to this dissertation. iv STATEMENT OF INTEGRITY I hereby declare having conducted this academic work with integrity. I confirm that I have not used plagiarism or any form of undue use of information or falsification of results along the process leading to its elaboration. I further declare that I have fully acknowledged the Code of Ethical Conduct of the University of Minho. v Resumo Os Advanced Driver Assistance Systems (ADAS) estão relacionados a vários sistemas nos veículos que se destinam a melhorar a segurança do tráfego rodoviário, ajudando os condutores a terem melhor consciência da estrada e dos seus perigos inerentes, bem como de outros motoristas nas proximidades. A deteção e o reconhecimento de sinais de trânsito são parte integrante do ADAS. Os sinais de trânsito fornecem informações sobre as regras de trânsito, condições das estradas, direções de rotas e auxiliam os motoristas para uma condução segura. Isto garante que o limite de velocidade atual e outros sinais de trânsito sejam exibidos para o motorista continuamente. O projeto de reconhecimento de sinais de trânsito tem sido um problema desafiador por muitos anos e, portanto, tornou-se um tópico de pesquisa importante e ativo na área de sistemas de transporte inteligentes. Esta tecnologia está a ser desenvolvida por uma variedade de fornecedores automóveis. Esta pode usar técnicas de processamento de imagem para detetar sinais de trânsito. Uma abordagem do problema de deteção e reconhecimento de sinais/semáforos usando Supervised Learning é apresentada nesta dissertação para dois cenários diferentes. Nesta dissertação, para cada objetivo, são apresentadas duas abordagens diferentes de Supervised Learning , bem como um estudo estendido dos hiperparâmetros. Para os dois primeiros objetivos, foram desenvolvidas as abordagens para um robô produzido pela equipa de Condução Autónoma do Laboratório de Automação e Robótica. O robô deve detetar e classificar corretamente o sinal/semáforo de trânsito apresentado. O terceiro objetivo é para a via pública e a abordagem desenvolvida deve detetar e classificar corretamente o sinal/semáforo de trânsito apresentado no conjunto de dados restritos. Palavras-chave: Supervised Learning, RoboCup, YOLOV3 , Detecção de sinais de trânsito, Robô móvel autónomo, Robótica, Robô simulado. vi Abstract Advanced Driver Assistance Systems (ADAS) relate to various in-vehicle systems that are intended to improve road traffic safety by supporting and improve drivers awareness of the road and its dangers as well as other drivers in the vicinity. Traffic sign detection and recognition is part of ADAS. Traffic signs give knowledge about the traffic rules, road conditions, route directions and assist drivers for safe driving. This ensures that the current speed limit and other road signs are displayed to the driver on an ongoing basis. The design of traffic sign recognition has been a challenging problem for many years and therefore became an important and active research topic in the area of intelligent transport systems. This technology is being developed by a variety of automotive suppliers. Typically it uses classical image processing techniques to detect traffic signs. An approach to the problem of Traffic Sign/Light detection and recognition using Supervised Learning is presented in this dissertation for two different scenarios. Two different Supervised Learning approaches are presented for each objective as well as an extended hyperparameter study. For the first two objectives, the approaches were developed for a robot produced by the Autonomous Driving team from the Laboratório de Automação e Robótica fom University of Minho. The robot must correctly detect and classify the presented Traffic Sign/Light. The third objective is for the public road and the developed approach must correctly detect and classify the presented Traffic Sign/Light in the restrained dataset. Keywords: Supervised Learning, Traffic Sign Detection, Autonomous Mobile Robot, Robotics, Simulated Robot, RoboCup, YOLOV3 vii Table of Contents Resumo vi Abstract vii Table of Contents viii List of Figures xi List of Tables xvii List of Listings xviii Acronyms xviii 1 Introduction 1 1.1 Motivation ................................. 2 1.2 Objectives ................................. 3 1.3 Dissertation Structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 2 Literature Review 5 2.1 Artificial Intelligence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 2.2 MachineLearning.............................. 6 2.3 DeepLearning ............................... 7 2.4 Convolutional Neural Networks . . . . . . . . . . . . . . . . . . . . . . . 10 2.4.1 Convolutional Layer . . . . . . . . . . . . . . . . . . . . . . . . 11 2.4.2 PoolingLayer............................ 13 2.4.3 Fully-connected Layer . . . . . . . . . . . . . . . . . . . . . . . . 14 2.5 Theoretical foundations . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 2.5.1 Supervised and unsupervised learning . . . . . . . . . . . . . . . 14 viii 5.1 YOLOV3 mAP for the Autonomous Driving Competition of the RoboCup Portuguese Open in simulation . . . . . . . . . . . . . . . . . . . . . . . . . 108 5.2 YOLOV3 loss for the Autonomous Driving Competition of the RoboCup Portuguese Open in simulation . . . . . . . . . . . . . . . . . . . . . . . . . 108 5.3 YOLOV3_tiny mAP for the Autonomous Driving Competition of the RoboCup Portuguese Open in simulation . . . . . . . . . . . . . . . . . . . . . . . 109 5.4 YOLOV3_tiny loss for the Autonomous Driving Competition of the RoboCup Portuguese Open in simulation . . . . . . . . . . . . . . . . . . . . . . . . . 109 5.5 YOLOV3 mAP for the Autonomous Driving Competition of the RoboCup PortugueseOpen................................ 110 5.6 YOLOV3 loss for the Autonomous Driving Competition of the RoboCup PortugueseOpen................................ 110 5.7 YOLOV3_tiny mAP for the Autonomous Driving Competition of the RoboCup PortugueseOpen.............................. 111 5.8 YOLOV3_tiny loss for the Autonomous Driving Competition of the RoboCup Portuguese Open in simulation . . . . . . . . . . . . . . . . . . . . . . . . . 111 5.9 YOLOV3 mAP for the Public Road . . . . . . . . . . . . . . . . . . . . . . 112 5.10 YOLOV3 loss for the Public Road . . . . . . . . . . . . . . . . . . . . . . 112 5.11 YOLOV3_tiny mAP for the Public Road . . . . . . . . . . . . . . . . . . . . 113 5.12 YOLOV3_tiny loss for the Public Road . . . . . . . . . . . . . . . . . . . . 113 5.13 Two frames with high confidence detection from the YOLOV3 network . . . . . 115 5.14 Two frames with high confidence detection from the YOLOV3_tiny network . . 115 5.15 Frame where the robot reacted to the Left Obligation sign . . . . . . . . . . 116 5.16 Frame where the robot reacted to the Left Obligation sign . . . . . . . . . . 117 5.17 Frame where the robot reacted to the 30km/h sign . . . . . . . . . . . . . 117 5.18 Frame where the robot reacted to the STOP sign . . . . . . . . . . . . . . . 118 5.19 Detection from the YOLOV3 (left) and YOLOV3_tiny (right) . . . . . . . . . . 119 5.20 Detection distance from the YOLOV3 (left) and YOLOV3_tiny (right) . . . . . . 119 5.21 Bounding Boxes from the YOLOV3 (left) and YOLOV3_tiny (right) . . . . . . . 120 5.22 Detection with less stability from the YOLOV3 (left) and YOLOV3_tiny (right) . . 120 5.23 Detection with extreme distortion from the YOLOV3 (left) and YOLOV3_tiny (right) 121 xv 5.24 Detection with ideal conditions from the YOLOV3 (left) and YOLOV3_tiny (right) 122 5.25 Detection with rain from the YOLOV3 (left) and YOLOV3_tiny (right) . . . . . . 122 5.26 Detection of a sign not facing the camera from the YOLOV3 (left) and YOLOV3_tiny (right) ................................... 123 5.27 Traffic Light detection from the YOLOV3 (left) and YOLOV3_tiny (right) . . . . 123 5.28 Detection of an unknown sign from the YOLOV3 (left) and YOLOV3_tiny (right) 124 5.29 Results of the YOLOV3 network in poor conditions . . . . . . . . . . . . . . 124 5.30 Results of the YOLOV3_tiny network in poor conditions . . . . . . . . . . . . 125 xvi List of Tables 2.1 ActivationFunction ............................. 10 2.2 Correspondence between each of the functions and the information (table adapted from [8] with permission from Manuel Silva) . . . . . . . . . . . . . . . . . 23 2.3 Performace of Real-Time Detectors . . . . . . . . . . . . . . . . . . . . . 26 2.4 The path from YOLO to YOLOv2 (table adapted from [12] with permission from AliFarhadi)................................. 28 2.5 Darknet-19 (table adapted from [12] with permission from Ali Farhadi) . . . . 29 2.6 Darknet-53 (figure from [10] displayed with permission from Ali Farhadi) . . . 30 3.1 The Traffic Signs/Lights in the Dataset . . . . . . . . . . . . . . . . . . . . 41 3.2 The Traffic Signs/Lights in the Dataset . . . . . . . . . . . . . . . . . . . . 42 3.3 Number of images per Traffic Sign/Light . . . . . . . . . . . . . . . . . . . 50 3.4 Number of images per Traffic Sign/Light . . . . . . . . . . . . . . . . . . . 70 3.5 The Traffic Signs/Lights in the Road Dataset . . . . . . . . . . . . . . . . . 78 3.6 Number of images per Traffic Sign/Light . . . . . . . . . . . . . . . . . . . 80 4.1 The hyperparameters and the corresponding tested values . . . . . . . . . . 87 4.2 Time each train lasted . . . . . . . . . . . . . . . . . . . . . . . . . . . 88 4.3 Time each train lasted . . . . . . . . . . . . . . . . . . . . . . . . . . . 93 4.4 Time each train lasted . . . . . . . . . . . . . . . . . . . . . . . . . . . 97 4.5 The hyperparameters and the corresponding tested values . . . . . . . . . . 97 4.6 Time each train lasted . . . . . . . . . . . . . . . . . . . . . . . . . . . 99 4.7 Time each train lasted . . . . . . . . . . . . . . . . . . . . . . . . . . . 103 4.8 Time each train lasted . . . . . . . . . . . . . . . . . . . . . . . . . . . 105 4.9 Optimized hyperparameters for the YOLOV3 network . . . . . . . . . . . . . 106 4.10 Optimized hyperparameters for the YOLOV3_tiny network . . . . . . . . . . 106 xvii Acronyms ADAS - Advanced Driver Assistance Systems AI - Artificial Intelligence API - Application Programming Interface ANN - Artificial Neural Networks CNN - Convolutional Neural Network CUDA - Compute Unified Device Architecture DNN - Deep Neural Networks FP - False Positives FN - False Negatives FPS - Frames Per Second GPU - Graphics Processing Unit IDE - Integrated Development Environment IoU - Intersection over Union mAP - Mean Average Precision NN - Neural Network OpenCV - Open Source Computer Vision Library RGB - Red Green Blue ROI - Region Of Interest ROS - Robot Operating System TP - True Positives TSR - Traffic Sign Recognition TS/L - Traffic Signs/Lights URDF - Unified Robot Description Format YOLO - You Only Look Once xviii Chapter 1 Introduction The world is under continuous development and with this, vehicles have become an essential means of transportation for people every day. This advancement in society has produced many traffic safety predicaments, like traffic congestion and regular road accidents, which are mainly caused by subjective reasons related to the driver. These are mostly a result of negligence, inappropriate driving manoeuvre and disregarding traffic signs [13]. The previously mentioned human factors can be evaded using self-driving technology which can support or control the driving operation. This can reduce the occurrence of accidents [14]. Traffic sign detection and recognition is essential in the evolution of intelligent vehicles. This dissertation presents the development of traffic sign detection and recognition using Machine Learning (ML). ML provides systems with the ability to automatically learn and improve from experience without being explicitly programmed. It examines the development and research of algorithms that can learn from and make predictions on data. These algorithms operate by building a model from inputs in order to make data-driven predictions rather than following strict guidelines [15]. ML is used to provide knowledge and understanding of the environment around the vehicle. This mainly involves the use of camera-based systems to detect and classify objects but can also be based on LiDAR or radar. The main concern around this technology is the possibility of a misclassification of an object that can be caused by just a few pixels of difference in an image produced by a camera system. This can lead to an incorrect action from the car causing a road accident. Using enhanced and more generalized training of the ML models can lead to an improvement in the network’s accuracy to correctly detect the object. Giving the system more diverse inputs on key parameters also improves performance [16]. 1 1.1 Motivation According to [17], the future of the automotive industry consists of connected and autonomous vehicles which will outnumber vehicles controlled by humans. This technological progress will lead to advancements in Advanced Driver Assistance Systems (ADAS). Adjacent to ADAS is Traffic Sign Recognition (TSR), a system that discovers and classifies traffic signs captured by a camera positioned in the vehicle. With this system, self-driving cars have a better perception of the road and its surroundings. Unlike most systems that are contained in ADAS, TSR only requires a camera and the corresponding ML software. The software associated with TSR demands cutting edge hardware to perform properly. One option is NVIDIA’s General Purpose Graphics Processing Unit (GPGPU), the DRIVE PX which is an Artificial Intelligence (AI) car computer that empowers vehicle manufacturers to stimulate the production of automated and autonomous vehicles [18]. It can facilitate real-time heavy machine learning software to run in the vehicle, supporting more accurate TSR systems. TSR systems face two major technical challenges [19]: •The broad disparity in illumination: two elements can contribute to this disparity, the time of day in which the incidence of sunlight will change and be reduced by nighttime, and the weather in which it will fluctuate depending on the presented conditions; •Low image resolution when the vehicle is at high speeds: the frames light from the camera and the other vehicles captures at high speeds are more blurred and the resolution affected. To overcome these challenges pre-processing methods are proposed. The TSR system is usually divided into two phases, the first one has the goal of discovering the Region of Interest (ROI) where the traffic sign is located in the input frame, the second one has the aim of classifying the sign in the ROI. The first phase of the TSR system can consist of colour or shape detection but have limited accuracy and can be affected by illumination. Another option is using Haar-like features [20]. For the second phase methods Convolutional Neural Network (CNN) can be used. 2 1.2 Objectives For autonomous driving vehicles, like the ones developed by Laboratório de Automação e Robótica to be able to move in a highly dynamic and regulated environment such as vehicle traffic, it is necessary to develop a strategy that complies with the Traffic Signaling Regulation. For this, it is necessary to develop a identification and classification system for traffic signs (for vertical and light signs). The developed algorithm will be based on Supervised Learning and Deep Learning (DL) strategies, where it has to search in non-standardized images, the occurrence of one or several signals and their respective classification. In the first instance, the algorithm will be tested in two environments: 1. the RoboCup Portuguese Open – Autonomous Driving competition; 2. an autonomous driving environment on real road and vehicle traffic. The first environment refers to the Autonomous Driving test of the RoboCup Portuguese Open, represented in figure 1.1. This event is characterized by creating a reduced and more controlled version of a two-lane road, crosswalks, obstacles, traffic lights, traffic signals, tunnels, parking spaces and construction sites with temporary signage. The robot has to be completely autonomous and respect all traffic rules in a more controlled version of public roads. In this track, the vertical traffic signs are the same as those regulated in the highway code, while the traffic light has some variations, as it also indicates the direction that the vehicle has to go. Before testing in the event a simulation must be developed that reproduces this competition. Participation in this event will serve as a benchmark of the developed algorithm as well as its application in a real environment. Figure 1.1: Autonomous Driving Test Track (left) and An example of autonomous driving vehicles developed by Laboratório de Automação e Robótica (right) 3 The second test environment consists of driving a real vehicle on a public road, with cameras mounted on the vehicle running the algorithm in real-time identifying and classifying vertical and light signals. 1.3 Dissertation Structure To better understand and describe the essential steps and features to the progression of the project, this dissertation is divided into five chapters. The first one is the Introduction chapter, where the motivation and objectives are described, followed by the Literature Review, where a study of Machine Learning was performed to understand its specificities and a state of the art is present and analysed, in which the used architecture is deepened. The third chapter, System Specification is the fundamental chapter of this dissertation since all the used methodologies to solve the objectives are explained. In this chapter all the used software, environments and procedures are disclosed. The following chapter, Tests, firstly presents a study to uncover the most optimized hyperparameters for the chosen networks. The next chapter, Results, describes the results for each network of each objective. Finally, in the Conclusions chapter, the core outcomes are mentioned as well as further work. 4 Chapter 2 Literature Review This chapter provides essential context about Artificial Intelligence (Section 2.1) and its subfields Machine Learning (Section 2.2) and Deep Learning (Section 2.3), as shown in figure 2.1. The concepts regarding the structure and layers of Convolutional Neural Networks are deepened (Section 2.4) and theoretical concepts will be presented (Section 2.5). A context is presented concerning Traffic Sign Recognition, as well as the levels of driving automation and examples of how recognition is carried out in car companies currently (Section 2.6). One of the objectives of the project is to participate in the RoboCup Portuguese Open Autonomous Driving Competition so its rules and information will be presented in Section 2.7. In Section 2.8, the You Only Look Once (YOLO) model is explained and an implementation of Traffic Sign Recognition for Autonomous Driving Robot is presented. Figure 2.1: Artificial intelligence, Machine Learning, and Deep Learning 5 2.1 Artificial Intelligence Artificial Intelligence (AI) is the development of computer systems that are able to perform tasks that would require human intelligence. Defined by John McCarthy in 1956 as ”science engineering of making intelligent machines”, AI is a branch in computer science that studies and designs intelligent agents that recognise its context and takes actions which maximize its chances of success [21]. AI should possess the following traits [22]: •Should be able to predict and adapt: AI uses multiple algorithms from huge amounts of data to discover patterns; •To make decisions on its own: It can augment human intelligence, deliver insights and enhance productivity; •It is continuous learning: in order to construct analytical models AI uses algorithms. From countless rounds of trial and error AI finds out how to perform tasks; •It is forward-looking: AI allows people to reevaluate how data is analyzed and information integrated, and then use these insights to make more reliable judgments; •It is capable of motion and perception. 2.2 Machine Learning In 1959, Arthur Samuel defined Machine Learning (ML) as a “Field of study that gives computers the ability to learn without being explicitly programmed” [23]. A concise definition of the field would be the effort to automate intellectual tasks normally performed by humans [24]. ML is a form of AI that enables a system to learn from data rather than through specific programming, it emerged from the study of pattern recognition and computational learning theory. As seen in figure 2.2, this form of AI is trained rather than explicitly programmed. 6 In examples in which the input is an image composed by Red Green Blue (RGB) colour spectrum, the kernel has the same depth, that is, the same number of channels as the input matrix. Each kernel channel does its multiplication process in its respective input matrix channel, then the results are added to the bias that results in a value that will fill a cell in the output matrix, that is, in the feature map. Thus, at the end of the process, it results in a compressed depth feature map [6]. Figure 2.11: Convolution operation (figure from [5] displayed with permission from Arden Dertat) 2.4.2 Pooling Layer The pooling layer is responsible for reducing the dimensions of the feature maps, thereby reducing computational costs and the complexity of the model. Thus allowing to better control the phenomenon of overfitting [6]. This layer also uses filters, similar to the convolutional layer, but instead of performing the weighted sum, it performs arithmetic functions in all previous activation maps [32]. The most common methods of the pooling layer are maximum pooling (Figure 2.12), which looks for the maximum value of all pixels within the pool, and average pooling, in which the average value is calculated [33]. 13 Figure 2.12: MaxPooling example (figure from [6] displayed with permission from Shadman Sakib) 2.4.3 Fully-connected Layer In a basic CNN model, the characteristics generated by the last convolution layer represent a portion of the input image since its receptive field does not integrate the entire spatial dimension of the image. Each node in a fully-connected layer is directly connected to every node in both the previous and in the next layer [6] as shown in figure 2.13. Figure 2.13: Fully-Connected Layers (figure from [6] displayed with permission from Shadman Sakib) 2.5 Theoretical foundations 2.5.1 Supervised and unsupervised learning ML is a vast field with a complex subfield taxonomy. This algorithms can be supervised and unsupervised learning based on the way the data is fed to. 14 Supervised learning is the most common type of learning. It is based on training data with correct classification attached to approximate the mapping function [34]. Supervised learning problems can be further grouped into classification or regression problems. Unlike the previous type of learning, Unsupervised Learning has the purpose to model the underlying structure or distribution in the data in order to learn more about the data. It is applied in problems where it is either impossible or unrealistic for a human to propose patterns in the data [35]. 2.5.2 Classification and regression Two major prediction problems which are usually dealt with in ML are Classification and Regression. Classification is the process of determining a model that distributes data into various classes. Data is designated under distinct labels conforming to the parameter given as an input. This process deals with problems where the data can be sorted into binary or multiple discrete labels [36]. Regression is the process of determining a model that detects the data into continuous real values and can also recognize the distribution movement depending on the historical data [36]. 2.5.3 Intersection over Union Used to measure the accuracy of an object detector, Intersection over Union (IoU) is an evaluation metric. If an algorithm produces predicted bounding boxes as output, IoU can judge it [37]. Concerning the evaluation using IoU in object detection, two inputs are required: •The predicted bounding boxes of the model; •The ground-truth bounding boxes (where the object in the image is specified) IoU =Area_of_Overlap Area_of_Union (2.1) 15 In equation 2.1, the numerator is the area of overlap connecting both inputs and the denominator is the area of union among both inputs, both represented in Figure 2.14. The IoU is a ratio that symbolizes how much the predicted box is confined by the ground truth [32]. Figure 2.14: Area of Overlap and Area of Union 2.5.4 Average Precision Mean Average Precision (mAP) is an object detection evaluation metric used mainly in computer vision. It assesses object detection systems, such as YOLO [9] [12] [10] and Faster R-CNN [38]. mAP gives developers a view of the way the model is performing versus other models on the same test dataset or if improvements were made when advancements were implemented. Precision quantifies how accurate is the prediction and is represented by equation 2.2. Precision =T P TP +FP (2.2) TP = True Positives (Predicted as positive and was correct) FP = False Positives (Predicted as positive but was incorrect) Recall quantifies the correctly detected positives in the input and is represented by equation 2.3. Recall =TP TP +FN (2.3) FN = False Negatives (Failed to predict an object that was there) 16 Precision recall curve allows the visualization of the performance of the model when the confidence threshold of the predictions is decreased. If the model is in a situation where avoiding false positives is more important than dodging false negatives, it can set its confidence threshold higher to encourage the model to only deliver high precision predictions at the risk of lowering its amount of coverage (recall) 2.5.5 Convolutional neural network architectures 2.5.5.1 LesNET-5 Developed by Yann LeCun in 1998, the LesNET-5 [7] is one of the most simple architectures. It was significant for the deep learning for image recognition because it was the first to state that image features are distributed across the image and the best way to extract them was with the employment of convolutions [32]. As demonstrated in figure 2.15, it is composed of: •Input layer: where the image is provided; •Two convolutional layers: which are the result of the convolution applied at the previous matrix; •Subsampling: which results on the Pooling Layer; •Three fully-connected layers: which allow each neuron to receive data from all the neurons in the previous layers; •Output layer: where the classification is conferred. Figure 2.15: LesNet architecture (figure from [7] displayed with permission from Yann LeCun) 17 2.5.5.2 ResNet The ResNet consists of a deep net with 152 layers of depth and acquainted the idea of residual learning to deep neural networks. In the ResNet, each layer of depth is called residual block [39]. (figure 2.16). Figure 2.16: Residual learning: a building block The residual block fundamentally employs convolutional filters to the input x , and the result of those filters is added to the original input. Differentiating it from other architectures, instead of computing the whole transformation of the input, that is, x to F(x), this residual block allows to only compute the variation of the input [39]. This architecture makes backpropagation an easier process [32]. The Resnet architecture won the ImageNet Large Scale Video Recognition Competition (ILSVRC) in 2015. The ILSVRC is an image classification contest, where teams develop neural networks, to classify a set of images in 1000 classes. 2.6 Context Technological advances offer the opportunity to change transportation with the introduction of autonomous vehicles. These features may fluctuate from supporting the driver like adaptive cruise control or crash warning systems to autonomous driving. 18 The Society of Automotive Engineers specifies 6 levels of driving automation ranging from 0 (fully manual) to 5 (fully autonomous) in the context of motor vehicles [40]: •Level 0 - No Driving Automation: The driver provides the dynamic driving tasks. Today, most vehicles on the road are level 0. •Level 1 - Driver Assistance: To assist the driver the vehicle features a single automated system like adaptive cruise control that keeps the car at a safe distance behind the next car. •Level 2 - Partial Driving Automation: Systems with Advanced Driver Assistance Systems or ADAS [41] can avoid accidents that are caused by human error. The vehicle can control accelerating/decelerating and steering but the driver can take command of the car at any time. Teslas Autopilot and General Motors Super Cruise are in this level. •Level 3 – Conditional Driving Automation: The vehicle has environmental detection capabilities and can make knowledgeable decisions but requires the driver to remain alert and ready to take control if the system is incapable to execute the task. •Level 4 - High Driving Automation: These cars don’t require human interaction but the driver can still manually override. They can only operate in limited areas. •Level 5 - Full Driving Automation: In this level, the cars are fully autonomous. They don’t have steering wheels or acceleration/braking pedals and can go anywhere that an experienced human driver can go. In the following sub chapters 2.6.1 and 2.6.2, examples of how recognition is carried out in car companies currently like Tesla and General Motors. 2.6.1 Tesla Led by Elon Musk, Tesla is one of the early pioneers and leaders within the self-driving cars market. Every Tesla car made since October 2016 is equipped with the necessary sensor suite for full self-driving, each of these cars also support our autonomous driving development [42] Tesla’s traffic lights and signs detection is possible by the combination of two modules [43]: 19 •Radars - located in the front bumpers, they can detect cars and objects from a substantial distance and all around; •Cameras - which have an average wide angled camera, are located in the front or on the roof of the car. They can detect various objects such as cars, cyclists, pedestrians and road markings. With the combination of these modules, Mobileye, the manufacturer of Tesla’s processor, was able to deploy the first Digital Neural Network on the road in Tesla’s vehicles. The DNN was trained with cars in the surroundings until it was able to detect them with consistent accuracy and consequently create a 3D model of the same. It is responsible for [43]: •Free Space Pixel Labeling - recognition of the area on-camera with no obstructions in which the car is allowed to perform; •Holistic Path Planning - an attribute that tells the car where to drive with little visual hints; •General Object Detection and Sign Detection - the software can recognize over 250 traffic signs in more than 50 countries which include turn signs, speed limits and traffic lights. It also detects debris and other undesirables situationssuch as potholes on the road. 2.6.2 General Motors Traffic Sign Recognition is a General Motors active safety technology that aids the driver. The software identifies speed limit and many other sorts of signs and displays them on the vehicle’s instrument panel. It also detects LED changing road signs. The current system can recognize signs at a distance of 60 meters. The camera mounted used for the Traffic Sign Recognition, the Opel Eye, is mounted in the front of the vehicle [44]. The Opel Eye camera allows the vehicle to detect the road, vehicles, pedestrians and traffic signs. Continually monitoring the scene around the car this technology processes distinct amounts of data providing the driver with warnings and information [44]. 20 2.7 RoboCup Portuguese Open The RoboCup Portuguese Open promotes Science and Technology among young people, teachers and researchers, through competitions for autonomous robots. Robotic competitions are designed to promote an innovative, entrepreneurial spirit in children and young people through active teaching methods, as well as the acquisition of transversal skills. The 2022 edition is in Santa Maria da Feira from April 27 to May 1 [45]. 2.7.1 Autonomous Driving Competition In this challenge, the robot must be a completely autonomous vehicle in which all decisions must be taken by the systems included in it. Competitors can not include communication between the robot and other external devices. The track is installed within an area of 6.95 x 16.7 m. This has the format of a traffic road. In figure 2.17, a representative view of this route is shown. The track floor is dark and infrared absorbing with lines painted in white [8]. Figure 2.17: Track overview (figure from [8] displayed with permission from Manuel Silva) Two Thin Film Transistor (TFT) panels are mounted right above the zebra crossing, one in each direction of the track. These panels will show indicating signals to the competing vehicles. The base of the panels are 87 cm above the ground as seen in figure 2.18 [8]. 21 Figure 2.18: Signaling panels (figure from [8] displayed with permission from Manuel Silva) 2.7.2 Indicating Signals Six possible indicating signals presented on the TFT are shown in figure 2.19. The symbols are displayed over a black background and bounded by a square box [8]. Figure 2.19: Signaling panel drawings (figure from [8] displayed with permission from Manuel Silva) The purpose of the indicating signals is to conduct the robots trials, giving them orders to give some randomness. The correspondence between each of the functions and the information presented in the indicating panels is represented in the following table [8]: 22 Table 2.5: Darknet-19 (table adapted from [12] with permission from Ali Farhadi) 2.8.2.3 Stronger This method uses labelled images to learn the location and the classification of the objects. In order to improve classification, hierarchical classification is introduced. Using Wordnet [47], a language dataset that arranges concepts and relates them, the network associates labels with their type. For example, ”Norfolk terrier” and ”Yorkshire terrier” are both hyponyms of ”terrier”, which is a type of ”dog”, this creates a structured tree. Figure 2.24 demonstrates the application of a WordTree hierarchy in comparison to ImageNet and COCO label structures [12]. Figure 2.24: Combining datasets using WordTree hierarchy (figure from [10] displayed with permission from Ali Farhadi) 29 2.8.2.4 Yolo9000 Using the previous explained YOLOv2 architecture and combining datasets using WordTree, YOLO9000 was created using 9418 classes. YOLO9000 gets 19.7 mAP overall [12]. 2.8.3 YOLOv3: An Incremental Improvement In 2018 the third iteration of YOLO, YOLOv3 was published [10]. This version is more accurate, features multi-scale detection and a stronger feature extractor network. This iteration can be divided into two elements: Feature Extractor and Detector. Both of these elements work with multiple scales. When an input is presented to the network, it goes through the Feature Extractor first and with the obtained features, the Detector gets the bounding boxes and predicts the correspondent class. 2.8.3.1 Darknet-53 A new network was presented to perform feature extraction, Darknet-53. This is an upgrade from the previous YOLOv2 version, Darknet-19. This network uses successive 3x3 and 1x1 convolutional layers, has shortcut connections and has 53 convolutional layers [10]. The network is represented in the following Table 2.6. Table 2.6: Darknet-53 (figure from [10] displayed with permission from Ali Farhadi) 30 2.8.3.2 Bounding Box Prediction The system uses dimension clusters as anchor boxes to predict bounding boxes [10]. Four coordinates are predicted for each bounding box: tx,ty,tw,th. If an offset is detected from the top left corner of the image (cx,cy) and the bounding box has width and height pw,ph, then the prediction corresponds to, as shown in Figure 2.25: bx=σ(tx) + cx by=σ(ty) + cy bw=pwetw bh=pheth The width and height of the box are predicted as offsets for cluster centroids. The centre coordinates are predicted using a sigmoid function. Figure 2.25: Bounding boxes with dimension priors and location prediction (figure from [10] displayed with permission from Ali Farhadi) Using logistic regression the network predicts an objectness score for each bounding box [10]. 2.8.3.3 Tiny Yolov3 YOLOv3 is one of the fastest deep learning-based object detectors but its computational demand for embedded devices such as the Raspberry Pi, for examples, is exceeding. In order to make 31 implementations in such embedded devices, the YOLO creators introduced Tiny YOLOv3, a variation of the YOLO architecture [48]. This architecture allows a network to be approximately 442% faster than YOLO. This network is comprised of 7 convolutional layers, 6 maxpool layers for image feature extraction and 2 scales of detection layers as demonstrated in Figure 2.26 [49]. Figure 2.26: Structure of Tiny YOLOv3 Since Tiny YOLOv3 is a smaller version, this means that it is unfortunately even less accurate. 2.8.4 Traffic Sign Recognition for Autonomous Driving Robot This project proposes a Traffic Sign Recognition software produced for a robot or a car. The aim of this project was the participation in the Autonomous Driving Competition in the Portuguese Festival of Robotics. The robot acquires the images for the detection and classification of traffic signs and traffic lights with a camera attached to its chassis. The system should be qualified to recognize nine different traffic signs that appear on the side of the track and five traffic lights displayed by a LED panel above the departure area [11]. Some examples are shown in figure 2.27 32 Figure 2.27: Examples of signs to recognize (figure from [11] displayed with permission from Vitor Filipe) The software was divided into three stages [11]: •Detection - using colour segmentation the image regions with signs are identified; •Pictogram extraction - in order to identify the type of signal the system obtains the pictogram; •Classification - a binary pattern is presented to a feed-forward neural network to classify the sign. 2.8.4.1 Traffic Signs Algorithm •Detection Using a digital camera the system acquires frames of the surrounding environment where the traffic signs and traffic lights are present. The frames collected are in the RGB colour space and are converted to HSV colour space in order to make the detection more simplified and robust to diverse illumination condition. The detection is achieved by thresholding the image in the HSV colour space. By thresholding the image, a binary image with several potential regions for sign location is achieved. This location has blue or red as the dominant colour. Every potential region is examined for a minimum and maximum value (region size filtering). Only roughly square regions are further evaluated as signs [11]. In this stage, the pictogram might be missing due to its colour (Figure 2.28). 33 Figure 2.28: a) Frame acquired with a visible traffic sign. b) Regions detected after thresholding. c) Detected external contour (figure from [11] displayed with permission from Vitor Filipe) •Pictogram Extraction Once the detection is performed, the area of the sign in the image is extracted, scaled and converted to grayscale. Two thresholdings are employed to segment the interior of the pictogram, one for dark pictograms and one for white pictograms. A region size filter is employed to eliminate the small noise regions of the resulting binary image [11]. (Figure 2.29) Figure 2.29: a) Sub-image (grayscale) with the sign extrated from the RGB frame. b) Binary image after threshold operation. c) Result of pictogram extraction (figure from [11] displayed with permission from Vitor Filipe) Applying a logical OR operation combines the external contour and the pictogram of the traffic sign and produces a 50x50 pixels binary pattern matrix [11]. (Figure 2.30) 34 Figure 2.30: a) External contour obtained in the detection stage. b) Internal symbol obtained in pictogram extraction stage. c) Binary pattern obtained after logical OR operation (figure from [11] displayed with permission from Vitor Filipe) •Classification The classifier is a feed-forward Neural Network (NN) that distinguishes the sign. The NN has 3 layers [11]: •An input layer with 2500 neurons, one for each 50*50 pixels in the input; •A hidden layer with 20 neurons; •An output layer with 9 possible outputs, one for each sign trained. Using a sigmoid activation function, the NN was trained in supervised mode by the backpropagation learning algorithm. To train the NN a dataset with 3150 real traffic sign images was used. In order to make the NN more robust, these images had different levels of perspective distortion and imprecision in the pictogram [11]. 2.8.4.2 Traffic Lights Algorithm The proposed traffic lights recognition algorithm is identical to the traffic sign algorithm however some steps are not needed because it is less complex. For the detection, thresholding is used to obtain colour segmentation to detect red, yellow and green regions. A conversion to HSV colour space is employed. This stage suppresses the pictogram extraction stage because it locates and obtains the binary pattern matrix. The NN for traffic lights is less complex than the traffic sign because there are fewer outputs and the patterns are simpler. It has the same 2500 neurons, one for each 50*50 pixels in the input but the hidden layer only has 15 neurons with a logistic regression activation function. The output layer is made up of 5 outputs [11]. 35 2.8.4.3 Results To assess the algorithms, tests were performed under the same conditions as the competition. In most signs, 100% precision was obtained in both algorithms. The traffic lights obtained over 96% recall and the traffic signs obtained between 52% and 88.2% recall. Most failures are in detection and pictogram extraction [11]. 36 Chapter 3 System Specification The software were implemented on Ubuntu 20.04 operating system on an ASUS Vivobook Pro N580VD with an Intel Core i7 7th Gen 7700HQ CPU and an Nvidia GeForce GTX 1050. Utilizing Python as the main programming language, the chosen Integrated Development Environment (IDE) was Pycharm Community 2021.1. The library that was essentially used was OpenCV. OpenCV (Open Source Computer Vision Library) is an open-source computer vision and machine learning software library and was used in multiple image processing tasks through the software development. Google Colab was used in every train that was performed in order to cut down the time that each train took. Colab provides free access to powerful GPUs (Graphics Processing Unit) that empowers users to write and execute arbitrary python code through the browser and is particularly well suited for machine learning. 3.1 Model Framework Two software were developed for analysis of traffic signs and traffic lights for each one of the following objectives: •Autonomous Driving Competition of the RoboCup Portuguese Open in simulation •Autonomous Driving Competition of the RoboCup Portuguese Open •Public road In order to develop the algorithms to solve each of the three objectives, it was necessary to divide the problem into three sequential parts: 37 •Acquisition •Training •Testing Figure 3.1: Model Framework In the acquisition part, different frames were acquired and labelled in order to train the network, and one example is shown in figure 3.2. Figure 3.2: Image from one of the Datasets and its correspondent text file with the bounding box coordinates, dimensions and label In the previous image, a bounding box is around the STOP sign and in the associated text file the sign location and label, which is explained in the following sections. After gathering a sufficient amount of labelled images for the network, a train is performed. 38 e Robótica and using SolidWorks, a 3D model of the chassis was built. In figure 3.8, a comparison between the real robot and the 3D model can be seen. Figure 3.8: Comparison between the real robot (left) and its 3D model (right) Afterwards, the model was imported as a URDF file to CoppeliaSim. Also according to the real robot, four wheels with the same dimensions were added. To simulate the motors from the rear wheels and to bind the base to the wheels, revolute joints were placed in each wheel. Revolute joints are a CoppeliaSim object that allows the user to emulate a motor by rotating the object that is attached to it. In the rear wheels, the joints were used to move the car forward or backwards, depending on the direction that the joints were rotating. The front wheels have joints in which the main purpose is to control the car steering. Depending on the direction the joints were rotating, the robot can control the steering and the angle that it’s heading. In figure 3.9, the addition of the wheels and the joints can be visualized. The dimensions of the joints were increased to capture this figure, usually, they are not visible. Figure 3.9: The joints that allow the movement of the car 45 A CoppeliaSim camera was added to the car in the same position as in the real robot to detect and categorize each traffic sign/light that is in the presence of the robot. This camera grabs frames that are processed in the network to detect which traffic sign/light is present. Figure 3.10: Camera in the car The frames that the car camera records are later processed in a python script, which is explained later, that outputs the signs/lights that are in the frame. To make a connection between CoppeliaSim, where the car and the signs/lights are simulated, and Pycharm, where the frames are processed, ROS was used. ROS (Robot Operating System) is a framework for writing robot software. It has multiple tools, libraries, and conventions that simplify the task of creating complex and robust robot behaviour across multiple robotic platforms. ROS can have multiple processes running in parallel, each one is called a node. Every node should be responsible for one task. Nodes communicate with each other using messages passing via logical channels called topics. Each node can send or get data from another node using the publish/subscribe model. If data needs to be transferred from one node to another, one node publishes it into a topic and another node subscribes to the topic to receive the data. The ROS Master provides naming and registration services to the rest of the nodes in the ROS system. It traces publishers and subscribers to topics as well as services. The role of the Master is to enable individual ROS nodes to locate one another. Once these nodes have located each other they can communicate among them peer-to-peer. 46 Figure 3.11: ROS Master All the supervised learning algorithms software were developed on a Python script external to CoppeliaSim. The system that controls the robot depending on the identified traffic signs/lights, consisted of two ROS nodes, one for the simulator, named /sim_ros_interface and another for the control script, named /talker. In figure 3.12, the implemented ROS rqt_graph is presented illustrating how both of the ROS nodes communicate. /image (image) - Frame from the car’s camera /chatter (array) - Label from the traffic sign/light that is detected in the image. If no sign/light is detected, this array is empty. Figure 3.12: Graphical representation of how both of the ROS nodes communicate To create a node in the CoppeliaSim environment to communicate with PyCharm, Non threaded associated child scripts are used to customise the simulation. This method consists of a Lua script. Lua is powerful, fast and designed to be a lightweight embeddable scripting language. It is used for all sorts of applications, including image processing. It is completely compatible with CoppeliaSim. This method allows users to control every element of the simulation. Two embedded scripts were developed, one regarding the image from the camera and the other one regarding the control of the car. 47 The following part of the script is the creation of the subscriber in CoppeliaSim. This subscriber will receive the labels of the detected signs and will subscribe to the ”chatter” node created in Pycharm. This creation is specified in function syCall_init() which is the function that runs at the start of the simulation. The subscriber_callback(msg) is the function in which CoppeliaSim receives the label of the detected sign and changes the robot speed and direction according to the detection. Figure 3.13: Subscriber creation in CoppeliaSim In the other script, the publisher that sends the frame from the robot camera is created. In the syCall_init() this initialization is performed. Figure 3.14: Publisher creation in CoppeliaSim Using this system the frames from the car’s camera are sent to Pycharm, which processes them and if there is any traffic sign/light present, a command is returned to CoppeliaSim so that the car can react to the sign/light. 48 3.2.1 Acquisition The acquisition phase of the first objective has the goal of creating a dataset with images from all the traffic signs and traffic lights in order to train the network. A great number of images with signs/lights in different positions, distances to the camera make the network more robust. The first step to create the dataset, after creating all the signs and lights, is to record videos. This method was chosen because it allows the capture of the signs/lights in different situations that will be later explained. Multiple videos were created diversifying the following traits: •Background: changing the colour, inserting CoppeliaSim objects such as robots, tables, chairs and other objects; •Number of signs/lights: videos with one or more; •The angle of the camera: ranging from a top, bottom or side view; •Distance to the camera: recorded from near or far away. After this step was finished, all the videos were combined creating a longer video. Just like the individual videos, the final video had 60 fps (frames per second). Using Python as the main language and Pycharm Community 2021.1 as the IDE, a script was developed in order to convert the video frames into images. Instead of converting every frame, only one in every six frames was converted. This decision was made to prevent having extremely similar images. Since the video camera is always in motion, converting a frame every 6 frames allows the user to have mostly different images. The final video has 4 minutes and 47 seconds and using the previously mentioned script, 2873 images were created. In the following table 3.3, the number of images in which the signs/lights appear is presented. 49 Traffic Sign/Light Number of images Left Direction 397 Roundabout 241 Lights 406 Public Transport 323 60km/h 322 Hospital 303 Car Park 391 Crosswalk 290 Animals 366 Road Depression 287 Narrow Passage 243 Dangers 334 Yield 323 STOP 345 U-Turn 324 Overtaking 362 Prohibited Direction 350 30km/h 370 40km/h 351 Parking 381 End Overtaking 358 End Prohibitions 354 Right Obligation 361 Left Obligation 330 Traffic Light Right 149 Traffic Light Front 73 Traffic Light Left 41 Traffic Light Park 83 Traffic Light Stop 69 Traffic Light Finish 66 Table 3.3: Number of images per Traffic Sign/Light The inputs for the YOLOv3 algorithm are images, each one with a correspondent text file. This text file has the following format for each object in an image: Label_ID X_CENTER_NORM Y_CENTER_NORM WIDTH_NORM HEIGHT_NORM where, Label_ID = Correspondent label X_CENTER_NORM = X_CENTER/IMAGE_WIDTH Y_CENTER_NORM = Y_CENTER/IMAGE_HEIGHT 50 WIDTH_NORM = WIDTH_OF_BOUNDING_BOX/IMAGE_WIDTH HEIGHT_NORM = HEIGHT_OF_BOUNDING_BOX/IMAGE_HEIGHT •The Label_ID is a number from 0 to (number of classes - 1) and expresses the class of the sign/light that is in the bounding box; •X_CENTER_NORM: the X coordinate in percentage from 0 to 1 of the centre of the bounding box relative to the image width; •Y_CENTER_NORM: the Y coordinate in percentage from 0 to 1 of the centre of the bounding box relative to the image height; •WIDTH_NORM: width in percentage from 0 to 1 of the bounding box relative to the image width; •HEIGHT_NORM: height in percentage from 0 to 1 of the bounding box relative to the image height. If an image has more than one sign/light, the correspondent text file will have each sign/light in each line. Unlike the acquisition phase from the following objectives, in order to label every image, LabelImg software was chosen. LabelImg is a graphical image annotation tool and allows users to manually save annotations as XML files in PASCAL VOC format or text in YOLO format. Every object from every image was manually labelled. This methodology is not the most time and work efficient because it could be made using a python script as will be explained in the following objectives. It was also used because the project was at an early stage when the images were gradually being labelled to also learn more about the network described in section 3.2.2. With the goal of evaluating the network the dataset must be divided into two datasets: •Training - the dataset that is used to train the model (weights and biases in the case of a Neural Network). The model sees and learns from this data. •Validation - the sample of data used to provide an unbiased evaluation of a model fit on the training dataset while tuning model hyperparameters. The evaluation becomes more biased as skill on the validation dataset is incorporated into the model configuration. 51 In this objective, the chosen ratio was validation 20% and training 80%. To achieve this ratio a Python script was developed to divide the dataset. In the script, one in every five images was sent to the validation dataset and the remaining were sent to the training dataset. Every image that was sent to every dataset was sent with its corresponding text file. 3.2.2 Training Following the acquisition phase, where the Dataset was created, the data is ready to be entered into the network to train itself. For the training phase, two networks were used to develop two possible solutions for the objective, YOLOV3 and YOLOV3_tiny. These networks were chosen due to their high FPS and accuracy. As mentioned in section 2.8.3 where both of the networks were described, YOLOV3_tiny is a variation of the YOLOV3 architecture that was built for devices that have limited computational power. In this chapter, the steps to train the dataset in each of the networks are explained and in the following chapter 5 a conclusion of which is the best solution for the objective is presented. While training, a neural network takes in inputs, which are then processed in hidden layers using weights that are adjusted to find patterns in order to make better predictions. These operations are essentially matrix multiplications and can be trained faster in DL models by running all operations at the same time instead of one after the other. This can be achieved by using a GPU to train the model. The more powerful the GPU is, the faster the training will be. With the aim of lowering the training time and avoiding having the network training in the computer for multiple days at full power, Google Colab Pro was chosen to perform every train. Colab enables training in state of the art GPU, mainly the Nvidia Tesla P100. YOLOV3 and YOLOV3_tiny have similar training steps which are explained. All the commands here presented were executed in Colab. The first step to train the network, after the dataset is completed, is to clone the Darknet repository. Darknet is an open source neural network framework that allows users to use its two optional dependencies: OpenCV if users want a wider variety of supported image types or CUDA if they want GPU computation. CUDA (Compute Unified Device Architecture) is a parallel computing platform and application programming interface (API) model developed by Nvidia. It allows users to use a CUDA-enabled graphics processing unit (GPU) for general purpose processing which accelerates the processes. 52 This framework features, among other options, YOLOv3 and YOLOV3_tiny which are used in this phase. Darknet is cloned with the following command: !git clone https://github.com/AlexeyAB/darknet The following step is configuring the Makefile in order to enable OpenCV, CUDA and GPU. OpenCV is enabled because it will be used after the train to process the input images to predict if any traffic signs/lights are present. CUDA and GPU are enabled to take full advantage of Nvidia’s CUDA enabled GPU training. Initially, the directory is changed to darknet. This is obtained with the following commands: %cd darknet !sed -i 's/OPENCV=0/OPENCV=1/' Makefile !sed -i 's/GPU=0/GPU=1/' Makefile !sed -i 's/CUDNN=0/CUDNN=1/' Makefile After the Makefile has every desired Darknet setup option, the repository is compiled to be used with all these new options. This step is made by running the following command: !make After Darknet is compiled, every available functionality can be used. In this step, the first difference between YOLOv3 and YOLOV3_tiny appears. Darknet provides a config file for each YOLO option. For YOLOv3, the yolov3.cfg is used and for the tiny version, yolov3-tiny.cfg is used. This step is the most important because every hyperparameter for the network is established. These hyperparameters are in both config files. The most general hyperparameters are: •batch; •subdivisions; •width; •height; •channels; 53 •classes; •max_batches. The batch parameter sets the batch size used during training. The train involves iteratively refreshing the weights of the network based on the detection/location errors made on the training dataset. The training dataset contains a great number of images and instead of using all these images at once to update the weights, a small division of images is used in one iteration, this is called the batch size. For example, if the batch size is 64 then 64 images are employed in one iteration to update the weights. Even with small batch sizes, the computational power required can be too demanding for a GPU with low memory. Subdivision is used in Darknet to allow users to process a fraction of the batch size at one time on the GPU. The GPU will process batch subdivision images each time but the iteration would be complete only after all the batch images are processed. The width and height parameters define the image dimensions in which the images will be resized. The maximum value is 608x608 and with these dimensions, the results should improve but it takes longer to train. The channels parameter indicates the number of channels in the input, every image will be converted to this number of channels during Training and Detection. In every train performed the network had 3 channels because RGB images were used. Classes represent the number of different classes contained in the dataset. For example, if the dataset has 10 traffic lights and 2 traffic signs, the classes parameter should be 12. Finally, it needs to be specified how many iterations should the training last. For multi-class object detectors, the max_batches number is higher, the train requires to run for more batches. For this class object detector, it is advisable to run the training for at least (average number of images per class)*(number of classes). The following are hyperparameters for optimization: •momentum; •decay; •learning_rate; •policy; 54 With the value of the time that the script takes to process the frame, the fps, that corresponds to the frequency, can be calculated with the following formula: fps =1000(ms) time_of_process(ms)(3.2) In the previous formula, the 1000 value corresponds to the conversion from milliseconds to seconds. This script is used in the three objectives described in section 3.1. Its application changes in each one with the three corresponding files that are loaded: the weights and config file for the network and the labels file. For each objective three types of frame inputs are accepted: image, video and camera. This script will be referenced in the following two objectives. 3.2.3.2 Testing Framework The network was tested twice: •First Test The first test had the purpose of evaluating how the network correctly detected the TS/L (traffic signs/lights). Videos were made in different scenarios from the dataset ones, referred to in section 3.2.1, so that it was possible to know how the network would behave in different conditions and to examine its robustness. This test did not involve the car or the ROS communication, only the TS/L (traffic signs/lights) detection and classification. Using the testing script, previously explained in section 3.2.3.1, and using the resulting weights, config and labels files from the train, the video frames were processed and the output analysed in order to visualize if the objects were correctly identified. In figure 3.16, the output of the testing script is presented in two different frames from the videos. 61 Figure 3.16: Images from two different frames from the test •Second Test The second test is a more realistic simulation of the competition, in which the network detects the possible TS/L that would be on the track. The main goal is the detection of the TS/L and with this information, the car will act depending on the traffic sign. Not all TS/L were chosen to assign an action because their sign/light meaning does not involve movement. The S/L and their correspondent actions are placed in the following list: •STOP, Yield and Prohibited Direction: stop the motion •Traffic Light Front: start of forwarding motion •30km/h, 40km/h and 60km/h: change speed to traffic signal value •Left Obligation: turn 45 degrees left •Right Obligation: turn 45 degrees right •U-Turn: turn 180 degrees In figures 3.17, 3.18 and 3.19, the output of PyCharm and CoppeliaSim are presented. All the communications in the software are carried out using ROS as previously explained. On each imagesleft side can be seen the processed frame that is sent from CoppeliaSim to PyCharm and processed using the script referenced in section 3.2.3.1. This script detects and classifies the present sign in the frame. The frame label is returned to CoppeliaSim so that the robot can perform the corresponding action. 62 On the figures right side there is CoppeliaSim simulator in which it is possible to view the robot movement. At the bottom of this simulator, the commands are presented in which the speed is presented in real time. Initially, the robot waits for the detection of the front traffic light to start its forward motion, as seen in figure 3.17. Figure 3.17: Start of the test From this point, the robot moves forward until it detects another signal, according to which it reacts. When it detects a speed signal the robot changes its speed depending on the sign value. Figure 3.18: Speed detection 63 The robot continues to react to the detected S/L until it reaches a sign with a stop action. Figure 3.19: End of the test after the STOP detection 3.3 Autonomous Driving Competition of the RoboCup Portuguese Open The second objective of the project is the development of a detection and classification algorithm for the Autonomous Driving Competition of the RoboCup Portuguese Open. Unlike the objective described in section 3.2 this objective is not in simulation but in a real competition context. In this section, the three phases are described in order to accomplish the objective. As in section 3.2, for the autonomous driving competition of the RoboCup Portuguese Open, the network has the goal of correctly identifying 24 traffic signs and 6 traffic lights which are listed in table 3.2. Following the competition guidelines and to obtain the images for the dataset, identical TS/L were constructed to mimic the environment that would be in the competition. Using images from Autoridade Nacional de Segurança Rodoviária , which has the mission of planning and coordination at a national level to support the Government’s policy in the field 64 of road safety, as well as the application of the administrative law on road trafficthe TS/L were created. These images were printed and attached to a K-line sheet. These were cut with the right dimensions. In the following figure 3.20, some examples of the TS/L are presented. Figure 3.20: Constructed Signs Also to mimic the environment that would be in the competition, a track with similar dimensions to the competition one was constructed in the Laboratório de Automação e Robótica facilities by its members. This track can be visualized in figure 3.21 Figure 3.21: Constructed Track 65 3.3.0.1 Distance to the Traffic Sign/Light software During the competition, the robot needs to know how far away are the TS/L so that with this information it knows the ideal moment to react to the sign. Given this problem, a Python script was developed that, starting from the Bounding Boxes area detected in the frame, could determine the distance to the TS/L. The first step was to create a relationship between the Bounding Box area and the distance to the camera. For this, a large number of values of the Bounding Box area resulting from the neural network were registered while increasing the distance between the signal and the camera. After obtaining this data, it was entered into Microsoft Excel and from this, a regression was performed to determine the relationship between the Bounding Box area and the distance to the camera. The equation resulting from this regression was: 39757.381 x−0.597 Distance = 39757.381 ∗Area−0.597 (3.3) where, Distance = Distance from the TS/L to the camera Area = Area of the Bounding Box In figure 3.22, it is possible to verify in black the line of the determined equation and in orange the points used to perform the regression. Figure 3.22: Graph with the distance from the T S/L to the camera 66 Using the Bounding Box area proved to be the best method compared to other tested variables such as Bounding Box height and width individually. 3.3.1 Acquisition The acquisition phase of the second objective (section 3.2.1) has the goal of creating a dataset with images from all the traffic signs and traffic lights in order to train the networks. Using the same methods referenced in the acquisition phase from the previous objective, videos of the TS/L were performed with the same goal. Although the previous objective is in simulation, in the second objective the diversity of traits is approximately the same: •Background: Using different scenarios in the Laboratório de Automação e Robótica facilities, varying the colour, textures and objects such as robots, tables and chairs; •Number of signs/lights: videos with one or multiple; •The angle of the camera: ranging from a top, bottom or side view; •Distance to the camera: recorded from near or far away. In the following figures, the variety of previously explained traits is demonstrated for the U-Turn sign in the images from the dataset. Figure 3.23: First Example 67 In figure 3.23 the image only has the U-Turn sign, it is at an average distance to the sign and the brightness is lower than normal. Figure 3.24: Second Example Unlike figure 3.23, in figure 3.24 the image has more than one traffic sign, has a shorter distance and is brighter. Figure 3.25: Third Example The main difference between figure 3.24 and 3.25 is that figure 3.25 has more objects in the background and is blurrier. 68 Figure 3.26: Fourth Example In figure 3.26 there is a greater amount of signs and these are further away. This figure has the most objects in the background. The Xiaomi Mi 9 SE smartphone was used to record the videos with 1080 resolution and with 30 fps. The smartphone was used due to its camera stabilization. The smartphone was used to obtain the videos because it would emulate the conditions in which the network would be tested. This method was also chosen because not all frames would be in focus and this would help the final network to be more robust in case the input has the same conditions. If the network is trained with some blurier images it will perform better in the testing phase. Using Python as the main language and Pycharm Community 2021.1 as the IDE, the same script described in section 3.2.1 was used to convert the video frames into images. A different ratio of 1 in every 3 frames was used due to the videos having half of the FPS. This allowed the user to have always different images. With this frame rate, the script generated 5 images every second instead of the 10 obtained with a 60 fps input. The final video has 9 minutes and 54 seconds and using the script as in the previous objective, 5949 images were created. In table 3.4, the number of images in which the signs/ lights appear is presented. 69 Traffic Sign/Light Number of images Left Direction 362 Roundabout 372 Lights 440 Public Transport 436 60km/h 459 Hospital 440 Car Park 433 Crosswalk 435 Animals 395 Road Depression 390 Narrow Passage 371 Dangers 361 Yield 420 STOP 387 U-Turn 415 Overtaking 383 Prohibited Direction 367 30km/h 450 40km/h 377 Parking 386 End Overtaking 454 End Prohibitions 354 Right Obligation 418 Left Obligation 399 Traffic Light Right 281 Traffic Light Front 424 Traffic Light Left 461 Traffic Light Park 262 Traffic Light Stop 363 Traffic Light Finish 378 Table 3.4: Number of images per Traffic Sign/Light Regarding the associated text file to the images, the same format referred to in section 3.2.1 is used since the network used in this objective will be the same as the last one. Unlike the acquisition phase in the previous objective in which the labels were manually performed, in this one, most of the labels were deployed using a developed Python script using the Template Matching function from OpenCV. Template Matching is a method for searching and finding the location of a template image in a larger image, it slides the template image over the input image returns all the detections with the corresponding confidence. For each TS/L, a template was generated. To improve detection, this template must have a black background with 70 ID Traffic Sign/- Light Image ID Traffic Sign/- Light Image 14 Prohibition Turn 15 Roundabout 16 Lights 17 Overtaking 18 Parking 19 STOP and Park 20 End Prohibitions 21 Danger Crosswalk 22 Traffic Congestion 23 Danger Roundabout 24 Junction 25 Curve 26 Narrow Passage 27 Road Hump 77 ID Traffic Sign/- Light Image ID Traffic Sign/- Light Image 28 Crosswalk 29 Car Park 30 Tunnel 31 STOP 32 Traffic light 33 Yield 34 Directional Sign 35 50km/h Obligation Table 3.5: The Traffic Signs/Lights in the Road Dataset 3.4.1 Acquisition With the goal of acquiring images for the Road Dataset, the methodology chosen was similar to the one used in the previous objective (section 3.3.1). The main goal of the acquisition phase was obtaining multiple images from the TS/L from different scenarios. To make the network more robust the TS/L needed to be in different locations in which the following characteristics would change: •Background: the TS/L are placed in different environments and subsequently all of them have different backgrounds; •Angle: the TS/L can have different angles; •Distance: the TS/L can be at different distances; 78 •Brightness: over the course of the day, the incidence of brightness will vary on TS/L. Videos were recorded to accomplish the acquisition of the images with the previously referred characteristics. These videos were recorded from the front passenger seat in the car as demonstrated in figure 3.33. In this figure, the blue shape corresponds to the zone that was being recorded. Figure 3.33: Zone recorded The videos were recorded on a Xiaomi Mi 9 SE smartphone. This smartphone was chosen due to the camera stabilization to reduce the blur of videos. These videos were recorded on multiple days and at different hours. The videos also were recorded during nighttime to ensure the network performs all day. The paths also vary in order to find the least common TS/L in the dataset and to have different characteristics in the images. The videos were joined and a final video was created. The final video has 15 minutes and 30 seconds and using the script from the other objectives, 4648 images were created. In the following table 3.6, the number of images in which the T S/L appear is presented. Traffic Sign/Light Number of images 10km/h 6 30km/h 58 40km/h 248 50km/h 476 60km/h 155 70km/h 286 79 80km/h 67 90km/h 135 100km/h 38 120km/h 110 30km/h Obligation 38 40km/h Obligation 67 Prohibited Direction 401 Obligation Arrow 1057 Prohibition Turn 414 Roundabout 652 Lights 91 Overtaking 132 Parking 51 STOP and Park 80 End Prohibitions 58 Danger Crosswalk 299 Traffic Congestion 15 Danger Roundabout 182 Junction 281 Curve 186 Narrow Passage 52 Road Hump 140 Crosswalk 1044 Car Park 113 Tunnel 41 STOP 367 Traffic light 700 Yield 1162 Directional Sign 490 50km/h Obligation 153 Table 3.6: Number of images per Traffic Sign/Light In table 3.6 the values from each TS/L fluctuate a lot which is representative of their quantity on the roads where the videos were recorded. In the following images examples from the dataset are presented with different characteristics. 80 Figure 3.34: Rainy day image In figure 3.34 the image has a low level of brightness due to the rain and there are multiple traffic signs at different distances. Figure 3.35: Bright day image Figure 3.35 is an example of an image in which the traffic signs are at a greater distance and the image is bright. Figure 3.36: Night image 81 Figure 3.36 is an example in which the images have the lowest brightness because they were recorded at night. After the images were generated, the corresponding labels needed to be produced. The method of using the Template Matching function from the OpenCV library, as described in section 3.3.1, was not used to generate images because the distance of the images to the TS/L was too high for the function to properly work. The LabelImg software, as in section 3.2.1 was used to label every image. This method was not the most time-efficient because the process took several days but it was the one that guaranteed the most accuracy from the ones that were tested. 3.4.2 Train Labelling Network Due to the immense time needed to label images with the methodology used in section 3.4.1, a network was created with the already labelled images as input. This network will label more images to make the final network more complete and robust. To create this network for labelling, a YOLOV3 network was created just like it was for the first objective, all the steps remain the same as explained in section 3.2.2. The only difference is the dataset described in section 3.4.1, in which the train is performed. For this network, only the YOLOV3 network was created, and not the YOLOV3_tiny, because the accuracy from YOLOV3 was required to label the images. 3.4.3 Acquisition Using The Labelling Network The number of videos used to train the previously described network is not enough because the network is not as accurate as it should be. This network is used to speed up the labelling process for the final network. After acquiring more videos in the same conditions referenced in section 3.4.1, the videos were joined, which resulted in a 12 minutes and 23 seconds video. Using the script explained in the previous objectives, 3686 images were created. These images were processed in the network that was developed in section 3.4.2 and the labels for each sign that the network outputs are saved with the YOLO format in a text file with the same name of the image. 82 Since this network is not the final one, the accuracy is not the desired one and every label needed to be checked. Figure 3.37 is an example of this process, the network correctly detects 3 traffic lights and 2 traffic signs but 1 traffic light is not detected. This is why every image needs to be checked. Even though all the labels were checked and the Bounding Boxes adjusted, this method proved to be faster than doing the labels manually. With this method, all the labels for the final dataset were performed. Figure 3.37: Example of Acquisition using The Labelling Network 3.4.4 Train Final Network With the final dataset ready for the Public Road network the train was performed. As in previous sections two networks were developed (YOLOV3 and YOLOV3_tiny). All the steps remain the same as explained in section 3.2.2. The only difference is the dataset, described in section 3.4.3, in which the train is performed. 83 3.4.5 Testing The main objective is the correct detection and classification of the traffic signs that are being captured by a camera from a moving car. To replicate this process and has was performed in the acquisition phases, images were recorded on a Xiaomi Mi 9 SE smartphone and then these frames were processed in the testing script, explained in section 3.2.3.1, to evaluate how accurate the network is. 84 Chapter 4 Tests This chapter introduces the numerous tests performed on the two networks (YOLOV3 and YOLOV3_tiny) that were carried out to determine the best parameters that represent the best results in terms of performance and accuracy. 4.1 Frame Rate The first tests that were performed consisted of testing the performance with different input dimensions. The tests were performed on both networks (YOLOV3 and YOLOV3_tiny) using the testing script referenced in section 3.2.3.1 and the value of the FPS was monitored. As was mentioned in section 2.8.2.1, the YOLO architecture can perform with multiple scales which can adopt an image dimension size that is multiple of 32 to make it more robust. The smallest choice is 320x320 and the most extensive 608x608. Figure 4.1 corresponds to the performance results for the YOLOV3 network and figure 4.2 corresponds to the performance results for the YOLO_tiny network. Figure 4.1: YOLOV3 performance with different input dimensions 85 Figure 4.2: YOLO_tiny performance with different input dimensions In both of the networks, the frame rate decreases when the input size increases. With a higher resolution, the networks have to process more pixels and this results in a higher processing time which lowers the FPS. As was explained in section 2.8.3.3, the YOLO_tiny network ensures a higher frame rate compared to YOLOV3. With the lowest input value (width = 320 and height = 320), the YOLO_tiny network has 29,03 fps in comparison to the 4,19 fps obtained in YOLOV3. On the other hand, accuracy is where YOLO_tiny has lower values but this was tested in the following sections. Following the values provided by the YOLOV3 creators, the resolution used in the following tests are width = 416 and height = 416. This way there is a balance between the accuracy and the FPS. 4.2 Hyperparameters For each of the three objectives explained in section 3, two networks were developed (YOLOV3 and YOLOV3_tiny). In each training phase the hyperparameters needed to be optimized to obtain the best performance. The YOLO architecture provides initial values for these hyperparameters and these are tested, in this section, in comparison to higher or lower values. In the following sections, the output of three tests with different values from the corresponding hyperparameter is going to be compared. The output values compared are the mAP and the Loss 86 is reached in iteration 12880 in the burn_in=1000 line. Although one of the lines had the best result, as the training progressed, the three lines tended to the same value, which leads to the conclusion that this hyperparameter does not affect mAP. Figure 4.10: Influence of Burn_in on Loss The same conclusion can be made regarding the loss since the three values, despite the fluctuations and the beginning of the train, tend towards the same value during training, analyzing figure 4.10. Burn in value Time the train lasted 500 17 hours 33 minutes 1000 16 hours 9 minutes 1500 16 hours 3 minutes Table 4.3: Time each train lasted Unlike previously analyzed metrics, mAP and loss, training times are shown in table 4.3. It is possible to establish that the lower the burn_in value, the longer the training time. This leads to the conclusion that burn_in affects training time and not the metrics already referenced. These tests were performed in Google Colab Pro to save time using the provided GPUs. The computational power available can fluctuate throughout the tests and this can lead to a slightly different training time for two equal trains. 93 4.2.1.5 Decay To prevent overfitting in situations where the model learns too much about the training data, the decay hyperparameter is used to penalize large values for weights. In these circumstances, the model is useful in reference only to its training dataset, and not to any other datasets, such as the testing one. In this case, the model proves to be fitting only for the training data, as if the model had only memorized the training data and was not able to generalize to other data never seen before. In figure 4.11 and 4.12 the two generated graphs are presented. Figure 4.11: Influence of decay on mAP The main conclusion that can be reached by analysing figure 4.11 is that the decay=0.005 line is the worst one of the ones tested because it constantly has inferior mAP values. There is not a great difference between the other two values for decay since they constantly have approximately the same values throughout the train. The highest mAP value, mAP=0.98857 or 98.857, was achieved in iteration 14200 for the decay=0.00005 value. 94 Figure 4.12: Influence of decay on Loss Interpreting figure 4.12, the same conclusion can be reached regarding the decay values. The decay=0.005 line has the highest loss values in comparison to the other two. The other two values have similar values throughout the train. It can be concluded that 0.0005 and 0.00005 vales are the most optimized from the ones tested for the decay hyperparameter. 4.2.1.6 Width and Height The width and height parameters define the image dimensions in which the images will be resized in order to be used in the training process. The maximum value is 608x608 and with these dimensions, the results should improve but it takes longer to train. In figure 4.13 and 4.14 the two generated graphs are presented and table 4.4 presents the training times. 95 Figure 4.13: Influence of width and height on mAP Performing an examination on figure 4.13, the highest mAP value, mAP= 0.99514 or 99.514%, is achieved in iteration 18380 when the width and height have a value of 544. The other two lines throughout the train also reach near mAP values. The line with width and height=416 reaches mAP=0.98804 or 98.804% and the line with width and height=320 reaches mAP=0.98392 or 98.392%. Figure 4.14: Influence of width and height on Loss 96 Width and Height value Time the train lasted 320 8 hours 56 minutes 416 18 hours 6 minutes 544 28 hours 13 minutes Table 4.4: Time each train lasted Analysing figure 4.14, the resulting lines for the three values for width and height are similar. The three lines evolve at the same rate and have a similar fluctuation. Since the resolution of the image influences the number of pixels each image has the higher the input resolution, the longer the train time for each iteration. This is proven in table 4.4 where train time is higher for higher resolutions. 4.2.2 YoloV3_tiny In this section, the values for the hyperparameters in Table 4.5 are tested in the YOLOV3_tiny network and the results are presented. Hyperparameter Test 1 Test 2 Test 3 Max Batches 19000 50000 72000 Learning Rate 0.01 0.001 0.0001 Momentum 0.2 0.45 0.9 Burn in 500 1000 1500 Decay 0.005 0.0005 0.00005 Width x height 320x320 416x416 544x544 Table 4.5: The hyperparameters and the corresponding tested values 4.2.2.1 Max Batches The max_batches hyperparameter was tested with the most optimized value tested in section 4.2.1.1, max_batches=1900, but analysing the mAP results the conclusion was reached that the network could increase its accuracy with higher iterations. The following values tested were max_batches=50000 and max_batches=72000 and the two generated graphs are granted in figure 4.15 and 4.16. 97 Figure 4.15: Influence of max_batches on mAP The max_batches=19000 line is the one whose max mAP value, mAP=0.86536 or 86.536%, is the lowest in comparison to the other two networks. With more iterations, max_batches=50000, the network is able to reach higher mAP values and strikes an mAP of 0.92611 or 92.611%. Increasing, even more, the number of iterations, max_batches=72000, the network only increases its mAP slightly to 0.92895 or 92.895%. Figure 4.16: Influence of max_batches on Loss Regarding the loss, the same conclusion can be reached concerning the max_batches=19000 line where it does not have enough iterations to reach lower values. In comparison to the 98 max_batches=50000 line, the extra iterations from the max_batches=72000 line does not offer an improvement in terms of loss. Max_batches value Time the train lasted 19000 4 hours 24 minutes 50000 14 hours 7 minutes 72000 17 hours 51 minutes Table 4.6: Time each train lasted As concluded in section 4.2.1.1, the train takes longer if it has more iterations. This outcome can be visualised by interpreting table 4.6. In conclusion, the max_batches=19000 line does not have enough iterations to reach the values that the other two lines obtain. The extra computational power for the max_batches=72000 line in comparison to the max_batches=50000 line does not result in significant upgrades. Consequently, the best result and the one that is going to be used in the following sections for the YOLO_tiny network is max_batches=50000. 4.2.2.2 Learning Rate The same values used in section 4.2.1.2 were used to test the optimal learning_rate value for the YOLOV3_tiny network. The output graphs are represented in figure 4.17 and 4.18. Figure 4.17: Influence of learning_rate on mAP 99 Interpreting figure 4.17, the effect that this hyperparameter has is visible. With a lower learning_rate value, as conceived in the lr=0.0001 line, the mAP does not reach the desired value. With a higher value, as shown in the lr=0.01 line, the curve has more fluctuation and reaches early its maximum value. This early maximum value can lead to overfitting. The line that reaches the highest value is the learning_rate=0.001 with mAP=0.92611 or 92.611%. Figure 4.18: Influence of learning_rate on Loss The loss of the lr=0.0001 line is the worst result obtained stabilizing with a value that doubles the achieved in the other two tests. The best result was accomplished with the lr=0.001 line that despite having a higher loss value for half of the train, obtained an inferior and more stable loss value at the second half of the train, compared to the lr=0.01 line. Analyzing the results obtained in figure 4.17 and 4.18 it was concluded that the best value for the learning_rate hyperparameter is lr=0.001. 4.2.2.3 Momentum Regarding the momentum hyperparameter, the values tested are the same as the ones used in section 4.2.1.3. Also, as stated in the previously mentioned section, values above 0.9 led to a train crash because the loss value was too high, leading to a variable overflow. The output graphs are reproduced in figure 4.19 and 4.20. 100 Figure 4.19: Influence of momentum on mAP The differences between the three output lines allow an easy conclusion regarding which value is the most optimized for the hyperparameter tested. Over the three trains, the momentum=0.9 line consistently contains the highest mAP values by a solid margin. Figure 4.20: Influence of momentum on Loss Analysing figure 4.20, the same conclusion can easily be made regarding the loss in which the momentum=0.9 reaches a lower value by a great margin. Concluding, the best value for the momentum is 0.9. 101 4.2.2.4 Burn in To test the optimal value for the burn_in hyperparameter, the values used in section REF are used for the YOLOV3_tiny network. The results are displayed in figure 4.21 and 4.22. Figure 4.21: Influence of Burn_in on mAP Figure 4.22: Influence of Burn_in on Loss 102 5.1.1.2 YOLOV3_tiny Figure 5.3: YOLOV3_tiny mAP for the Autonomous Driving Competition of the RoboCup Portuguese Open in simulation Figure 5.4: YOLOV3_tiny loss for the Autonomous Driving Competition of the RoboCup Portuguese Open in simulation The best mAP, mAP=0.98791 or 98.791%, was achieved at iteration 46617. 109 5.1.2 Autonomous Driving Competition of the RoboCup Portuguese Open 5.1.2.1 YOLOV3 Figure 5.5: YOLOV3 mAP for the Autonomous Driving Competition of the RoboCup Portuguese Open Figure 5.6: YOLOV3 loss for the Autonomous Driving Competition of the RoboCup Portuguese Open The best mAP, mAP=0.99086 or 99.086%, was achieved at iteration 14024. 110 5.1.2.2 YOLOV3_tiny Figure 5.7: YOLOV3_tiny mAP for the Autonomous Driving Competition of the RoboCup Portuguese Open Figure 5.8: YOLOV3_tiny loss for the Autonomous Driving Competition of the RoboCup Portuguese Open in simulation The best mAP, mAP=0.98479 or 98.479%, was achieved at iteration 30600. 111 5.1.3 Public Road 5.1.3.1 YOLOV3 Figure 5.9: YOLOV3 mAP for the Public Road Figure 5.10: YOLOV3 loss for the Public Road The best mAP, mAP= 0.98914 or 98.914%, was achieved at iteration 17548. 112 5.1.3.2 YOLOV3_tiny Figure 5.11: YOLOV3_tiny mAP for the Public Road Figure 5.12: YOLOV3_tiny loss for the Public Road The best mAP, mAP=0.95584 or 95.584%, was achieved at iteration 49462. 113 5.2 Comparison between networks per objective In this section, the output of the developed YOLOV3 and YOLOV3_tiny networks is compared in each objective. The outputs consists of the videos that were clarified in the corresponding objectives from section 3. These videos were created and later processed by the corresponding networks. 5.2.1 Autonomous Driving Competition of the RoboCup Portuguese Open in simulation To test the simulation environment for the RoboCup Portuguese Open, the two tests referenced in section 3.2.3 will be used. These tests allow network performance visualization and a more detailed comparison between the two developed networks. 5.2.1.1 First Test For this test, the signs were placed in two parallel lines to verify their correct classification and detection. The minimum confidence for which signals were detected was 80%. Comparing both networks in the first test located in appendix A.1, it is possible to verify that the computational power required for the YOLOV3 network is superior to what the computer offers. This means that only a fraction of the frames is processed. Regarding classification, the fraction of frames that are processed unveils that the signs are all classified and detected with confidence over 95%. In Figure 5.13, two examples of this high confidence in detection are demonstrated. 114 Figure 5.13: Two frames with high confidence detection from the YOLOV3 network On the other hand, the YOLOV3 tiny network manages to process the frames in such a way that the robot is constantly identifying and classifying the signs. Compared to the other network, YOLOV3_tiny does not have such high detection confidence or stable Bounding Boxes but its ability to process frames overcomes these limitations. In Figure 5.14, this detection is displayed. Figure 5.14: Two frames with high confidence detection from the YOLOV3_tiny network 115 5.2.1.2 Second Test For the second test, the signs were placed so that the robot would detect them sequentially while performing the respective movements. The robot reacted to the signs when confidence was over 98%. Comparing the outputs of the two networks in appendix A.1 from the second test, it is possible to verify the difference between the two networks. As concluded in the previous test, the output for the YOLOV3 network cannot process all the frames so that the robot reacts to the TS/L promptly to perform the corresponding movements. This lack of performance implied that the robot only reacted to the Left Obligation sign when it was already extremely close to it. This meant that the robot could not turn in time to move to the other signs. Regarding classification, the network was able to correctly classify with confidence above 99% the signs it obtained in the frames. The moment the robot reacted to the Left Obligation sign is presented in Figure 5.15. Figure 5.15: Frame where the robot reacted to the Left Obligation sign Unlike the YOLOV3 network, the YOLOV3_tiny network was able to timely react to all traffic signs. In Figure 5.16, it is possible to visualize the frame in which the robot reacted to the Left Obligation sign. It reacted earlier than the YOLOV3 network. This difference can be visualized by comparing the position of the robot in figure 5.15 and 5.16. 116 Figure 5.16: Frame where the robot reacted to the Left Obligation sign As explained in section 3.2.3.2 the robot has the ability to change its speed if it detects one of the three available speed signs. This change of speed can be visualized in the videos, in the robot or the CoppeliaSim command line. In figure 5.17, the robot correctly identifies the 30km/h sign and changes its velocity to achieve the velocity designated by the sign. Figure 5.17: Frame where the robot reacted to the 30km/h sign 117 In this test, the YOLOV3_tiny network enabled the robot to go all the way until it stopped at the STOP sign, frame displayed in figure 5.18. Before it achieved the previously mentioned sign, the robot also accurately classified the U-Turn sign and performed its corresponding manoeuvre. Along its path, the robot was able to detect all the signs in time and this led to being able to perform the corresponding movements to reach the end of the path. Figure 5.18: Frame where the robot reacted to the STOP sign 5.2.2 Autonomous Driving Competition of the RoboCup Portuguese Open To test the environment that was built to emulate the Autonomous Driving Competition of the RoboCup Portuguese Open, since the competition did not take place due to the pandemic, two videos were recorded. Both videos demonstrate the robot driving along the track and observing the signs/lights that were randomly placed along. The only difference between these is that the first video was recorded with as much stability as possible and the second was recorded on the robot prototype while it was remotely controlled manually, this led to a lower stabilization. In appendix A.2, these videos are presented sequentially for each of the networks. In these videos, the signs are only detected when confidence is over 80%. In these videos, it is not possible to control the processing time of each frame as they were recorded before and processed afterwards. The YOLOV3 network managed to process an average of 2 fps while the YOLOV3_tiny 17 FPS. 118 Some frames of the output of the YOLOV3_tiny network, using the same input, are presented in figure 5.30. The network correctly identifies the Direction Signs and the Obligation Arrow but does not recognize the 50km/h, Danger Crosswalk or the Junction signs. Figure 5.30: Results of the YOLOV3_tiny network in poor conditions 5.3 Results Discussion In this chapter, the results for the three dissertation objectives are described in section 3.1. For each objective, a comparison is made between the YOLOV3 and the YOLOV3_tiny networks. For the Autonomous Driving Competition of the RoboCup Portuguese Open in simulation, the results presented proved that the YOLOV3_tiny network is the most suitable for this environment. The YOLOV3 network exhibited a slightly higher detection and classification confidence but the required computational power to run the simulation in real-time was higher than the computational provided by the used computer. The tiny version was able to correctly detect the TS/L in a timely manner and this provided enough time for the robot to perform the corresponding manoeuvres. For the Autonomous Driving Competition of the RoboCup Portuguese Open, the same conclusion as the previous objective was reached. The device controlling the robot would not have 125 enough computational power to provide the required YOLOV3 network power. This would lead the robot to only detect the TS/L when it has already passed it or perform the corresponding manoeuvres belatedly. The tiny version would allow timely detection and the robot would be able to perform the manoeuvers correctly. The Public Road is the objective where the YOLOV3 network is the most suitable. The detection in this environment must be the most accurate since its application is destined for autonomous cars. Cars that have TSR software have enough computational power to allow a low processing time that allows the network to be able to identify the signs without delays. From the provided results, the YOLOV3_tiny does not have enough accuracy to ensure a correct and safe detection. 126 Chapter 6 Conclusions and Future Work The work presented in this dissertation describes the development of neural networks capable of detecting and classifying traffic signs and traffic lights in three different scenarios. Initially, some introductory concepts are presented regarding fundaments from Artificial Intelligence to Convolutional Neural Networks that leads to a better understanding of the networks mentioned. A context of how two automotive companies perform Traffic Sign Recognition is presented to better understand how this technology is performed in an industrial environment. Follows an analysis surrounding the Autonomous Driving Competition from the RoboCup Portuguese Open to acquaint how the robot must perform concerning the Vertical Traffic Signs Detection Challenge. A detailed explanation about the architecture used to develop the final networks is shown, which includes its evolution, starting in the first version until the third one. Finally, an application of a TSR for the same competition performed by another team is exhibited using different methods for the same purpose. The second part of this dissertation focus is on the used methodologies to develop the final networks. For each of the presented objectives, an explanation of the adopted steps is presented. In the first objective, initially, the steps for the creation of all the parts for the simulation are presented. Then, an explanation regarding how these parts were used to develop the dataset is made. The following step had the greatest importance because it was used in all the objectives. It consisted of making a detailed explanation of how the training plan in the YOLO architecture was performed, this included every command line and clarification of the hyperparameters and network structure. Next, the testing phase of each objective was presented to judge if any improvements were required for the networks. The detailed explanation of the YOLO architecture throughout the dissertation and all the adjacent steps to create all networks can allow further projects of Laboratório de Automação e Robótica to use this work. 127 The third part of this project consisted of testing the hyperparameters for each of the networks. With this study, for the YOLOV3 and the YOLOV3_tiny networks, the optimized values for the hyperparameters were discovered. Testing three values per hyperparameter, the mAP and the loss were used to select which one was the most optimized for each network. Using the most optimized values for each hyperparameter, concerning each objective, two final networks were developed (YOLOV3 and YOLOV3_tiny), in the fourth part of the dissertation. The mAP and loss graphs for each of the six networks were displayed. Videos representing the results of these networks were created using the corresponding input videos explained in the third part, which allow a study of how these performed. Regarding the first objective, the YOLOV3_tiny network guarantees results that allow the robot to detect and classify the Traffic Signs/Lights in time to perform the necessary manoeuvres along its path. Despite having a higher accuracy, the processing time of each frame on the V3 network does not allow the robot to work correctly. In the second objective, the processing time led to the choice to commit to the YOLOV3_tiny network since the computational power of the onboard computer for the competition is usually limited and with the YOLOV3 network, it would not be able to react to the Traffic Signs/Lights in time. Unlike the networks chosen for the previous objectives, for Public Road, the YOLOV3 network provides the most accurate results and usually, cars have onboard computers that provide enough computational power. Despite correct detection in most signs, the results displayed classification and detection errors when faced with Traffic Lights and with Signs not contained in the dataset. 6.1 Future Work The work carried out in this dissertation is available to be used by the Autonomous Driving team of Laboratório de Automação e Robótica at the RoboCup Portuguese Open when it occurs. The future of this project aims to enhance the accuracy of the detection and classification for the three objectives. This can be performed by introducing more images in different scenarios into the dataset and adding new signs making the networks more complete. In the Public Road example, the use of a camera from the car or one that enables a higher stabilization will most 128 likely improve the results. The application of these datasets for the YOLOV4 [50] and YOLOV5 networks can also be analyzed. 129 Appendix A Videos for each objective A.1 Autonomous Driving Competition of the RoboCup Portuguese Open in simulation https://youtu.be/oaBd6Ub-o7E A.2 Autonomous Driving Competition of the RoboCup Portuguese Open https://youtu.be/T2USKNakM9w A.3 Public road https://youtu.be/zzIkw8suny4 130 References [1] A. Friedman, “Introduction to neurons,” pp. 1–20, jan 2005. [2] Ayşegül Uçar, “The-structure-of-an-artificial-neuron.png.” [Online]. Available: https://www.researchgate.net/profile/Ayseguel{_}Ucar/publication/327260166/figure/ fig3/AS:664442090049538@1535426747843/The-structure-of-an-artificial-neuron.png [3] K. O’Shea and R. Nash, “An Introduction to Convolutional Neural Networks,” nov 2015. [Online]. Available: http://arxiv.org/abs/1511.08458 [4] Thomas Christopher, “An introduction to Convolutional Neural Networks | by Christopher Thomas BSc Hons. MIAP | Towards Data Science,” 2019. [Online]. Available: https:// towardsdatascience.com/an-introduction-to-convolutional-neural-networks-eb0b60b58fd7 [5] D. Arden, “Applied Deep Learning - Part 4: Convolutional Neural Networks | by Arden Dertat | Towards Data Science,” 2019. [Online]. Available: https://towardsdatascience. com/applied-deep-learning-part-4-convolutional-neural-networks-584bc134c1e2 [6] S. Sakib, Ahmed, A. Jawad, J. Kabir, and H. Ahmed, “An Overview of Convolutional Neural Network: Its Architecture and Applications,” ResearchGate , no. November, nov 2018. [Online]. Available: www.preprints.orghttps://www.researchgate.net/publication/ 329220700 [7] Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2323, 1998. [8] SPR, “Robótica 2019 - Rules for Autonomous Driving,” Tech. Rep. [9] J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” in Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition , vol. 2016-Decem. IEEE Computer Society, jun 2016, pp. 779–788. [Online]. Available: http://arxiv.org/abs/1506.02640 131 REFERENCES 132 [10] J. Redmon and A. Farhadi, “YOLOv3: An incremental improvement,” 2018. [Online]. Available: https://pjreddie.com/yolo/. [11] T. Moura, A. Valente, A. Sousa, and V. Filipe, “Traffic sign recognition for autonomous driving robot,” in 2014 IEEE International Conference on Autonomous Robot Systems and Competitions, ICARSC 2014 , 2014, pp. 303–308. [12] J. Redmon and A. Farhadi, “YOLO9000: Better, faster, stronger,” in Proceedings - 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017 , vol. 2017-Janua, 2017, pp. 6517–6525. [Online]. Available: http://pjreddie.com/yolo9000/ [13] J. Cao, C. Song, S. Peng, F. Xiao, and S. Song, “Improved traffic sign detection and recognition algorithm for intelligent vehicles,” Sensors (Switzerland) , vol. 19, no. 18, sep 2019. [14] J. Yang and J. F. Coughlin, “In-vehicle technology for self-driving cars: Advantages and challenges for aging drivers,” International Journal of Automotive Technology , vol. 15, no. 2, pp. 333–340, 2014. [15] P. Dönmez, Introduction to Machine Learning The Wikipedia Guide , 2013, vol. 19, no. 2. [16] A. Osman, “The Role of Machine Learning in Autonomous Vehicles | Electronic Design,” 2020. [Online]. Available: https://www.electronicdesign.com/markets/automotive/article/21147200/ nxp-semiconductors-the-role-of-machine-learning-in-autonomous-vehicles [17] S. Chapman and N. Agashe, “Traffic Signs in the Evolving World of Autonomous Vehicles,” Tech. Rep., 2019. [18] Nvidia, “Autonomous Car Development Platform from NVIDIA DRIVE PX2.” [Online]. Available: https://www.nvidia.com/content/nvidiaGDC/sg/en_SG/self-driving-cars/drive-px/ http://www.nvidia.com/object/drive-px.html [19] K. Lim, Y. Hong, Y. Choi, and H. Byun, “Real-time traffic sign recognition based on a general purpose GPU and deep-learning,” PLoS ONE , vol. 12, no. 3, mar 2017. [Online]. Available: /pmc/articles/PMC5338798//pmc/articles/PMC5338798/?report= abstracthttps://www.ncbi.nlm.nih.gov/pmc/articles/PMC5338798/ REFERENCES 133 [20] P. Viola and M. Jones, “Rapid object detection using a boosted cascade of simple features,” Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition , vol. 1, pp. 11–18, 2001. [21] M. Y. Zhou and W. F. Lawless, “An Overview of Artificial Intelligence in Education,” in Encyclopedia of Information Science and Technology, Third Edition , 2014, pp. 2445–2452. [Online]. Available: https://www.researchgate.net/publication/ 236346414{_}AN{_}OVERVIEW{_}OF{_}ARTIFICIAL{_}INTELLIGENCE [22] Z. Mohammed, “Artificial Intelligence Definition, Ethics and Standards,” Tech. Rep. [Online]. Available: https://www.researchgate.net/publication/ 332548325{_}Artificial{_}Intelligence{_}Definition{_}Ethics{_}and{_}Standards [23] P. W. and Simon, “Too Big to Ignore : The Business Case for Big Data,” Journal of Chemical Information and Modeling , vol. 53, no. 9, pp. 1689–1699, 2013. [Online]. Available: https://www.goodreads.com/book/show/17134962-too-big-to-ignore [24] F. Chollet, Deep Learning with Python . [25] D. K. Judith Hurwitz, Machine Learning, IBM Limited Edition , 2018, vol. 35, no. 5. [Online]. Available: http://www.wiley.com/go/permissions. [26] O. Pentakalos, “Introduction to machine learning,” in CMG IMPACT 2019 , vol. 19, no. 2, 2019, pp. 285–288. [27] P. Boucher, “How artificial intelligence works,” STOA | Panel for the Future of Science and Technology , vol. 3, no. March, p. 10, 2019. [Online]. Available: https://www.globalme.net/blog/the-present-future-of-speech-recognition [28] A. Zayegh and N. Al Bassam, “Neural Network Principles and Applications,” in Digital Systems . IntechOpen, nov 2018. [29] R. E. Brown, “Donald O. Hebb and the Organization of Behavior: 17 years in the writing,” apr 2020. REFERENCES 134 [30] S. Sagar, “Activation Functions in Neural Networks | by SAGAR SHARMA | Towards Data Science,” 2017. [Online]. Available: https://towardsdatascience.com/ activation-functions-neural-networks-1cbd9f8d91d6 [31] A. Mathew, P. Amudha, and S. Sivakumari, “Deep learning techniques: an overview,” in Advances in Intelligent Systems and Computing , vol. 1141. Springer, 2021, pp. 599–608. [32] T. Pinto, “Object detection with artificial vision and neural networks for service robots,” Tech. Rep., 2018. [Online]. Available: http://repositorium.sdum.uminho.pt/handle/1822/62251http://repositorium.sdum. uminho.pt/http://repositorium.sdum.uminho.pt/handle/1822/62251{%}0Ahttp: //repositorium.sdum.uminho.pt/ [33] S. Albawi, T. A. Mohammed, and S. Al-Zawi, “Understanding of a convolutional neural network,” in Proceedings of 2017 International Conference on Engineering and Technology, ICET 2017 , vol. 2018-Janua. Institute of Electrical and Electronics Engineers Inc., mar 2018, pp. 1–6. [34] R. Sathya and A. Abraham, “Comparison of Supervised and Unsupervised Learning Algorithms for Pattern Classification,” International Journal of Advanced Research in Artificial Intelligence , vol. 2, no. 2, 2013. [35] S. Devin, “Supervised vs. Unsupervised Learning | by Devin Sony | Towards Data Science,” 2018. [Online]. Available: https://towardsdatascience.com/ supervised-vs-unsupervised-learning-14f68e32ea8d [36] A. Bisht, “ML | Classification vs Regression - GeeksforGeeks.” [Online]. Available: https://www.geeksforgeeks.org/ml-classification-vs-regression/https: //www.geeksforgeeks.org/ml-classification-vs-regression/{%}0Ahttps://www. geeksforgeeks.org/ml-classification-vs-regression/?ref=rp [37] H. Rezatofighi, N. Tsoi, J. Gwak, A. Sadeghian, I. Reid, and S. Savarese, “Generalized intersection over union: A metric and a loss for bounding box regression,” in Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition , vol. 2019-June, 2019, pp. 658–666.