Full text
Universit` a degli Studi di Padova Dipartimento di Ingegneria dell’Informazione Master Degree in ICT for Internet and Multimedia Title Seeing with Sound: Object Detection and Localization by YOLOv8 and Audio Feedback for Blind Individuals. Supervisor: Prof.ssa Federica Battisti Co-Supervisor: Prof.ssa Yeongmi Kim Candidate: Ali Tavakoli Yaraki 2040500
Academic Year 2023-2024 II
Abstract Object detection is a challenging Computer Vision (CV) application, particularly for assisting blind individuals. With the rapid advancement of Deep Learning (DL), algorithms like Convolutional Neural Network (CNN) have significantly improved video analysis and image understanding for this purpose. Blind individuals face substantial challenges when navigating indoor or outdoor environments, underscoring the pressing need for assistive technologies. In this thesis, a system has been developed to address this need, integrating the You Only Look Once (YOLO) object detection algorithm with audio guidance to aid blind users. The solution utilizes YOLOv8’s State-Of-The-Art (SOTA) deep convolutional neural network architecture to detect objects in the user’s environment, providing spatial information and counting processes through audio feedback. The system, equipped with a text-to-speech engine, converts all the information into verbal instructions, in some cases acting as a virtual assistant shape program. This context-aware feedback, available in multiple languages, has been optimized for webcams as a real-time scenario, images, and videos. The system has shown promising results, enhancing the autonomy and quality of life for blind users, a significant step towards addressing the challenges they face in daily environments. Keywords: Object detection, blind people, audio feedback, spatial location, YOLO, computer vision, bounding box. III
Acknowledgements I express my deepest gratitude to my supervisors at the University of Padova and the Management Center Innsbruck (MCI). I thank Professor. Federica Battisti for her kind and helpful support and feedback throughout this journey. I am equally grateful to Professor Yeongmi Kim for assigning me a title that perfectly fits my interests and guiding me toward this research direction. I am incredibly thankful to my family, partner, and friends, whose unwavering support and encouragement have been my endless source of strength. I also extend my sincere appreciation to the University of Padova for the comprehensive education and knowledge imparted to me and MCI for providing an impactful internship experience that demanded hard work and dedication. Lastly, I would like to thank Cybathlon for inspiring me with the idea for this thesis and their motivating influence throughout the process. V
Contents Abstract III Acknowledgements V 1 Introduction 1 1.1 Background and Motivation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1 1.2 ProblemStatement ................................. 2 1.3 ObjectiveofTheStudy ............................... 3 1.4 Deep learning for computer vision application . . . . . . . . . . . . . . . . . . 3 1.4.1 ConvolutionalLayer ............................ 4 1.4.2 PoolingLayer................................ 4 1.4.3 Networktraining.............................. 4 1.4.4 Backpropagation .............................. 4 1.4.5 Fully Connected Layers . . . . . . . . . . . . . . . . . . . . . . . . . . 5 1.5 ObjectDetection................................... 5 1.5.1 Single stage detection . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 1.5.2 Twostagedetection ............................ 6 1.6 Object detection models performance evaluation metrics . . . . . . . . . . . . 6 1.6.1 Precision................................... 7 1.6.2 Recall .................................... 7 1.6.3 F1-Score................................... 7 1.6.4 AveragePrecision.............................. 7 1.7 YOLO(YouOnlyLookOnce)............................ 8 1.7.1 Step-by-Step Insights into YOLO . . . . . . . . . . . . . . . . . . . . . 9 1.7.2 Residualblocks............................... 9 1.7.3 Bounding box regression . . . . . . . . . . . . . . . . . . . . . . . . . . 9 1.7.4 Intersection over Union . . . . . . . . . . . . . . . . . . . . . . . . . . 10 VII
1.7.5 Non-Maximum suppression . . . . . . . . . . . . . . . . . . . . . . . . 11 2 Related Works 13 2.1 A Comprehensive Review of YOLO Architectures in Computer Vision: From YOLOv1 to YOLOv8 and YOLO-NAS . . . . . . . . . . . . . . . . . . . . . . . . 13 2.2 Real-Time Object Detection Using Yolo Algorithm for Blind People . . . . . . 14 2.3 An Real Time Object Detection Method for Visually Impaired Using Machine Learning ....................................... 15 2.4 Object Detection for Blind People Using Yolov3 . . . . . . . . . . . . . . . . . 15 2.5 Understanding of Object Detection Based on CNN Family and YOLO . . . . . 16 2.6 Real-Time Object Detection with YOLO . . . . . . . . . . . . . . . . . . . . . . 16 2.7 An Assistive Model for Visually Impaired People using YOLO and MTCNN . . 17 2.8 Object Detection System For The Blind With Voice guidance . . . . . . . . . . 17 2.9 Real Time Object Detection with Audio Feedback using Yolo vs. Yolov3 . . . . 18 2.10 Real Time Object Detection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 2.11 Object Detection Using Machine Learning for Visually Impaired People . . . . 19 2.12 Object Detection Featuring 3D Audio Localization for Microsoft HoloLens: A Deep Learning Based Sensor Substitution Approach for the Blind . . . . . . . 19 2.13 Deep learning based object detection and surrounding environment description for visually impaired people . . . . . . . . . . . . . . . . . . . . . . . . . . 20 2.14 Object Detection with Voice Guidance to Assist Visually Impaired Using YOLOv7 20 2.15 Android Based Object Recognition for Visually Impaired . . . . . . . . . . . . 21 2.16 Object Detection and Recognition for a Pick and Place Robot . . . . . . . . . . 21 2.17 Assistive Technologies for Obstacle Detection and Identification . . . . . . . . 22 2.18 Absolute Distance Prediction Based on Deep Learning Object Detection and Monocular Depth Estimation Models . . . . . . . . . . . . . . . . . . . . . . . 22 2.19 ExistingSystems................................... 23 2.19.1 WearableDevices.............................. 23 2.19.2 Smartphone Applications . . . . . . . . . . . . . . . . . . . . . . . . . 24 2.19.3 Robotics................................... 24 2.19.4 Websites................................... 24 3 Proposed Method 25 3.1 ToolsandLibraries ................................. 25 VIII
3.2 Dataset and Data Augmentation . . . . . . . . . . . . . . . . . . . . . . . . . . 25 3.3 YOLOv8 ....................................... 26 3.4 Training ....................................... 30 3.5 Phase1........................................ 33 3.5.1 ObjectDetection .............................. 33 3.6 Phase2........................................ 35 3.6.1 ObjectCounting .............................. 35 3.7 Phase3........................................ 35 3.7.1 ObjectDistance............................... 35 3.7.2 Object Spatial Location . . . . . . . . . . . . . . . . . . . . . . . . . . . 36 3.8 Phase4........................................ 38 3.8.1 TexttoSpeech ............................... 38 3.8.2 Generating the Audio Feedback . . . . . . . . . . . . . . . . . . . . . . 40 3.8.3 Delivering the Audio Feedback . . . . . . . . . . . . . . . . . . . . . . 40 3.8.4 LanguageModifying............................ 40 3.9 SpeechRecognition................................. 41 4 Results and Discussion 45 4.1 SystemImplementation............................... 45 4.2 ExperimentalSetup ................................. 46 4.2.1 Framework ................................. 46 4.2.2 Datasetlimitation.............................. 46 4.3 Performance Metrics and Evaluation Criteria . . . . . . . . . . . . . . . . . . . 47 4.3.1 Accuracy .................................. 47 4.3.2 TheLoss................................... 49 4.3.3 Speed .................................... 50 4.4 Analyzing of Model Performance . . . . . . . . . . . . . . . . . . . . . . . . . 50 4.4.1 ObjectDetection .............................. 50 4.4.2 Object’s Spatial Location . . . . . . . . . . . . . . . . . . . . . . . . . . 54 4.5 Audiofeedbackanalysing.............................. 55 4.6 A Prototype of Virtual Assistant . . . . . . . . . . . . . . . . . . . . . . . . . . 56 5 Limitations and Future Work 59 6 Conclusions 61 IX
1.3 Objective of The Study This thesis proceeds into integrating the latest version of YOLO, YOLOv8, with audio guidance systems to develop an assistive device specifically designed for blind individuals. The system aims to identify objects in the surroundings and deliver real-time audio descriptions, effectively converting visual information into auditory feedback. By using YOLO’s enhanced precision and processing speed, this project aims to improve real-time applications in various and rapidly varying circumstances. The implementation involves setting up a framework where a webcam captures scenes, which is then processed by the YOLO algorithm to detect and identify objects. The implementation of the same procedure on images and videos was also examined. The detected objects are subsequently relayed to a text-to-speech system that converts the information into audio signals, guiding the user about the location and counting of nearby objects. This setup not only enhances spatial awareness for blind users but also promotes their independence by reducing reliance on human assistance. 1.4 Deep learning for computer vision application Computational models inspired by the composition and operations of the human brain are known as neural networks. They are made up of layer-organized, networked nodes known as neurons. Every neuron generates an output signal after applying an activation function (such as Relu, Tanh) to receive input signals. By varying the weights of the connections between neurons, neural networks can learn from data. DL allows computational models composed of multiple processing layers to learn data representations with numerous levels of abstraction. These methods have dramatically improved the SOTA in speech recognition, visual object recognition, object detection, and many other domains, such as drug discovery and genomics. Deep learning discovers complex structures in large data sets using the backpropagation algorithm to indicate how a machine should change its internal parameters used to compute each layer’s representation from the previous layer’s. Deep convolutional nets have brought about breakthroughs in processing images, video, speech, and audio, whereas recurrent convolutional networks nets have shone light on sequential data such as text and speech[13] In exploring DL architectures, particularly CNN , several critical components and processes define the effectiveness and efficiency of model training and operation. Among the 3
foundational elements of CNNs are the convolutional layer, pooling layer, fully connected layer, network training, and the backpropagation algorithm. Each of these aspects are essential for building robust and efficient neural network models. 1.4.1 Convolutional Layer This is the first significant layer in CNNs. It involves filters that move across the input image and capture important features like edges, colors, patterns and textures. These filters help the network focus on specific parts of the image. 1.4.2 Pooling Layer There’s often a pooling layer after the convolutional layer. Pooling helps reduce the size of the data the network needs to process. It does this by summarizing the features captured in small areas of the image. For example, it might take the most significant number (max pooling) or the average (average pooling) from a small square in the picture. This makes the network faster and more efficient. 1.4.3 Network training involves multiple steps: Initialization: Set initial weights and biases. These could be random or follow a specific distribution. Forward Propagation: Input data passes through the network, layer by layer, until the output layer is reached. Loss Calculation: Use a loss function to compute the difference between the network output and the actual values. This measures the network’s performance. 1.4.4 Backpropagation To optimize the neural network, it is essential to calculate the gradient of the loss function concerning each weight. This involves tracing the impact of the loss function through each layer of the network and then adjusting the weights to minimize the loss. This process, known as backpropagation, allows for the iterative refinement of the weights, ultimately leading to improved network performance. 4
1.4.5 Fully Connected Layers Often follow convolutional and pooling layers in CNN architectures and play a crucial role in decision-making processes. These layers integrate learned features into final outcomes such as class scores, Where every neuron in one layer is connected to every neuron in the subsequent layer. They are critical in synthesizing the features extracted by convolutional and pooling layers into outputs that pertain to the task at hand. 1.5 Object Detection Object detection is a CV task that involves identifying and locating objects in images or videos. It is essential in many applications, such as surveillance, self-driving cars, or robotics. The process will combine elements of an image by classification, which will represent what object categories are present in an image, and localization, which brings out the location of the objects in an image by drawing some bounding boxes around them. These features will typically be done by using either two-stage detectors, which first propose regions and then classify them, or one-stage detectors like, which simplify this process by doing both tasks in one shot Figure 1.1: Object detection example 1.5.1 Single stage detection Single-shot object detection uses a single pass of the input image to make predictions about the presence and location of objects in the image. It processes an entire image in a single pass, making them computationally effective. However, single-shot object detection is generally less 5
accurate than other methods, and it’s less effective in detecting small objects. Such algorithms can be used to detect objects in real time in resource-constrained environments.YOLO is a single-shot detector that uses a fully CNN to process an image. 1.5.2 Two stage detection Two-shot object detection uses two passes of the input image to make predictions about the presence and location of objects. The first pass is used to generate a set of proposals or potential object locations, and the second pass is used to refine these proposals and make final predictions. This approach is more accurate than single-shot object detection but is also more computationally expensive. In this thesis, single-shot object detection was deployed to better suit real-time applications. Figure 1.2: Object detection stages 1.6 Object detection models performance evaluation metrics To determine and compare the predictive performance of different object detection models, we need standard quantitative metrics. Below the three most common evaluation metrics are represented. 6
1.6.1 Precision It is a necessary metric in model evaluation as it serves to quantify the accuracy of the positive predictions made by the model. It represented how well the model distinguishes true objects from false positives. Precision provides insight into the model’s ability to make positive predictions that are indeed accurate. A high precision score indicates that the model is skilled at avoiding false positives and provides reliable positive predictions. Precision =T P T P +F P 1.6.2 Recall Recall, also known as perceptiveness or true positive rate, is another essential metric to evaluate model performance. It measures the model’s capability to capture all appropriate objects in the image and evaluates the model’s completeness in identifying objects of interest. A high recall score suggests that the model effectively identifies most of the relevant objects in the data. Recall =T P T P +F N 1.6.3 F1-Score Another metric related to the trade-off mean of precision and recall, the F1-score, provides a measure of the model’s performance, considering both false positives and false negatives. This metric is particularly useful when there is an imbalance between positive and negative classes in the dataset. F1= 2 ·Precision ·Recall Precision +Recall 1.6.4 Average Precision Average Precision (AP) is calculated as the area under a precision vs. recall curve for a set of predictions. Recall and precision offer a trade-off that is graphically represented into a curve by varying the classification threshold. The area under this precision vs. recall curve gives us the Average Precision per class for the model. The average of this value, taken over all classes, is called Mean Average Precision (mAP). In object detection, precision and recall are not used for class predictions. Instead, they serve as predictions of boundary boxes for 7
measuring the decision performance. An IoU value > 0.5. is taken as a positive prediction, while an IoU value < 0.5is a negative prediction. 1.7 YOLO (You Only Look Once) YOLO is a popular object detection algorithm that has transformed field of CV. It is fast and structured, making it a great choice for almost all object detection tasks. Moreover, It has achieved state-of-the-art performance on diverse standards and has been widely adopted in various real-world applications. YOLO is widely known as unified network and very fast compared to Faster Region-based Convolutional Neural Network (RCNN) and runs using single convolutional neural network.[14] Despite limitations such as struggling with small objects and the inability to perform fine-grained object classification, YOLO gains more than twice the mAP of other real-time systems. It has been confirmed to be a beneficial tool for object detection and has opened up many new possibilities for researchers and practitioners. YOLO, or You Only Look Once, introduced a novel approach that unified object detection and classification into a single neural network model, enabling real-time performance without compromising accuracy. Unlike traditional methods that relied on sliding window approaches or region proposal networks, YOLO divided the input image into a grid and predicted bounding boxes and class probabilities directly from this grid. By considering the entire image at once, YOLO achieved impressive detection speed while maintaining competitive accuracy levels. This approach streamlined the object detection pipeline, making it suitable for a wide range of applications requiring fast and efficient object detection capabilities. Our detection network has 24 convolutional layers followed by 2 fully connected layers. Alternating 1*1 convolutional layers reduce the features space from preceding layers. We pretrain the convolutional layers on the ImageNet classification task at half the resolution (224 * 224 input image) and then double the resolution for detection.[1] Figure 1.3: YOLO architecture[1] 8
1.7.1 Step-by-Step Insights into YOLO 1.7.2 Residual blocks By dividing the original image into N*N grid cells of equal shape YOLO will start the processing. Each cell in the grid is responsible for localizing and predicting the class of the object that it covers, Onwards with the probability/confidence value. Figure 1.4: Grid cells 1.7.3 Bounding box regression After dividing image, the algorithm will determine the bounding boxes which correspond to rectangles highlighting all the objects in the image. There is no limitations regarding the bounding boxes it can be as many object as in the image. YOLO controls the attributes of these bounding boxes using a single regression module in the following format, where Y is the final vector representation for each bounding box. Y = [pc, bx, by, bh, bw, c1, c2] This is especially important during the training phase of the model. pc corresponds to the probability score of the grid containing an object. For instance, all the grids in red will have a probability score higher than zero. The image on the right is the simplified version since the probability of each yellow cell is zero (insignificant). bx, by are the x and y coordinates of the center of the bounding box with respect to the enveloping grid cell. bh, bw correspond to the height and the width of the bounding box. c1 and c2 correspond to the two classes. 9
Figure 1.5: Bounding box regression 1.7.4 Intersection over Union IoU is a metric to measure localization accuracy and calculate localization errors regarding object detection models. To calculate the IoU between the predicted and the ground truth bounding boxes, first, the intersecting area between the two corresponding bounding boxes for the same object is examined, then the calculation of the total area covered by the two bounding boxes— also known as the “Union” and the area of overlap between them called the “Intersection.” The metric mentioned above gives the ratio of the overlap to the total area, providing a good estimate of how close the prediction bounding box is to the original bounding box. Figure 1.6: IoU 10
1.7.5 Non-Maximum suppression NMS As a post-processing technique used in YOLO as the last. It is used to remove redundant detections and retain only the most confident ones. It compares the predicted bounding boxes of the detected objects and discards the ones that have a high overlap or intersection. NMS involves two main steps, Thresholding and Suppression. The thresholding step filters out detections with low confidence scores. The suppression step removes redundant detections that have a high overlap with each other, retaining only the most confident ones.By performing NMS, YOLO can reduce the number of positives and improve the accuracy of object detection. Figure 1.7: NMS 11
2.11 Object Detection Using Machine Learning for Visually Impaired People Integrating advanced machine learning techniques has shown promising potential to enhance object detection abilities. This paper significantly contribute to the mentioned area by implementing a system that utilizes RetinaNet and neural network technologies to facilitate accurate navigation aids for indoor and outdoor settings. The system looked at different environments, crowded and empty, and obtained great accuracy during the challenging area by looking at different algorithms and exploring them deeply theoretically, It identifies objects with high accuracy and communicates this information through audio outputs, thus aiding visually impaired individuals in understanding their surroundings effectively. While the system shows potential, its performance across diverse environmental conditions and its dependence on specific hardware configurations highlight areas for future improvement. These findings underscore the need for ongoing research to refine and adapt machine learning applications in assistive technologies, ensuring they are robust, versatile, and widely accessible[23]. 2.12 Object Detection Featuring 3D Audio Localization for Microsoft HoloLens: A Deep Learning Based Sensor Substitution Approach for the Blind The integration of augmented reality (AR) and artificial intelligence (AI) offers novel tracks to enhance navigation and interaction within the environment. An innovative application of this integration in their development of a system that combines YOLOv2-based object detection with 3D audio localization via Microsoft HoloLens was studied. Figure 2.4: User interaction Through system[4] This system detects objects in real-time and also provides spatial audio feedback,a speech recognition program was setup to interact with the user. In order to test how well voice com19
mands are recognized and processed, the debug output for each recognized voice command has been used. The voice commands that should be recognized were limited to “Scan”, “Start”, “Stop” and “Chair”. Such technological advancements underscore the potential of AR and AI to transform everyday experiences for individuals with visual impairments, providing them with greater independence and mobility. The reliance on continuous network connectivity and the processing demands on external servers highlight areas for further development, particularly in enhancing the system’s autonomy and operational efficiency[4]. 2.13 Deep learning based object detection and surrounding environment description for visually impaired people Deep learning for both object detection and environmental interaction represents a forward leap in functionality and user experience.regarding this paper , It introduces a comprehensive system that detects objects using a TensorFlow-based SSDLite MobileNetV2 model and then describes the surrounding environment through an innovative ’ambiance mode.’ Trained on a diverse dataset, including weather conditions, this system offers auditory feedback that enhances spatial and situational awareness for visually impaired users. the system will demonstrate robust performance metrics on standard benchmarks relying on Raspberry Pi hardware that presents scalability and environmental adaptability challenges.[24]. 2.14 Object Detection with Voice Guidance to Assist Visually Impaired Using YOLOv7 Innovative applications of deep learning models in assistive technologies for the visually impaired have shown marked progress, particularly with integrating YOLOv7 in object detection systems. Authors exemplify this advancement by combining YOLOv7’s precise detection capabilities with voice guidance technology, offering a practical solution for enhancing the independence of visually impaired individuals. This system identifies objects with high accuracy and communicates this information through voice, thus making spatial navigation more accessible. 20
Figure 2.5: Design diagram[5] Future developments could focus on reducing the system’s computational demands and extending its applicability to more complex environments[5]. 2.15 Android Based Object Recognition for Visually Impaired The research by Saeed presents a feasible solution for real-time, offline object recognition on Android devices. This solution significantly benefits visually impaired users by enhancing their environmental awareness. The system was evaluated using a dataset of 600 images across various conditions. It achieved the highest accuracy, at 95.36%. The primary challenge was ensuring accurate object recognition under varying conditions without excessive computational load. Optimization of classifier parameters and feature selection were critical to balancing performance and resource usage. Future improvements include expanding the object database, incorporating more complex recognition tasks, and further optimizing the system for better performance and lower energy consumption. The proposed Android-based object recognition system will be used to implement a money reader for the Egyptian currency. No such money reader is known among the blind and visually impaired[25]. 2.16 Object Detection and Recognition for a Pick and Place Robot Recent studies have demonstrated that YOLO has an impressive processing speed of 155 frames per second while achieving double the mean Average Precision (mAP) compared to other realtime detectors. Despite this notable efficiency, the methodology has its limitations. For exam21
ple, while contemporary detection techniques have effectively reduced the occurrence of false positives, YOLO is inclined to exhibit more localization errors. Anyhow, YOLO demonstrates exceptional generalization capabilities across diverse domains. For instance, it outperforms approaches such as the Deformable Parts Model (DPM) and R-CNN in transferring knowledge from natural images to other domains, including art, exhibiting adaptability across various applications.[26]. 2.17 Assistive Technologies for Obstacle Detection and Identification Applications centered on sensing obstacles near the user and alerting them through alarms or beeping sounds implemented on sensory devices such as ultrasonic, smart sticks with obstacle detectors, mobile phones, and navigators. These devices are expensive and could sometimes be a barrier to the user. In some situations, tactile signs and Braille texts are labeled at the top of the items for identification. Moreover, High-tech systems such as Radio Frequency Identification Devices (RFID), barcodes, or talking labels can be used to find objects in near distance[27][28]. 2.18 Absolute Distance Prediction Based on Deep Learning Object Detection and Monocular Depth Estimation Models The You Only Look Once (YOLOv5) model, which is renowned for its ability to detect objects in real time with high accuracy, specifically enhances speed and performance over its predecessors, making it an effective tool for detecting and localizing objects within an image. The challenge of depth estimation using a single (monocular) camera is established since monocular cameras capture 2D images, making it challenging to intimate depth directly. To address this, the authors use a deep autoencoder network, trained in a self-supervised form, to estimate depth from the image, providing relative distances within the scene. By integrating these two techniques, the framework effectively predicts the absolute distance to objects, achieving an accuracy of 96% with a Root Mean Square Error (RMSE) both the method and result out of this study is presented in Figure [6]. 22
Figure 2.6: Visual process of whole network[6] Figure 2.7: Estimated distance vs. absolute distance[6] 2.19 Existing Systems 2.19.1 Wearable Devices Wearable devices are widely used in object detection and recognition for visually impaired individuals. These devices typically incorporate sensors and cameras to detect objects in the user’s surroundings, providing audio or tactile feedback. An example is the Orcam MyEye, which utilizes a camera and Optical character recognition (OCR) technology to read text and recognize faces, products, and more. 23
Figure 2.8: Orcam MyEye device 2.19.2 Smartphone Applications Numerous smartphone applications assist visually impaired individuals in object recognition by leveraging the smartphone’s camera and computer vision algorithms for real-time detection and classification. Examples include the Be My Eyes app, which connects blind users with sighted volunteers for object identification, and Seeing AI by Microsoft, which can recognize faces, read text, and identify objects. 2.19.3 Robotics Robotics represents another approach to object detection and recognition for visually impaired users. Researchers have developed robotic systems that can detect and recognize objects and navigate environments to provide assistance. An example is the Blind Explorer, a robotic platform capable of detecting obstacles and offering audio feedback to help users navigate through complicated environments. 2.19.4 Websites Considerable websites have been developed to assist visually impaired individuals with object detection and other tasks by employing web-based technologies and computer vision algorithms for real-time object detection and classification. These websites typically utilize useruploaded images and live streams to analyze objects, text, and scenes on the website or to provide remote individual assistance. One example is a web captioner, which offers realtime captioning and text recognition. 24
Chapter 3 Proposed Method This chapter provides a comprehensive overview of the methodology employed in each part of the project, outlining the precise procedures and approaches utilized to ensure successful execution. 3.1 Tools and Libraries Various libraries, such as CV Numpy and the others, Torch as the framework, are installed and used to succeed in different tasks. The YOLO model imported from ultralytics libraries who has produced the intended YOLO version[7]. The Pygame Python module is used to play the audio. Pyttsx3 and Google Text-To-Speech (gTTS) are libraries that provide a text-tospeech conversion engine in Python offline and online, respectively. The second one, gTTS, was presented only to support a wide range of languages in the algorithm to tackle concerning distinct languages. The main algorithm will work on the Pyttsx3 to provide more accessible feedback without worrying about an internet connection. It is also faster than gTTS providing both male and female voices. 3.2 Dataset and Data Augmentation The dataset included 2282 images and 3523 annotations around the specific objects, which were completely manually and carefully selected to achieve the best result. Only square shape box was determined for the labeling object that can be used related to bounding boxes. While there were two other options, such as smart polygon and smart labeling, that had a misunderstanding between segmentation and detection tasks for the algorithm. The detection, labeling, and annotating data was used on RoboFlow[29], a web-based platform. 25
The dataset covers both indoor and outdoor navigation objects, with 17 classes related to supermarkets, shopping, walking tours, and home surfing. After labeling and annotating objects, data augmentation was considered for almost all variations, such as flipping horizontal and vertical, to create a robust model for the training process and tackle the algorithm’s limitation of not detecting objects in diverse environments through bad lightening areas, noisy data, and disparities in color. Hence, data transformations such as Rotation about 90°, Clockwise, Counter-Clockwise, Upside Down, Grayscaling was considered for 17% of images, Hue and saturation Between -19°and +19°, Between -29% and +29% respectively, Brightness Between -17% and +17%, Exposure Between -13% and +13% , Blur Up to 2.5 pixels and Noise Up to 1.29% of pixels were adjusted and added to the training data. Furthermore, The Mosaic augmentation was automatically applied during training by YOLOv8, but it is disabled before the last 10 epochs. Data augmentation has been accomplished using the RoboFlow platform as well[29]. Figure 3.1: Sample of data augmentation 3.3 YOLOv8 YOLOv8 architecture is divided into three essential parts: Backbone, Neck, and Head. The backbone acts as a feature extractor, capturing patterns in the first layers, such as edges 26
and textures. The neck part will perform as a bridge between the head and the backbone. It gathers feature maps from different stages of the backbone and provides information related to the accuracy and speed of the model. Another functionality of this part is Performing concatenation or fusion of features of different scales to ensure the network can detect objects of different sizes. The head is the final part of the network and is responsible for generating the network’s outcomes, such as bounding boxes and confidence scores for object detection. The YOLO architecture utilizes a local feature analysis approach instead of processing the entire image at once, aiming to reduce computational efforts and facilitate real-time detection. The algorithm employs convolutions multiple times to extract feature maps. Convolution, a mathematical operation that combines two functions to produce a third one, is commonly used to apply filters to images or signals in computer vision and signal processing, thereby emphasizing specific patterns. Additionally, it extracts features from inputs such as images in CNN. Convolutions are characterized by Kernels, Strides, and Paddings. The kernel, also known as the filter, is a small array of numbers that moves slowly across the input (image or signal) during the convolution operation. The goal of Kernels is to apply local operations to the input to detect specific characteristics. Each element in the kernel represents a weight multiplied by the corresponding value in the input during convolution. 27
Figure 3.2: Kernel operation Another factor related to Convolution is stride, which is the amount of displacement the kernel considers as it moves across the input. A stride of 1 means the kernel moves one position at a time. Stride directly influences the spatial dimensions of the convolution output. Larger of them can decrease the dimensionality of the output and reduce computational effort, thereby increasing the speed of the operation, which can directly impact quality. while smaller strides retain more spatial information. Padding refers to adding extra pixels around the edges of the input image, also known as zero padding. It occurs before applying convolution operations. This ensures that information at the image’s edges is treated the same way as information in the center during convolution 28
3.6 Phase 2 3.6.1 Object Counting The number of the detected objects is followed based on their labels and spatial locations. Adefaultdict named detected objects is initialized to store the counts and locations of each object type. Each object’s count is incremented in detected objects [label][’count’]. The feedback is created by repeating through the detected objects and generating illustrative sentences based on their numbers and locations. 1# Initialize defaultdict to store counts and locations 2detected objects = defaultdict(lambda: –'count': 0, 'locations': defaultdict(int) ) 3 4for i in range(len(labels)): 5 label = model.names[int(labels[i])] 6 x1, y1, x2, y2 = int(coords[i][0]), int(coords[i][1]), int(coords[i ][2]), int(coords[i][3]) 7 8# Access count for a specific label 9detected objects[label]['count'] 10 11 # Adding count 12 detected objects[label]['count'] += 1 Figure 3.9: Code snippets for counting objects 3.7 Phase 3 3.7.1 Object Distance Three distinct methods were studied to detect the object spatial location. The simplest one was measuring the distance of an object from the camera by knowing its actual width/height and the focal length of the camera, in addition to the sensor size either determining an estimated width/height for each class and measuring the distance with the help of a third-party app to join the mobile phone’s camera to the laptop. This strategy was not as accurate as needed since It is not easy to assign only one size for the images in a class, for example, one size for all images related to class CAR, since the size of the classes is different image by image. Distance =Actual W/H of the Object ×Focal Length W/H of the Object in Image Another experienced method was looking at the size of the bounding boxes and setting 35
a threshold for them to determine three stages such as Close, Far, Very Far. However, the algorithm had some errors regarding this technique; for example, it looked at the bounding boxes of the objects and estimated the ones with low Confidence thresholds with the small size of the bounding boxes around them as a ”far” distance since it was not always true. 3.7.2 Object Spatial Location The mentioned methods were not available to completely accomplish through user experience, Therefore, best approach was instead of information related to distance of the object , dividing the images into a 3x3 grid to categorize the location of each detected objects. It will use the center point of the object’s bounding box to determine its position. x1: Left boundary of the bounding box. y1: Top boundary of the bounding box. x2: Right boundary of the bounding box. y2: Bottom boundary of the bounding box. img width, img height represented the image dimensions, These provide the width and height of the image, necessary for converting normalized coordinates into pixel values and for location categorization. 1img width = img.shape[1] 2img height = img.shape[0] Figure 3.10: Image dimensions. The center of the bounding box is calculated by averaging the x-coordinates and ycoordinates of the bounding box: center x = (x1 + x2) / 2 center y = (y1 + y2) / 2 This center point is used to determine the location category. The image is conceptually divided into a grid with nine regions. The divisions are: Horizontal: Divided into three equal columns: left, center, right. Vertical: Divided into three equal rows: top, middle, bottom. The boundaries of these divisions are: 36
Horizontally: img width 3: First column boundary. 2·img width 3: Second column boundary. Vertically: img height 3: First row boundary. 2·img height 3: Second row boundary. 1def get location category(x1, y1, x2, y2, img width, img height): 2 center x = (x1 + x2) / 2 3 center y = (y1 + y2) / 2 4 if center x < img width / 3: 5 if center y < img height / 3: 6 return ”top-left” 7 elif center y < 2 * img height / 3: 8 return ”left” 9 else: 10 return ”bottom-left” 11 elif center x < 2 * img width / 3: 12 if center y < img height / 3: 13 return ”up” 14 elif center y $<$ 2 * img height / 3: 15 return ”center” 16 else: 17 return ”down” 18 else: 19 if center y < img height / 3: 20 return ”top-right” 21 elif center y < 2 * img height / 3: 22 return ”right” 23 else: 24 return ”bottom-right” Figure 3.11: Location category based on coordinates The function get location category takes the bounding box coordinates and the image dimensions as inputs. It uses the following logic to categorize the location in nine equal sections. 37
Column Formula Spatial location Left Column centerx<img width 3 Top centery<img height 3top-left Middle img height 3≤centery< 2·img height 3 left Bottom centery≥2·img height 3bottom-left Middle Column img width 3≤centerx< 2·img width 3 Top centery<img height 3up Middle img height 3≤centery< 2·img height 3 center Bottom centery≥2·img height 3down Right Column centerx≥2·img width 3 Top centery<img height 3top-right Middle img height 3≤centery< 2·img height 3 right Bottom centery≥2·img height 3bottom-right Table 3.1: Categorizing the location. top-left up top-right left center right bottom-left down bottom-right Table 3.2: Visualization of the 3x3 grid 3.8 Phase 4 3.8.1 Text to Speech The process of providing audio feedback for detected objects involves aggregating information such as their names, numbers and locations, converting this data into spoken feedback using a text-to-speech engine. This step enhances user interaction by allowing them to receive immediate, audible updates about the objects detected in a scene. 1# Initializing the text-to-speech engine 2engine = pyttsx3.init() Figure 3.12: initializing text-to-speech engine The results contains the detection results from the YOLO model, labels contain the detected class indices and coords contains the normalized coordinates (x1, y1, x2, y2) and confidence 38
scores for the bounding boxes. Defaultdict is a specialized dictionary from the collections module. It initializes entries with a default structure with keys count and locations. count keeps track of the total number of each type of detected object. locations uses another defaultdict to count occurrences of each object in specific spatial categories. The function provide audio feedback encapsulates the entire process in order to generating and delivering audio feedback. 1def provide audio feedback(results, img width, img height): 2 labels = results[0].boxes.cls 3 coords = results[0].boxes.xyxy 4 detected objects = defaultdict(lambda: –'count': 0, 'locations': defaultdict(int) ) 5 6 for i in range(len(labels)): 7 label = model.names[int(labels[i])] 8 x1, y1, x2, y2 = int(coords[i][0]), int(coords[i][1]), int(coords[i ][2]), int(coords[i][3]) 9 location category = get location category(x1, y1, x2, y2, img width , img height) 10 detected objects[label]['count'] += 1 11 detected objects[label]['locations'][location category] += 1 Figure 3.13: Code Snippets Through a Loop Regarding the function ”def provide audio feedback (results)” a looping procedure implemented for each detected object updating it’s counts and spatial location by extract its label and bounding box coordinates and converts the normalized coordinates to pixel values by multiplying by the image dimensions. It also employs the label index to get the actual class name from the YOLO model names array. get location category calls the function to determine the spatial category of the object. Increment the count for the object’s label and the count for its location in the detected objects dictionary. 39
3.8.2 Generating the Audio Feedback 1if detected objects: 2 feedback parts = [] 3 for label, info in detected objects.items(): 4 count = info['count'] 5 if count == 1: 6 location = next(iter(info['locations'])) 7 feedback parts.append(f”one –label at the –location ”) 8 else: 9 locations = [f”–loc count at the –loc ” for loc, loc count in info['locations'].items()] 10 feedback parts.append(f”–count –label s: ” + ”, ”.join( locations)) 11 feedback = ”I see: ” + ”, ”.join(feedback parts) 12 else: 13 feedback = ”No objects detected.” Figure 3.14: Feedback Generation By inspecting the detected objects dictionary, algorithm checks if any objects were detected, Creating a list to store parts of the feedback message. finally it Joins all parts into a single feedback message for example : ”I see one object at the location”. If no objects are detected, the feedback message set to ”No objects detected”. 3.8.3 Delivering the Audio Feedback 1engine.say(feedback) 2engine.runAndWait() Figure 3.15: Delivering the Audio Feedback Audio feedback will be provided to the user immediately after processing. The output will be represented on the console as well. The engine will activate audio feedback to announce the result, ”first specifying the total number of classes, the ones with more detection, separated each with a matching location, and then the objects with less detection”. The feedback will be stored in MP3 or any other supported format. 3.8.4 Language Modifying gTTS will allow the Algorithm to perform on additional languages to have compatibility with various users. This process impacted future work focusing on the virtual assistant shape pro40
gram. 1# Translating audio to Italian 2translated feedback = translator.translate(feedback, src='en', dest='it'). text 3 4# Print translated feedback 5print(translated feedback) 6 7# modify language to supported ones in google translate : de,fr etc 8tts = gTTS(translated feedback, lang='it') 9tts.save('audio feedback.mp3') 10 audio = AudioSegment.from mp3('audio feedback.mp3') 11 play(audio) 12 os.remove('audio feedback.mp3') Figure 3.16: Language Modifying Script By Modifying the dest string to accessible languages from Google Translate, the feedback will also provide the represented language. Figure 3.16 demonstrated feedback in console in Italian language. 3.9 Speech Recognition A semi-virtual assistant program has been incorporated into the system to improve user experience and engagement. This speech recognition algorithm will specifically address the streaming functionality, particularly the webcam feature. Upon activation, the algorithm will initiate by asking, ”How can I help you?” and then remain in a listening mood for the user’s command. Once the user issues the command, ”What do you see?” the webcam will be activated, allowing it to provide feedback on any detected objects. 41
1 2# Recognizing voice commands 3def recognize command(): 4 recognizer = sr.Recognizer() 5 with sr.Microphone() as source: 6 print(”Listening...”) 7 audio = recognizer.listen(source) 8 9 try: 10 command = recognizer.recognize google(audio) 11 print(f”User said: –command ”) 12 return command.lower() 13 except sr.UnknownValueError: 14 engine.say(”Sorry, I did not understand that.”) 15 engine.runAndWait() 16 return None 17 18 # manage the interaction 19 def main(): 20 cap = cv2.VideoCapture(0) # Open the webcam 21 if not cap.isOpened(): 22 print(”Error: Could not open video stream.”) 23 return 24 25 while True: 26 ret, frame = cap.read() 27 if not ret: 28 break 29 30 # Recognize voice commands 31 command = recognize command() 32 if command: 33 if ”what do you see” in command: 34 detect and announce(frame) 35 elif ”exit” in command or ”quit” in command: 36 engine.say(”Exiting the program.”) 37 engine.runAndWait() 38 break 39 else: 40 engine.say(”Command not recognized. Please try again.”) 41 engine.runAndWait() 42 43 cap.release() 44 45 if name == ” main ”: 46 main() Figure 3.17: Speech recognition code snippets On the other hand, the user can input the command ”capture image”, and the algorithm will utilize the webcam to capture an image, process it, and generate audio feedback. This functionality is currently limited to these features. However, plans are in place to expand its capabilities by integrating two buttons into the hardware. The first button would work when the user holds it and request the algorithm to describe the environment in real-time. The other button would act as a second choice, allowing user to press it, capture the environment, and 42
then process it by the algorithm to provide feedback. This approach, which involves integrating hardware and camera sensors, aims to offer extensive assistance for blind individuals. 43
Figure 4.7: Detection result the model evaluates its predicted bounding boxes by comparing them with the true boxes using IoU to measure how accurately the model has located the objects. Higher IoU scores indicate more accurate localization. After IoU determined, YOLO will use NMS as a post processing technique to eliminate duplicate bounding boxes for the same object, keeping only the most confident predictions. The workload of IoU, NMS and Confidence threshold could be significant, especially in blind assistance scenarios where providing feedback could become challenging due to the presence of many objects. For this reason, an IoU of 0.7 was determined to safely detect ordinary amount of objects. 51
Figure 4.8: Model first performance In Figure 4.8, there are two buses with a confidence threshold above 0.90 and two cars with different confidence thresholds, one 0.89 and another 0.39. First, by adjusting this parameter to 0.90, the results in Figure 4.9 will demonstrate that one object in class Car and another in class Bus are wholly removed from the detection. The other parameter was set to 0.33 This will keep the ones above this ratio, removing the car with a confidence threshold of about 0.32. Figure 4.9: Adjusting parameters Moreover by indicating IoU very high it will detect some extra detections out of each object which NMS will discard the ones with lowest confidence. However, the model also makes some errors regarding detection. For example, the figure 4.10 illustrates that the model incorrectly predicted a cup as a soft drink because of the 52
similarity between the bounding boxes. Figure 4.10: Evaluation detection accuracy Ignoring these few false detections about the model showed significant results in both speed and accuracy throughout the whole procedure, such as streaming, image processing, and video analysis. Regarding the video and real-time application, the challenging part is detecting frame by frame to implement it with the best impact on blind people in terms of accuracy and speed. Furthermore, it was investigated in different lighting environments. 53
Figure 4.11: Real time example 4.4.2 Object’s Spatial Location Due to the lack of availability of the actual sizes and sensors, such as the camera, The best approach was, instead of the distance, having feedback about the spatial location of the objects. The image is divided into nine equal sections in all directions regarding spatial location. The algorithm accuracy was displaying significant results out of detection in any direction. Figure 4.12: Practical result Figure 4.12 represents a practical result of all the procedures, especially spatial location, 54
which is printed on the console beside the audio playback. It printed the spatial location of each detected object related to the class Car, describing them at the down, bottom left, and traffic light at the up, that are entirely correct. 4.5 Audio feedback analysing In the case of an image, the practical result is providing feedback after capturing and processing the data. For video sources, the video is analyzed frame by frame, available for adjusting to detect objects in each frame, storing the result separately, and providing feedback during analysis. In a real-time application, the webcam processes the environment as a live stream and sends feedback frame by frame. Figure 4.13: Printing audio on console In order to have the application in different languages, implementing on gTTS was considered integrating with a Google translate library to translate simply to the other languages, enhancing user experience. Figure 4.14, corresponds to the practical result out of language modification in Italian. It is reasonable to point out that in all the processes, the response time of the audio was almost immediate except for the video part, which was slightly challenging. 55
Figure 4.14: Italian Language 4.6 A Prototype of Virtual Assistant as a prototype for virtual assistant, a speech recognition algorithm separately was studied to detect objects. This feature will allow users to interact with a semi-virtual assistant through a few talking procedures. This algorithm is focused on working only on a webcam, which can be the primary procedure and satisfactory in all tasks. However, future work would focus on an entire virtual assistance program. represents a practical result of the it. the algorithm started by saying to the user, ”How can I help you?” Then, it remains in listening mode to acquire the user’s command. In this example, after the user provides the command ”Hello,” the algorithm will deliver feedback as ”Command not recognized. Please try again.” Finally, the correct command was implemented, as well as the result. 56
Figure 4.15: Speech recognition practical result In the Figure 4.16, Another interaction occurred between the user and the algorithm. In this case, instead of capturing, the algorithm would describe all the objects that can be seen and provide feedback. This form of analysis will improve real-time applications. Figure 4.16: Speech recognition another practical result 57
Chapter 5 Limitations and Future Work Due to the lack of accessibility to hardware such as Arduino or Raspberry Pi, the current program remained theoretical-based. On the other hand, future developments will focus on improving the detection algorithm with better metrics regarding the accuracy and speed involved in implementing it on hardware. Besides the improvements of the current algorithm, text recognition using the OCR and a color sensor can be used for a wide range of tasks. The concept for future hardware involves eyeglasses with attached hardware, including a color sensor and a microphone, to provide feedback. The proposed wearable system would have two buttons: one dedicated to capturing and sending processed information and another for real-time recording and feedback. The hardware would be compact and user-friendly, and the power supply would be a rechargeable battery. Furthermore, the audio feedback will be enhanced to provide a form of virtual assistance, offering users a more interactive, supportive experience. During the composition of this thesis, newer versions of the YOLO algorithm, YOLOv9[31] and YOLOv10[32], were announced. Regrettably, YOLOv9 was not easily accessible and did not readily accommodate audio guidance. Shortly after the completion of the thesis, YOLOv10 was released. It is anticipated that these impressive versions with outstanding transformations in YOLO object detection era, will be used in upcoming attempts. 59
[17] Juan Du. Understanding of object detection based on cnn family and yolo. Journal of Physics: Conference Series, 1004(1):012029, apr 2018. [18] Geethapriya. S, N. Duraimurugan, and S. P. Chokkalingam. Real-time object detection with yolo. International Journal of Engineering and Advanced Technology (IJEAT), 8(3S):2249–8958, February 2019. ISSN: 2249-8958. [19] Ferdousi Rahman, Israt Jahan Ritun, Nafisa Farhin, and Jia Uddin. An assistive model for visually impaired people using yolo and mtcnn. In Proceedings of the 3rd International Conference on Cryptography, Security and Privacy, ICCSP ’19, pages 225–230, New York, NY, USA, 2019. Association for Computing Machinery. [20] Rajeshvaree Karmarkar and Vikas Honmane. Object detection system for the blind with voice guidance. International Journal of Engineering Applied Sciences and Technology, 6(2), June 2021. [21] Mansi Mahendru and Sanjay Kumar Dubey. Real time object detection with audio feedback using yolo vs. yolo v3. In Proceedings of the 2021 11th International Conference on Cloud Computing, Data Science & Engineering (Confluence), pages 734–740, 2021. [22] Simranjeet Kaur, Anup Lal Yadav, and Abhishek Joshi. Real time object detection. In Proceedings of the 2022 International Conference on Cyber Resilience (ICCR), pages 1–5, 2022. [23] Venkata Mandhala, Debnath Bhattacharyya, Vamsi Bandi, and N. Thirupathi Rao. Object detection using machine learning for visually impaired people. International Journal of Current Research and Review, 12(2):157–167, January 2020. [24] Raihan Bin Islam, Samiha Akhter, Faria Iqbal, Md. Saif Ur Rahman, and Riasat Khan. Deep learning based object detection and surrounding environment description for visually impaired people. Heliyon, 9(6):e16924, 2023. [25] Nada N. Saeed, Mohammed A.-M. Salem, and Alaa Khamis. Android-based object recognition for the visually impaired. In 2013 IEEE 20th International Conference on Electronics, Circuits, and Systems (ICECS), pages 645–648, 2013. [26] Rahul Kumar, Sanjesh Kumar, Sunil Lal, and Praneel Chand. Object detection and recognition for a pick and place robot. In Proceedings of the 2014 International Conference on Computer Vision and Robotics, November 2014. 67
[27] Samkit Shah. Cnn based auto-assistance system as a boon for directing visually impaired person. In 3rd International Conference on Trends in Electronics and Informatics (ICOEI), pages 1–5. IEEE, 2019. [28] Jack M. Loomis. Digital map and navigation system for the visually impaired. Technical report, Department of Psychology, University of California, Santa Barbara, 1985. [29] B. Dwyer, J. Nelson, T. Hansen, et al. Roboflow (version 1.0) [software], 2024. Computer vision software. [30] Glenn Jocher. Ultralytics yolov5, 2020. [31] Chien-Yao Wang, I-Hau Yeh, and Hong-Yuan Mark Liao. Yolov9: Learning what you want to learn using programmable gradient information, 2024. [32] Ao Wang, Hui Chen, Lihao Liu, Kai Chen, Zijia Lin, Jungong Han, and Guiguang Ding. Yolov10: Real-time end-to-end object detection, 2024. [33] F. Liu, Z. Lu, and X. Lin. Vision-based environmental perception for autonomous driving. Proceedings of the Institution of Mechanical Engineers, Part D: Journal of Automobile Engineering, 0(0), 2023. [34] Navneet Dalal and Bill Triggs. Histograms of oriented gradients for human detection. 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), 1:886–893 vol. 1, 2005. [35] Matthew Blaschko and Christoph Lampert. Learning to localize objects with structured output regression. In Proceedings of the European Conference on Computer Vision (ECCV) 2008, volume 5302 of Lecture Notes in Computer Science, pages 2–15, Berlin, Heidelberg, October 2008. Springer. [36] Marek Vajgl, Petr Hurtik, and Tom´ aˇ s Nejezchleba. Dist-yolo: Fast object detection with distance estimation. Applied Sciences, 12(3), 2022. 68
70