Generative AI and neural networks towards advanced robot cognition
Full text
Generative AI and neural networks towards advanced robot cognition Christoforos Aristeidou, Nikos Dimitropoulos, George Michalos (2)* Laboratory for Manufacturing Systems and Automation, Department of Mechanical Engineering and Aeronautics, University of Patras, Patras, 26504, Greece ARTICLE INFO Article history: Available online 19 April 2024 ABSTRACT Enhancing autonomy and applicability of robotic systems across diverse scenarios, requires efficient environment perception. Conventional vision systems are highly performing but limited to simple tasks, while AI based ones require extensive data collection, processing and training. This paper presents a framework leveraging generative AI and Neural Networks to implement a dynamically updateable perception system. A multimodal conditional Generative Adversarial Network generates large image datasets which are automatically annotated by a Large Multimodal Model. A Convolutional Neural Network performs further dataset augmentation. A case study on the inspection of aircraft fuel tanks is used to showcase the potential of the approach. © 2024 The Author(s). Published by Elsevier Ltd on behalf of CIRP. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/) Keywords: Cognitive robotics Neural networks Generative artificial intelligence 1. Introduction Since its inception, machine vision has been a main technological flexibility, enabling multiple applications in the production shopfloor including quality control and part/area inspection which are often encountered in production systems [1,2]. The technological evolution from simple two-dimensional RGB sensors all the way to time of flight and RGBD cameras able to process massive amounts of data has further enabled functionalities such as spatial mapping and real time perception of the sensors operating environment [3,4]. Nevertheless, the quality of any perception system is as good as the underlying algorithms that are used to interpret the raw data and convert them to usable information and eventually knowledge about the perceived environment [5]. In this context methods such as Machine Learning have been used to provide information-rich interpretation of the acquired data [6]. These approaches have proven to be high performing however they prerequisite a lot of effort on preparing and processing the required training datasets. On the other hand, recent advances in AI have led to great technological leaps on the way that data can be generated and processed. Specifically, large language models (LLMs) provide a highly performing way to process, summarize, generate and predict new content based on extremely large datasets [7]. Particularly for the case of robotic perception, the availability of relevant images that can be used as training datasets is what usually limits the flexibility of ML applications. To overcome such limitation Large Multimodal Models (LMMs) can be used to generate information from multiple data modalities or sources, such as text, images, audio, and video, which in turn can be used to dynamically provide scalable and updatable training datasets [8]. However and due to the partially stochastic outcome of AI algorithms, proper techniques need to be employed to ensure relevance and accuracy of these datasets which in turn affects the performance of the perception system itself [9]. In this context, this paper proposes an approach for using LMMs in the synthesis of perception systems for robotic applications with minimal user intervention, technical knowledge and data input. Section 2 presents the architecture for synthesizing LMM based perception systems by automating a large portion of the dataset generation, classification, annotation and training process. The implementation of the approach in the form of a prototype is described in Section 3 while Section 4 provides the evaluation on a case study of the aeronautics sector. Section 5 provides a discussion on the findings and outlines areas of future work. 2. Approach The contradicting objective of having both flexible and accurate robot vision for detecting multiple objects during inspection tasks, is addressed by a continuous learning approach based on LMMs. Unlike simple visual recognition, a cognition enabled system can generate new information and process it to update its inferencing capability on object/ feature detection, similarly to meta-learning [10]. Likewise, humans can use their memories/ experiences to analyse a newly presented situation and reason on how to act. Of course, human abilities for sensor fusion, planning, execution adaptation, and collaboration are far more advanced and cannot be fully duplicated in a single system yet. In our case the inspection is carried out by a robotic arm which is equipped with an RGBD camera that is used to detect objects or features as well as to avoid obstacles. Especially for the process of detecting foreign objects, fast real-time object detectors (Single Shot multibox Detector SSD [11], Region-based Convolutional Neural Networks RCNN [12], YOLO [13]) that run on edge and achieve high accuracy in computer vision benchmark datasets are used. However, such models can only detect the set of objects that were included in the training datasets. The enhancement of the dataset with new images is a manual and time-consuming process, requiring * Corresponding author at: Laboratory for Manufacturing Systems & Automation, Mechanical Engineering and Aeronautics Department, University of Patras, University Campus, 26504, Patras, Greece. E-mail address: [email protected] (G. Michalos). https://doi.org/10.1016/j.cirp.2024.04.013 0007-8506/© 2024 The Author(s). Published by Elsevier Ltd on behalf of CIRP. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/) CIRP Annals - Manufacturing Technology 73 (2024) 2124 Contents lists available at ScienceDirect CIRP Annals - Manufacturing Technology journal homepage: https://www.editorialmanager.com/CIRP/default.aspx
extensive search for representative image collection, annotation and programming/technical skills to carry them out. To simplify and shorten the duration and effort required for this task, an intuitive approach has been adopted (Fig. 1). The operator supervising the operation of the robot is provided with a live view of the camera feed. Upon identifying an object that is not included in the training dataset, he/she may change the robot position to capture different viewing angles (at least 5). This can be done by robot hand guiding or pinpointing the area of interest in the image and letting the robot plan its trajectory by using its Digital Twin. Following, he/ she inserts a simple text description (e.g. red wire cutter) for the object using natural language. The dual input (text and images) is consequently fed to a conditional Generative Adversarial Network (cGAN) to generate images based on it (e.g. images of completely different wire cutter). Input images represent the initial knowledge to the model and text prompts semantically describe features of the generated image. An API is used to automatically download the newly generated images which are then added to the pool of original images. Following, the extended dataset is used as input to an LMM zeroshot object detection model which supports multi modal input. Natural language text prompts, either the same as in the previous step or new ones, are used to guide the detection of objects/ characteristics in the image and the output consists of localization coordinates with class labels for each detected object. The LMM model enables the full automation of the annotation process. At this point human input may also be engaged in order to verify the quality of the annotation results and rectify possible mistakes such as false positives or negatives. This corresponds to a very small percentage of the effort that would otherwise be required for a fully manual image annotation process which could involve hundreds or even thousands of images. Having added the new objects in the recognition process, the proposed approach proceeds to ensure that further variations of these objects can be accommodated by the inspection system. For this purpose, image augmentation techniques are applied based on spatial and pixel transformations, to create new and diverse versions with respect to colors, size, shape, and angle of perspective. As a final step, the reference images, the augmented images and the derived labels are added in the training dataset and automatically split into train, validation and test sets using predefined percentage thresholds. The derived dataset is used to retrain the model and acquire a recognition system which can now detect multiple versions and instances of the newly added object. 3. Implementation 3.1. Image generative adversarial network Generative Adversarial Networks have emerged in the last years and together with the rapid progress of GPU acceleration methods for model training and inference, have enabled the implementation of complex architectures for generating high fidelity realistic images. In this study, the Leonardo API [14] is used which encompasses a variety of image generative models such as Stable Diffusion, and also fine-tuned models based on specific image generation tasks (photorealistic, artistic, landscape and 3D animation). For the purposes of this study, a model with multimodal capabilities is used, meaning that it is able to receive both image and text inputs. Such models are called conditional Generative Adversarial Networks (cGAN) [15]. Usually, image generators start with random pixel values in the range [0,255], for each RGB channel. Thus, a matrix for each colour channel is created, having the dimensions of the final image (e.g. 640 £480). Through several iteration cycles they are updated to achieve results converging to the requested outcome. In the case of cGAN, the input is a reference image and not random pixels, allowing the model to generate images which are closer to the input prompts. Parameters such as text and image prompt weighting, image input and output size, number of inference steps and negative prompts are some of many that can determine the output results of generated images. 3.2. GroundingDINO zero-shot object detector Object detection models normally require training on images and corresponding labels in order to predict localization coordinates for specific features inside an image. In this work the LMM model GroundingDINO [16] is used thanks to its ability to accept both image and corresponding text prompts that describe details of the object. It is based on GLIP [17], a transformer-based vision model that unifies object detection with phrase prompts that describe the candidate classes of the detection task. GroundingDINO advances further, and instead of passing the class names in text prompting along with image input, it uses expression comprehensions with specific attributes and characteristics of each object to be detected. Such models are called Base Models, because they have wide knowledge for various basic objects and attributes. A text prompt that could be used as example is ‘red gloves’or ‘gray stainless nut’. Specified characteristics such ‘red’and ‘gray stainless’affect the detection result and omit objects not fitting these criteria. Apart from Prompt Engineering methods which can be used to design prompts in order to influence the model’s response, a user defined threshold which determines the model prediction confidence is also used. Fig. 2, presents examples of predictions using different text prompts, tested on a scene with various tools. 3.3. Pixel and spatial image augmentation Deep neural networks require a lot of training data and are sensitive to overfitting [18]. To prevent such occurrence, image augmentation is used to artificially create new images from existing ones. Augmentation Fig. 1. Flowchart for continuous vision learning procedure. Fig. 2. Prediction using text prompts: (a) ‘stainless nut. Red glove on the right’, (b) ‘screwdriver. Black tape’, (c) ‘rivet. Gray gloves’, (d) ‘gloves’. 22 C. Aristeidou et al. / CIRP Annals - Manufacturing Technology 73 (2024) 2124
is based on two categories, pixel-level and spatial-level transformations. The first category includes transformations that change the values of each pixel in the image such as RGB channel shuffling with can mix the colours of the image, increase or decrease of brightness, add blur, sharpen, etc. The second category includes transformations that change the point of view and position of the image such as rotation, crop, resize, vertical flip, etc. Fig. 3 presents indicative augmentation techniques using the open source library Albumentations. Augmentation is used in almost all classification, detection and segmentation models that achieve very high recognition accuracy and robustness thanks to the model generalization it offers. By diversifying the dataset through augmentation, the model adapts better in handling variations from real-world scenarios. 3.4. Real-time object detector model training In robotics vision tasks, fast and accurate models that run in realtime and achieve high fps are used for simultaneous prediction quality and light deployment that can run on the edge. In the proposed approach an object detector with proven speed and accuracy is used. Specifically, YOLOv8 [19] is a CNN model with five different architecture sizes, able to run in moderately powered mobile devices leveraging runtime optimization methods such as 8-bit integer quantization (int8). Regarding the training datasets notation, we assume D 0 an initial state of a dataset which contains captured images from real objects used to train a model. A training set for this case is T0¼fðxi;yiÞg where xiis an image and yiis its corresponding labels. The generated and augmented image datasets Dgen and Daug are combined with the initial to produce Dcomb ¼D0[Dgen [Daug.While training goal is to minimize the loss function (1): JuðÞ¼1 mX m i¼1 Lf x i;uðÞ;yi ðÞ ð1Þ where mis the number of samples, Lis the loss function measuring the difference of predicted fðxi;uÞand the actual yiby finding the model weights u¼argminuJðuÞthat minimize the error. Our goal is to maximize detection predictive accuracy by minimizing model weights uwhen training using Dcomb. 3.5. Transfer learning and fine-tuning When training a deep learning model, it is crucial to achieve fast and accurate convergence of model weights to a local or global minima. Transfer learning is an approach that involves leveraging gained knowledge from a pre-trained model on a specific vision task [20].Insteadof starting the training from random generated weights or using another weight initialization method, transfer learning initializes the model weights from an already trained model. Except the output layers, which are the final layers of a neural network, all other layer weights get transferred to the new model that needs to be trained. Then new output layers which are constructed for the vision task may involve multiple sub-layers responsible for each aspect of the prediction. In situations where a new object needs to be added in the dataset and further learned by a model, transfer learning can reduce the training time utilizing already available knowledge from a previous training process. Thus, the training process is reduced to a slight tuning of the weights for the updated dataset with newly introduced features. 4. Case study To validate our implementation a task of aeronautics sector was selected and specifically the maintenance and inspection process during the conversion of passenger aircrafts to freighter. One of the tasks included is the visual inspection of the airplane’s fuel tanks by technicians who physically enter the tank and check for discrepancies or foreign objects using a flashlight and mirror. Cramped spaces, very complex piping/ cabling, fuel fumes and hard to reach areas are the main challenges for such inspection tasks. Due to hazardous environment, safety precautions are taken before a technician can enter the fuel tanks, leading to a long process time and unwanted exposure to health hazards. Additionally, the inspected tank may undergo significant modification during its lifetime without updating the CAD models. In such cases, and unlike robotic manipulation via simulation, there is a strong need for inspection methods that can handle unknown environments, geometries and objects. A collaborative robot (Fig. 4) is used to inspect the tanks and detect foreign objects which may come from wear/damage of the tank internal equipment or mistakenly forgotten parts and tools by the technicians. The robot must also extract them by using a gripper or a vacuum tube and in cases where it is not possible, notify the operator about the position of the object to be extracted. Following the approach described in Section 2 an object detection model was trained on a dataset containing common tools and mechanical parts (screwdrivers, rivets, screws, washers, wires, flashlights, gloves, etc.) that can be found inside the tanks during the inspection process, totaling to 68 variants. While the operator monitors the visual inspection process performed by the robot, she/he can interfere when noticing a foreign object that is not included in the training set. Three new objects that were not included in the dataset were placed inside the tank including: a wrench tool, a wire cutting tool and a stainless nut. Following the operator used manual guidance to position the camera of the robot in suitable positions to capture 5 reference images per object. All 15 captured images were fed as initial image prompts to the Generative AI model via the Leonardo API. Through each request 4 new images were generated, resulting in a total of 60. Samples of the results are presented in Fig. 5. The generated images were fed to GroundingDINO and after automatic annotation, minimal rectification (<10 % of the total predictions) was required. Following, the images and annotations were augmented creating 3x new variations, i.e. 180 images in total. The combined dataset was split into train validation test folders, with 0.7, 0.2, 0.1 percentage thresholds respectively. To evaluate the performance regarding the accuracy and execution times, three different methods of model training were performed: a) with captured images only, b) with captured and augmented, and c) with generated, captured and augmented combined. In order to have a common baseline the total number of Fig. 3. Image augmentation transformations: (a) rotation, (b) horizontal flip, (c) zoom in, (d) hue change, (e) mix of cases b and d, (f) saturation. Fig. 4. Airplane tank setup. C. Aristeidou et al. / CIRP Annals - Manufacturing Technology 73 (2024) 2124 23
images was capped to 180 and their mix was allowed to be variable. Each method was executed in five iteration cycles by a different operator and the results are presented in Table 1. The proposed method manages to reduce the execution time by »75 %, while achieving a 0.95 mean Average Precision at an Intersection Over Union (IoU) of 0.5 (mAP50). The robustness for high accuracy prediction (mean of mAPs for IoU 0.5 to 0.95) is similar. The accuracy (a) convergence rate has been evaluated against the number and type of images used. Fig. 6 (left) denotes that the examined method can reach a>0.9 when at least 10 captured images are used. The right image shows that even using only 5 captured images (Group A) and adding 20 AI generated ones (Total: 25 - Group B) lead to a>0.8. Further addition of 50 augmented images (Total: 75 - Group C) further allow for a>0.9. 5. Conclusions and future work This paper presents an AI-based approach that enhances the capabilities of machine perception by utilizing and combining multiple models and processes to reduce the time and effort for collecting, annotating and preparing the training dataset. Novel models with multimodal capabilities were used to create diverse and realistic images from simplified input including only captured imaged and text prompts. The annotation process was also automated using an LMM to predict localization coordinates. The application in an aeronautics case study revealed that a reduction in execution time of the model’s dataset expansion can be achieved together with higher accuracy compared to traditional methods. Future work will aim to exploit the latest GenAI models for automating the review of generated annotations as well as the generation of more complex scenes with multiple objects and enhanced variability regarding size, colours and points of view. Combinations of objects and diversified image backgrounds can be explored to test the range of generalization learning that can be automatically achieved in a higher scale, towards more flexible machine vision capabilities. The use of LMMs in quality inspection tasks could also be investigated, depending on their capability to generate images with defects (e.g. cracks, torn cable, etc.). Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. CRediT authorship contribution statement Christoforos Aristeidou: Writing review & editing, Writing original draft, Visualization, Validation, Software, Investigation, Formal analysis, Data curation, Conceptualization. Nikos Dimitropoulos: Writing review & editing, Writing original draft, Validation, Supervision, Project administration, Methodology, Investigation, Conceptualization. George Michalos: Writing review & editing, Writing original draft, Methodology, Funding acquisition, Formal analysis, Conceptualization. Acknowledgement This work has been supported by “CONVERGING - Social industrial collaborative environments integrating AI, Big Data and Robotics for smart manufacturing”project, funded by the European Commission Horizon Europe research and innovation programme under grant agreement No 101058521. References [1] Chryssolouris G (2006) Manufacturing Systems: Theory and Practice, 2nd Edition Springer-VerlagNew York. [2] Weimer D, Scholz-Reiter B, Shpitalni M (2016) Design of Deep Convolutional Neural Network Architectures For Automated Feature Extraction In Industrial Inspection. CIRP Annals 65/1:417–420. [3] Zheng P, Li S, Xia L, Wang L, Nassehi A (2022) A Visual Reasoning-Based Approach For Mutual-Cognitive Human-Robot Collaboration. CIRP Annals 71/1:377–380. [4] Zheng P, Li S, Fan J, Li C, Wang L (2023) A Collaborative Intelligence-Based Approach For Handling Human-Robot Collaboration Uncertainties. CIRP Annals 72/1:1–4. [5] ElMaraghy H, Monostori L, Schuh G, ElMaraghy W (2021) Evolution and Future of Manufacturing Systems. CIRP Annals 70/2:635–658. [6] Dimitropoulos N, Togias T, Zacharaki N, Michalos G, Makris S (2021) Seamless HumanRobot Collaborative Assembly Using Artificial Intelligence and Wearable Devices. Applied Sciences 11/12:5699. [7] Wang X, Anwer N, Dai Y, Liu A (2023) ChatGPT for Design, Manufacturing, and Education. Procedia CIRP 119:7–14. [8] Lichtenwalter D, Burggr€ af P, Wagner J, Weißer T (2021) Deep Multimodal Learning for Manufacturing Problem Solving. Procedia CIRP 99:615–620. [9] Whang SE, Roh Y, Song H, Lee JG (2023) Data Collection And Quality Challenges In Deep Learning: A Data-Centric AI Perspective. The VLDB Journal 32/4:791–813. [10] Li Y, Liu C, Hua J, Gao J, Maropoulos P (2019) A Novel Method For Accurately Monitoring And Predicting Tool Wear Under Varying Cutting Conditions Based On Meta-Learning. CIRP Annals 68/1:487–490. [11] Liu W, Anguelov D, Erhan D, Szegedy C, Reed S, Fu CY, Berg AC (2016) SSD: Single Shot Multibox Detector. Computer VisionECCV 2016,21–37. [12] Girshick R, Donahue J, Darrell T, Malik J (2015) Region-Based Convolutional Networks for Accurate Object Detection and Segmentation,IEEE Transactions on Pattern Analysis and Machine Intelligence, 142–158. [13] Jiang P, Ergu D, Liu F, Cai Y, Ma B (2022) A Review of Yolo Algorithm Developments. Procedia Computer Science 199:1066–1073. [14] Leonardo A.P.I., https://docs.leonardo.ai, last accessed on 14/1/24. [15] Mirza M., Osindero S. (2014) Conditional Generative Adversarial Nets. arXiv preprint [16] Liu S., Zeng Z., Ren T., Li F., Zhang H., Yang J., Li C., Yang J., Su H., Zhu J., Zhang L. (2023) Grounding DINO: Marrying DINO with Grounded Pre-Training for OpenSet Object Detection. arXiv preprint [17] Li L., Zhang P., Zhang H., Yang J., Li C., Zhong Y., Wang L., Yuan L., Zhang L., Hwang J., Chang K. (2021) Grounded Language-Image Pre-training. arXiv preprint: 2112.03857. [18] Shorten C, Khoshgoftaar TM (2019) A Survey on Image Data Augmentation for Deep Learning. Journal of Big Data 6/1:1–48. [19] YOLOv8, https://docs.ultralytics.com/, last accessed on 14/1/24. [20] Zhuang F, Qi Z, Duan K, Xi D, Zhu Y, Zhu H, Xiong H, He Q (2020) A Comprehensive Survey on Transfer Learning. Proceedings of the IEEE 109/1:43–76. Fig. 5. API Generated images: (a) cutter tool, (b) wrench, (c) stainless nut. Table 1 Metrics table of accuracy and process time. All results represent mean values across five iterations cycles. No of images: Captured: 180 Captured: 45 Augmented: 135 Captured: 15 Generated: 60 Augmented: 105 mAP50 (higher is better) 0.96 0.945 0.954 mAP50-95 (higher is better) 0.816 0.786 0.829 Time (lower is better) »34 min »16 min »8 min Fig. 6. Left: Accuracy vs. number of captured images, Right: Accuracy enhancement vs. type and number of a) Captured Images only, b) Captured and AI Generated Images, c) Captured, AI Generated & Augmented images. 24 C. Aristeidou et al. / CIRP Annals - Manufacturing Technology 73 (2024) 2124