Estrategias de visión por computador para la estimación de pose en el contexto de aplicaciones robóticas industriales: avances en el uso de modelos tanto clásicos como de Deep Learning en imágenes 2D
Abstract
184 p.
Full text
Departamento de Ciencia de la Computación e Inteligencia Artificial Universidad del País Vasco (UPV/EHU) Estrategias de visión por computador para la estimación de pose en el contexto de aplicaciones robóticas industriales: avances en el uso de modelos tanto clásicos como de Deep Learning en imágenes 2D y nubes de puntos Ibon Merino Bermejo Dirección Basilio Sierra Anthony Remazeilles 18 de julio de 2023 (cc)2023 IBON MERINO BERMEJO (cc by 4.0)
Agradecimientos Con estas palabras voy a dar comienzo al final de esta larga etapa. Una etapa que ha tenido sus buenos y malos momentos, pero que ambos me han permitido desarrollarme en diferentes aspectos. Este recorrido que ha durado algo más de 4 años lo inicié siendo aún más ignorante de lo que soy ahora. No tenía ni idea de lo que me venía por delante. Desde un inicio he tenido la suerte de encontrarme en el camino a personas maravillosas que me han ayudado a crecer como persona y como profesional. Primero quiero agradecer a Tecnalia por la oportunidad que me ha dado y a todos con los que he coincidido. Sobre todo quiero agradecer a la gente de mi grupo de robótica flexible. Gracias por todo lo que me habéis enseñado, por toda la ayuda que me habéis ofrecido y sobre todo por aguantar mis chistes malos. Quiero agradecer a las 3 personas que han sido fundamentales en este recorrido: a mis directores de tesis, Basi y Anthony, por su paciencia y por su ayuda; y a Jon, por su dedicación y todo el apoyo que me ha dado durante todo este tiempo. Gracias a todos por hacer que este camino haya sido más llevadero. Como las raíces son importantes y más aún si no te despegas de ellas, quiero agradecer al colegio donde estudié, Salesianos Deusto, y a toda la obra de los Salesianos que me han formado como persona. Han sido muchos años en ese colegio y posteriormente como monitor. Gracias por todo lo que me habéis enseñado y por todo lo que me habéis dado. Y gracias a todos los monitores y chavales que habéis compartido camino conmigo porque las experiencias adquiridas en esos años son invaluables. También quiero agradecer a todos mis amigos que me han apoyado siempre. Por un lado a mis amigos de la universidad en especial a Judity e Itziar. Por otro lado a mi querida “Villa Lola” que me ha permitido distraerme, reírme y disfrutar en los momentos de más tensión. A mi “Zulo” que siempre han estado ahí presentes. A mis queridas “Txandaleras” que tienen un lugar especial en mi corazón. A mi pareja Gonzalo que me ha dado ánimos y fuerzas en los momentos más duros. Y a todos los demás amigos que han estado ahí siempre. Como los postres, lo más dulce es el final. Quiero agradecer especialmente a mi
ii familia: a mi aita, que muy bien me ha enseñado que “más vale maña que fuerza” y me ha inspirado a ser resolutivo y perspicaz; a mi ama, que siempre ha estado ahí con su amor incondicional, tanto en lo bueno como en lo malo y de la que he aprendido que la vida son dos días y hay que sacarle el jugo a cada momento; a mi hermano, el pequeño de la casa, al que considero uno de mis mejores amigos; a mi aitite, que me ha demostrado que siempre hay que luchar y seguir hacia adelante; a mi amama, que es como mi segunda madre y todo lo he aprendido de ella; y al resto de mi familia por estar siempre ahí. Tras todo esto, no me queda otra cosa que seguir en este camino de la eterna ignorancia que espero que me siga nutriendo de conocimientos, de experiencias y de personas maravillosas como lo ha hecho hasta ahora.
Resumen La visión por computador es una tecnología habilitadora que permite a los robots y sistemas autónomos percibir su entorno. Dentro del contexto de la industria 4.0 y 5.0, la visión por computador es esencial para la automatización de procesos industriales. Entre las técnicas de visión por computador, la detección de objetos y la estimación de la pose 6D son dos de las más importantes para la automatización de procesos industriales. Para dar respuesta a estos retos, existen dos enfoques principales: los métodos clásicos y los métodos de aprendizaje profundo. Los métodos clásicos son robustos y precisos, pero requieren de una gran cantidad de conocimiento experto para su desarrollo. Por otro lado, los métodos de aprendizaje profundo son fáciles de desarrollar, pero requieren de una gran cantidad de datos para su entrenamiento. En la presente memoria de tesis se presenta una revisión de la literatura sobre técnicas de visión por computador para la detección de objetos y la estimación de la pose 6D. Además se ha dado respuesta a los siguientes retos: (1) estimación de pose mediante técnicas de visión clásicas, (2) transferencia de aprendizaje de modelos 2D a 3D, (3) la utilización de datos sintéticos para entrenar modelos de aprendizaje profundo y (4) la combinación de técnicas clásicas y de aprendizaje profundo. Para ello, se han realizado contribuciones en revistas de alto impacto que dan respuesta a los anteriores retos. iii
Laburpena Ordenagailu bidezko ikusmena robotei eta sistema autonomoei beren ingurunea hautemateko aukera ematen dien teknologia gaitzailea da. 4.0 eta 5.0 industriaren testuinguruan, ordenagailuen bidezko ikusmena funtsezkoa da prozesu industrialak automatizatzeko. Ordenagailuen bidezko ikusmen-tekniken artean, objektuen detekzioa eta 6D posearen estimazioa dira prozesu industrialen automatizaziorako garrantzitsuenetarikoak. Erronka horiei erantzuteko, bi ikuspegi nagusi daude: metodo klasikoak eta ikasketa sakoneko metodoak. Metodo klasikoak sendoak eta zehatzak dira, baina ezagutza aditu ugari behar dute garatzeko. Bestalde, ikasketa sakoneko metodoak erraz garatzen dira, baina datu asko behar dira entrenatzeko. Tesi-memoria honetan, objektuak detektatzeko eta 6D posea estimatzeko ordenagailuen bidezko ikusmen-teknikei buruzko literaturaren berrikuspena aurkezten da. Gainera, honako erronka hauei erantzuna eman zaie: (1) ikuspegi-teknika klasikoen bidez posea estimatzea, (2) 2D ereduen ikaskuntza 3Dra transferitzea, (3) datu sintetikoak erabiltzea ikasketa sakoneko ereduak entrenatzeko, eta (4) teknika klasikoak eta ikasketa sakonekoak konbinatzea. Horretarako, eragin handiko aldizkarietan ekarpenak egin dira, aurreko erronkei erantzuteko. v
Índice de contenidos Índice de contenidos vii Índice de figuras x Índice de tablas xii I Ámbito de investigación 1 1 Introducción y motivación 3 1.1. Introducción ................................. 3 1.2. Motivación .................................. 5 1.3. Contexto ................................... 6 1.3.1. TECNALIA ............................. 6 1.3.2. UPV/EHU .............................. 6 1.3.3. Proyectos .............................. 7 1.4. Hipótesis de la investigación ........................ 17 1.5. Estructura de la memoria .......................... 18 2 Técnicas de visión por computador para la detección de objetos. 19 2.1. Introducción ................................. 19 2.2. Técnicas clásicas 2D ............................. 23 2.2.1. Descriptores globales ........................ 23 2.2.2. Descriptores locales ........................ 27 2.2.3. Búsqueda de correspondencias y clasificación .......... 30 2.3. Técnicas 3D ................................. 34 2.3.1. Métodos locales basados en parches ............... 35 2.3.2. Métodos basados en matcheo de nubes de puntos ........ 35 2.3.3. Métodos basados en plantillas ................... 36 2.4. Aprendizaje profundo ............................ 36 vii
CAPÍTULO 1 Introducción y motivación 1.1. Introducción En los últimos años la industria ha experimentado muchos cambios a un ritmo vertiginoso. Ha pasado una década desde que se acuñó el término Industria 4.0 [ 1 ] y ya se contempla la siguiente revolución industrial, la Industria 5.0. Si la cuarta revolución industrial se centra en la automatización y el intercambio de datos mediante IoT, sistemas ciber-físicos o computación en la nube, la Industria 5.0 se centra en la sostenibilidad, resiliencia y en el bienestar humano mientras mantiene los criterios que marca la Industria 4.0. La Comisión Europea, quien bautizó esta nueva era, expone en su informe [2]: En lugar de preguntarnos qué podemos hacer con la nueva tecnología, nos preguntamos qué puede hacer la tecnología por nosotros. En lugar de pedir al trabajador de la industria que adapte sus habilidades a las necesidades de una tecnología que evoluciona rápidamente, queremos utilizar la tecnología para adaptar el proceso de producción a las necesidades del trabajador. La visión por computador es un campo importante dentro de este paradigma. Es importante que los robots sepan identificar su entorno. Por ejemplo, de nada sirve que un robot sepa coger objetos si no sabe donde está el objeto. De la misma forma, un robot no puede ejecutar movimientos de forma amigable para el operario si no sabe donde está para poder evitarlo o actuar de forma más sutil cuando se aproxime a él. Esto es importante para el bienestar del operario y ayuda a adaptar la robótica al operario. 3
1. Introducción y motivación En la Figura 1.1 se muestra un brazo robótico colaborativo Doosan con una cámara 3D de luz estructurada Zivid. Este es un caso de uso típico de visión por computador en la industria. La visión por computador permite realizar tareas como localización de objetos [ 3 , 4 , 5 ], detección de objetos [ 6 , 7 , 8 ], tracking de objetos [ 9 , 10 , 11 ], detección de humanos [ 12 , 13 , 14 ], reconstrucción de objetos [ 15 , 16 , 17 ], segmentación semántica [ 18 , 19 ], entre muchas otras, las cuales encajan perfectamente en este paradigma de Industria 4.0 e Industria 5.0. Figura 1.1: Brazo robótico colaborativo Doosan con una cámara 3D de luz proyectada Zivid. El trabajo presentado tiene como motivación la preparación para esta Industria 5.0 (ver Sección 1.2), ya que se ha realizado en el centro de investigación orientado a las necesidades de la industria TECNALIA (ver Sección 1.3.1) en colaboración con la UPVEHU (ver Sección 1.3.2) debido al trabajo de colaboración entre estos dos y con el autor. El trabajo realizado ha sido aplicado en varios proyectos (ver Sección 1.3.3) los cuales intentan desarrollar tecnologías habilitantes dentro de la Industria 5.0. Finalmente, se han definido unas hipótesis de investigación (ver Sección 1.4) las cuales se han intentado responder con las contribuciones realizadas (ver Capítulo 3). 4
1.2. Motivación 1.2. Motivación En las últimas décadas el sector industrial ha experimentado un cambio radical en su forma de producción. La automatización de procesos ha sido una de las claves para que la industria pueda competir en un mercado globalizado. Para ello, se han desarrollado nuevas tecnologías que permiten automatizar procesos de forma más eficiente y segura. Originalmente esta programación era predefinida, es decir, cada robot realizaba un movimiento o acción concreta repetidamente. Gracias a los avances en inteligencia artificial y sensórica, se ha dotado a los robots de mayor autonomía. Esto es, permitimos al robot actuar en concordancia con su entorno. Por ejemplo, si un robot tiene que realizar una tarea en un espacio de trabajo, es capaz de detectar los objetos que hay en ese espacio y planificar su trayectoria para realizar la tarea. Para realizar esto son necesarias varias tecnologías como una capacidad de adaptación para gestionar las órdenes dinámicas del robot o una visión artificial para detectar qué hay en nuestro espacio de trabajo. En el contexto de este proyecto de tesis, dentro de la visión por ordenador se ha centrado en la detección de objetos y la estimación de su pose 6D (translación y rotación). La Figura 1.2 muestra un ejemplo de este problema. Figura 1.2: Estimación de la pose de los objetos del dataset Linemod proyectada sobre una imagen de test. 5
1. Introducción y motivación Dentro de la detección de objetos y estimación de pose 6D, se pueden diferenciar dos tipos de técnicas: las técnicas clásicas y las de aprendizaje profundo. Llamamos técnicas o métodos clásicos a aquellas técnicas que predominaban antes del auge de los métodos de aprendizaje profundo. Estos métodos no requieren de una gran cantidad de datos para entrenar y han sido utilizados durante décadas de forma satisfactoria. Aun así, en esta última década los métodos de aprendizaje profundo han demostrado obtener unos mejores resultados para muchos casos y su uso se ha extendido. Es por ello que el uso de ambos tipos de técnicas todavía tiene cabida y se pueden utilizar en conjunto para obtener mejores resultados. Por eso la principal motivación de este proyecto de tesis ha sido estudiar las técnicas existentes de ambos tipos y extraer el máximo potencial de ambas combinándolas. 1.3. Contexto 1.3.1. TECNALIA TECNALIA es un centro privado de investigación aplicada con varias sedes en todo el planeta. Es el centro de investigación aplicada y desarrollo tecnológico más grande de España, un referente en Europa y miembro del Basque Research and Technology Alliance (BRTA). TECNALIA tiene dos premisas principales: transformar investigación tecnológica en prosperidad y ser agentes de transformación de las empresas y de la sociedad para su adaptación a los retos de un futuro en continua evolución. TECNALIA dispone de 5 unidades operativas: Transición energética, climática y urbana;Industria y movilidad;Salud;Digital; y, Lab Services. Dentro de cada unidad operativa hay áreas de negocio. El autor pertenece al área de negocio Medios de Producción y Robótica de la unidad operativa Industria y movilidad. Más concretamente, a la plataforma de Robótica para flexibilidad industrial. Esta plataforma está especializada en el desarrollo de tecnologías robóticas inteligentes, autónomas, adaptativas y fáciles de programar. Gracias a la plataforma de Robótica para flexibilidad industrial el autor ha podido colaborar en proyectos reales (Sección 1.3.3) para añadir habilidades de visión a los robots. La Figura 1.3 muestra el stand de TECNALIA en la Bienal de Máquina Herramienta (BIEMH) de 2022. En este stand se muestran las tecnologías habilitadoras de TECNALIA para la Industria 5.0. El autor participó en la creación de las herramientas de visión por computador para la detección de las piezas de los demostradores. 1.3.2. UPV/EHU El autor ha estado ligado a la Universidad del País Vasco / Euskal Herriko Unibertsitatea (UPV/EHU) desde que comenzó sus estudios de Grado. Durante los últimos años ha colaborado con el grupo de Robótica y Sistemas Autónomos (RSAIT) de la UPV/EHU en varios proyectos de investigación. 6
1.3. Contexto Figura 1.3: Stand de TECNALIA de la BIEMH 2022 1.3.3. Proyectos Durante el transcurso del doctorado del autor, éste ha participado activamente en proyectos de TECNALIA, tanto directamente ligados con su proyecto de tesis como otros que no lo están. Estos proyectos, que son reflejo de las necesidades tecnológicas que se demandan en entornos industriales, han dado sentido a esta proyecto de tesis doctoral. Los proyectos que han permitido desarrollar este proyecto de tesis, en su mayoría de financiación pública competitiva, han sido financiados por la Comisión Europea (Programa Horizon 2020) y el Gobierno Vasco (Programa Elkartek). La Figura 1.4 muestra el orden cronológico de los proyectos y el origen de la subvención. Los proyectos ligados a este proyecto de tesis son los siguientes: Elkarbot (Figura 1.5): El proyecto Elkarbot con n º de proyecto KK-2020/00092 pertenece al programa de ayudas a la investigación colaborativa Elkartek. El principal objetivo de este proyecto es el de dotar al País Vasco de herramientas para su robotización. Para realizar esto se tienen en cuenta tres premisas a lograr: mitigar el riesgo percibido al desarrollo de tecnología de robótica flexible propia por los integradores; habilitar el uso de la robótica no solo en los sectores de automoción y aeronáutica; y coordinar mejor las importantes capacidades tecnológicas de la RVCTI (Red Vasca de Ciencia, Tecnología e Innovación) en torno a la robótica, en estrecha relación con el BDIH (Basque Digital Innovation Hub) y dentro del 7
1. Introducción y motivación Figura 1.4: Orden cronológico de los proyectos y origen de la subvención. En negrita están aquellos proyectos ligados al proyecto de tesis. contexto de BRTA (Basque Research and Technology Alliance). Este proyecto ha tenido una duración de dos años empezando en 2020. La principal aportación del autor en este proyecto ha sido diseñar algoritmos de clasificación de piezas industriales usando arquitecturas de aprendizaje profundo 3D preentrenadas con modelos 2D. Esto ha permitido identificar piezas de forma efectiva usando cámaras 3D. Como resultado de este proyecto se ha publicado el siguiente artículo: Merino, I., Azpiazu J., Remazeilles A., Sierra B., 2021. 3D Convolutional Neural Networks Initialized from Pretrained 2D Convolutional Neural Networks for Classification of Industrial Parts . Sensors 21, 1078. MDPI. https://doi. org/10.3390/s21041078 Sherlock - Seamless and safe human centred robotic applications for novel collaborative workplaces (Figura 1.6): Es un proyecto europeo bajo el programa Horizon 2020 cuyo indicador del acuerdo de subvención (Grant Agreement) es 820689. Sherlock tiene como objetivo introducir las últimas tecnologías en seguridad robótica en entornos de producción, habilitándolas con mecatrónica inteligente y cognición basada en IA, creando, de esta forma, estaciones HRC (Human Robot Collaboration) eficientes que están diseñadas para ser seguras y garantizar la aceptación y bienestar de los operarios. Este proyecto tiene una du8
1.3. Contexto Figura 1.5: Paquetes de trabajo del proyecto Elkarbot para responder a los retos de la robótica industrial del futuro. ración de cuatro años entre 2018 y 2022. En este proyecto el autor ha contribuido diseñando técnicas de reconocimiento de piezas industriales usando descriptores locales y globales. Inicialmente estas técnicas han permitido clasificar piezas industriales y una vez identificadas obtener la pose 6D de la pieza para poder cogerla. Durante este proyecto se han publicado los siguientes artículos: Merino, I., Azpiazu, J., Remazeilles, A., Sierra, B., 2019. 2d Image Features Detector and Descriptor Selection Expert System . Computer Science & Information Technology (CS & IT), AIRCC Publishing Corporation. https://doi.org/ 10.5121/csit.2019.91206 Merino, I., Azpiazu, J., Remazeilles, A., Sierra, B., 2019. 2D Features-based detector and descriptor selection system for hierarchical recognition of industrial parts . International Journal of Artificial Intelligence & Applications (IJAIA), AIRCC Publishing Corporation. https://doi.org/10.5121/ijaia.2019.10601 9
1. Introducción y motivación Figura 1.6: Paquetes de trabajo del proyecto Sherlock con sus respectivas tareas. Merino, I., Azpiazu, J., Remazeilles, A., Sierra, B., 2020. Histogram-Based Descriptor Subset Selection for Visual Recognition of Industrial Parts . Applied science 10, 3701. MDPI. https://doi.org/10.3390/app10113701 J.L. Outón, I. Merino, I. Villaverde, A. Ibarguren, H. Herrero, P. Daelman, B. Sierra, 2021. A Real Application of an Autonomous Industrial Mobile Manipulator within Industrial Context. Electronics 10, 1276. MDPI. https://doi.org/10.3390/ electronics10111276 PROFLOW (Figura 1.7): PROFLOW (Producción Fluida para la Industria Inteligente) es un proyecto perteneciente al programa Elkartek con n º de proyecto KK-2022/00024. En la producción fluida, no-lineal o matricial basada en células de fabricación flexibles y configurables, las estaciones de trabajo no están físicamente conectadas, como ocurre con las cintas transportadoras en las líneas convencionales de producción, sino que se trata de islas interconectadas mediante robots 10
1.3. Contexto Figura 1.7: Descripción gráfica del proyecto PROFROW móviles. La adopción de un paradigma de producción fluida o matricial como la propuesta en PROFLOW favorece que las empresas manufactureras vascas mejoren su capacidad de respuesta y adecuación a un contexto marcado por un elevado y creciente nivel de complejidad en la demanda de productos (i.e., muchas variantes de producto a fabricar en series cada vez más cortas) permitiéndoles mejorar su competitividad en un mercado globalizado. Este proyecto tiene una duración de dos años comprendidos entre 2022-2023. En este proyecto se ha diseñado un algoritmo de estimación de pose 6D fusionando las predicciones de modelos de aprendizaje profundo. Los resultados de la investigación se han publicado en: Merino, I., Azpiazu, J., Remazeilles, A., Sierra, B, 2023. Ensemble of 6 dof pose estimation from state-of-the-art deep methods . Neurocomputing, Volume 541. Elsevier. https://doi.org/10.1016/j.neucom.2023.126270 El autor ha participado también en proyectos de investigación que han contribuido a la carrera investigadora del autor, pero que no están ligados directamente con la tesis. Estos proyectos son los siguientes: 11
1. Introducción y motivación 1.5. Estructura de la memoria Esta memoria está dividida en dos partes. En la Parte Ise expone la investigación realizada. Ésta a su vez está dividida en Capítulos. En total hay 4 Capítulos: Introducción y motivación, Técnicas de visión para la detección de objetos, Contribuciones y Conclusiones. El Capítulo 1es el capítulo actual y en él se expone la motivación y las hipótesis de la investigación. En el Capítulo 2se hace un estado del arte sobre las diferentes técnicas de visión por computador, primero se introduce y se presentan los tipos de tareas que existen en visión por computador (2.1), después se continúa con técnicas clásicas 2D (2.2), técnicas clásicas 3D (2.3), técnicas de aprendizaje profundo (2.4), y por último, sobre datasets y benchmarks (2.5). En el Capítulo 3se exponen las contribuciones que se han realizado para dar respuesta a las hipótesis presentadas en la Sección 1.4. Finalmente, en el Capítulo 4se exponen las conclusiones que se han obtenido durante el proyecto de tesis. Por último, en la Parte II están como anexos los artículos que presentan las contribuciones expuestas en el Capítulo 3. 18
CAPÍTULO 2 Técnicas de visión por computador para la detección de objetos. 2.1. Introducción Uno de los problemas clave tratados en este proyecto de tesis es la estimación de pose 6D de objetos observados por un sensor de visión. Esta consiste en estimar la posición y orientación de un objeto en el espacio. Esta pose está compuesta por 6 dimensiones, que se pueden expresar de distintas maneras. Una de ellas y la más común es: x , y , z , θ , φ , ψ (los tres primeros para la posición y los 3 últimos para la rotación). Estas dimensiones se pueden representar mediante una matriz de rotación y una matriz de traslación o mediante una matriz de transformación homogénea. Hay que tener en cuenta que esta pose 6D es relativa a un sistema de coordenadas concreto. Como norma general, se suele utilizar el sistema de coordenadas de la cámara, ya que es el sistema de coordenadas que se utiliza para obtener las imágenes. Para entender mejor la estimación de pose 6D, es importante tener en cuenta los siguientes conceptos. Esa estimación de pose 6D se realiza a partir de la información visual adquirida por la cámara. Estos datos adquiridos por la cámara se pueden representar de diferentes maneras, formatos o tipologías. Por un lado tenemos los datos 2D, que pueden ser imágenes RGB, imágenes térmicas, imágenes de profundidad, etc; y por otro lado tenemos los datos 3D, que pueden ser nubes de puntos, voxels, etc. Entre los datos 2D las 19
2. Técnicas de visión por computador para la detección de objetos. imágenes RGB son las más comunes. Estas son imágenes 2D que contienen información de color. Cada pixel de una imagen RGB contiene información de color en los canales Rojo (R: Red), Verde (G: Green) y Azul (B: Blue). Cualquier imagen 2D puede estar compuesta de uno o varios canales. En el caso de las RGB son 3 canales (R, G y B), pero también existen imágenes 2D en escala de grises que solo tienen un canal, como por ejemplo las imágenes térmicas o las de profundidad. Las imágenes de profundidad son imágenes 2D que contienen información de la profundidad. Cada pixel de una imagen de profundidad contiene información de la distancia a la que se encuentra el objeto que se está observando. Por eso, aunque son imágenes 2D, contienen información 3D e incluso se pueden transformar en nubes de puntos. El uso de datos 3D es menos común que el uso de datos 2D, debido a que los datos 3D son más costosos de obtener y procesar. Aun así, existen sensores que pueden obtener datos 3D de forma relativamente sencilla, como por ejemplo los sensores de profundidad estilo Kinect. Las nubes de puntos son estructuras de datos 3D que contienen un conjunto de puntos en el espacio, cada uno con su coordenada x , y y z . Además, cada punto también puede disponer de información de color y de su normal. El vóxel es la contraparte 3D del píxel. Es la unidad cúbica que compone una matriz tridimensional. Al igual que los píxeles, no tienen información de la posición en el espacio, sino que viene dada por su posición en la estructura de datos. Muchas veces los vóxeles se utilizan para representar una nube de puntos, ya que es una forma de reducir la cantidad de datos. A mayor tamaño de vóxel mayor cantidad de puntos abarcará y será más fácil procesarla, pero menor será la resolución y, por consiguiente, la precisión que se obtendrá. La Figura 2.1 muestra un ejemplo de un objeto del dataset YCB-Video en formato malla, voxel y nube de puntos. (a) Malla (Mesh). (b) Voxel. (c) Nube de puntos. Figura 2.1: Pinza del dataset YCB-Video. Además, también es importante definir el resto de tareas típicas de la visión por computador. La clasificación es una de las tareas más básicas, pero a su vez más utilizadas, de la visión por computador. Consiste en dar una clase a una imagen. La segmentación semántica es una extensión de la clasificación. Consiste en etiquetar la imagen a nivel de 20
2.1. Introducción píxel o punto, de esta forma se divide la imagen en diferentes regiones, cada una de ellas perteneciente a una clase diferente. La segmentación de instancia consiste en localizar en la imagen todas las instancias diferentes. La segmentación panóptica es una combinación de las dos anteriores, la cual consiste en dividir la imagen en diferentes regiones, cada una de ellas perteneciente a una instancia diferente, y además a cada instancia se le asigna una clase. La detección 2D consiste en encontrar objetos en imágenes. Estas detecciones se dan en forma de Bounding Box 2D (caja delimitadora). La Figura 2.2 [ 20 ] muestra un ejemplo de detección 2D, segmentación semántica, segmentación de instancia y segmentación panóptica para ver sus diferencias. Otra tarea importante es el tracking o seguimiento de objetos. Consiste en seguir un objeto durante una secuencia (en un vídeo por ejemplo). Figura 2.2: Comparativa de una detección 2D, segmentación semántica, segmentación de instancia y segmentación panóptica a la misma imagen. (Origen de la imagen: lamaquinaoraculo.com) Muchas de las tareas de visión por computador que se han mencionado anteriormente se pueden realizar con datos 2D o 3D. Además, la salida de una tarea puede ser 21
2. Técnicas de visión por computador para la detección de objetos. de diferentes tipologías de datos. Cuando hacemos una detección 2D obtenemos un Bounding Box 2D y cuando hacemos una detección 3D obtenemos un Bounding Box 3D. Aun así, existen métodos que permiten hacer una correspondencia entre datos 2D y 3D, como por ejemplo PnP [ 21 ] (Perspective-n-Point). Este método permite obtener la pose 6D de un objeto a partir de una detección 2D y un modelo 3D del objeto. La Figura 2.3 muestra el conjunto de las tareas que se pueden realizar con visión por computador. Figura 2.3: Tareas más comunes de visión por computador con el output que generan. En este capítulo, se van a presentar técnicas clásicas 2D (Sección 2.2), que se pueden 22
2.2. Técnicas clásicas 2D dividir entre descriptores globales (Subsección 2.2.1), descriptores locales (Subsección 2.2.2), y la busqueda de correspondencias y clasificación (Subsección 2.2.3). Posteriormente, se van a presentar técnicas 3D (Sección 2.3) para la estimación de pose 6D: métodos locales basados en parches (Subsección 2.3.1), métodos locales basados en nubes de puntos (Subsección 2.3.2) y métodos basados en plantillas (Subsección 2.3.3). Además, se van a introducir las bases del aprendizaje profundo (Sección 2.4) y las arquitecturas y técnicas más utilizadas como son las redes neuronales convolucionales (Subsección 2.4.1), las redes neuronales recurrentes (Subsección 2.4.2), los autoencoders (Subsección 2.4.3), los modelos generativos adversariales (Subsección 2.4.4), Transformers (Subsección 2.4.5), modelos de difusión (Subsección 2.4.6) y aprendizaje transferido (Subsección 2.4.7). Por último, se va a presentar el estado del arte de los datasets más utilizados para la estimación de pose 6D (Sección 2.5), como el benchmark BOP (Subsección 2.5.1) y herramientas para la generación de datasets sintéticos (Subsección 2.5.2). 2.2. Técnicas clásicas 2D Dentro de la visión por ordenador más tradicional o visión clásica, el pipeline de detección de objetos ha seguido un patrón determinado, el cual ha consistido en buscar características concretas o importantes en las imágenes, describir o codificar esas características y buscar correspondencias entre esas características codificadas. La Figura 2.4 muestra los pipelines o flujos que siguen la gran mayoría de algoritmos de visión clásica. El primer flujo, la descripción global, está recogido en 2.4a. En este contexto de descripción global, se busca extraer características de la imagen en su totalidad formando un vector de características de tamaño concreto que depende del tipo de descriptor. Este vector de características se usa para clasificar la imagen. El segundo flujo, la descripción local, se resume en 2.4b. En este contexto primero se detectan características concretas las cuales se llaman keypoints o puntos de interés que pueden ser esquinas, bordes, blobs o parches de la imagen concretos. Una vez detectados estos puntos de interés, se describen cada uno de ellos mediante un descriptor local. De esta forma obtenemos un vector de características por cada punto de interés. Estos vectores de características se pueden utilizar para clasificar la imagen en su totalidad, similar a lo que se hace con los descriptores globales, o matchear (emparejar) esos vectores con los de modelos conocidos para lograr obtener la localización de esos objetos conocidos (estimación de pose). La Figura 2.5 muestra la relación entre los diferentes métodos clásicos locales y globales del estado del arte. Las flechas indican que un método se basa en otro, que es una mejora respecto al anterior o que es una combinación de varios métodos. 2.2.1. Descriptores globales Los descriptores globales nos permiten describir una imagen en su totalidad y formar un vector de características único para toda la imagen. 23
2. Técnicas de visión por computador para la detección de objetos. (a) Descriptores globales (b) Descriptores locales Figura 2.4: Flujos de visión por ordenador 2D clásico Entre ellos podemos encontrarnos HOG (Histograms of Oriented Gradients) [ 22 ] que es un histograma de gradientes orientados normalizado localmente, HSOG (Histograms of the Second-Order Gradients) [ 23 ] que calcula los gradientes de segundo orden para capturar la curvatura relacionada con propiedades geométricas, GIST [ 24 ] que utiliza una representación de la estructura espacial dominante de una escena que integra un conjunto de dimensiones perceptuales que llaman Spatial Envelope, y MPEG-7 [25,26] que está basado en el estándar con el mismo nombre e incluye diferentes descriptores basados en histogramas como CLD (Color Layout Descriptor), EHD (Edge Histogram Descriptor) y SCD (Scalable Color Descriptor). Existe un gran número de descriptores locales que se pueden utilizar como descriptores globales calculando su histograma. A este grupo de algoritmos se les llama histogramas de patrones equivalentes [27]. Entre ellos podemos encontrar: LBP (Local Binary Pattern) [ 28 ] es el más conocido de ellos el cual es robusto frente a variaciones de luz. LBP da pie a una gran variedad de diferentes descriptores locales que describen texturas y tienen un nivel de caracterización bajo para calcular su histograma. STU (Simplified Texture Unit) [ 29 ] consigue reducir el rango de posibles valores sin una pérdida significativa en el poder de caracterización. 24
2.2. Técnicas clásicas 2D Figura 2.5: Relación entre los diferentes métodos clásicos locales y globales del estado del arte 25
2. Técnicas de visión por computador para la detección de objetos. MTS (Modified Texture Spectrum) [ 30 ] es una versión simplificada del LBP utilizando un subset de los píxeles. CLBP (Completed Local Binary Pattern) [ 31 ] combina tres diferentes versiones del LBP modificado para mejorar la invarianza a la rotación. LTP (Local Ternary Pattern) [ 32 ] es una generalización del LBP más discriminativa y menos sensible al ruido en regiones uniformes pero deja de ser estrictamente invariante a transformaciones a nivel de grises. ELTP (Enhanced Local Ternary Pattern) [ 33 ] ataca el problema que tiene LTP utilizando una estrategia adaptativa para seleccionar el umbral o threshold. CELTP (Completed Enhanced Local Ternary Pattern) [ 33 ] utiliza la misma estrategia que el CLBP usando ELTP. LTrP (Local Tetra Pattern) [ 34 ] extrae mayor información utilizando derivadas de mayor orden. BGC (Binary Gradient Contours) [ 35 ] presenta 3 formas diferentes de calcular el contorno del gradiente binario de forma parecida a LBP. Añadiendo filtros Gabor a LBP obtenemos el descriptor LGBPHS (Local Gabor Binary Pattern Histogram Sequence) [36]. LQP (Local Quantized Pattern) [ 37 ] es una generalización del LBP que utiliza un vector de cuantificación y una tabla de búsqueda para utilizar más pixeles para la caracterización del patrón local y un nivel de caracterización mayor sin sacrificar simplicidad y eficiencia computacional, GLCM (Gray Level Coocurrences Matrices) [ 38 ] está basado en describir las dependencias espaciales en tonos grises. Es un descriptor de texturas muy simple pero muy utilizado en muchas aplicaciones. WLD (Weber Local Descriptor) [ 39 ] es un descriptor local muy simple pero la vez muy potente y robusto basado en la ley de Weber [40]. Como está definido en el pipeline de los descriptores globales (Figura 2.4a), una vez obtenido el vector de características global (como descriptor global o histograma de patrones equivalentes), se puede utilizar para clasificar la imagen en su totalidad. Para ello, se utilizan algoritmos de aprendizaje automático o de matcheo (ver Sección 2.2.3.1). 26
2.2. Técnicas clásicas 2D 2.2.2. Descriptores locales Para hacer una correspondencia (matching) de características, es necesario primero encontrar puntos o regiones en la imagen que tengan cierta significancia, es decir, que sean características del objeto. Para ello, se utilizan detectores de puntos de interés (keypoint detectors); estos pueden ser detectores de bordes, de esquinas, de regiones (blobs) o de otro tipo. 2.2.2.1. Detectores Un detector de puntos de interés localiza puntos que tengan ciertas características para que sean reconocibles, como por ejemplo, bordes o esquinas. La Figura 2.6 muestra un ejemplo de detección de puntos de interés en una imagen de una pieza industrial. Figura 2.6: Detección de puntos de interés en una pieza industrial. Entre los detectores de bordes más conocidos están los Steerable filters [ 41 ], Sobel [ 42 ] y Canny [ 43 , 44 , 45 ]. Los filtros orientables o Steerable filters permiten sintetizar filtros de orientaciones arbitrarias a partir de filtros básicos. El detector de bordes Sobel se basa en un operador isotrópico para obtener el gradiente en un vecindario de 3x3. El detector Canny, en cambio, detecta bordes buscando máximos en el gradiente de imágenes con un suavizado gaussiano. 27
2. Técnicas de visión por computador para la detección de objetos. Figura 2.9: Un árbol de decisión para clasificación de tres clases. Como el clasificador SVM es binario, se utilizan estrategias como uno contra todos (one-vs-all) [ 84 ] o uno contra uno (one-vs-one) [ 85 ] para transformar un problema multiclase en múltiples problemas binarios. SVM es un clasificador muy utilizado para clasificar descriptores HOG (anteriormente mencionados), como por ejemplo para el reconocimiento de caras o detección de objetos. Los árboles de decisión buscan fabricar un sistema de predicción basado en una serie de condiciones de forma recursiva para la resolución de un problema. Este tipo de estructuras se parecen a las que se generan con los sistemas basados en agentes o en reglas. La Figura 2.9 muestra un árbol de decisión que tiene en cuenta 3 de las N variables ( X0 , X1 y X3 ) y 3 clases ( C1 , C2 y C3 ). Se puede apreciar en el ejemplo que no es necesario que el árbol utilice todas las variables disponibles, porque no son necesarias o porque se ha realizado una poda. También existen métodos que utilizan múltiples algoritmos de aprendizaje en conjunto para refinar el resultado que da cada algoritmo por separado. Estos métodos se les conoce como ensemble y algunas de esas técnicas son: bagging [ 86 ], boosting [ 87 ], stacking [88] y voting [89]. 2.3. Técnicas 3D Dentro de la visión por computador, existen muchos sensores que capturan nubes de puntos o imágenes de profundidad en vez de imágenes 2D de color. Estos sensores pueden ser cámaras por triangulación láser, de luz estructurada, estéreo o basadas en tiempo de vuelo. La adquisición de esos sensores produce una nube de puntos o un mapa de profundidad. Con los parámetros intrínsecos de los sensores se pueden obtener las nubes de puntos con los mapas de profundidad. Trabajar con nubes de puntos en vez de con imágenes 2D tiene mayor complejidad. 34
2.3. Técnicas 3D Esto se debe a diversas razones. La primera es que la adquisición de la nube de puntos genera datos dispersos, es decir, no es capaz de adquirir información de toda la escena 3D, la resolución del sensor limita los datos obtenidos. Segundo, las técnicas para trabajar con imágenes son más eficientes que las que trabajan con nubes de puntos. Añadir una dimensión a los datos incrementa su complejidad para tratarlos. Debido a la complejidad de los datos, se han abordado la estimación de pose 6D desde diferentes perspectivas. 2.3.1. Métodos locales basados en parches Estos métodos utilizan sistemas de votos basados en árboles (como random forest) sobre parches locales para la detección, localización y estimación de pose. Entre estos métodos nos podemos encontrar el método diseñado por Brachmann et al. que estima directamente las coordenadas de los objetos [ 90 ], Tejani et al. que utilizan un parche invariante a la escala para predecir la localización y la pose [ 91 ] o Bonde et al. que utilizan una ventana deslizante volumétrica [92]. 2.3.2. Métodos basados en matcheo de nubes de puntos Estos métodos son los más utilizados para la detección, localización y registro en nubes de puntos. Tienen bastantes similitudes con los descriptores 2D, pero en vez de describir el color describen la geometría de los objetos. Spin Images [ 93 ]: es un descriptor de la forma a nivel de dato para matchear superficies junto con un esquema de compresión para reconocer simultaneamente varios objetos. Point Feature Histogram (PFH) [ 94 ]: es un descriptor que captura información sobre la geometría al rededor de un punto analizando las direcciones de las normales de su vecindario. Fast Point Feature Histogram (FPFH) [ 95 ]: una versión más rápida de PFH manteniendo prácticamente todo su poder discriminativo. Radius-based Surface Descriptor (RSD) [ 96 ]: busca describir la forma de la superficie alrededor de un punto. Por cada de punto de interés con su vecino se calcula la diferencia entre sus normales y la distancia entre ellos. Se ajusta una esfera de forma que encaje con las normales y la distancia entre los puntos. De esta manera si la superficie es plana el radio de la esfera será infinito, y en caso contrario, el radio de la esfera será parecido a la de la superficie curva. 3D Shape Context (3DSC) [ 97 ]: es un descriptor que se basa en el Shape Context [73] pero en 3D. 35
2. Técnicas de visión por computador para la detección de objetos. Unique Shape Context (USC) [ 98 ]: extiende el descriptor 3DSC definiendo un frame de referencia para tener orientaciones únicas reduciendo la comp-lejidad del descriptor y mejorando su precisión. Signature of Histograms of Orientations (SHOT) [ 99 ]: encapsula la información topologica de la superficie de cada punto de interés con una estructura de soporte parecida a la usada en 3DSC. Las divisiones de la esfera son fijas (32) y se calcula un histograma por cada volumen. Es un método invariante a la rotación y robusto ante el ruido y al desorden. Point Pair Features (PPF) [ 100 ]: es un descriptor formado por información a pares de puntos que incluye la distancia entre los puntos, el ángulo que forman cada normal con el vector que resulta de la diferencia entre las normales y el ángulo que forman las dos normales. Este descriptor es invariante a transformaciones rígidas. Iteratite Closest Point (ICP) [ 101 ]: es un método iterativo que se utiliza para alinear dos nubes de puntos. ICP necesita una estimación inicial por lo que se suele utilizar para refinar las poses obtenidas por otros métodos. La mayoría de estos métodos se pueden encontrar en la librería PCL [ 102 ], la cual es una librería de código abierto C++ para el procesamiento de nubes de puntos. 2.3.3. Métodos basados en plantillas Estos métodos primero crean unas plantillas de los objetos a reconocer que capturan las diferentes formas de los objetos desde diferentes perspectivas. Los objetos son detectados cuando una plantilla encaja en la imagen y su pose es la que da esa plantilla. Uno de los primeros métodos se llama Snakes [ 103 ] que consiste en una Spline que se ajusta a la forma del objeto. A raíz de este método han surgido muchos otros que siguen la misma idea [ 104 , 105 ]. Uno de los autores que más ha trabajado en este campo es Stefan Hinterstoisser. En 2010 [106], propuso un método basado en plantillas que buscaba las orientaciones de gradiente dominantes en un pequeño subconjunto de píxeles. Otros métodos parecidos propuestos por el mismo autor [ 107 ] han dado pie a uno de los métodos basado en plantillas más conocidos LINEMOD [ 108 ], el cual posteriormente también ha sido refinado [ 109 ]. Otros métodos que utilizan plantillas son [ 110 ] y [ 111 ]. 2.4. Aprendizaje profundo Las bases de las redes neuronales se han afianzado en el siglo pasado [ 112 , 113 ], pero no ha sido hasta esta última década cuando se han podido utilizar de manera eficiente. El aprendizaje profundo ha ido ganando cada vez más popularidad debido a los méritos 36
2.4. Aprendizaje profundo conseguidos con ellas. En 2012, el modelo de aprendizaje profundo AlexNet [ 114 ] ganó el challenge ImageNet [ 115 ] con una ventaja del 10,8 % del top-5 error respecto al siguiente método. Esto fue posible gracias a la profundidad de la red y a que se utilizaron unidades de procesamiento gráfico (GPUs) durante el entrenamiento. Este momento ha sido un punto de inflexión para las redes neuronales profundas. A partir de ese momento y gracias la facilidad de obtener una GPU, han surgido una infinidad de arquitecturas y modelos que se han utilizado en diversos sectores [116,117,118]. El challenge ImageNet ha sido bastante importante en la generación de nuevas arquitecturas de aprendizaje profundo. Los métodos que conseguian liderar el podium han conseguido relevancia y popularidad. Entre ellos se encuentran: VGG [ 119 ], Inception [ 120 , 121 ], ResNet [ 122 ], Inception-ResNet [ 123 ], ResNeXt [ 124 ], NASNET [ 125 ], EfficientNet [ 126 ] y ViT [ 127 ]. ImageNet es un dataset para clasificación, por lo que la mayoría de métodos que se han presentado son para clasificación. Aun así, la parte de extracción de características de esos modelos se puede utilizar para otras tareas. En cuanto a estimación de pose 6D, el BOP challenge [ 128 ] ha definido una referencia común para comparar técnicas. La mayoría de las técnicas 2D de aprendizaje profundo primero hacen la detección 2D y obtienen la pose 6D de esa detección [ 129 , 130 , 131 ] o directamente buscan correspondencias 2D a 3D para obtener la pose 6D [ 132 , 133 , 134 , 135 , 136 ]. Profundizando más en estos métodos, los métodos como [ 135 ]o[ 136 ] siguen un planteamiento muy parecido a los métodos tradicionales. Existen también muchos métodos semi-supervisados que utilizan datos no etiquetados para mejorar el rendimiento de las redes neuronales. Por ejemplo, Pseudo Labels [ 137 ], Noisy Student [ 138 ] o Meta Pseudo Labels [ 139 ]. O incluso métodos que permiten manejar entradas y salidas arbitrarias de datos [140,141]. Para entender el funcionamiento de las redes neuronales, se van a explicar las partes más comunes que las componen y las arquitecturas más utilizadas. Las partes más comunes son las siguientes: Capa densa (Fully-connected): es la capa más básica. Consiste en un conjunto de funciones no lineales que encapsulan una neurona (perceptron). Estas neuronas aplican una transformación lineal a una entrada. A esta salida es a la que se le aplica una transformación no lineal a traves de la anteriormente mencionada función de no linealidad. Por ello, la salida de esta capa y es el resultado de aplicar la función de activación f (no linear) al producto escalar entre la matriz de pesos Wy el vector de entrada x, más el vector de bias b(Ecuación 2.5). y=f(Wx +b)(2.5) Las funciones de activación posibles se explican más adelante. 37
2. Técnicas de visión por computador para la detección de objetos. Capa convolucional: las capas convolucionales son capas que aplican una convolución a la entrada. La convolución es una operación matemática que consiste en un producto escalar, donde el núcleo (kernel) se desplaza a lo largo de la matriz de entrada, y tomamos el producto escalar entre ambos como si fueran vectores (Figura 2.10). Figura 2.10: Ilustración del comportamiento de una capa convolucional. Imagen: Diego Unzueta. Como en las capas densas, a la salida de la convolución se le aplica una función de activación no lineal. Capa de pooling: Las capas de pooling sirven para reducir la dimensionalidad de un mapa de características (matriz de salida de una capa convolucional). Para un kernel de tamaño k y un stride s , el tamaño de mapa de características de nh×nw×nc se reduce a un tamaño de (nh−f+ 1)/s ×(nw−f+ 1)/s ×nc , donde nh es la altura, nw es el ancho y nc es el número de canales del mapa de características. Existen diferentes tipos de pooling, como por ejemplo el max pooling (selecciona el valor máximo para cada desplazamiento del kernel), el average pooling (computa la media de los valores que abarca el kernel) o el global pooling (hace un max o average pooling que reduce el mapa de características a 1×1×nc). Capa de normalización: Son capas que permiten reducir el desplazamiento interno de covariables [ 142 ] que es un error que se da en la actualización de los pesos de los modelos cuando se asume que los pesos anteriores a la actualización son fijos. Para ello se existen métodos como el Batch normalization [ 142 ] o el Layer normalization [ 143 ] que ayudan a coordinar la actualización de pesos en multiples capas. Función de activación: Las funciones de activación son funciones no lineales que se aplican a la salida de una capa para introducir no linealidad en la red. Algunas de las funciones de activación más utilizadas son: ReLU [ 144 ], Leaky ReLU [ 145 ], ELU [146], SELU [147], tanh [148], sigmoid [148] y softmax [149]. 38
2.4. Aprendizaje profundo Dropout: Para evitar que las redes neuronales se sobreajusten (overfitting) a los datos de entrenamiento, se utilizan técnicas de regularización. Una de las técnicas más utilizadas es el dropout [ 150 ]. Esta técnica consiste en desactivar aleatoriamente un porcentaje de neuronas de la red durante el entrenamiento. De esta manera, la red no puede depender de una neurona en concreto para realizar la predicción. Loss: Las redes neuronales se tratan como un problema de optimización en la que los pesos de las capas son los parámetros a optimizar. Para ello se utiliza una función de coste que se minimiza mediante un algoritmo de optimización (por ejemplo, Stochastic Gradient Descent). Esta función de coste se conoce como función de pérdida (loss function). Existen diferentes tipos de funciones de pérdida en función del problema que se quiera resolver. Por ejemplo, para problemas de clasificación se suele utilizar la entropía cruzada (cross entropy) y para problemas de regresión se suele utilizar el error cuadrático medio (mean squared error). Los grupos de arquitecturas más relevantes (aunque existen muchos más) son: Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Autoencoders (AE), Generative Adversarial Networks (GAN), Transformers y Modelos de Difusión (Difussion models). En las siguientes secciones se explican brevemente cada uno de ellos. 2.4.1. Convolutional Neural Networks (CNN) Las CNN son las redes neuronales más utilizadas hasta la fecha en el aprendizaje profundo para visión por computador. Estas consisten en una serie de capas convolucionales y capas de pooling que extraen características de la imagen. La Figura 2.11 muestra un ejemplo de una arquitectura convolucional. Algunas de las CNN más conocidas y utilizadas son: VGG [ 119 ], Inception [ 120 , 121 ], ResNet [ 122 ], Inception-ResNet [ 123 ], ResNeXt [ 124 ], NASNET [ 125 ] y EfficientNet [ 126 ]. Muchas de estas arquitecturas han definido unos bloques básicos que se pueden repetir para crear arquitecturas más complejas. Por ejemplo, en ResNet se definen el bloque residual, el cual añade una conexión de salto entre bloques y Inception module el bloque de Inception, que combina varias capas de multiples tamaños de kernel que luego pasa a la siguiente capa concatenados. 2.4.2. Recurrent Neural Networks (RNN) Las redes neuronales recurrentes son redes neuronales que tienen una memoria interna que les permite recordar información de pasos anteriores. Por ello, son muy útiles para procesar datos secuenciales como texto o audio, aunque también se pueden utilizar 39
2. Técnicas de visión por computador para la detección de objetos. Figura 2.11: Ejemplo de una arquitectura convolucional que toma como entrada una imagen 2D, y está compuesta por 3 capas convolucionales y 2 capas de pooling intercalados. Su salida se aplana para formar el vector de características, que se pasa a la capa densa para su clasificación. Figura 2.12: Ejemplo simplificado de una arquitectura recurrente. combinadas con arquitecturas convolucionales para procesar vídeos. La Figura 2.12 muestra un ejemplo de una arquitectura recurrente. Algunas de las RNN más conocidas y utilizadas son: LSTM [151] y GRU [152]. 2.4.3. Autoencoders (AE) Los Autoencoders [ 153 ] son un tipo de arquitectura que se utiliza para aprender representaciones eficientes de datos. Estas arquitecturas se componen de dos partes: una codificadora y una decodificadora. La codificadora es una red neuronal que reduce la dimensionalidad de los datos de entrada y la decodificadora es una red neuronal que reconstruye los datos de entrada a partir de la codificación. La Figura 2.13 muestra un ejemplo de una arquitectura Autoencoder. Estas arquitecturas son muy utilizadas para tareas de clasificación (Sparse, Denoising y Contractive autoencoders) que incluyen 40
2.4. Aprendizaje profundo Figura 2.13: Ejemplo simplificado de una rquitectura Autoencoder que se compone de una primera parte de encoding y su posterior decoding. reconocimiento de caras, detección de características o detección de anomalias, o como modelos generativos (Variational Autoencoders) para generar datos aleatorios similares a los datos de entrada. Los Sparse autoencoders como [ 154 , 155 ] añaden restricciones de dispersión en las unidades ocultas. Los Denoising Autoencoders [ 156 , 157 ], en cambio, corrompen la entrada de datos a propósito con ruido para evitar que la red aprenda la función identidad y consiga reducir la dimensionalidad o aprender representaciones útiles. Los Contractive Autoencoders [ 158 , 159 ] añaden una penalización de en la función de coste de reconstrucción. Los Variational Autoencoders [ 160 ] son una variante de los autoencoders que permiten generar muestras aleatorias de los datos de entrada, el cual se suele utilizar como modelo generativo. 2.4.4. Generative Adversarial Networks (GAN) Los modelos generativos adversariales o Generative Adversarial Network [ 161 , 162 , 163 ] son un tipo de arquitectura generativa que entrenan simultaneamente dos modelos uno que genera nuevos ejemplos y otro que discrimina entre ejemplos reales y generados. La Figura 2.14 muestra un ejemplo de una arquitectura GAN. Algunos ejemplos de modelos GAN para generar imagenes 2D son [ 164 ], [ 165 ] y [ 166 ]; también se han utilizado para generar datos 3D como [167]y[168]. 2.4.5. Transformers Los Transformers [ 169 , 170 ] son un tipo de arquitectura de red neuronal que se basa en la atención. La atención es un mecanismo que permite a una red neuronal prestar atención a diferentes partes de la entrada para realizar una tarea. La Figura 2.15 muestra 41
2. Técnicas de visión por computador para la detección de objetos. Figura 2.14: Representación de un sistema generativo GAN. un ejemplo de una arquitectura Transformer y la Figura 2.16 muestra la arquitectura de la capa de atención. En la arquitectura Transformer de [ 170 ] se utilizan capas de atención multi-cabeza, las cuales consisten en concatenar varias capas de atención en paralelo con pesos diferentes. Estas arquitecturas son muy utilizadas en tareas de NLP (Natural Language Processing) como traducción, resumen de texto o generación de texto [ 171 ]. Algunas de estas arquitecturas son genéricas y se pueden utilizar para resolver muchas de las tareas de NLP, como por ejemplo, BERT [ 172 ], GPT [ 173 , 174 , 175 ] y XLNet [ 176 ]. Aun así estos últimos años también se estan utilizando en tareas de visión por computador [177,178,127]. 2.4.6. Modelos de difusión Los modelos de Difusión son unos modelos generativos que han ganado mucho interés en los últimos años [ 179 , 180 , 181 ]. Los modelos difusos se basan en la difusión de una distribución inicial para generar una distribución final. Es decir, en añadir ruido iterativamente a una imagen para generar datos de entrenamiento y entrenar un modelo (el modelo de difusión) el cual sea capaz de deshacer cada paso de ruido añadido a las imagenes de entrenamiento. La Figura 2.17 muestra un esquema del proceso de entrenamiento de una arquitectura de difusión. Estos modelos se pueden utilizar para 42
2.4. Aprendizaje profundo Figura 2.15: Ejemplo de una arquitectura Transformer compuesto de capas de atención multicabeza. 43
3. Contribuciones [218] Merino, I., Azpiazu, J., Remazeilles, A., Sierra, B., 2020. Histogram-Based Descriptor Subset Selection for Visual Recognition of Industrial Parts . Applied science 10, 3701. MDPI. https://doi.org/10.3390/app10113701 [ 219 ] J.L. Outón, I. Merino, I. Villaverde, A. Ibarguren, H. Herrero, P. Daelman, B. Sierra, 2021. A Real Application of an Autonomous Industrial Mobile Manipulator within Industrial Context. Electronics 10, 1276. MDPI. https://doi.org/10. 3390/electronics10111276 [ 220 ] Merino, I., Azpiazu, J., Remazeilles, A., Sierra, B., 2021. 3D Convolutional Neural Networks Initialized from Pretrained 2D Convolutional Neural Networks for Classification of Industrial Parts . Sensors 21, 1078. MDPI. https://doi.org/10.3390/s21041078 [ 221 ] Merino, I., Azpiazu, J., Remazeilles, A., Sierra, B, 2023. Ensemble of 6 dof pose estimation from state-of-the-art deep methods . Neurocomputing, Volume 541. Elsevier. https://doi.org/10.1016/j.neucom.2023.126270 3.2. Contribuciones a la visión por computador clásica La visión por computador clásica 2D no ha tenido avances significativos en la última década. Esto se debe al gran éxito que han tenido los algoritmos de aprendizaje profundo. Aun así, como se refleja en la hipótesis H1 , creemos que este tipo de algoritmos aún no está obsoleto y para ciertos casos es mejor utilizar modelos de detección o clasificación clásicos en vez de algoritmos de aprendizaje profundo. Muchos de los algoritmos de detección como SIFT o SURF han sido líderes entre este tipo de métodos durante muchos años para buscar correspondencias entre texturas; por otro lado HOG es muy utilizado para detectar y clasificar caras. Es por eso que muchos de estos algoritmos tienen un uso concreto, es decir, que funcionan muy bien para ciertos casos de uso mientras que para otros no tan bien. Por eso hasta el auge de los algoritmos de aprendizaje profundo era necesario un experto o experta en visión para poder determinar cuál era el algoritmo que se tenía que usar para el caso de uso concreto que se quería resolver. Los algoritmos de aprendizaje profundo han resuelto esta problemática aprendiendo cómo codificar o filtrar la imagen para luego clasificarla. Por lo tanto, existen algunos casos de uso concretos en los que el modelo de aprendizaje profundo requiere menor presencia del experto o experta en visión. Aun así, para poder configurar bien el modelo de aprendizaje profundo sigue siendo necesario que alguien con conocimientos en el tema lo realice. Partiendo de esta premisa nos preguntamos si existía alguna forma de replicar este comportamiento de los algoritmos de aprendizaje profundo: buscar un método que aprendiese cuál es el algoritmo de visión clásica o qué combinación de ellos es la mejor para el caso concreto al que nos enfrentámos. 50
3.2. Contribuciones a la visión por computador clásica Para ello desarrollamos el método presentado en [ 216 , 217 ]. El caso de uso de esos artículos es clasificar imágenes de ciertas piezas industriales. Para ello se define un clasificador jerárquico que se entrena en dos fases. En la primera fase se busca separar las piezas en K subgrupos diferentes en función de cómo se comportan con cada descriptor. Para ello, se calculan los valores F1 , por objeto para cada pipeline (detector + descriptor + matching). El F1 es la media armónica entre la precisión o precision y la exhaustividad o recall (Ecuación 3.1). El pipeline que obtenga mayor F1 es el que se utilizará como clasificador de tipologías. Una vez obtenidos los F1 por objeto y pipeline, clusterizamos estos objetos en K ( K < número de objetos ) tipologías utilizando K-means. Una vez definidas las llamadas tipologías, para cada una de ellas se utiliza el descriptor que mejor funciona. F1=2 recall−1+precision−1(3.1) Con la primera fase terminada, hemos definido 3 elementos: El pipeline que se utiliza para separar las tipologías. Este es el pipeline que obtiene mayor media de valores F1. Las diferentes tipologías. Estas tipologias pueden estar compuestas por uno o varios objetos. Los pipelines que mejor funcionan para cada tipología. Esos pipelines son los que mayor media tengan para cada tipología y servirán para diferenciar cada objeto dentro de la tipología. En la segunda fase se entrena el pipeline que clasifica las tipologías y los pipelines de cada una de las tipologías. Para evaluar el clasificador jerárquico se ha utilizado el valor F1 obtenido con un número de imágenes y número de objetos variable. Se han utilizado sets de 10, 20, 30, 40 y 50 imágenes por objeto y 3, 4, 5, 6 y 7 objetos, dando un total de 25 casos diferentes. Debido al bajo número de imágenes que hay en el dataset y que para algunas de las combinaciones solo se tienen 10 imágenes por objeto, se utiliza un Nested Leave-OneOut Cross-Validation (Nested LOOCV [ 222 , 223 ]). Primero se calculan los vectores de características para todas las imágenes con cada pipeline (utilizando solo el detector y descriptor sin hacer el matching). El LOOCV interior sirve para definir los sets de entrenamiento y validación de la primera fase. El LOOCV exterior, en cambio, sirve para entrenar los clasificadores en la segunda fase y testearlos. Los resultados obtenidos con el clasificador jerárquico para nuestro dataset son mejores que con los pipelines definidos. Se puede apreciar que a mayor cantidad de imágenes y número de objetos, la mejora del clasificador jerárquico es más significativa. 51
3. Contribuciones Otro método que sigue la premisa que se ha presentado al inicio y da respuesta a la hipótesis H1 es el que se presenta en [ 218 ]. En este artículo se presenta una técnica de selección de características de descriptores globales para clasificación de piezas industriales. Este método define la forma de buscar qué conjunto de descriptores globales funciona mejor para un caso de uso concreto. Para ello, se utilizan técnicas de selección de características (Feature Selection), más concretamente una selección de características secuencial hacia delante y hacia atrás (Sequential Forward Subset Selection ySequential Backward Subset Selection, respectivamente). La SFSS, de sus siglas en ingles, parte del conjunto vacío y va incluyendo en cada iteración el descriptor que maximice la métrica a evaluar hasta que no se obtenga mejora. La SBSS, en cambio, parte del conjunto de todos los descriptores y va retirando descriptores siempre que mejore la métrica. En el caso del artículo la métrica es el F1. En el artículo se realizan experimentos con diferentes descriptores y formas de aplicar los descriptores tanto con SFSS y SBSS. Los experimentos muestran que al utilizar SFSS aplicando el descriptor a toda la imagen, en vez de a trozos uniformes y concatenarlos, se obtienen los mejores resultados. De hecho esta versión es también la más rápida en computar y la más sencilla de implementar. Debido al bajo número de imágenes que se tiene en el dataset (mismo dataset que en [ 216 , 217 ]) este método supera a los resultados obtenidos con redes neuronales como Xception [224] o Siamese [225]. Con todo esto, se ha dado una respuesta a la hipótesis H1 . Se puede confirmar que existen formas de mejorar los métodos clásicos actuales y que todavía tienen un gran margen de mejora. Aplicar refinamientos o fusión de características a los métodos clásicos permite incrementar la precisión de estos. En nuestros casos, conseguimos unas mejoras en torno al 10 % respecto de los métodos base. 3.3. Contribuciones a la transferencia de aprendizaje 2D-3D Como se ha explicado en la Sección 2.2 los métodos de aprendizaje profundo han revolucionado la visión por computador. Las redes convolucionales, los datasets con una gran cantidad de imagenes y los avances en hardware han permitido desarrollar técnicas de visión muy precisas y, aunque lentas entrenando, rápidas en la inferencia. Este gran avance es especialmente notorio para imágenes 2D pero los métodos que incorporan datos 3D como nubes de puntos o vóxeles 3D no obtienen resultados tan buenos como los obtenidos con los métodos de aprendizaje profundo 2D. Esto se puede deber a la mayor dificultad de adquisición de este tipo de datos (sensores más caros si se quiere una mayor precisión) y la dificultad de etiquetarlos (por ejemplo definir las poses 6D de los objetos). En la Sección 3.4 se profundizará más sobre esto último. El artículo [ 220 ] presenta un método para dar respuesta a este problema ( H2 ). Este método consiste en transformar los pesos de redes convolucionales 2D preentrenadas para inicializar los pesos de redes convolucionales 3D. Para ello se presentan dos 52
3.3. Contribuciones a la transferencia de aprendizaje 2D-3D transformaciones: la extrusión de pesos y la rotación de pesos. La extrusión consiste en replicar los pesos 2D a lo largo de un eje (X, Y o Z). En la rotación se rotan los pesos respecto de un eje como al generar un cuerpo de revolución. Las Figuras 3.1 y3.2, extraidas de [220], muestran un ejemplo de extrusión y rotación, respectivamente. Figura 3.1: Extrusión de pesos 2D para obtener los pesos 3D: se replica una la matriz de pesos 2D a lo largo de nueva dimensión añadida para inicializar arquitecturas 3D. Figura 3.2: Rotación de pesos 2D para obtener los pesos 3D: para cada valor de la matriz de dimensionalidad agrandada se le asigna un valor de la matriz de pesos 2D siguiendo el mapeo de la Ecuación 3.2 para inicializar arquitecturas 3D. 53
3. Contribuciones T(x, y, z) = M(x, min(bpy2+z2c, H)),(3.2) donde T es la matriz de pesos 3D, M es la matriz de pesos 2D, H es la altura de la matriz de pesos 2D y x,yyzson las indices de la posición de las matrices de pesos. Se compara el comportamiento de las versiones 3D de 4 arquitecturas de redes convolucionales (VGG16 [ 119 ], ResNet [ 122 ], Inception-ResNet v2 [ 123 ] y EfficientNet [ 126 ]) inicializadas con los pesos de sus versiones 2D preentrenados con ImageNet [ 114 ] transformados con las técnicas antes mencionadas. Las versiones 3D de estas arquitecturas sustituyen todas las capas 2D por sus respectivas versiones 3D, y se ajusta el padding () y las dimensiones de entrada: VGG16: 224 x 224 x 3 (input versión 2D) a 96 x 96 x 96 x 3 (input versión 3D). ResNet: 224 x 224 x 3 (input versión 2D) a 96 x 96 x 96 x 3 (input versión 3D). Inception-ResNet v2: 299 x 299 x 3 (input versión 2D) a 139 x 139 x 139 x 3 (input versión 3D). EfficientNet: sigue la siguiente ecuación: r3D=b3 pr2 2Dc , donde r2D es la resolución en 2D y r3D es la resolución en 3D. Por ejemplo, para EfficientNet-B0 pasa de224x224x3a36x36x36x3. Estas arquitecturas inicializadas con los pesos 2D transformados consiguen obtener un F1 mejor que sus versiones sin preentrenamiento. En la Tabla 3.1 se puede ver que ambos preentrenamientos funcionan mejor para todos los casos que la versión sin preentrenar. Para algunos casos funciona mejor la extrusión y para otros la rotación por lo que no es posible decir con los experimentos actuales que uno de los dos preentrenamientos sea mejor que el otro. Además, se han comparado con una arquitectura (PointNet [ 226 ]), que toma como entrada nubes de puntos, perteneciente al estado del arte y el método propuesto obtiene mejores resultados. Arquitectura Sin preentrenamiento Extrusión Rotación ResNet 0,8272 0,8612 0,8512 Inception ResNet 0,7558 0,8887 0,8939 EfficientNet B0 0,8605 0,9217 0,9052 EfficientNet B1 0,8372 0,8420 0,8422 PointNet 0,9048 - - Tabla 3.1: Comparativa de los F1 obtenidos en los experimentos con cada arquitectura y preentrenamiento 54
3.4. Contribuciones a la generación de datasets sintéticos Como se ha podido apreciar sigue existiendo un margen de mejora para redes convolucionales 3D. En comparación con sus versiones 2D los datasets que se disponen para entrenar son mucho más grandes, debido a la complejidad de adquisición y etiquetado de los datos 3D. Aun así, con la transferencia de conocimiento 2D a 3D se ha conseguido hasta un 6 % de mejora para algunas arquitecturas, validando la hipótesis H2. 3.4. Contribuciones a la generación de datasets sintéticos Como hemos comentado en la anterior sección, uno de los mayores problemas para trabajar con datos 3D es la dificultad de definir las posiciones reales de los objetos. En imágenes 2D, normalmente, se suelen dar como posición de los objetos el bounding box 2D de estos. Definir este bounding box es bastante sencillo y existen muchas herramientas que facilitan su etiquetado [ 227 , 228 , 229 ]. En el caso 3D este proceso es bastante más complicado ya que añadimos una dimensión extra. Equivocarse a la hora de definir la posición de los objetos es más común y puede complicar el entrenamiento. Para intentar solucionar esta problemática, en los últimos años se han empezado a utilizar datos sintéticos ultra-realistas basados en físicas. Las herramientas de modelado 3D, diseño de videojuegos y entornos virtuales han avanzado mucho tanto que muchas veces cuesta diferenciar una imagen real de una generada por ordenador. Es por ello que este tipo de herramientas se han utilizado en la generación de datasets sintéticos para entrenar diferentes métodos. El método propuesto en [ 220 ] ha sido entrenado con un dataset sintético para dar respuesta a H3 . Este dataset se ha generado usando Unreal Engine 4 (UE4) [ 207 ] y el plugin NVIDIA Deep Learning Dataset Synthesizer (NDDS) [ 208 ]. Este dataset está compuesto por 7 piezas industriales que pertenecen a varios proyectos en los que el autor ha participado. Se ha hecho una reconstrucción de esas piezas utilizando una cámara 3D de luz proyectada y CloudCompare [230]. Una vez generados los modelos, éstos se importan a UE4. Gracias al plugin NDDS, se definen unos distractores, objetos geométricos aleatorios, cuya posición se va adaptando en cada iteración; se define una iluminación con cambios en la intensidad y dirección; se definen cambios en el fondo y se aleatorizan las posiciones de los modelos. Por cada iteración se realizan cambios en todos aspectos antes comentados y se obtiene una captura RGBD junto con su segmentación semántica, la segmentación de instancia y por cada objeto la pose 6D, la visibilidad, el bounding box 2D, el bounding box 3D y la proyección 2D del bounding box 3D. Las Figuras 3.3a,3.3b,3.3c y3.3d muestran un ejemplo de una captura de una iteración de UE4 con el plugin NDDS. La Figura 3.3a es la captura RGB, la Figura 3.3b es la imagen de profundidad, la Figura 3.3c es la segmentación semántica (Ground truth) y la Figura 3.3d es la segmentación de instancia (Ground truth). 55
3. Contribuciones (a) Captura RGB obtenida mediante UE4 y el plugin NDDS (b) Imagen de profundidad obtenida mediante UE4 y el plugin NDDS (c) Segmentación semántica obtenida mediante UE4 y el plugin NDDS (d) Segmentación de instancia obtenida mediante UE4 y el plugin NDDS En [ 221 ] también se propone el uso de imágenes sintéticas para mejorar los resultados obtenidos con redes neuronales. En este caso la herramienta utilizada es BlenderProc [ 214 ]. BlenderProc es un pipeline procedural de Blender [ 213 ], un entorno abierto y gratuito para la creación de contenido 3D, para generar renderizados fotorealistas. El pipeline utilizado para generar las imágenes sintéticas para este caso es el que define el BOP challenge [231]. Las imágenes obtenidas con BlenderProc (Figura 3.4) son mucho más realistas que las obtenidas con UE4 y además la forma de definir el pipeline es mucho más sencilla. De hecho gracias a BlenderProc muchas de las técnicas de aprendizaje profundo que han participado en el BOP challenge que previamente no lograban superar a los métodos PPF (Point Pair Features), han conseguido liderar la tabla de clasificación. En el caso de nuestro artículo ocurre lo mismo, entrenar los métodos utilizados en él (PVN3D, FFB6D y Cosypose) con solo datos reales obtiene peores resultados que entrenando con datos sintéticos. 56
3.5. Contribuciones de ensamblaje de métodos (ensemble) (a) Imagen sintética del dataset Linemod (b) Imagen sintética del dataset YCBV (c) Imagen sintética del dataset TLESS Figura 3.4: Imagenes sintéticas generadas con BlenderProc para los datasets Linemod, YCBV y TLESS Utilizando este tipo de herramientas de generación de datasets sintéticos, se consigue disminuir la cantidad de datos reales necesarios para entrenar modelos de aprendizaje profundo que generalicen bien y permitan clasificar, localizar o detectar objetos de forma eficiente. Esto es especialmente necesario para datos 3D debido a su complejidad de adquisición y etiquetado. Por eso, la H3 es correcta y no solo ayuda a disminuir los datos reales necesarios, sino que también mejora los resultados que se obtienen utilizando solo imágenes reales. 3.5. Contribuciones de ensamblaje de métodos (ensemble) Al igual que se han explicado las contribuciones realizadas en visión clásica mediante fusión o mejora de estos algoritmos en la Sección 3.2, en esta Sección vamos a hacer lo mismo para algoritmos de aprendizaje profundo. Una de las grandes diferencias entre los algoritmos clásicos y los de aprendizaje profundo es que los primeros por lo general estaban integrados por diferentes herramientas dentro del pipeline (detectores, descriptores y matching), las cuales se definen en función del caso del uso y del objetivo que se quiere lograr. Mientras que en los algoritmos de aprendizaje profundo se intenta evitar esto y por lo general se buscan o diseñan métodos end-to-end, es decir, que una única arquitectura permita realizar todo lo que antes hacia cada parte y entrenarlo en su totalidad. Por un lado, esto permite que el entrenamiento de las diferentes fases se retroalimente, esto es, que las características o features que se extraen estén pensadas para que el clasificador o estimador de pose que viene después sea capaz de utilizarlas de forma eficiente para maximizar su propósito, o lo que es lo mismo minimizar el la función de perdida o Loss function. Esta función sirve para evaluar la desviación entre las predicciones de la red neuronal y los valores reales. Por otro lado, los métodos clásicos permiten definir exactamente qué es lo que queremos extraer (bordes, esquinas, puntos de interés...). Aun así definir estos extractores de características y descriptores es mucho más laborioso y se necesita un experto o experta en visión para poder definir 57
3. Contribuciones que extractor y descriptor concreto funciona bien para el tipo de datos que tenemos (formas de los objetos, tipo de cámara, iluminación...). Como se puede apreciar cada uno tiene sus ventajas y desventajas por eso queremos extraer el potencial de ambos tipos de métodos fusionando métodos de aprendizaje profundo con técnicas más clásicas. En [ 221 ] hemos presentado varios métodos de ensemble para fusionar varias técnicas de aprendizaje profundo para estimar poses 6D y de esta forma dar respuesta a la hipotesis H4 . Para ello, tomamos como base 3 modelos de aprendizaje profundo: PVN3D[ 135 ], FFB6D[ 136 ] y Cosypose[ 131 ]. Los dos primeros métodos se basan en encontrar ciertos puntos de interés de los objetos (parte de aprendizaje profundo) y luego se aplica un ajuste de mínimos cuadrados (Least Square Fitting) con los puntos de interés del modelo. Los dos métodos difieren en la parte de aprendizaje profundo. El primero, PVN3D, saca información geométrica de la nube de puntos e información de color de la imágen RGB y luego fusiona esta información. Luego estos features se pasan a 3 módulos que sirven para obtener los centroides, la segmentación semántica y los puntos de interés. Gracias a los centroides y la segmentación semántica se pueden obtener los puntos de interés para cada instancia presente en la escena. La principal diferencia con el segundo modelo, FFB6D, es la parte de extracción de features. En vez de obtener las características geométricas y de color por separado, en cada capa de la red se fusionan ambos tipos de características para aprovechar la información geométrica a la hora de extraer features de color y viceversa. Cosypose, en cambio, es un método que solo utiliza imágenes RGB para obtener la pose 6D. Primero hace una detección gruesa de las poses 6D con una red convolucional gruesa (CNN coarse) y después refina esas poses 6D con otra red convolucional (CNN refiner). Cosypose incluye un refinamiento multivista que no hemos utilizado. Una vez definidos los 3 modelos, vamos a definir cómo se han fusionado. Se han propuesto 2 estrategias: la estrategia de unión (merge) y la estrategia de apilado (stacking). Por un lado, la estrategia de unión permite juntar los resultados de los modelos base de forma geométrica. Esta estrategia consiste en hacer una media de las poses de los modelos bases. La forma más sencilla es hacer la media sin contemplar nada más. Esto es lo que llamamos simple merge. Otra propuesta es añadir la confianza que dan los métodos sobre sus predicciones. De forma que se hace una media ponderada de las poses (weighted merge). Esto sigue teniendo sus inconvenientes cuando las poses 6D que da alguno de los modelos sea de alguna otra instancia. Para solucionar esto, se propone la unión por agrupación (clustering merge). Este método pretende primero agrupar las poses 6D obtenidas de los modelos base en grupos de poses cercanas y luego aplicar un simple merge oweighted merge al grupo más grande. En el caso de que tengan la misma cantidad de elementos se elige el grupo con mayor confianza media. De esta forma tenemos 4 métodos dentro de la estrategia de unión: simple merge,weighted merge, clustering simple merge yclustering weighted merge. 58
3.5. Contribuciones de ensamblaje de métodos (ensemble) Por otro lado, la estrategia de apilado utiliza un modelo de machine learning para fusionar las poses 6D que estiman los modelos base. Estos modelos tienen como objetivo tomar como entrada los valores de las poses 6D y hacer una regresión de la nueva pose refinada. Se han probado 6 diferentes modelos: SVR, Arboles de decisión, Regresión lineal de Ridge, Regresión lineal de mínimos cuadrados ordinarios, Regresor KNN y MLP. Los resultados obtenidos demuestran que la estrategia de unión funciona mejor para regresión que las estrategias de apilado. Además, se logra mejorar para algunos datasets las poses 6D obtenidas con los modelos base. Esto valida la hipótesis H4. 59
Bibliografía [1] Heiner Lasi, Peter Fettke, Hans-Georg Kemper, Thomas Feld, and Michael Hoffmann. Industry 4.0. Business & information systems engineering, 6(4):239–242, 2014. Ver página 3. [2] European Commission, Directorate-General for Research, Innovation, Maija Breque, Lars De Nul, and Athanasios Petridis. Industry 5.0 : towards a sustainable, human-centric and resilient European industry. Publications Office, 2021. Ver página 3. [3] Samir Yitzhak Gadre, Mitchell Wortsman, Gabriel Ilharco, Ludwig Schmidt, and Shuran Song. Clip on wheels: Zero-shot object navigation as object localization and exploration. arXiv preprint arXiv:2203.10421, 2022. Ver página 4. [4] Zhengxue Zhou, Leihui Li, Alexander Fürsterling, Hjalte Joshua Durocher, Jesper Mouridsen, and Xuping Zhang. Learning-based object detection and localization for a mobile robot manipulator in sme production. Robotics and Computer-Integrated Manufacturing, 73:102229, 2022. Ver página 4. [5] Guoguang Du, Kai Wang, Shiguo Lian, and Kaiyong Zhao. Vision-based robotic grasping from object localization, object pose estimation to grasp estimation for parallel grippers: a review. Artificial Intelligence Review, 54(3):1677–1734, 2021. Ver página 4. [6] E Shreyas, Manav Hiren Sheth, and Mohana. 3d object detection and tracking methods using deep learning for computer vision applications. In 2021 International Conference on Recent Trends on Electronics, Information, Communication & Technology (RTEICT), pages 735–738. IEEE, 2021. Ver página 4. [7] Zhuoqi Cheng and Thiusius Rajeeth Savarimuthu. A novel robot-assisted electrical impedance scanning system for subsurface object detection. Measurement Science and Technology, 32(8):085902, 2021. Ver página 4. [8] Márton Szemenyei and Vladimir Estivill-Castro. Fully neural object detection solutions for robot soccer. Neural Computing and Applications, 34(24):21419–21432, 2022. Ver página 4. [9] Imran Ahmed, Sadia Din, Gwanggil Jeon, Francesco Piccialli, and Giancarlo Fortino. Towards collaborative robotics in top view surveillance: A framework for multiple object tracking by detection using deep learning. IEEE/CAA Journal of Automatica Sinica, 8(7):1253– 1270, 2021. Ver página 4. 67
Bibliografía [10] Jiayao Shan, Sifan Zhou, Zheng Fang, and Yubo Cui. Ptt: Point-track-transformer module for 3d single object tracking in point clouds. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 1310–1316. IEEE, 2021. Ver página 4. [11] Yizhe Wu, Oiwi Parker Jones, Martin Engelcke, and Ingmar Posner. Apex: Unsupervised, object-centric scene segmentation and tracking for robot manipulation. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 3375–3382. IEEE, 2021. Ver página 4. [12] Andrea Bonci, Pangcheng David Cen Cheng, Marina Indri, Giacomo Nabissi, and Fiorella Sibona. Human-robot perception in industrial environments: A survey. Sensors, 21(5):1571, 2021. Ver página 4. [13] Janis Arents, Valters Abolins, Janis Judvaitis, Oskars Vismanis, Aly Oraby, and Kaspars Ozols. Human–robot collaboration trends and safety aspects: A systematic review. Journal of Sensor and Actuator Networks, 10(3):48, 2021. Ver página 4. [14] Óscar G Hernández, Vicente Morell, José L Ramon, and Carlos A Jara. Human pose detection for robotic-assisted and rehabilitation environments. Applied Sciences, 11(9):4183, 2021. Ver página 4. [15] Matheus Zorawski Silva, Thadeu Brito, José L Lima, and Manuel F Silva. Industrial robotic arm in machining process aimed to 3d objects reconstruction. In 2021 22nd IEEE International Conference on Industrial Technology (ICIT), volume 1, pages 1100–1105. IEEE, 2021. Ver página 4. [16] Mohamed Tahoun, Omar Tahri, Juan Antonio Corrales Ramón, and Youcef Mezouar. Visualtactile fusion for 3d objects reconstruction from a single depth view and a single gripper touch for robotics tasks. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 6786–6793. IEEE, 2021. Ver página 4. [17] William Agnew, Christopher Xie, Aaron Walsman, Octavian Murad, Yubo Wang, Pedro Domingos, and Siddhartha Srinivasa. Amodal 3d reconstruction for robotic manipulation via stability and connectivity. In Conference on Robot Learning, pages 1498–1508. PMLR, 2021. Ver página 4. [18] Aleksandar Jokić, Milica Petrović, and Zoran Miljković. Semantic segmentation based stereo visual servoing of nonholonomic mobile robot in intelligent manufacturing environment. Expert Systems with Applications, 190:116203, 2022. Ver página 4. [19] Jiehao Li, Yingpeng Dai, Junzheng Wang, Xiaohang Su, and Ruijun Ma. Towards broad learning networks on unmanned mobile robot for semantic segmentation. In 2022 International Conference on Robotics and Automation (ICRA), pages 9228–9234. IEEE, 2022. Ver página 4. [20] Rubén Rodríguez Abril. Segmentación panóptica. https://lamaquinaoraculo.com/ computacion/segmentacion-panoptica/. Accessed: 2023-02-13. Ver página 21. [21] Vincent Lepetit, Francesc Moreno-Noguer, and Pascal Fua. Ep n p: An accurate o (n) solution to the p n p problem. International journal of computer vision, 81:155–166, 2009. Ver página 22. 68
Bibliografía [22] Navneet Dalal and Bill Triggs. Histograms of oriented gradients for human detection. In IEEE computer society conference on computer vision and pattern recognition (CVPR’05), volume 1, pages 886–893. Ieee, 2005. Ver página 24. [23] Di Huang, Chao Zhu, Yunhong Wang, and Liming Chen. Hsog: a novel local image descriptor based on histograms of the second-order gradients. IEEE Transactions on Image Processing, 23(11):4680–4695, 2014. Ver página 24. [24] Aude Oliva and Antonio Torralba. Modeling the shape of the scene: A holistic representation of the spatial envelope. International journal of computer vision, 42(3):145–175, 2001. Ver página 24. [25] Bangalore S Manjunath, J-R Ohm, Vinod V Vasudevan, and Akio Yamada. Color and texture descriptors. IEEE Transactions on circuits and systems for video technology, 11(6):703–715, 2001. Ver página 24. [26] Thomas Sikora. The mpeg-7 visual standard for content description-an overview. IEEE Transactions on circuits and systems for video technology, 11(6):696–702, 2001. Ver página 24. [27] Antonio Fernández, Marcos X Álvarez, and Francesco Bianconi. Texture description through histograms of equivalent patterns. Journal of mathematical imaging and vision, 45(1):76–102, 2013. Ver página 24. [28] Timo Ojala, Matti Pietikainen, and Topi Maenpaa. Multiresolution gray-scale and rotation invariant texture classification with local binary patterns. IEEE Transactions on pattern analysis and machine intelligence, 24(7):971–987, 2002. Ver página 24. [29] Francisco J Madrid-Cuevas, R Medina Carnicer, M Prieto Villegas, NL García, and A Carmona Poyato. Simplified texture unit: A˜ new descriptor of the local texture in gray-level images. In Iberian Conference on Pattern Recognition and Image Analysis, pages 470–477. Springer, 2003. Ver página 24. [30] Bing Xu, Peng Gong, Edmund Seto, and Robert Spear. Comparison of gray-level reduction and different texture spectrum encoding methods for land-use classification using a panchromatic ikonos image. Photogrammetric Engineering & Remote Sensing, 69(5):529–536, 2003. Ver página 26. [31] Zhenhua Guo, Lei Zhang, and David Zhang. A completed modeling of local binary pattern operator for texture classification. IEEE transactions on image processing, 19(6):1657–1663, 2010. Ver página 26. [32] Xiaoyang Tan and Bill Triggs. Enhanced local texture feature sets for face recognition under difficult lighting conditions. IEEE transactions on image processing, 19(6):1635–1650, 2010. Ver página 26. [33] Jing-Hua Yuan, Hao-Dong Zhu, Yong Gan, and Li Shang. Enhanced local ternary pattern for texture classification. In International conference on intelligent computing, pages 443–448. Springer, 2014. Ver página 26. 69
Bibliografía [34] Subrahmanyam Murala, RP Maheshwari, and R Balasubramanian. Local tetra patterns: a new feature descriptor for content-based image retrieval. IEEE transactions on image processing, 21(5):2874–2886, 2012. Ver página 26. [35] Antonio Fernández, Marcos X Álvarez, and Francesco Bianconi. Image classification with binary gradient contours. Optics and Lasers in Engineering, 49(9-10):1177–1184, 2011. Ver página 26. [36] Wenchao Zhang, Shiguang Shan, Wen Gao, Xilin Chen, and Hongming Zhang. Local gabor binary pattern histogram sequence (lgbphs): A novel non-statistical model for face representation and recognition. In Tenth IEEE International Conference on Computer Vision (ICCV’05) Volume 1, volume 1, pages 786–791. IEEE, 2005. Ver página 26. [37] Sibt Ul Hussain, Thibault Napoléon, and Fréderic Jurie. Face recognition using local quantized patterns. In British machive vision conference, pages 11–pages, 2012. Ver página 26. [38] Robert M Haralick, Karthikeyan Shanmugam, and Its’ Hak Dinstein. Textural features for image classification. IEEE Transactions on systems, man, and cybernetics, (6):610–621, 1973. Ver página 26. [39] Jie Chen, Shiguang Shan, Chu He, Guoying Zhao, Matti Pietikäinen, Xilin Chen, and Wen Gao. Wld: A robust local image descriptor. IEEE transactions on pattern analysis and machine intelligence, 32(9):1705–1720, 2009. Ver página 26. [40] Gustav Theodor Fechner. Elemente der psychophysik, volume 2. Breitkopf u. Härtel, 1860. Ver página 26. [41] William T. Freeman and Edward H. Adelson. The design and use of steerable filters. IEEE Transactions on Pattern analysis and machine intelligence, 13(9):891–906, 1991. Ver página 27. [42] Irwin Sobel and Gary Feldman. A 3x3 isotropic gradient operator for image processing. a talk at the Stanford Artificial Project in, pages 271–272, 1968. Ver página 27. [43] John Canny. A computational approach to edge detection. IEEE Transactions on pattern analysis and machine intelligence, (6):679–698, 1986. Ver página 27. [44] Ramesh Jain, Rangachar Kasturi, and Brian G. Schunck. Machine vision, volume 5. McGrawhill New York, 1995. Ver página 27. [45] Emanuele Trucco and Alessandro Verri. Introductory techniques for 3-D computer vision, volume 201. Prentice Hall Englewood Cliffs, 1998. Ver página 27. [46] Hans Peter Moravec. Obstacle avoidance and navigation in the real world by a seeing robot rover. Stanford University, 1980. Ver página 28. [47] Chris Harris and Mike Stephens. A combined corner and edge detector. In Alvey vision conference, volume 15, pages 10–5244, 1988. Ver página 28. [48] Jianbo Shi and Tomasi. Good features to track. In IEEE Conference on Computer Vision and Pattern Recognition, pages 593–600, 1994. Ver página 28. 70
Bibliografía [49] Stephen M Smith and J Michael Brady. Susan—a new approach to low level image processing. International journal of computer vision, 23(1):45–78, 1997. Ver página 28. [50] Edward Rosten and Tom Drummond. Fusing points and lines for high performance tracking. In Tenth IEEE International Conference on Computer Vision (ICCV’05) Volume 1, volume 2, pages 1508–1515. Ieee, 2005. Ver página 28. [51] Elmar Mair, Gregory D Hager, Darius Burschka, Michael Suppa, and Gerhard Hirzinger. Adaptive and generic corner detection based on the accelerated segment test. In European conference on Computer vision, pages 183–196. Springer, 2010. Ver página 28. [52] Stefan Leutenegger, Margarita Chli, and Roland Y Siegwart. Brisk: Binary robust invariant scalable keypoints. In 2011 International conference on computer vision, pages 2548–2555. Ieee, 2011. Ver páginas 28,29. [53] Ethan Rublee, Vincent Rabaud, Kurt Konolige, and Gary Bradski. Orb: An efficient alternative to sift or surf. In 2011 International conference on computer vision, pages 2564–2571. Ieee, 2011. Ver páginas 28,29. [54] James L Crowley and Alice C Parker. A representation for shape based on peaks and ridges in the difference of low-pass transform. IEEE transactions on pattern analysis and machine intelligence, (2):156–170, 1984. Ver página 28. [55] Tony Lindeberg. Detecting salient blob-like image structures and their scales with a scalespace primal sketch: A method for focus-of-attention. International Journal of Computer Vision, 11(3):283–318, 1993. Ver página 28. [56] D.G. Lowe. Object recognition from local scale-invariant features. In Seventh IEEE International Conference on Computer Vision, volume 2, pages 1150–1157 vol.2, 1999. Ver páginas 28,29. [57] Tony Lindeberg. Scale-space theory: A basic tool for analyzing structures at different scales. Journal of applied statistics, 21(1-2):225–270, 1994. Ver página 28. [58] Herbert Bay, Tinne Tuytelaars, and Luc Van Gool. Surf: Speeded up robust features. In European conference on computer vision, pages 404–417. Springer, 2006. Ver páginas 28,29. [59] Motilal Agrawal, Kurt Konolige, and Morten Rufus Blas. Censure: Center surround extremas for realtime feature detection and matching. In European conference on computer vision, pages 102–115. Springer, 2008. Ver página 29. [60] Krystian Mikolajczyk and Cordelia Schmid. An affine invariant interest point detector. In European conference on computer vision, pages 128–142. Springer, 2002. Ver página 29. [61] Krystian Mikolajczyk and Cordelia Schmid. Scale & affine invariant interest point detectors. International journal of computer vision, 60(1):63–86, 2004. Ver página 29. [62] Jiri Matas, Ondrej Chum, Martin Urban, and Tomás Pajdla. Robust wide-baseline stereo from maximally stable extremal regions. Image and vision computing, 22(10):761–767, 2004. Ver página 29. [63] Tai Sing Lee. Image representation using 2d gabor wavelets. IEEE Transactions on pattern analysis and machine intelligence, 18(10):959–971, 1996. Ver página 29. 71
Bibliografía [64] John Illingworth and Josef Kittler. A survey of the hough transform. Computer vision, graphics, and image processing, 44(1):87–116, 1988. Ver página 29. [65] Michael Calonder, Vincent Lepetit, Christoph Strecha, and Pascal Fua. Brief: Binary robust independent elementary features. In European conference on computer vision, pages 778–792. Springer, 2010. Ver página 29. [66] Yan Ke and Rahul Sukthankar. Pca-sift: A more distinctive representation for local image descriptors. In 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2004. CVPR 2004., volume 2, pages II–II. IEEE, 2004. Ver página 29. [67] Svetlana Lazebnik, Cordelia Schmid, and Jean Ponce. A sparse texture representation using local affine regions. IEEE transactions on pattern analysis and machine intelligence, 27(8):1265–1278, 2005. Ver página 30. [68] Bin Fan, Fuchao Wu, and Zhanyi Hu. Rotationally invariant descriptors using intensity order pooling. IEEE transactions on pattern analysis and machine intelligence, 34(10):2031– 2045, 2011. Ver página 30. [69] Krystian Mikolajczyk and Cordelia Schmid. A performance evaluation of local descriptors. IEEE transactions on pattern analysis and machine intelligence, 27(10):1615–1630, 2005. Ver página 30. [70] Engin Tola, Vincent Lepetit, and Pascal Fua. Daisy: An efficient dense descriptor applied to wide-baseline stereo. IEEE transactions on pattern analysis and machine intelligence, 32(5):815–830, 2009. Ver página 30. [71] Zhenhua Wang, Bin Fan, and Fuchao Wu. Local intensity order pattern for feature description. In 2011 International Conference on Computer Vision, pages 603–610. IEEE, 2011. Ver página 30. [72] Matthew Brown, Richard Szeliski, and Simon Winder. Multi-image matching using multiscale oriented patches. In IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), volume 1, pages 510–517. IEEE, 2005. Ver página 30. [73] Serge Belongie, Jitendra Malik, and Jan Puzicha. Shape matching and object recognition using shape contexts. IEEE transactions on pattern analysis and machine intelligence, 24(4):509–522, 2002. Ver páginas 30,35. [74] Eli Shechtman and Michal Irani. Matching local self-similarities across images and videos. In 2007 IEEE Conference on Computer Vision and Pattern Recognition, pages 1–8. IEEE, 2007. Ver página 30. [75] Jingneng Liu, Guihua Zeng, and Jianping Fan. Fast local self-similarity for describing interest regions. Pattern Recognition Letters, 33(9):1224–1235, 2012. Ver página 30. [76] Alexandre Alahi, Raphael Ortiz, and Pierre Vandergheynst. Freak: Fast retina keypoint. In IEEE conference on computer vision and pattern recognition, pages 510–517. Ieee, 2012. Ver página 30. [77] Li Fei-Fei and Pietro Perona. A bayesian hierarchical model for learning natural scene categories. In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), volume 2, pages 524–531. IEEE, 2005. Ver página 31. 72
Bibliografía [78] Ajmal S Mian, Mohammed Bennamoun, and Robyn Owens. Three-dimensional modelbased object recognition and segmentation in cluttered scenes. IEEE transactions on pattern analysis and machine intelligence, 28(10):1584–1601, 2006. Ver página 31. [79] Karl Pearson. Liii. on lines and planes of closest fit to systems of points in space. The London, Edinburgh, and Dublin philosophical magazine and journal of science, 2(11):559–572, 1901. Ver página 31. [80] Harold Hotelling. Analysis of a complex of statistical variables into principal components. Journal of educational psychology, 24(6):417, 1933. Ver página 31. [81] Sam T Roweis and Lawrence K Saul. Nonlinear dimensionality reduction by locally linear embedding. science, 290(5500):2323–2326, 2000. Ver página 31. [82] Aapo Hyvärinen and Erkki Oja. Independent component analysis: algorithms and applications. Neural networks, 13(4-5):411–430, 2000. Ver página 31. [83] Corinna Cortes and Vladimir Vapnik. Support-vector networks. Machine learning, 20(3):273–297, 1995. Ver página 32. [84] Ryan Rifkin and Aldebaro Klautau. In defense of one-vs-all classification. The Journal of Machine Learning Research, 5:101–141, 2004. Ver página 34. [85] Stefan Knerr, Léon Personnaz, and Gérard Dreyfus. Single-layer learning revisited: a stepwise procedure for building and training a neural network. In Neurocomputing: algorithms, architectures and applications, pages 41–50. Springer, 1990. Ver página 34. [86] Leo Breiman. Bagging predictors. Machine learning, 24:123–140, 1996. Ver página 34. [87] Robert E Schapire. A brief introduction to boosting. In Ijcai, volume 99, pages 1401–1406, 1999. Ver página 34. [88] David H Wolpert. Stacked generalization. Neural networks, 5(2):241–259, 1992. Ver página 34. [89] Artittayapron Rojarath, Wararat Songpan, and Chakrit Pong-inwong. Improved ensemble learning for classification techniques based on majority voting. In 2016 7th IEEE international conference on software engineering and service science (ICSESS), pages 107–110. IEEE, 2016. Ver página 34. [90] Eric Brachmann, Alexander Krull, Frank Michel, Stefan Gumhold, Jamie Shotton, and Carsten Rother. Learning 6d object pose estimation using 3d object coordinates. In Computer Vision – ECCV 2014, pages 536–551. Springer, 2014. Ver página 35. [91] Alykhan Tejani, Danhang Tang, Rigas Kouskouridas, and Tae-Kyun Kim. Latent-class hough forests for 3d object detection and pose estimation. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, pages 462–477. Springer, 2014. Ver páginas 35,46. [92] Ujwal Bonde, Vijay Badrinarayanan, and Roberto Cipolla. Robust instance recognition in presence of occlusion and clutter. In Computer Vision – ECCV 2014, pages 520–535. Springer, 2014. Ver página 35. 73
Bibliografía [93] Andrew E Johnson and Martial Hebert. Using spin images for efficient object recognition in cluttered 3d scenes. IEEE Transactions on pattern analysis and machine intelligence, 21(5):433–449, 1999. Ver página 35. [94] Radu Bogdan Rusu, Zoltan Csaba Marton, Nico Blodow, and Michael Beetz. Persistent point feature histograms for 3d point clouds. In 10th International Conference on Intelligent Autonomous Systems (IAS-10), Baden-Baden, Germany, pages 119–128, 2008. Ver página 35. [95] Radu Bogdan Rusu, Nico Blodow, and Michael Beetz. Fast point feature histograms (fpfh) for 3d registration. In IEEE international conference on robotics and automation, pages 3212–3217. IEEE, 2009. Ver página 35. [96] Zoltan-Csaba Marton, Dejan Pangercic, Nico Blodow, Jonathan Kleinehellefort, and Michael Beetz. General 3d modelling of novel objects from a single view. In 2010 IEEE/RSJ International Conference on Intelligent Robots and Systems, pages 3700–3705. IEEE, 2010. Ver página 35. [97] Andrea Frome, Daniel Huber, Ravi Kolluri, Thomas Bülow, and Jitendra Malik. Recognizing objects in range data using regional point descriptors. In Computer Vision-ECCV 2004: 8th European Conference on Computer Vision, Prague, Czech Republic, pages 224–237. Springer, 2004. Ver página 35. [98] Federico Tombari, Samuele Salti, and Luigi Di Stefano. Unique shape context for 3d data description. In ACM workshop on 3D object retrieval, pages 57–62, 2010. Ver página 36. [99] Samuele Salti, Federico Tombari, and Luigi Di Stefano. Shot: Unique signatures of histograms for surface and texture description. Computer Vision and Image Understanding, 125:251–264, 2014. Ver página 36. [100] Bertram Drost, Markus Ulrich, Nassir Navab, and Slobodan Ilic. Model globally, match locally: Efficient and robust 3d object recognition. In 2010 IEEE computer society conference on computer vision and pattern recognition, pages 998–1005. Ieee, 2010. Ver página 36. [101] Paul J. Besl and Neil D. McKay. Method for registration of 3-D shapes. In Paul S. Schenker, editor, Sensor Fusion IV: Control Paradigms and Data Structures, volume 1611, pages 586 – 606. International Society for Optics and Photonics, SPIE, 1992. Ver página 36. [102] Radu Bogdan Rusu and Steve Cousins. 3D is here: Point Cloud Library (PCL). In IEEE International Conference on Robotics and Automation (ICRA), Shanghai, China, May 9-13 2011. IEEE. Ver página 36. [103] Michael Kass, Andrew Witkin, and Demetri Terzopoulos. Snakes: Active contour models. International journal of computer vision, 1(4):321–331, 1988. Ver página 36. [104] Daniel P Huttenlocher, Gregory A. Klanderman, and William J Rucklidge. Comparing images using the hausdorff distance. IEEE Transactions on pattern analysis and machine intelligence, 15(9):850–863, 1993. Ver página 36. [105] Carsten Steger. Similarity measures for occlusion, clutter, and illumination invariant object recognition. In Pattern Recognition: 23rd DAGM Symposium Munich, Germany, pages 148–154. Springer, 2001. Ver página 36. 74
Bibliografía [106] Stefan Hinterstoisser, Vincent Lepetit, Slobodan Ilic, Pascal Fua, and Nassir Navab. Dominant orientation templates for real-time detection of texture-less objects. In IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pages 2257–2264. IEEE, 2010. Ver página 36. [107] Stefan Hinterstoisser, Cedric Cagniart, Slobodan Ilic, Peter Sturm, Nassir Navab, Pascal Fua, and Vincent Lepetit. Gradient response maps for real-time detection of textureless objects. IEEE transactions on pattern analysis and machine intelligence, 34(5):876–888, 2011. Ver página 36. [108] Stefan Hinterstoisser, Stefan Holzer, Cedric Cagniart, Slobodan Ilic, Kurt Konolige, Nassir Navab, and Vincent Lepetit. Multimodal templates for real-time detection of texture-less objects in heavily cluttered scenes. In International conference on computer vision, pages 858–865. IEEE, 2011. Ver página 36. [109] Stefan Hinterstoisser, Vincent Lepetit, Slobodan Ilic, Stefan Holzer, Gary Bradski, Kurt Konolige, and Nassir Navab. Model based training, detection and pose estimation of texture-less 3d objects in heavily cluttered scenes. In Asian conference on computer vision, pages 548–562. Springer, 2013. Ver página 36. [110] Reyes Rios-Cabrera and Tinne Tuytelaars. Discriminatively trained templates for 3d object detection: A real time scalable approach. In IEEE international conference on computer vision, pages 2048–2055, 2013. Ver página 36. [111] Paul Wohlhart and Vincent Lepetit. Learning descriptors for object recognition and 3d pose estimation. In IEEE conference on computer vision and pattern recognition, pages 3109–3118, 2015. Ver página 36. [112] Warren S McCulloch and Walter Pitts. A logical calculus of the ideas immanent in nervous activity. The bulletin of mathematical biophysics, 5:115–133, 1943. Ver página 36. [113] Frank Rosenblatt. The perceptron: a probabilistic model for information storage and organization in the brain. Psychological review, 65(6):386, 1958. Ver página 36. [114] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. Ver páginas 37,54. [115] Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3):211–252, 2015. Ver página 37. [116] Andreas Kamilaris and Francesc X Prenafeta-Boldú. Deep learning in agriculture: A survey. Computers and electronics in agriculture, 147:70–90, 2018. Ver página 37. [117] Juan Izquierdo Tomás. Reliability engineering advancements to enhance fleet asset management in service-oriented business models. 2020. Ver página 37. [118] Francesco Piccialli, Vittorio Di Somma, Fabio Giampaolo, Salvatore Cuomo, and Giancarlo Fortino. A survey on deep learning in medicine: Why, how and when? Information Fusion, 66:111–137, 2021. Ver página 37. 75
Bibliografía [197] Tomáš Hodaň, Pavel Haluza, Štěpán Obdržálek, Jiří Matas, Manolis Lourakis, and Xenophon Zabulis. T-LESS: An RGB-D dataset for 6D pose estimation of texture-less objects. IEEE Winter Conference on Applications of Computer Vision (WACV), 2017. Ver página 46. [198] Bertram Drost, Markus Ulrich, Paul Bergmann, Philipp Hartinger, and Carsten Steger. Introducing mvtec itodd-a dataset for 3d object recognition in industry. In IEEE international conference on computer vision workshops, pages 2200–2208, 2017. Ver página 46. [199] Roman Kaskman, Sergey Zakharov, Ivan Shugurov, and Slobodan Ilic. Homebreweddb: Rgb-d dataset for 6d pose estimation of 3d objects. In IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019. Ver página 46. [200] Berk Calli, Arjun Singh, Aaron Walsman, Siddhartha Srinivasa, Pieter Abbeel, and Aaron M Dollar. The ycb object and model set: Towards common benchmarks for manipulation research. In 2015 international conference on advanced robotics (ICAR), pages 510–517. IEEE, 2015. Ver página 46. [201] Colin Rennie, Rahul Shome, Kostas E Bekris, and Alberto F De Souza. A dataset for improved rgbd-based object detection and pose estimation for warehouse pick-and-place. IEEE Robotics and Automation Letters, 1(2):1179–1185, 2016. Ver página 46. [202] Andreas Doumanoglou, Rigas Kouskouridas, Sotiris Malassiotis, and Tae-Kyun Kim. Recovering 6d object pose and predicting next-best-view in the crowd. In IEEE conference on computer vision and pattern recognition, pages 3583–3592, 2016. Ver página 46. [203] Tomáš Hodaň, Jiří Matas, and Štěpán Obdržálek. On evaluation of 6d object pose estimation. In Computer Vision–ECCV 2016, pages 606–619. Springer, 2016. Ver página 46. [204] Eric Brachmann, Frank Michel, Alexander Krull, Michael Ying Yang, Stefan Gumhold, and Carsten Rother. Uncertainty-driven 6d pose estimation of objects and scenes from a single rgb image. In IEEE conference on computer vision and pattern recognition, pages 3364–3372, 2016. Ver página 46. [205] Celso M de Melo, Antonio Torralba, Leonidas Guibas, James DiCarlo, Rama Chellappa, and Jessica Hodgins. Next-generation deep learning based on simulators and synthetic data. Trends in cognitive sciences, 2021. Ver página 46. [206] John K Haas. A history of the unity game engine. 2014. Ver página 46. [207] Epic Games. Unreal engine. Ver páginas 46,55. [208] Thang To, Jonathan Tremblay, Duncan McKay, Yukie Yamaguchi, Kirby Leung, Adrian Balanon, Jia Cheng, William Hodge, and Stan Birchfield. NDDS: NVIDIA deep learning dataset synthesizer, 2018. https://github.com/NVIDIA/Dataset_Synthesizer . Ver páginas 46,55. [209] Unreal gt. https://unrealgt.github.io/. Accessed: 2022-12-20. Ver página 46. [210] James Fort Anthony Navarro. Supercharge your computer vision models with synthetic datasets built by unity. https://blog.unity.com/technology/superchargeyour-computer-vision-models-with-synthetic-datasets-built-by-unity , 2021. Accessed: 2023-01-24. Ver página 46. 82
Bibliografía [211] Nvidia. Nvidia omniverse. https://www.nvidia.com/es-es/omniverse/ , 2021. Accessed: 2023-01-26. Ver página 47. [212] Nvidia. Nvidia isaac sim. https://developer.nvidia.com/isaac-sim , 2021. Accessed: 2023-01-26. Ver página 47. [213] Blender Online Community. Blender - a 3D modelling and rendering package. Blender Foundation, Stichting Blender Foundation, Amsterdam, 2018. Ver páginas 47,56. [214] Maximilian Denninger, Martin Sundermeyer, Dominik Winkelbauer, Youssef Zidan, Dmitry Olefir, Mohamad Elbadrawy, Ahsan Lodhi, and Harinandan Katam. Blenderproc. arXiv preprint arXiv:1911.01911, 2019. Ver páginas 47,56. [215] Angel X. Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu. ShapeNet: An Information-Rich 3D Model Repository. Technical Report arXiv:1512.03012 [cs.GR], Stanford University — Princeton University — Toyota Technological Institute at Chicago, 2015. Ver página 47. [216] Ibon Merino, Jon Azpiazu, Anthony Remazeilles, and Basilio Sierra. 2d features-based detector and descriptor selection system for hierarchical recognition of industrial parts. international Journal of Artificial Intelligence and Applications (IJAIA), 10:1–13, 2019. Ver páginas 49,51, and 52. [217] Ibon Merino, Jon Azpiazu, Anthony Remazeilles, and Basilio Sierra. 2d image features detector and descriptor selection expert system. arXiv preprint arXiv:2006.02933, 2020. Ver páginas 49,51, and 52. [218] Ibon Merino, Jon Azpiazu, Anthony Remazeilles, and Basilio Sierra. Histogram-based descriptor subset selection for visual recognition of industrial parts. Applied Sciences, 10(11):3701, 2020. Ver páginas 50,52. [219] Jose Luis Outón, Ibon Merino, Iván Villaverde, Aitor Ibarguren, Héctor Herrero, Paul Daelman, and Basilio Sierra. A real application of an autonomous industrial mobile manipulator within industrial context. Electronics, 10(11):1276, 2021. Ver página 50. [220] Ibon Merino, Jon Azpiazu, Anthony Remazeilles, and Basilio Sierra. 3d convolutional neural networks initialized from pretrained 2d convolutional neural networks for classification of industrial parts. Sensors, 21(4):1078, 2021. Ver páginas 50,52,53, and 55. [221] Ibon Merino, Jon Azpiazu, Anthony Remazeilles, and Basilio Sierra. Ensemble of 6 dof pose estimation from state-of-the-art deep methods. Neurocomputing, page 126270, 2023. Ver páginas 50,56, and 58. [222] Mervyn Stone. Cross-validatory choice and assessment of statistical predictions. Journal of the royal statistical society: Series B (Methodological), 36(2):111–133, 1974. Ver página 51. [223] Sudhir Varma and Richard Simon. Bias in error estimation when using cross-validation for model selection. BMC bioinformatics, 7(1):1–8, 2006. Ver página 51. [224] François Chollet. Xception: Deep learning with depthwise separable convolutions. In IEEE conference on computer vision and pattern recognition, pages 1251–1258, 2017. Ver página 52. 83
Bibliografía [225] Iaroslav Melekhov, Juho Kannala, and Esa Rahtu. Siamese network features for image matching. In 2016 23rd international conference on pattern recognition (ICPR), pages 378–383. IEEE, 2016. Ver página 52. [226] Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In IEEE conference on computer vision and pattern recognition, pages 652–660, 2017. Ver página 54. [227] Kentaro Wada. labelme: Image polygonal annotation with python. https://github.com/ wkentaro/labelme, 2018. Ver página 55. [228] Boris Sekachev, Nikita Manovich, Maxim Zhiltsov, Andrey Zhavoronkov, Dmitry Kalinin, Ben Hoff, TOsmanov, Dmitry Kruchinin, Artyom Zankevich, DmitriySidnev, Maksim Markelov, Johannes222, Mathis Chenuet, a andre, telenachos, Aleksandr Melnikov, Jijoong Kim, Liron Ilouz, Nikita Glazov, Priya4607, Rush Tehrani, Seungwon Jeong, Vladimir Skubriev, Sebastian Yonekura, vugia truong, zliang7, lizhming, and Tritin Truong. opencv/cvat: v1.1.0, August 2020. Ver página 55. [229] Maxim Tkachenko, Mikhail Malyuk, Andrey Holmanyuk, and Nikolai Liubimov. Label Studio: Data labeling software, 2020-2022. Open source software available from https://github.com/heartexlabs/label-studio. Ver página 55. [230] Cloudcompare. https://www.danielgm.net/cc/. Accedido: 21-12-2022. Ver página 55. [231] Tomáš Hodaň, Martin Sundermeyer, Bertram Drost, Yann Labbé, Eric Brachmann, Frank Michel, Carsten Rother, and Jiří Matas. BOP challenge 2020 on 6D object localization. European Conference on Computer Vision Workshops (ECCVW), 2020. Ver página 56. [232] Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv, 2022. Ver página 64. [233] Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution text-to-3d content creation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. Ver página 64. [234] Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3: Zero-shot one image to 3d object, 2023. Ver página 64. [235] Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, dec 2021. Ver página 65. 84
Parte II Publicaciones obtenidas 85
CAPÍTULO 5 2d Image Features Detector and Descriptor Selection Expert System Título: 2d Image Features Detector and Descriptor Selection Expert System Autores: I. Merino, J. Azpiazu, A. Remazeilles, B. Sierra Conferencia: Computer Science & Information Technology (CS & IT) Editor: AIRCC Publishing Corporation DOI: 10.5121/csit.2019.91206 Año: 2019 Cuartil (Scimago/WoS): -/- 87
2D IMAGE FEATURES DETECTOR AND DESCRIPTOR SELECTION EXPERT SYSTEM Ibon Merino1, Jon Azpiazu1, Anthony Remazeilles1, and Basilio Sierra2 1Industry and Transport, Tecnalia Research and Innovation, Donostia-San Sebastian, Spain {ibon.merino, jon.azpiazu, anthony.remazeilles}@tecnalia.com 2Computer Science and Artificial Intelligence, University of the Basque Country UPV/EHU, Donostia-San Sebastian, Spain b[email protected] ABSTRACT Detection and description of keypoints from an image is a well-studied problem in Computer Vision. Some methods like SIFT, SURF or ORB are computationally really efficient. This paper proposes a solution for a particular case study on object recognition of industrial parts based on hierarchical classification. Reducing the number of instances leads to better performance, indeed, that is what the use of the hierarchical classification is looking for. We demonstrate that this method performs better than using just one method like ORB, SIFT or FREAK, despite being fairly slower. KEYWORDS Computer vision, Descriptors, Feature-based object recognition, Expert system 1. INTRODUCTION Object recognition is an important branch of computer vision. Its main idea is to extract important data or features from images in order to recognize which object is present on it. Many different techniques are used in order to achieve this. In recent computer vision literature, it has been a widely spread tendency to use deep learning due to their benefits throwing out many techniques of previous literature that, actually, have a good performance in many cases. Our aim is to recover those techniques in order to boost them and increase their performance or use their benefits that neural networks may not have. The classical methods in computer vision are based in pure mathematical operations were images are used as matrices. These methods look for gradient changes, patterns... and try to find similarities in different images or build a machine learning model to try to predict the objects that are present in the image. Our use case is the industrial area were many similar parts are to be recognized. Those parts vary a lot from one to another (textures, size, color, reflections,...) so an expert is needed for choosing which method is better for recognizing the objects. We propose a method that simulates the expert role. This is achieved learning a model that classifies the objects in groups that behave similarly to different recognition methods. This leads to a hierarchical classification that first classifies the object to be recognized in one of the previously obtained groups and inside the group the method that works better in that group is used to recognize the object. The paper is organized as follows. In Section 2 we present a state of art of the most used 2D feature-based methods, including detectors, descriptors and matchers. The purpose of Section 3 is to present the method that we propose and how we evaluate it. The experiments done and their Natarajan Meghanathan et al. (Eds) : NLP, ARIA, JSE, DMS, ITCS - 2019 pp. 51-61, 2019. © CS & IT-CSCP 2019 DOI: 10.5121/csit.2019.91206 5. 2d Image Features Detector and Descriptor Selection Expert System 88
results are shown in section 4. Section 5 summarizes the conclusions that can be drawn from our work. 2. BACKGROUND There are several methods for object recognition. In our case, we have focused on feature-based methods. These methods look for points of interest of the images (detectors), try to describe them (descriptors) and match them (matchers). The combination of different detectors, descriptors and matchers vary the perfomance of the whole system. This is a fast growing area in image processing field. The following short and chronologically ordered review presents the gradual improvements in feature detection (Subsection 2.1), description (Subsection 2.2) and matching (Subsection 2.3). 2.1. 2D features detectors One of the most used methods was proposed in 1999 by Lowe [13]. This method is called SIFT, which stands for Scale Invariant Feature Transform. The main idea is to use the Difference-ofGaussian function (a close approximation to the Laplacian-of-Gaussian proposed by Lowe) to search for extrema in the scale space. Even if SIFT was relatively fast, a new method, SURF (Speeded Up Robust Features) [3], outperforms it in terms of repeatability, distinctiveness and robustness, although it can be computed and compared much faster. In addition, FAST (Features from Accelerated Segment Test) [24] proposed by Rosten and Drummond introduce a fast detector. FAST outperforms previous algorithms (like SURF and SIFT) in both computational performance and repeatability. AGAST [16] is based on the FAST, but it is more efficient as well as generic. BRISK [11] is a novel method for keypoint detection, description and matching which has a low computational cost (as stated in the corresponding article, an order of magnitude faster than SURF in some cases). Following the same line of FAST based mehods, we find ORB Rublee et al. [25], an efficient alternative to SIFT or SURF. This method’s detector is based on FAST but it adds orientation in order to obtain better results. In fact, this method performs at two orders of magnitude faster than SIFT, in many situations. 2.2. 2D features descriptors Lowe also proposed a descriptor called SIFT. As mentioned above, is one of the most popular feature detector and descriptor. The descriptor is a position-dependent histogram of local image gradient directions around the interest point and is also scale invariant. It has numerous extensions such as PCA-SIFT [9], that mixes PCA with SIFT; CSIFT [1], Color invariant SIFT; GLOH [17]; DAISY [26], a dense descriptor inspired in SIFT and GLOH; and so on. SURF descriptor [3] relies on integral images for image convolutions in order to obtain its speed. BRIEF [4] is a highly discriminative feature descriptor that is fast both to build and to match. BRISK [11] descriptor is composed as a binary string by concatenating the results of simple brightness comparison tests. ORB descriptor is BRIEF-based and adds rotation invariance and resistance to noise. LBP (Local Binary Patterns) [21] is a two-level version of the texture spectrum method [27]. This methods has been really popular and many derivatives has been proposed. Based on this, the CSLBP (Center-Symmetric Local Binary Pattern) [7] combines the strengths of SIFT and LBP. Later in 2010, the LTP (Local Ternary Pattern) [12] appeared, a generalization of the LBP that is more discriminant and less sensitive to noise in uniform regions. Same year, ELTP (Extended local ternary pattern) [20] improved this by attempting to strike a balance by using a clustering method to group the patterns in a meaningful way. In 2012, LTrP (Local Tetra Patterns) [19] encoded the relationship between the referenced pixel and its neighbors, based on the directions that are Computer Science & Information Technology (CS & IT) 52 89
calculated using the first-order derivatives in vertical and horizontal directions. In [22] there are gathered other methods that are based on the LBP. Other descriptor called FREAK [2] is a keypoint descriptor inspired by the human visual system and more precisely the retina. It is faster, usess less memory and more robust than SIFT, SURF and BRISK. They are thus competitive alternatives to existing descriptors in particular for embedded applications. 2.3. Matchers The most widely used method for matching is Nearest Neighbor (NN). Many algorithms follow this method. One of the most used is the kd-tree [23] which works well with low dimensionality. For dealing with higher dimensionalities many researchers have proposed diverse methods such as the Approximate Nearest Neighbor (ANN) by Indyk and Motwani [8] or the Fast Approximate Nearest Neighbors of Muja and Lowe [18] which is implemented in the well known open source library FLANN (Fast Library for Approximate Nearest Neighbors). 3. PROPOSED APPROACH As we have stated before, the issue we are dealing with is the recognition of industrial parts for pick-and-placing. The main problem is that the accurate recognition of some kind of parts are highly dependant on the recognition pipeline used. This is because parts’ characteristics like texture (presence or absence), forms, colors, brightness; make some detectors or descriptors work differently. We are thus proposing a systematic approach for selecting the best recognition pipeline for a given object (Subsection 3.2). We also propose in Subsection 3.3 an expert system that identifies groups of parts that are recognized similarly to improve the overall accuracy. The recognition pipeline is explained in Subsection 3.1. We start defining some notations. An industrial part, or object, is named instance. The images captured of each part are named views. Given the set of views X, the set of instance labels Y and the set of recognition pipelines Ψ, the function ωΨ X,Y (y)returns for each y∈Ythe best pipeline ψ∗ ∈ Ψaccording to a metric F1that is later discussed. We call ψ∗∗ to the pipeline that on average performs better according to the evaluation metric, this is, that maximizes the average of the scores per instance (2). ωΨ X,Y (y) = argmax ψ∈Ψ F1ψ y(X, Y ) = ψ∗(1) ψ∗∗ = argmax ψ∈Ψ X y∈Y F1ψ y(X, Y ) |Y|(2) 3.1. Recognition Pipeline A recognition pipeline Ψis composed of 3 steps: detection, description and matching. Detectors, Γ, localize interesting keypoints in the view (gradient changes, changes in illumination,...). Descriptors, Φ, are used to represent those keypoints in order to locate them in other views. Matchers, Ω, find the closest features between views. So, a pipeline ψis composed by a keypoint detector γ, a feature descriptor φand a matcher ω. Figure 1 shows the structure of the recognition pipeline. The keypoints detection and description are described previously in the background section. In the matching, are two groups of features: the ones that form the model (train) and the ones that Computer Science & Information Technology (CS & IT) 53 5. 2d Image Features Detector and Descriptor Selection Expert System 90
FLANN Brute force L2 Brute force Hamming etc Harris SIFT ORB etc Images (views) Keypoints detector Keypoints Features descriptor Features Harris SIFT ORB etc Harris SIFT ORB etc Harris SIFT ORB etc SIFT ORB FREAK etc SIFT ORB FREAK etc SIFT ORB FREAK etc SIFT ORB FREAK etc Test Train Part 0 Part 1 Part 2 Part ¿? Matching Result Part y Figure 1: Recognition pipeline need to be recognized (test). Different kind of methods could be used to match features, but, mainly, distance based techniques are used. This techniques make use of different distances (L2, hamming,...) to find the closest feature to the one that needs to be labeled. Those two features (the test feature and the closest to this one) are considered a match. In order to discard ambiguous features, we use the Lowe’s ratio test [14] to define whether two features are a ”good match”. Assuming ftis the feature to be recognized, and fl1and fl2its two closest features from the model, then (ft,fl1) is a good match if: d(ft, fl1) d(ft, fl2)< r (3) where d(fA, fB)is the distance (L2, Hamming,...) between features A and B, and ris a threshold that is used to validate if two features are similarly close to the test feature and discard it. This threshold is set at 0.8. Now a simple voting system is used for labeling the view. For each view from the model (train) the number of good matches are counted. The good matches of each instances are summed and the test view is labeled as the instance with more good matches. 3.2. Recognition Evaluation As we have said, we have the input views X, the instance labels Yand the pipelines Ψ. To evaluate the pipelines we have to separate the views in train and test. The evaluation method used for it is Leave-One-Out Cross-Validation (LOOCV) [10]. It consists of |X|iterations, that for each iteration i, the train dataset is (X−xi)and the test sample is xi. With this separation train-test we can generate the confusion matrix. Table 1 is an example of a confusion matrix for 3 instances. As mentioned in the introduction of Section 3, we use the metric F1value [6] for scoring the performance of the system. The score is calculated for the tests views from the LOOCV. F1 score, or value, is calculated per each instance (4). This metric is an harmonic mean between the Computer Science & Information Technology (CS & IT) 54 91
variant texture descriptor for classifying pain states. Expert Systems with Applications, 37 (12):7888–7894, 2010. [21] T. Ojala, M. Pietikinen, and D. Harwood. A comparative study of texture measures with classification based on featured distributions. Pattern Recognition, 29(1):51–59, January 1996. [22] M. Pietikinen, A. Hadid, G. Zhao, and T. Ahonen. Local Binary Patterns for Still Images. In Computer Vision Using Local Binary Patterns, Computational Imaging and Vision, pages 13–47. Springer London, 2011. [23] John T. Robinson. The k-d-b-tree: A search structure for large multidimensional dynamic indexes. In Proceedings of the 1981 ACM SIGMOD International Conference on Management of Data, SIGMOD ’81, pages 10–18, 1981. [24] E. Rosten and T. Drummond. Fusing points and lines for high performance tracking. In Tenth IEEE International Conference on Computer Vision (ICCV’05) Volume 1, pages 1508–1515 Vol. 2. IEEE, 2005. [25] E. Rublee, V. Rabaud, K. Konolige, and G. Bradski. ORB: An efficient alternative to SIFT or SURF. In 2011 International Conference on Computer Vision, pages 2564–2571, November 2011. [26] E. Tola, V. Lepetit, and P. Fua. DAISY: An Efficient Dense Descriptor Applied to WideBaseline Stereo. IEEE Transactions on Pattern Analysis and Machine Intelligence, 32(5): 815–830, May 2010. [27] L. Wang and D.C. He. Texture classification using texture spectrum. Pattern Recognition, 23(8):905–910, 1990. Computer Science & Information Technology (CS & IT) 61 5. 2d Image Features Detector and Descriptor Selection Expert System 98
CAPÍTULO 6 2D Features-based detector and descriptor selection system for hierarchical recognition of industrial parts Título: 2D Features-based detector and descriptor selection system for hierarchical recognition of industrial parts Autores: I. Merino, J. Azpiazu, A. Remazeilles, B. Sierra Revista: International Journal of Artificial Intelligence & Applications (IJAIA) Volumen: 10 Número: 6 Editor: AIRCC Publishing Corporation DOI: 10.5121/ijaia.2019.10601 Año: 2019 Cuartil (Scimago/WoS): -/- 99
2D FEATURES-BASED DETECTOR AND DESCRIPTOR SELECTION SYSTEM FOR HIERARCHICAL RECOGNITION OF INDUSTRIAL PARTS Ibon Merino1, Jon Azpiazu1, Anthony Remazeilles1, and Basilio Sierra2 1Industry and Transport, Tecnalia Research and Innovation, Donostia-San Sebastian, Spain {ibon.merino, jon.azpiazu, anthony.remazeilles}@tecnalia.com 2Computer Science and Artificial Intelligence, University of the Basque Country UPV/EHU, Donostia-San Sebastian, Spain b[email protected] ABSTRACT Detection and description of keypoints from an image is a well-studied problem in Computer Vision. Some methods like SIFT, SURF or ORB are computationally really efficient. This paper proposes a solution for a particular case study on object recognition of industrial parts based on hierarchical classification. Reducing the number of instances leads to better performance, indeed, that is what the use of the hierarchical classification is looking for. We demonstrate that this method performs better than using just one method like ORB, SIFT or FREAK, despite being fairly slower. KEYWORDS Computer vision, Descriptors, Feature-based object recognition, Expert system 1. INTRODUCTION Object recognition is an important branch of computer vision. Its main idea is to extract important data or features from images in order to recognize which object is present on it. Many different techniques are used in order to achieve this. In recent computer vision literature, it has been a widely spread tendency to use deep learning due to their benefits throwing out many techniques of previous literature that, actually, have a good performance in many cases. Our aim is to recover those techniques in order to boost them and increase their performance or use their benefits that neural networks may not have. The classical methods in computer vision are based in pure mathematical operations were images are used as matrices. These methods look for gradient changes, patterns... and try to find similarities in different images or build a machine learning model to try to predict the objects that are present in the image. Our use case is the industrial area were many similar parts are to be recognized. Those parts vary a lot from one to another (textures, size, color, reflections,...) so an expert is needed for choosing which method is better for recognizing the objects. We propose a method that simulates the expert role. This is achieved learning a model that classifies the objects in groups that behave similarly to different recognition methods. This leads to a hierarchical classification that first classifies the object to be recognized in one of the previously obtained groups and inside the group the method that works better in that group is used to recognize the object. International Journal of Artificial Intelligence & Applications (IJAIA) Vol.10, No.6, November 2019 DOI: 10.5121/ijaia.2019.10601 1 6. 2D Features-based detector and descriptor selection system for hierarchical recognition of industrial parts 100
The paper is organized as follows. In Section 2 we present a state of art of the most used 2D feature-based methods, including detectors, descriptors and matchers. The purpose of Section 3 is to present the method that we propose and how we evaluate it. The experiments done and their results are shown in section 4. Section 5 summarizes the conclusions that can be drawn from our work. 2. BACKGROUND There are several methods for object recognition. In our case, we have focused on feature-based methods. These methods look for points of interest of the images (detectors), try to describe them (descriptors) and match them (matchers). The combination of different detectors, descriptors and matchers vary the perfomance of the whole system. This is a fast growing area in image processing field. The following short and chronologically ordered review presents the gradual improvements in feature detection (Subsection 2.1), description (Subsection 2.2) and matching (Subsection 2.3). 2.1. 2D features detectors One of the most used methods was proposed in 1999 by Lowe [14]. This method is called SIFT, which stands for Scale Invariant Feature Transform. The main idea is to use the Difference-ofGaussian function (a close approximation to the Laplacian-of-Gaussian proposed by Lowe) to search for extrema in the scale space. Even if SIFT was relatively fast, a new method, SURF (Speeded Up Robust Features) [3], outperforms it in terms of repeatability, distinctiveness and robustness, although it can be computed and compared much faster. In addition, FAST (Features from Accelerated Segment Test) [25] proposed by Rosten and Drummond introduce a fast detector. FAST outperforms previous algorithms (like SURF and SIFT) in both computational performance and repeatability. AGAST [17] is based on the FAST, but it is more efficient as well as generic. BRISK [12] is a novel method for keypoint detection, description and matching which has a low computational cost (as stated in the corresponding article, an order of magnitude faster than SURF in some cases). Following the same line of FAST based mehods, we find ORB Rublee et al. [26], an efficient alternative to SIFT or SURF. This method’s detector is based on FAST but it adds orientation in order to obtain better results. In fact, this method performs at two orders of magnitude faster than SIFT, in many situations. Figure 1 shows some detectors and the relation between them chronologically ordered. 2.2. 2D features descriptors Lowe also proposed a descriptor called SIFT. As mentioned above, is one of the most popular feature detector and descriptor. The descriptor is a position-dependent histogram of local image gradient directions around the interest point and is also scale invariant. It has numerous extensions such as PCA-SIFT [10], that mixes PCA with SIFT; CSIFT [1], Color invariant SIFT; GLOH [18]; DAISY [27], a dense descriptor inspired in SIFT and GLOH; and so on. SURF descriptor [3] relies on integral images for image convolutions in order to obtain its speed. BRIEF [4] is a highly discriminative feature descriptor that is fast both to build and to match. BRISK [12] descriptor is composed as a binary string by concatenating the results of simple brightness comparison tests. ORB descriptor is BRIEF-based and adds rotation invariance and resistance to noise. LBP (Local Binary Patterns) [22] is a two-level version of the texture spectrum method [28]. This methods has been really popular and many derivatives has been proposed. Based on this, the CSLBP (Center-Symmetric Local Binary Pattern) [8] combines the strengths of SIFT and LBP. Later International Journal of Artificial Intelligence & Applications (IJAIA) Vol.10, No.6, November 2019 2 101
Moravec (1976) Harris (1988) Shi-Tomasi (1994) SIFT (1999) SURF (2006) Hessian (1998) Hessian-affine (2009) ORB (2011) FAST (2005) MSER (2004) SUSAN (1997) BRISK (2011) Gabor-wavelet (1996) Steerable filters (1991) Canny (1986) (2001) Harris-laplaceHessian-laplace Harris-affine (2002) AGAST (2010) CenSurE (2008) Figure 1: Recognition pipeline in 2010, the LTP (Local Ternary Pattern) [13] appeared, a generalization of the LBP that is more discriminant and less sensitive to noise in uniform regions. Same year, ELTP (Extended local ternary pattern) [21] improved this by attempting to strike a balance by using a clustering method to group the patterns in a meaningful way. In 2012, LTrP (Local Tetra Patterns) [20] encoded the relationship between the referenced pixel and its neighbors, based on the directions that are calculated using the first-order derivatives in vertical and horizontal directions. In [23] there are gathered other methods that are based on the LBP. MTS (Modified texture spectrum) proposed by Xu et al. [29] can be considered as a simplified version of LBP, where only a subset of the peripheral pixels (up-left, up, up-right and right) is considered. The Binary Gradient Contours (BGC) [6] is a binary 8-tuple proposed by Fernndez et al. The simple loop form (BGC1) makes a closed path around the central pixel computing a set of eight binary gradients between pairs of pixels. Other descriptor called FREAK [2] is a keypoint descriptor inspired by the human visual system and more precisely the retina. It is faster, usess less memory and more robust than SIFT, SURF and BRISK. They are thus competitive alternatives to existing descriptors in particular for embedded applications. Figure 2 shows some detectors and the relation between them chronologically ordered. 2.3. Matchers The most widely used method for matching is Nearest Neighbor (NN). Many algorithms follow this method. One of the most used is the kd-tree [24] which works well with low dimensionality. For dealing with higher dimensionalities many researchers have proposed diverse methods such as the Approximate Nearest Neighbor (ANN) by Indyk and Motwani [9] or the Fast Approximate Nearest Neighbors of Muja and Lowe [19] which is implemented in the well known open source library FLANN (Fast Library for Approximate Nearest Neighbors). International Journal of Artificial Intelligence & Applications (IJAIA) Vol.10, No.6, November 2019 3 6. 2D Features-based detector and descriptor selection system for hierarchical recognition of industrial parts 102
(2011) LBP (1995) SIFT (1999) Shape context (2002) RIFT (2005) MOPS (2005) HOG (2005) SURF (2006) LSS (2007) CS-LBP (2009) DAISY (2009) ELTP (2010) BRIEF (2010) WLD (2010) ORB (2011) BRISK (2011) LIOP (2011) Derivated from LBP (2011) MROGHMRRID LSS, C (2012) FLSS, C (2012) LTrP (2012) FREAK (2012) HSOG (2014) GLOH (2005) PCA-SIFT (2004) LTP (2007) Figure 2: Recognition pipeline 3. PROPOSED APPROACH As we have stated before, the issue we are dealing with is the recognition of industrial parts for pick-and-placing. The main problem is that the accurate recognition of some kind of parts are highly dependant on the recognition pipeline used. This is because parts’ characteristics like texture (presence or absence), forms, colors, brightness; make some detectors or descriptors work differently. We are thus proposing a systematic approach for selecting the best recognition pipeline for a given object (Subsection 3.2). We also propose in Subsection 3.3 an expert system that identifies groups of parts that are recognized similarly to improve the overall accuracy. The recognition pipeline is explained in Subsection 3.1. We start defining some notations. An industrial part, or object, is named instance. The images captured of each part are named views. Given the set of views X, the set of instance labels Y and the set of recognition pipelines Ψ, the function ωΨ X,Y (y)returns for each y∈Ythe best pipeline ψ∗ ∈ Ψaccording to a metric F1that is later discussed. We call ψ∗∗ to the pipeline that on average performs better according to the evaluation metric, this is, that maximizes the average of the scores per instance (2). ωΨ X,Y (y) = argmax ψ∈Ψ F1ψ y(X, Y ) = ψ∗(1) ψ∗∗ = argmax ψ∈Ψ X y∈Y F1ψ y(X, Y ) |Y|(2) 3.1. Recognition Pipeline A recognition pipeline Ψis composed of 3 steps: detection, description and matching. Detectors, Γ, localize interesting keypoints in the view (gradient changes, changes in illumination,...). Descriptors, Φ, are used to represent those keypoints in order to locate them in other views. Matchers, Ω, find the closest features between views. So, a pipeline ψis composed by a keypoint detector γ, a feature descriptor φand a matcher ω. Figure 3 shows the structure of the recognition pipeline. The keypoints detection and description are described previously in the background section. In the matching, are two groups of features: the ones that form the model (train) and the ones that International Journal of Artificial Intelligence & Applications (IJAIA) Vol.10, No.6, November 2019 4 103
FLANN Brute force L2 Brute force Hamming etc Harris SIFT ORB etc Images (views) Keypoints detector Keypoints Features descriptor Features Harris SIFT ORB etc Harris SIFT ORB etc Harris SIFT ORB etc SIFT ORB FREAK etc SIFT ORB FREAK etc SIFT ORB FREAK etc SIFT ORB FREAK etc Test Train Part 0 Part 1 Part 2 Part ¿? Matching Result Part y Figure 3: Recognition pipeline need to be recognized (test). Different kind of methods could be used to match features, but, mainly, distance based techniques are used. This techniques make use of different distances (L2, hamming,...) to find the closest feature to the one that needs to be labeled. Those two features (the test feature and the closest to this one) are considered a match. In order to discard ambiguous features, we use the Lowe’s ratio test [15] to define whether two features are a ”good match”. Assuming ftis the feature to be recognized, and fl1and fl2its two closest features from the model, then (ft,fl1) is a good match if: d(ft, fl1) d(ft, fl2)< r (3) where d(fA, fB)is the distance (Euclidean or L2 distance: Equation 4, Hamming distance: Equation 5, where δis the kronecker delta (Equation 6),...) between features A and B, and ris a threshold that is used to validate if two features are similarly close to the test feature and discard it. This threshold is set at 0.8. Now a simple voting system is used for labeling the view. For each view from the model (train) the number of good matches are counted. The good matches of each instances are summed and the test view is labeled as the instance with more good matches. dE(P, Q) = v u u t n X i=1 (pi−qi)2(4) dH(P, Q) = n X i=1 δ(pi, qi)(5) δ(pi, qi) = 0if xi=yi 1if xi6=yi (6) International Journal of Artificial Intelligence & Applications (IJAIA) Vol.10, No.6, November 2019 5 6. 2D Features-based detector and descriptor selection system for hierarchical recognition of industrial parts 104
Table 1: Example of a confusion matrix for 3 instances. Actual instance object 1 object 2 object 3 Predicted instance object 1 40 10 0 50 object 2 0 30 25 55 object 3 10 10 25 45 50 50 50 150 3.2. Recognition Evaluation As we have said, we have the input views X, the instance labels Yand the pipelines Ψ. To evaluate the pipelines we have to separate the views in train and test. The evaluation method used for it is Leave-One-Out Cross-Validation (LOOCV) [11]. It consists of |X|iterations, that for each iteration i, the train dataset is (X−xi)and the test sample is xi. With this separation train-test we can generate the confusion matrix. Table 1 is an example of a confusion matrix for 3 instances. As mentioned in the introduction of Section 3, we use the metric F1value [7] for scoring the performance of the system. The score is calculated for the tests views from the LOOCV. F1 score, or value, is calculated per each instance (7). This metric is an harmonic mean between the precision and the recall. The mean of all the F1’s, ¯ F1(8) is used for calculating ψ∗∗. F1(y) = 2 ·precisiony∗recally precisiony+recally (7) ¯ F1=Py∈YF1(y) |Y|(8) The precision (Equation 9) is the ratio between the correctly predicted views with label y(tpy) and all predicted views for that given instance (|ψ(X) = y|). The recall (Equation 10), instead, is the relation between correctly predicted views with label y(tpy) and all views that should have that label (|label(X) = y|). precisiony=tpy |ψ(X) = y|(9) recally=tpy |label(X) = y|(10) 3.3. Expert system The function ωgives a lot of information about objects but it needs the instance to return the best pipeline for that instance which is not available a priori. Indeed, this is what we want to identify. We use the information that would provide ωto build a hierarchical classification based in a clustering of similar objects. Since some parts work better with some particular pipelines because of their shape, color or texture, we try to take advantage of this and make clusters of objects that are classified similarly well by each pipeline. For example, two parts that have textures may be better recognized by pipelines that use descriptors like SIFT or SURF rather than non textured parts. We call these clusters typologies. This clustering is made using the algorithm K-means [16], that aims to partition the International Journal of Artificial Intelligence & Applications (IJAIA) Vol.10, No.6, November 2019 6 105
test view ψ**+ k-NN Typology 2 (t=2) ψ*t=2 Instance 6 (y=6) Instance 4 (y=4) Instance 5 (y=5) Typology 3 (t=3) ψ*t=3 Instance 8 (y=8) Instance 7 (y=7) Typology 1 (t=1) ψ*t=1 Instance 1 (y=1) Instance 2 (y=2) Instance 3 (y=3) input Figure 4: Hierarchical classification objects into K clusters (where K < |Y|) in which each object belongs to the cluster with the nearest centroids. The input is a matrix with the instances as rows and for each row the F1value of each pipeline. The inputs for this algorithm are for each instance an array of the F1value obtained with every pipeline. The election of a good K may highly vary the result since if almost all the clusters are composed by 1 instance the result would be close to just using ψ∗∗. After obtaining the Ktypologies, the ψ∗ T’s (11) are calculated, i.e., the best pipeline for each typology. ψ∗ T= argmax ψ∈Ψ X y∈T F1ψ y(X, Y ) |T|(11) The first step of the hierarchical recognition is to recognize the typology with the ψ∗∗. Given the typology tas the typology predicted, the ψ∗ tis used to recognize the instance yof the object. We call the hierarchical recognition Υ. The Figure 4 shows an scheme of the hierarchical recognition for clarification. 4. EXPERIMENTS AND RESULTS Our initial hypothesis is that Υhas a better performance than ψ∗∗. In order to demonstrate this hypothesis we conducted some experiments. Moreover, we want to know in which way does the number of parts and the number of views per part affect the result. The pipelines used (detector, descriptor and matcher) are defined in Subsection 4.1. In Subsection 4.2, we explain the dataset we have created to evaluate the proposed method under the use case that is the industrial area and the results obtained. In order to compare these results with a wellknown dataset in Subsection 4.3 we present the Caltech dataset [5] and the results obtained. International Journal of Artificial Intelligence & Applications (IJAIA) Vol.10, No.6, November 2019 7 6. 2D Features-based detector and descriptor selection system for hierarchical recognition of industrial parts 106
Table 2: Pipelines composition. Pipeline Detector Descriptor Matcher ψ0SIFT SIFT FLANN ψ1SURF SURF FLANN ψ2ORB ORB Brute force Hamming ψ3—- LBP FLANN ψ4SURF BRIEF Brute force Hamming ψ5BRISK BRISK Brute force Hamming ψ6AGAST DAISY FLANN ψ7AGAST FREAK Brute force Hamming Table 3: F1’s of the ψ∗∗’s and Υfor each subset of our dataset. pstands for number of parts and tfor number of pictures per part. @@ @t10 20 30 40 50 p@@ @ψ∗∗ Υψ∗∗ Υψ∗∗ Υψ∗∗ Υψ∗∗ Υ 30.935 0.862 0.967 0.983 0.989 10.992 10.993 1 40.899 0.854 0.924 0.962 0.932 0.966 0.944 0.801 0.91 0.865 50.859 0.843 0.868 0.863 0.883 0.818 0.876 0.901 0.87 0.912 6 0.865 0.967 0.873 0.992 0.891 0.87 0.88 0.88 0.856 0.901 7 0.872 0.9 0.886 0.986 0.894 0.891 0.88 0.876 0.845 0.94 4.1. Pipelines The pipelines we have selected are shown in Table 2. Many combination could be done but it is not consistent to match binary descriptors with a L2 distance. The combinations chosen are compatible and may not be the best combination. Global descriptors, such as LBP, does not need a detector. 4.2. Our dataset We select 7 random industrial parts and on a white background we make 50 pictures per part from different angles randomly. That way, we have a dataset with 350 pictures. In Figure 5 are shown zoomed in examples of the pictures taken to the parts. We use subsets of the dataset to evaluate if changing the number of views per instance and the number of instance vary the performance. This subsets have from 3 to 7 parts and from 10 to 50 views (10 views step). In Table 3 are gathered the results for all the subsets using ψ∗∗ and Υ. The highest score for each subset is in bold. On average the hierarchical recognition performs better. The more parts or views per part, the better that performs the hierarchical recognition comparing with the best pipeline. Now we focus on the whole dataset. In Figure 6 are shown the F1’s of each instance using each pipeline for this particular case. The horizontal lines mark the ¯ F1for that pipeline. The score we obtain with our method (last column) is higher (0.94) than the best pipeline which is ψ2that corresponds to the pipeline that uses ORB (0.845). International Journal of Artificial Intelligence & Applications (IJAIA) Vol.10, No.6, November 2019 8 107
applied sciences Article Histogram-Based Descriptor Subset Selection for Visual Recognition of Industrial Parts Ibon Merino 1,2,* , Jon Azpiazu 1, Anthony Remazeilles 1and Basilio Sierra 2 1TECNALIA, Basque Research and Technology Alliance (BRTA), Paseo Mikeletegi 7, 20009 Donostia-San Sebastian, Spain; [email protected] (J.A.); anthony[email protected] (A.R.) 2Department of Computer Science and Artificial Intelligence, University of the Basque Country UPV/EHU, 20018 Donostia-San Sebastian, Spain; [email protected] *Correspondence: [email protected] Received: 1 April 2020; Accepted: 25 May 2020; Published: 27 May 2020 Abstract: This article deals with the 2D image-based recognition of industrial parts. Methods based on histograms are well known and widely used, but it is hard to find the best combination of histograms, most distinctive for instance, for each situation and without a high user expertise. We proposed a descriptor subset selection technique that automatically selects the most appropriate descriptor combination, and that outperforms approach involving single descriptors. We have considered both backward and forward mechanisms. Furthermore, to recognize the industrial parts a supervised classification is used with the global descriptors as predictors. Several class approaches are compared. Given our application, the best results are obtained with the Support Vector Machine with a combination of descriptors increasing the F1 by 0.031 with respect to the best descriptor alone. Keywords: computer vision; feature descriptor; histogram; feature subset selection; industrial objects 1. Introduction Computer vision, in the last years, has gained much interest in many fields, such as autonomous driving [ 1 ], medical [ 2 ], face recognition [ 3 ], object detection [ 4 ], and object segmentation [ 5 ]. Perception is also regarded as one of the key enabling technologies for extending the robot capabilities, preferentially targeting flexibility, adaptation, and robustness, as required for fulfilling the industry 4.0 paradigm [ 6 ]. Although in most fields large and complex datasets can be obtained, detection of industrial parts has a lack of datasets. One of the reasons is that most of the time in industrial context, the aim is to detect an object from which usually the CAD is available. However, sometimes there is a need of detecting diverse, complex, and tiny objects [ 7 ] and lack of time to generate a robust dataset (taking pictures and labeling). One of the solutions is to generate simulated data to train the models but usually there is a significant gap transferring that learned knowledge to reality. To make matter worse, industrial parts are usually texture-less. This means that many of the most used recognition methods cannot deal with them. One of the methods to deal with texture-less objects are Convolutional Neural Networks. Nowadays, computer vision researches are mainly focused on using Convolutional Neural Networks (CNN) [ 8 – 10 ]. One of the disadvantages of the CNNs is the need of a large dataset to train them. Even if it is possible to use the CNN trained on other fields in industry [ 11 ], there is still a need of a large enough training dataset to obtain good results. Feature descriptors based on classical methods have been very useful and thoroughly spread in the literature previous to CNN. One of the benefits of using this approach is that there is no need of a large training set to obtain good results. Actually, there are many image descriptors and each of them has its advantages and disadvantages. Appl. Sci. 2020,10, 3701; doi:10.3390/app10113701 www.mdpi.com/journal/applsci 115
Appl. Sci. 2020,10, 3701 2 of 17 Our approach is based in the idea that the combination of different descriptors leads to a better performance, taking advantage of the benefits of each descriptor to deal with the two problems mentioned before (lack of a large dataset and texture-less objects). The crux of the matter is to select the descriptors that contribute to achieve a better result and discard those that do not provide any improvement. Our method achieves a classification quality similar to state-of-the-art methods on the experiments done. In Section 2, we present a background of the description methods, classifiers, and features subset selection techniques. In Section 3, we explain the combination of the descriptors and the image classification. The experiments done and their results are gathered in Section 4. Finally, in Section 5, the conclusions are summarized. 2. Background The analysis of images usually relies on the extraction of visual features. Such an approach can be observed in classification [ 12 ], object detection [ 4 ], and segmentation [ 5 ]. In this section, we provide an overview of the main feature descriptors, together with some of the related classification techniques. 2.1. Features Descriptors Local features extractors are characteristic local primitives as points focusing on a close neighborhood. Some examples of those features are SIFT [ 13 ], SURF [ 14 ], and LBP [ 15 ]. Global descriptors, instead, extract information directly from the whole image by computing histograms for example. Local features are good for image recognition as each point is independent from the rest and the features are more discriminant. Global features instead are more used for classification and object detection as they achieve a more global representation. Nevertheless, small changes have a larger impact on global features and a better preprocessing is needed when using them. Extracting global features and their classification is usually faster. As a matter of a fact, combining both local and global features usually performs better [ 16 ]. Many researchers use histograms of local features to obtain benefits of both types. Doing so, we obtain a global representation of the local features. [ 16 ] present a taxonomy called Histogram of Equivalent Patterns (HEP) that gathers those histograms of local features. In order for a feature to be part of this framework, it needs to have a delimited quantification, that is, the number of possible values of the extracted feature must be small enough to obtain a relevant histogram. For example, LBP [ 15 ] is part of this framework as the possible values are 256 so the resulting histogram is of length 256, while HOG or SIFT are not part of the HEP framework as the number of possible values is high and the resulting histogram is not relevant. In [ 17 ], a combination of descriptors was also used, but limited to local descriptors. One of the first HEP methods was introduced in 1973. This method, called Gray Level Co-occurrences Matrices (GLCM) [ 18 ] measures the joint probability of the gray levels of two pixels standing in some predefined relative positions. Since 1973, it has been widely used in many texture analysis applications as a feature extractor in this context. In 1990, [ 19 ] proposed the texture spectrum (TS), which inspired many HEP methods. This texture descriptor is based in decomposing the image into a set of essential small units, called Texture Units (TUs). The occurrence distribution of TU is the TS. One of the first and most used TU-based descriptors is the Local Binary Pattern (LBP) [ 15 ]. This last one is a two-level TU, gray-scale invariant and easily combined with a simple contrast measure. One of the main characteristics is its robust invariant to light changes. Another method based in the TU is the Simplified Texture Unit (STU) [ 20 ]. This method use a more reduced range of values without a significant loss of the characterization power. This way, there are two options of STU: using the crosswide neighbors (up, right, down, and left) and using diagonal neighbors (up-left, up-right, down-right, and down-left); its reduced length is commonly used in real-time applications obtaining similar performance to LBP. 7. Histogram-Based Descriptor Subset Selection for Visual Recognition of Industrial Parts 116
Appl. Sci. 2020,10, 3701 3 of 17 The modified texture spectrum (MTS) [ 21 ] can be considered as a simplified version of LBP, where only a subset of the peripheral pixels (up-left, up, up-right, and right) are considered. Its TS is 16 elements in length, significantly improving the computation efficiency on classification. Similarly to STU, the reduction on the TS length leads to a faster classification while achieving similar performance. The GaborLBP [ 22 ] considers the advantages of the Gabor filters in computer vision and exploits them. It first applies a Gabor transformation and encodes the magnitude values with the LBP operator. Fusing both tools enables handling of illumination changes, viewpoint angle changes, and non-rigid bodies. Usually this combination is used for face recognition or person identification. The Local Ternary Pattern (LTP) [ 23 ] is a generalization of the LBP and it is more discriminant and less sensitive to noise in uniform regions. It is a local texture descriptor that uses a 3-value coding that thresholds around zero. Comparing to the LBP, LTP is more resistant to noise but no longer invariant to gray-level transformations. The Binary Gradient Contours (BGC) [ 24 ] is a binary 8-tuple. It relies on computing a set of eight binary gradients between pairs of pixels all along a closed path around the central pixel of a 3 × 3 grayscale image patch. They defined the closed path in three different ways: single-loop (BGC1), double-loop (BGC2), and triple-loop (BGC3). Another HEP descriptor, is the Local Quantized Patterns (LQP) [ 25 ]. This is a generalization of local pattern features that makes use of vector quantization. It uses large local neighbourhoods and/or deeper quantization with domain-adaptative vector quantization. The Weber’s Law Descriptor (WLD) [ 26 ] was proposed in 2010 as a simple, yet very powerful and robust descriptor. It is based on the fact that human pattern perception also depends on the original intensity of the stimulus and not only on the change of a stimulus (such as sound and lighting). It is composed of two components: differential excitation and orientation. The Histogram of oriented gradients (HOG) [ 27 ] is a feature descriptor that counts the occurrences of gradient orientation in localized portions of an image. Operating on local cells provides invariation to geometric and photometric transformations. The HOG descriptor is particularly suited for human detection in images. Even if HOG is not part of HEP, the way it generates the descriptor (calculating a histogram of gradients) works similar to HEP methods so it can be used similarly. 2.2. Classifiers Descriptors are used to obtain features from images. Those features are then used by the classifiers to predict which object is on each image. Many machine learning algorithms are used for classifying images, but some of the most popular ones are K-Nearest Neighbors, Naive Bayes, Random Forest, Support Vector machine, Random Committee, Bagging, and Multiclass Classifier. The Nearest Neighbor Rule is a well-known algorithm and the simplest nonparametric decision procedure that assigns to the uncategorized object the label of the closest sample of the training set. In 1967, a modification of this algorithm led to one of the most used classification algorithms, the K-Nearest Neighbors (KNN) [ 28 ]. It is based on looking for closest points and classifying them as the majority class. For a given set of n pairs (x1 , θ1) ,..., (xn , θn) , where xi is in a metric space X and θi is the category that xi belongs to from a subset { 1,2,..., M} , a new arriving instance x is analyzed to estimate its corresponding class θ . This estimation is done by looking for the nearest neighbor X0 n∈(x1,x2,..., xn): min d(xi,x) = d(x0 n,x)i=1, 2,..., n where d is a distance metric according to the space X . The new instance x will be assigned to the category θ0 n . This is the basic 1-NN. In general, KNN rule decides x belongs to the category of majority vote of the nearest kneighbors. The Naive Bayes [ 29 ], the simplest Bayesian classifier, is another classification algorithm that is often used for its simplicity. It is based on the Bayesian Rule and assumes that variables are independent 117
Appl. Sci. 2020,10, 3701 4 of 17 given the class. Despite this unrealistic assumption, it is successful in practice. The Bayesian rule states that the probability that a instance xbelongs to class Ckis P(Ck|x) = P(Ck)P(x|Ck) P(x)(1) where Ck is the class between the K possible classes and x the instance to be classified. Taking into account the independence assumption, the conditional distribution over the class variable C is p(Ck|x1,..., xn) = 1 Zp(Ck) n ∏ i=1 p(xi|Ck)(2) The instance is classified as the class with more p(Ck|x1,..., xn). The Random Forest (RF) [ 30 ] is a combination of decision trees that use random subsets of the features to be built. Figure 1shows an example of RF. tree 1 tree 2 tree 3 Majority voting X 1 121 Figure 1. Random Forest example where each tree classifies the new instance and the resulting class is decided by majority voting. Support Vector Machines (SVM) [ 31 ] are supervised learning models that look for optimal hyperplanes that separates classes. An optimal hyperplane is defined as the linear decision function with maximal margin between the vectors of the two classes (Figure 2). Figure 2. Support vector machine: maximum separation between two classes. 7. Histogram-Based Descriptor Subset Selection for Visual Recognition of Industrial Parts 118
Appl. Sci. 2020,10, 3701 5 of 17 Random Committee (RC) [ 32 ] is a committee of random classifiers. The base randomizable classifiers (that form the committee members) are built using different random number seeds based in the same data. The final prediction is a straight average of the predictions generated by the individual base classifiers. The Bagging [ 33 ] technique is called after Bootstrap aggregating. This machine learning ensemble that can be used to improve the stability of a model by improving the accuracy and reducing variance in order to reduce overfitting. 2.3. Feature Selection As stated before, the crux of the matter in this paper relies on how to select the different visual features to improve the individual score of each descriptor. Some authors have used different techniques to do this [ 34 , 35 ]. Feature Selection is a machine learning technique that is used in many fields and usually improves the accuracy of the model. In [ 34 ], the authors uses different feature selection techniques to improve the score in the Quantitative Structure–Activity Relationship (QSAR). In [ 35 ], instead, they use a similar approach for hand pose recognition. In [ 36 ], a view over the different feature selection techniques and its variations is described. Our approach is based in those methods and is used in a completely different context. 3. Proposed Approach In order to achieve a better performance than just using a single global descriptor, we propose using a Descriptor Subset Selector. That is, we try to find the combination of global descriptors that scores a better result. Among all available options of subset selection, we have used 2 for their greedy approach which achieve a significant performance: forward selection and backward selection. First, we present the classification of a single image, given a descriptor and a classifier. After that, we explain the feature selection techniques to choose the combination of descriptor to use. Next, we present the evaluation methods, in order to decide which is the best solution. Finally, we present the whole pipeline of the proposed approach. 3.1. Classification The first step in the pipeline is to classify a picture into the C different classes. Given a descriptor and a classifier, the classifier is trained with features obtained from the description of the set of images for training. Given a new image to be classified, the descriptor extracts the feature from the image and that feature is classified by the classifier (Figure 3). Image ClassifierDescriptor Ci Feature Figure 3. Classification of a new image given a descriptor and a classifier. 3.2. Feature Selection Techniques The feature selection techniques are used to chose the descriptors for the classification. An exhaustive search of best combination of descriptors is computationally inefficient, while it guarantees that the optimal solution is achieved. Nevertheless, a suboptimal solution can be achieved using a sequential search. This is an iterative search that once a stage of the search is reached, is impossible to go back. The complexity of the exhaustive search is exponential ( O( 2 n) ), while the sequential search remains polynomial ( O(nk+1) ), where k is the number of evaluated subsets in each stage. This last one does not guarantee an optimal solution. 119
Appl. Sci. 2020,10, 3701 6 of 17 Another important consideration in the feature selection techniques is the generation of the successors, i.e., how to select the next candidates for the following stage. The simplest and most used methods are Forward and Backward generation [ 36 ]. In forward generation, on each stage the element which makes J (the evaluation measure) greater is selected and added to the selected subset. For example, the first descriptor added to the subset would be the one with the best individual score. The next stage would add to the subset the one that concatenated with the previous one makes the score greater. We refer to this method as Sequential Forward Subset Selection (SFSS) [ 36 ], and its pseudocode is described in Algorithm 1. The backwards is the opposite behavior. The subset is initialized with all the elements and on each stage the element that that makes J greater when removed is done so. The stopping criteria in both cases can be that J is not increased in j steps or the subset achieves a desired length. We refer to this method as Sequential Backward Subset Selection (SBSS) [ 36 ], and its pseudocode is described in Algorithm 2. Algorithm 1: Sequential Forward Subset Selection Input : X—Set of elements J—evaluation metric Output: X0—solution found X0=∅ repeat x0:=argmax{J(X0∪x)|x∈(X\X0)} X0:=X0∪ {x0} until not improvement in J OR X0=X; where ∪ stands for union between two sets or an element and a set and \ operator stands for difference. Algorithm 2: Sequential Backward Subset Selection Input : X—Set of elements J—evaluation metric Output: X0—solution found X0=X repeat x0:=argmax{J(X0\x)|x∈X0} X0:=X0\ {x0} until not improvement in J OR X0=∅; 3.3. Evaluation Measure A classification quality can be quantified using measures such the one of Equation (3). This measure, named F-value [ 37 ] or F-score, is an evaluation measure that takes into account the precision and the recall. More precisely, the metric used is a particular case of the F-value where the precision and the recall are balanced. This is called F1 , an harmonic mean between the precision and the recall. F1(y) = 2·precisiony∗recally precisiony+recally(3) 7. Histogram-Based Descriptor Subset Selection for Visual Recognition of Industrial Parts 120
Appl. Sci. 2020,10, 3701 7 of 17 where y refers to a class (also referred in this paper as Ci ). F1 is class-dependent, so for each class, y , the precision and the recall are computed for that class. The precision (Equation (4)) is the ratio between the correctly predicted views with label y ( tpy or true positive) and all predicted views for that given instance ( |ψ(X) = y| ). The recall (Equation (5)), instead, is the relation between correctly predicted views with label y(tpyor true positive) and all views that should have that label (|label(X) = y|). precisiony=tpy |ψ(X) = y|(4) recally=tpy |label(X) = y|(5) To evaluate each stage of the feature selection we use the averaged F1 . This is the mean of the F1 ’s of all the classes (Equation (6)). F1=1 |Y|∑ y∈Y F1(y)(6) 3.4. Full Pipeline The dataset is divided in two sets: training and test. During the search of the best combination of descriptors, training set is used for training the classifiers and validate the feature selection technique. This separation is made by a Leave-One-Out Cross-Validation (LOOCV) [ 38 ]. Each image of the set is used as validation while the rest of the set is used to train the model. Figure 4shows the whole process. Given a descriptor and a classifier, both are tested using the LOOCV to set the training and validation sets. Once the best combination of descriptors is found, to test the quality of this combination, we use the test set to obtain a general evaluation metric. Figure 4. Full pipeline of the proposed method, including training, validation, and evaluation. 4. Experiments and Results As stated before, the aim of this paper is to present a method to improve the accuracy on reduced datasets of texture-less objects. In order to prove that our method improves the score of the descriptors by their own, we have created a small dataset composed by seven different random industrial parts (Figure 5). We took 50 pictures of each industrial part taken from different viewpoints and different illumination conditions. Objects are rotated and translated but all images are free from occlusion, and with an empty and white background. 121
Appl. Sci. 2020,10, 3701 8 of 17 Figure 5. Pictures of the parts used in the experiment. Our pool of descriptors D for discovering the best combination is made up of BGC1 BGC2, BGC3, LBP, GaborLBP, GLCM, HOG, LQP, LTP, MTS, STU+ (or STU1), STU × (or STU2), and WLD. All descriptors but HOG are computed on grids of different sizes: 1 × 1, 4 × 4, and 8 × 8. The length of gridded histograms is the length of the descriptor multiplied by the number of grids. The HOG is applied to the whole image directly. Figure 6shows a sample image from our database that has been described by each of the descriptors. (a) Sample image (b) BGC1 (c) BGC2 (d) BGC3 (e) GaborLBP (f) GLCM (g) HOG (h) LBP (i) LQP Figure 6. Cont. 7. Histogram-Based Descriptor Subset Selection for Visual Recognition of Industrial Parts 122
Appl. Sci. 2020,10, 3701 9 of 17 (j) LTP (k) MTS (l) STU1 (m) STU2 (n) WLD Figure 6. Histogram of all the used descriptors applied to a sample image. The vertical axis represents the number of occurrences of each texture unit normalized and the horizontal axis represents each of the texture units of the histograms. The descriptors are the ones that are part of D described at the beginning of Section 4. The classifiers used are KNN, NB, SVM 1-vs-1 trained with SMO (Sequential Minimal Optimization [ 39 ]), SVM 1-againt-all trained with SGD (Stochastic Gradient Descent [ 40 ]), RC, RF, and Bagging. To distinguish between the two SVM implementations, we call SVM to the one trained with SMO and SVM-SGD to the other one. In terms of performance, some of the classifiers are drastically affected by the parameters, but tuning the parameters makes a complex casuistry which is not the aim of this paper. Used parameters are standards and those are given in the Appendix A. The results are obtained for a Intel Xeon CPU of 3GHz and 16GB of RAM, and no GPU acceleration has been used. The following subsections explain the results obtained in the experiments. 4.1. Forward Subset Selection Forwards Subset Selection of descriptors applied to the whole image (from now on, FSS1 × 1) experiments results are shown in Table 1. In Table 1, the classifier that is between brackets is the one that achieves the highest mean score. If we would use the best descriptor alone, the F1 would be 0.94 with WLD. By combining it with BGC2 and MTS, and using SVM as classifier, we are able to augment quality of 3% to reach 0.971. On first iteration WLD outperforms the other descriptors with a difference of 0.1 comparing to the next best descriptor. The second iteration increases the overall accuracy and in almost all the cases improves the accuracy of the previous iteration best case. Table 2shows the results of the Forwards Subset Selection of descriptors applied to a 4 × 4 grid (FSS4 × 4). On average, the first iteration performs better than the non-gridded version FSS4 × 4, but the last iteration does not improve the results obtained with FSS4 × 4. The first iteration achieves an F1 of 0.934 and the final iteration 0.969. Therefore, an improvement of 3.5% is obtained. The final combination of descriptors, the one which achieves the highest score, is composed by STU1 and WLD. Table 3shows the results for the 8 × 8 gridded version (FSS8 × 8). The results are similar to the ones obtained in FSS4 × 4. The first iteration achieves an F1 of 0.94, while the last one achieves a score of 0.96. In this case, the improvement is 2%. The performance of the 3 options of the parameters are similar but the speed of the classification is much faster with the FSS1 × 1 version because the length of the final descriptor is shorter. Therefore, the 123
Appl. Sci. 2020,10, 3701 16 of 17 16. Fernández, A.; Álvarez, M.X.; Bianconi, F. Texture Description Through Histograms of Equivalent Patterns. J. Math. Imaging Vis. 2013,45, 76–102. [CrossRef] 17. Merino, I.; Azpiazu, J.; Remazeilles, A.; Sierra, B. 2D Features-based Detector and Descriptor Selection System for Hierarchical Recognition of Industrial Parts. IJAIA 2019,10, 1–13. [CrossRef] 18. Haralick, R.M.; Shanmugam, K.; Dinstein, I. Textural Features for Image Classification. IEEE Trans. Syst. Man Cybern. 1973,SMC-3, 610–621. [CrossRef] 19. Wang, L.; He, D.C. Texture classification using texture spectrum. Pattern Recognit. 1990,23, 905–910. 20. Madrid-Cuevas, F.J.; Medina, R.; Prieto, M.; Fernández, N.L.; Carmona, A. Simplified Texture Unit: A New Descriptor of the Local Texture in Gray-Level Images. In Pattern Recognition and Image Analysis; Springer Berlin Heidelberg: Berlin/Heidelberg, Germany, 2003; Volume 2652, pp. 470–477. 21. Xu, B.; Gong, P.; Seto, E.; Spear, R. Comparison of Gray-Level Reduction and Different Texture Spectrum Encoding Methods for Land-Use Classification Using a Panchromatic Ikonos Image. Photogramm. Eng. Remote Sens. 2003,69, 529–536. 22. Zhang, W.; Shan, S.; Gao, W.; Chen, X.; Zhang, H. Local Gabor binary pattern histogram sequence (LGBPHS): A novel non-statistical model for face representation and recognition. In Proceedings of the Tenth IEEE International Conference on Computer Vision (ICCV’05) Volume 1, Beijing, China, 17–21 October 2005; Volume 1, pp. 786–791. 23. Tan, X.; Triggs, W. Enhanced local texture feature sets for face recognition under difficult lighting conditions. IEEE Trans. Image Process. 2010,19, 1635–1650. 24. Fernández, A.; Álvarez, M.X.; Bianconi, F. Image classification with binary gradient contours. Opt. Lasers Eng. 2011,49, 1177–1184. 25. Hussain, S.U.; Napoléon, T.; Jurie, F. Face Recognition using Local Quantized Patterns. In Procedings of the British Machine Vision Conference 2012; British Machine Vision Association: Guildford, UK, 2012; pp. 99.1–99.11. 26. Chen, J.; Shan, S.; He, C.; Zhao, G.; Pietikainen, M.; Chen, X.; Gao, W. WLD: A Robust Local Image Descriptor. IEEE Trans. Pattern Anal. Mach. Intell. 2010,32, 1705–1720. [CrossRef] 27. Dalal, N.; Triggs, B. Histograms of oriented gradients for human detection. In Proceedings of the 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05), San Diego, CA, USA, 20–25 June 2005; IEEE: Piscataway, NJ, USA, 2005; Volume 1, pp. 886–893. 28. Cover, T.; Hart, P. Nearest neighbor pattern classification. IEEE Trans. Inf. Theory 1967,13, 21–27. 29. Rish, I. An empirical study of the naive Bayes classifier. In IJCAI 2001 Workshop on Empirical Methods in Artificial Intelligence; IBM: New York, NY, USA, 2001; Volume 3, pp. 41–46. 30. Breiman, L. Random forests. Mach. Learn. 2001,45, 5–32. [CrossRef] 31. Cortes, C.; Vapnik, V. Support-vector networks. Mach. Learn. 1995,20, 273–297. [CrossRef] 32. Lira, M.M.S.; de Aquino, R.R.B.; Ferreira, A.A.; Carvalho, M.A.; Neto, O.N.; Santos, G.S.M. Combining Multiple Artificial Neural Networks Using Random Committee to Decide upon Electrical Disturbance Classification. In Proceedings of the 2007 International Joint Conference on Neural Networks, Orlando, FL, USA, 12–17 August 2007; pp. 2863–2868. 33. Breiman, L. Bagging predictors. Mach. Learn. 1996,24, 123–140. [CrossRef] 34. Shahlaei, M. Descriptor Selection Methods in Quantitative Structure–Activity Relationship Studies: A Review Study. Chem. Rev. 2013,113, 8093–8103. [CrossRef] [PubMed] 35. Rasines, I.; Remazeilles, A.; Bengoa, P.M.I. Feature selection for hand pose recognition in human-robot object exchange scenario. In Proceedings of the 2014 IEEE Emerging Technology and Factory Automation (ETFA), Barcelona, Spain, 16–19 September 2014; pp. 1–8. 36. Molina, L.; Belanche, L.; Nebot, A. Feature selection algorithms: A survey and experimental evaluation. In Proceedings of the 2002 IEEE International Conference on Data Mining, Maebashi City, Japan, 9–12 December 2002; pp. 306–313. [CrossRef] 37. Chinchor, N. MUC-4 Evaluation Metrics. In Proceedings of the 4th Conference on Message Understanding; Association for Computational Linguistics: Stroudsburg, PA, USA, 1992; pp. 22–29. [CrossRef] 38. Forman, G.; Scholz, M. Apples-to-Apples in Cross-Validation Studies: Pitfalls in Classifier Performance Measurement. SIGKDD Explor. Newsl. 2010,12, 49–57. [CrossRef] 39. Platt, J.C. Sequential Minimal Optimization: A Fast Algorithm for Training Support Vector Machines; MIT Press: Cambridge, MA, USA, 1998. 7. Histogram-Based Descriptor Subset Selection for Visual Recognition of Industrial Parts 130
Appl. Sci. 2020,10, 3701 17 of 17 40. Robbins, H.; Monro, S. A Stochastic Approximation Method. Ann. Math. Statist. 1951 ,22, 400–407. [CrossRef] 41. Chollet, F. Xception: Deep Learning With Depthwise Separable Convolutions. In Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA, 20–25 June 2009; pp. 248–255. 42. Melekhov, I.; Kannala, J.; Rahtu, E. Siamese network features for image matching. In Proceedings of the 2016 23rd International Conference on Pattern Recognition (ICPR), Cancun, Mexico, 4–8 December 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 378–383. 43. Szegedy, C.; Vanhoucke, V.; Ioffe, S.; Shlens, J.; Wojna, Z. Rethinking the Inception Architecture for Computer Vision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 26 June–1 July 2016; pp. 2818–2826. c 2020 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/). 131
CAPÍTULO 8 3D Convolutional Neural Networks Initialized from Pretrained 2D Convolutional Neural Networks for Classification of Industrial Parts Título: 3D Convolutional Neural Networks Initialized from Pretrained 2D Convolutional Neural Networks for Classification of Industrial Parts Autores: I. Merino, J. Azpiazu, A. Remazeilles, B. Sierra Revista: Sensors Volumen: 21 Número: 1078 Editor: MDPI DOI: 10.3390/s21041078 Año: 2021 Cuartil (Scimago/WoS): Q2 (Electrical and electronic engineering) / Q2 (Engineering, electrical & electronic) 133
sensors Article 3D Convolutional Neural Networks Initialized from Pretrained 2D Convolutional Neural Networks for Classification of Industrial Parts Ibon Merino 1,2,* , Jon Azpiazu 1, Anthony Remazeilles 1and Basilio Sierra 2 Citation: Merino, I.; Azpiazu, J.; Remazeilles, A.; Sierra, B. 3D Convolutional Neural Networks Initialized from Pretrained 2D Convolutional Neural Networks for Classification of Industrial Parts. Sensors 2021,21, 1078. https:// doi.org/10.3390/s21041078 Academic Editor: Sheryl Berlin Brahnam Received: 28 December 2020 Accepted: 2 February 2021 Published: 4 February 2021 Publisher’s Note: MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations. Copyright: © 2021 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/). 1TECNALIA, Basque Research and Technology Alliance (BRTA), Mikeletegi Pasealekua 7, 20009 Donostia-San Sebastián, Spain; [email protected] (J.A.); anthony[email protected] (A.R.) 2Robotics and Autonomous Systems Group, Universidad del País Vasco/Euskal Herriko Unibertsitatea, 48940 Basque, Spain; [email protected] *Correspondence: [email protected] Abstract: Deep learning methods have been successfully applied to image processing, mainly using 2D vision sensors. Recently, the rise of depth cameras and other similar 3D sensors has opened the field for new perception techniques. Nevertheless, 3D convolutional neural networks perform slightly worse than other 3D deep learning methods, and even worse than their 2D version. In this paper, we propose to improve 3D deep learning results by transferring the pretrained weights learned in 2D networks to their corresponding 3D version. Using an industrial object recognition context, we have analyzed different combinations of 3D convolutional networks (VGG16, ResNet, Inception ResNet, and EfficientNet), comparing the recognition accuracy. The highest accuracy is obtained with EfficientNetB0 using extrusion with an accuracy of 0.9217, which gives comparable results to state-of-the art methods. We also observed that the transfer approach enabled to improve the accuracy of the Inception ResNet 3D version up to 18% with respect to the score of the 3D approach alone. Keywords: computer vision; deep learning; transfer learning; object recognition 1. Introduction Industrial processes are continuously changing and now digitalization and smart automation are the main focuses to improve performance and productivity of the industrial plants. Robotics, combined with computer science techniques, such as machine learning, have boosted exponentially the production and security. This industrial revolution has been named Industry 4.0. Many fields have been integrated to this paradigm. One of them is computer vision. Computer vision is a sub-field of machine learning consisting in acquiring, processing, analyzing and understanding images of the real world in order to generate information that a computer can deal with. As in all artificial intelligence fields, deep learning techniques are extensively used nowadays in computer vision. Deep learning approaches usually need a big dataset to obtain significant results. In order to meet this requirement, some techniques use deep learning nets that have been trained with huge datasets and transfer that learned knowledge to smaller or different datasets. Those techniques are called transfer learning [1,2]. In industrial application, and particularly in SMEs, or when small production batches are targeted, making huge datasets can be too expensive and arduous. In addition, objects to be recognized can be small and uncommon. Reducing the number of images needed for training is critical in this context. Using transfer learning methods can help to reduce computing time and the minimum dataset size that is needed to obtain significant results. Sensors 2021,21, 1078. https://doi.org/10.3390/s21041078 https://www.mdpi.com/journal/sensors 135
Sensors 2021,21, 1078 2 of 18 In parallel, in recent years, 3D cameras are gaining more and more popularity, specially in robotic applications. Working with 3D data is a relatively new paradigm that 2D convolutional networks cannot handle so easily. Therefore, new deep learning methods have been designed to deal with this paradigm [3,4], like the 3D convolutional networks. Our proposal is a transfer learning technique that relies on using 2D features learned by 2D convolutional nets to train a 3D convolutional net. The rest of this paper is organized as follows: Section 2outlines the related works, Section 3details the proposed approach, Section 4describes the training phase of the network, Section 5shows the experimental results obtained, and Section 6summarizes the conclusions. 2. Related Works Recent improvements in computing power and the rapid development of more affordable 3D sensors, have opened a new paradigm where 3D data, such as point clouds, are providing better understanding of the environment. Even if some advances have been done in deep learning on point clouds, this is still an underdeveloped field compared to 2D deep learning [3]. Dealing with 3D data in deep learning opens many new fronts. For example, 3D data is difficult to label, so that a significant time is required to label training data. Therefore, usually, the size of the training set for 3D approaches is notably smaller than the one used with 2D techniques. Many different Convolutional Neural Networks (CNN) have been possible and gained a great success due to the large amount of public image repositories, such ImageNet [ 5 , 6 ] and high-performance computing systems, like GPUs. Two-dimensional CNN have been widely studied, and there are many successful methods, but 3D CNN still need more research, as we will show in the next sections. 2.1. 2D CNN The most important deep learning architectures are identified through the ImageNet Large Scale Visual Recognition Competition (ILSVRC) [ 6 ]. One of the first ones winning this competition is Alexnet [ 5 ]. ZFNet [ 7 ] and OverFeat [ 8 ] followed Alexnet, improving the results they obtain for the ImageNet dataset. Understanding of convolutional layers is improved by Reference [ 7 ], thanks to their visualization. The following architectures focused on extracting features on low spatial resolutions. One of them is VGG [ 9 ], which is still being used as a base to many other architectures because of its simple and homogeneous topology. VGG scored the second place in the ILSVRC 2014. The first place was achieved by GoogLeNet [ 10 ], also known as Inception Network. This network was an improvement of the AlexNet, reducing the number of parameters while being much deeper. They introduced the Inception module, which enabled to recognize patterns of different sizes within the same layer, concurrently performing several convolutions of different receptive fields and combining the results. Another influential architecture was introduced by Reference [ 11 ] named the Residual blocks. The architecture called ResNet introduced those Residual blocks which include a skip connection on a convolution block that is merged by summation with the output of that block. This network won the ILSVRC 2015 localization and classification contests and also the COCO detection and segmentation challenges [12]. A modification of the GoogLeNet called Inception-v4 [ 13 ] included an improvement on the inception module and three different kinds of inception modules. In addition, this paper also presents a combination of the inception module with the residual connection, named Inception-ResNet, resulting in a more efficient network. Another network similar to the previous one, the ResNeXts [ 14 ], achieved the second place in the 2016 ILSVRC classification challenge. The first place in classification, localization, and detection challenges was achieved by ResNet101, Inception-v4, and Inception-ResNet-v1, respectively. 8. 3D Convolutional Neural Networks Initialized from Pretrained 2D Convolutional Neural Networks for Classification of Industrial Parts 136
Sensors 2021,21, 1078 3 of 18 Due to the success of the Inception and Residual modules, many subsequent networks have been derived from them. For example, DenseNets [ 15 ] combine the output of the residual connection and the output of the residual block by depth wise filter concatenation. The 2017 ILSVRC localization challenge’s first place and the top 3 in classification and detection categories were won by Dual Path Network (DPN) [ 16 ], a network that combines the architectures of DenseNets and ResNet. Since previous networks focus on achieving the highest possible accuracy, they are not prepared for real-time applications with restricted hardware, like mobile platforms. MobileNets [ 17 ] tackles this problem by replacing standard convolutions with Depthwise Separable Convolutions. Recently, Reference [ 18 ] proposed a novel scaling method that uniformly scales network’s depth, width, and resolution, obtaining a new family of models called EfficientNet. This family achieves much better accuracy with a 6.1 × gain factor in computation time and a 8.4×factor in size reduction compared to previous ConvNets. 2.2. 3D CNN Some researchers have taken advantage of the fact that 2D deep learning is more mature than 3D deep learning, trying to obtain a solution to 3D based on 2D deep learning. Recently, the arrival of RGB-D sensors, such as the Microsoft’s Kinect or the Intel’s Realsense, has enabled to acquire at a low cost 3D information. These sensors provide a 2D color image (RGB), along with a depth map (D), which provides the 3-dimensional information. Since both RGB and D are 2D images, 2D deep learning methods can be adapted to receive as input two images instead of one. Even if this representation is quite simple, they are effective for different tasks, such as human pose regression [ 19 ], 6D pose estimation [20], or object detection [21]. Despite representing 3D data, RGB-D images are composed by 2D data and no transformation is needed. One possible transformation as proposed in Reference [ 22 , 23 ], consists of projecting the 3D data into another 2D space while keeping some of the original 3D shape key properties. To keep 3D data without transforming it to 2D, some works, like Reference [ 24 , 25 ], propose a Voxel-based method. Voxels are used to describe how the 3D object is distributed in the three dimensions of the space. This representation is not always the best option since it stores both the occupied and non-occupied parts of the scene. Voxel-based methods are not recommended for high-resolution data since they store a huge unnecessary amount of data. To deal with this problem, octree-based methods with varying sized voxels [ 26 , 27 ] are proposed. In order to reduce the number of parameters, which is too high in voxel-based methods, some methods propose point-based methods that include point cloud as an unordered set of points as input [28,29]. Our proposal changes this perspective. We adapt the 2D deep learning architecture to 3D and transform the weights from 2D to 3D as initial weights for the newly generated 3D Convolutional model. This approach makes it possible to leverage on existing nets trained on 2D data and apply them on 3D data while maintaining the original data structure. 3. Proposed Approach Due to the great success of 2D CNN in computer vision, our proposal uses those nets as the base to train a 3D CNN for classification. Figure 1shows the overview of the proposed architecture. First, the weights of a pre-trained 2D CNN are transformed to 3D to be, therefore, used as the weights of the 3D CNN. The input point cloud is discretized by computing a voxel grid. That grid is the input tensor to the 3D CNN, which is an adapted form of the 2D CNN using 3D layers instead of 2D layers. That 3D CNN computes the 3D features that are then passed on to the classifier. 137
Sensors 2021,21, 1078 4 of 18 Figure 1. Overall architecture of the proposed method. The following subsections explains the different modules of the architecture. 3.1. 2D to 3D Transformations CNN weights can be represented as 2D matrices. Thus, we need to transform a 2D matrix into a 3D tensor, i.e., map the function M(x , y) = (r , g , b) to T(x , y , z) = (r , g , b) , where x , y∈N , and r , g , b∈R . For each value of x and y of the 2D matrix and each of the new possible z values the transformation function is: h(x , y , z , M(x , y)) = T(x0 , y0 , z0)=(r , g , b) . We have proposed 2 different transformation functions, the extrusion and the rotation. 3.1.1. Extrusion Extrusion of the plane consists of filling the tensor copying the RGB values along one axis. Given a matrix M of size (W×H) and the resulting tensor T of size (W×H×D) and the fact that we use inputs that have all the dimensions of the same length, this is, W=H=D, the Extrusion mapping is defined as: ∀x,y,z≤W:T(x,y,z) = M(x,y). The extrusion can be done along the three main axes: • Z axis: T(x,y,z) = M(x,y), • Y Axis: T(x,z,y) = M(x,y), • X Axis: T(z,x,y) = M(x,y). Figure 2shows how the extrusion along the Z axis is done for a matrix. Figure 2. Extrusion along Z axis of a 2D matrix to generate a 3D tensor. 8. 3D Convolutional Neural Networks Initialized from Pretrained 2D Convolutional Neural Networks for Classification of Industrial Parts 138
Sensors 2021,21, 1078 5 of 18 3.1.2. Rotation To add curvature to the 2D weights, a rotation from 0 to 90 degrees with respect to the Z axis is applied to the 2D weights. The mapping from 2D to 3D is defined as: T(x,y,z) = M(x,min(qy2+z2,W)). Figure 3shows an example of the 2D to 3D rotation transformation. Figure 3. Rotation on Z axis of a 2D matrix to generate a 3D tensor. 3.2. Discretization of the Point Cloud The input of the architecture is a point cloud. In order to use a CNN, we have to discretize/sample the point cloud to a gridded structure (tensor). The representation used is the voxel grid, in which a voxel is the three-dimensional equivalent of a pixel. This method generates a three-dimensional grid of shape ( nx , ny , nz ), where each point of the point cloud is assigned to a voxel. If more than one point is assigned to the same cell, an interpolation is used to calculate the RGB value of that voxel. As an illustration, Figure 4a shows a point cloud and Figure 4b presents its voxelization. The number of voxels on each dimension ( nx , ny , nz ) depends on the architecture used, and it is explained in the subsection of each architecture. (a) Point cloud example (b) Voxelization of the point cloud Figure 4. Transformation from a point cloud to a voxel grid. 139