Control de robots con realidad mixta multimodal: Programación a través del guiado manual de gemelos digitales
Abstract
139 p.
Full text
Universidad del País Vasco / Euskal Herriko Unibertsitatea (UPV/EHU) Tesis Doctoral Control de Robots con Realidad Mixta Multimodal: Programación a través del Guiado Manual de Gemelos Digitales Andoni Rivera Pinto 1. Supervisor Johan Kildal Okiñena 2. Supervisora Elena Lazkano Ortega 2025
Andoni Rivera Pinto Control de Robots con Realidad Mixta Multimodal: Programación a través del Guiado Manual de Gemelos Digitales 28 de febrero de 2025 Supervisores: Johan Kildal Okiñena y Elena Lazkano Ortega Universidad del País Vasco (UPV/EHU) Facultad de Informática de San Sebastián Ciencia de la Computación e Inteligencia Artificial Manuel Lardizabal pasealekua, 1 20018, Donostia - San Sebastián (cc) 2025 Andoni Rivera Pinto (cc by-sa 4.0)
Resumen En esta memoria se presenta el trabajo de investigación sobre el uso de realidad mixta y gemelos digitales para la programación de robots colaborativos, proponiendo un enfoque que busca superar los desafíos actuales en la programación de brazos robóticos en entornos industriales. El trabajo explora cómo la integración de tecnologías inmersivas, tales como dispositivos de realidad mixta y retroalimentación háptica, puede facilitar la programación intuitiva de robots mediante el guiado manual de sus gemelos digitales holográficos. Esta investigación consta de 3 fases experimentales en las que se evalúa desde la manipulación básica de hologramas hasta la programación de robots físicos a través de gemelos digitales en contextos de trabajo colaborativo. Los experimentos realizados abordan cuestiones clave, como la viabilidad de utilizar interfaces de realidad mixta para definir trayectorias y poses robóticas, el impacto de la retroalimentación táctil en la percepción de control y fisicalidad, y las limitaciones en la precisión alcanzable en comparación con métodos tradicionales. Los estudios de usuario indican una mejora en la percepción de control con retroalimentación háptica, pero también revelan retos importantes relacionados con la precisión y la adaptación a entornos complejos. Las conclusiones sugieren que la realidad mixta tiene un gran potencial para reducir la complejidad de la programación robótica, haciendo posible que operarios sin experiencia técnica avanzada programen robots de manera efectiva. Por otra parte, se identifican áreas para futuras investigaciones, incluyendo la optimización de la integración háptica y la mejora en la identificación de obstáculos del entorno real y su representación en el entorno virtual. iii
Agradecimientos Quiero comenzar expresando mi más sincero agradecimiento a mis directores de tesis, Johan y Elena, por guiarme a lo largo de todo el proceso del doctorado. Más allá de las correcciones habituales, han sido una fuente constante de inspiración en los momentos clave de la investigación y me han ayudado a superar los bloqueos con una paciencia que, en muchos momentos, me ha faltado. A Tekniker, también, por brindarme la oportunidad de desarrollar esta tesis junto a su equipo, donde he aprendido la mayor parte de lo que hoy sé sobre interacción, robótica y entornos virtuales. Quiero agradecer profundamente a mis padres, cuyo apoyo incondicional me ha dado fuerzas durante todo este recorrido. Su respaldo ha sido fundamental para mi crecimiento personal. A mi hermano, por ser siempre una referencia para mí, no solo en lo académico, sino también en los aspectos más personales y emocionales de este proceso. A mis compañeros de trabajo, quienes han sido una fuente constante de aprendizaje, apoyo y motivación. Gracias por compartir su conocimiento y experiencia, y, sobre todo, por hacer el día a día más ameno. Finalmente, pero no menos importante, a la cuadrilla y al resto de mis amigos, que durante este proceso han sido los mejores aliados para desconectar y recargar energías para llevar a cabo este trabajo. v
Índice general I Introducción 1 1 Introducción 3 1.1 Motivación . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 1.2 Realidad Mixta . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 1.3 Objetivo general . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 1.4 Resultados de la investigación . . . . . . . . . . . . . . . . . . . . . . 7 1.5 Entorno de investigación . . . . . . . . . . . . . . . . . . . . . . . . . 8 II Actividad Investigadora y aportación científica 11 2 Estado del arte 13 2.1 Robots colaborativos . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 2.2 Realidad Virtual, Mixta y Aumentada . . . . . . . . . . . . . . . . . . 14 2.3 Realidad Virtual, Mixta y Aumentada en la industria . . . . . . . . . 16 3 Impacto de un sistema de realidad mixta multimodal para el guiado manual de un holograma de cobot 23 3.1 Dispositivos y tecnología utilizada . . . . . . . . . . . . . . . . . . . . 24 3.2 Configuración de los dispositivos . . . . . . . . . . . . . . . . . . . . 25 3.3 Escena de entrenamiento: “Laboratorio virtual” . . . . . . . . . . . . 26 3.4 Escena de guiado manual del robot holográfico . . . . . . . . . . . . 27 3.5 Estudio de usuario . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29 3.6 Resultados del estudio . . . . . . . . . . . . . . . . . . . . . . . . . . 30 3.7 Discusión de los resultados . . . . . . . . . . . . . . . . . . . . . . . . 32 3.8 Conclusiones . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 4 Un entorno de realidad mixta para programar un robot a través de su gemelo digital 35 4.1 Requisitos de la aplicación . . . . . . . . . . . . . . . . . . . . . . . . 35 4.2 Dispositivos y tecnología utilizada . . . . . . . . . . . . . . . . . . . . 36 4.3 Configuración y comunicación del sistema . . . . . . . . . . . . . . . 38 4.4 Diseño de la interacción y la aplicación . . . . . . . . . . . . . . . . . 38 4.4.1 Definición de la trayectoria del extremo del robot . . . . . . . 40 4.4.2 Definición de la pose objetivo del robot . . . . . . . . . . . . . 42 vii
4.4.3 Escaneo del espacio circundante . . . . . . . . . . . . . . . . 43 4.5 Estudiodeusuario ............................ 44 4.6 Resultados del estudio . . . . . . . . . . . . . . . . . . . . . . . . . . 45 4.7 Discusión de los resultados . . . . . . . . . . . . . . . . . . . . . . . . 48 4.8 Conclusiones ............................... 49 5 Entorno de realidad mixta para teleoperar un robot físico en una tarea de inspección de alerones 51 5.1 Requisitos del sistema . . . . . . . . . . . . . . . . . . . . . . . . . . 51 5.2 Dispositivos y tecnología utilizada . . . . . . . . . . . . . . . . . . . . 52 5.3 Configuración y comunicación del sistema . . . . . . . . . . . . . . . 53 5.4 Diseño de la interacción y la aplicación . . . . . . . . . . . . . . . . . 54 5.5 Procedimiento de programación del robot para la tarea . . . . . . . . 55 5.6 Estudiodeusuario ............................ 59 5.7 Análisis y discusión de los resultados . . . . . . . . . . . . . . . . . . 60 5.8 Conclusiones ............................... 61 6 Conclusiones y trabajo futuro 63 6.1 TrabajoFuturo .............................. 64 Bibliografía 67 III Artículos publicados 73 7 Multimodal Mixed Reality Impact on a Hand Guiding Task with a Holographic Cobot 75 8 Toward Programming a Collaborative Robot by Interacting with Its Digital Twin in a Mixed Reality Environment 97 9 Collaborative Robot Teleoperation in Mixed Reality Environment for Inspection Tasks 111 viii
Índice de figuras 1.1 Representación del continuo de virtualidad con los distintos tramos de losqueestácompuesto. .......................... 6 2.1 Primer visor de realidad aumentada creado por Ivan Sutherland. . . . 15 3.1 Dispositivo de realidad mixta HoloLens 1. Permite mostrar hologramas sobre el mundo real e interactuar con ellos. . . . . . . . . . . . . . . . 24 3.2 Dispositivo háptico que funciona a través de emisores de ultrasonidos y una cámara que permite virtualizar la pose de la mano. . . . . . . . . . 25 3.3 Configuración de los distintos dispositivos utilizados en el desarrollo. . 26 3.4 Vista general de la escena de prueba para conocer las posibilidades de latecnologíautilizada............................ 27 3.5 Punto de vista del usuario de la escena de realidad mixta durante la manipulación de un brazo robótico holográfico. En la esquina inferior derecha se muestra la perspectiva del entorno real. . . . . . . . . . . . 28 4.1 Segunda versión del dispositivo de realidad mixta HoloLens 2. . . . . . 37 4.2 Esquema del flujo de comunicación entre la aplicación de HoloLens 2 y ROS...................................... 38 4.3 De izquierda a derecha, la evolución en el diseño del holograma que permite definir la pose del extremo del robot. Pintado en naranja el núcleo del objeto que define la posición objetivo y en color rojo la parte que define la orientación. . . . . . . . . . . . . . . . . . . . . . . . . . 40 4.4 Vista en primera persona del modo de programación describiendo la trayectoria. ................................. 41 4.5 Diseños de las esferas del robot que permiten definir el punto inicial y final objetivo del extremo del gemelo digital del robot. . . . . . . . . . 42 4.6 Holograma del gemelo digital en su posición actual, junto al holograma que muestra la planificación del la trayectoria. . . . . . . . . . . . . . . 43 4.7 Imagen del mallado del entorno real que realiza HoloLens 2. . . . . . . 44 5.1 Esquema de la comunicación entre los distintos elementos y servicios delaaplicación................................ 53 5.2 Placa de calibración utilizada para obtener la pose del alerón con respecto a la cámara colocada junto al extremo del brazo robótico. . . 56 ix
Fig. 1.1: Representación del continuo de virtualidad con los distintos tramos de los que está compuesto. 1.3 Objetivo general Este trabajo ha tenido como propósito principal dar respuesta a la pregunta de investigación presentada con anterioridad: ¿Es posible programar un brazo robótico real a través de un programa de realidad mixta? Es por ello que se ha realizado un estudio de tecnologías de realidad mixta para la programación de un brazo robótico a través de su gemelo digital. Dentro de las tecnologías de realidad mixta, se ha evaluado también el impacto del uso de dispositivos hápticos en la manipulación de elementos holográficos. Todo el proceso de investigación realizado se ha dividido en tres partes. Cada parte tenía un objetivo o hito que son los siguientes: 1. Explorar la posibilidad de mover un brazo robótico virtual mediante guiado manual haciendo uso de su representación holográfica en un entorno de realidad mixta. Evaluar el impacto del uso de dispositivos hápticos en la manipulación de hologramas. 2. Explorar la posibilidad de programar un brazo robótico virtual a través de su gemelo digital en un entorno de realidad mixta. 3. Programar un robot real a través su gemelo digital en un contexto colaborativo de inspección de un alerón usando tecnología de realidad mixta. Retomando la pregunta principal, en caso de que la respuesta sea positiva, será posible realizar esta tarea mientras el robot que se pretende programar se encuentra 6Capítulo 1 Introducción
realizando otra tarea. Además, podría acercar este tipo de labor a trabajadores que carecen de conocimientos de programación. 1.4 Resultados de la investigación A partir del trabajo realizado en el contexto de la tesis se han publicado los siguientes artículos en revistas indexadas: 1. Rivera Pinto A, Kildal J., Lazkano E. Multimodal Mixed Reality Impact on a Hand Guiding Task with a Holographic Cobot. Multimodal Technologies and Interaction. 2020; 4(4):78. (Factor de impacto: 2.4; Cuartil: Q3; Número de citas: 14) 2. Rivera Pinto A., Kildal J., Lazkano E. Toward Programming a Collaborative Robot by Interacting with Its Digital Twin in a Mixed Reality Environment. International Journal of Human–Computer Interaction. 2023;0(0):1-13. (Factor de impacto: 3.4; Cuartil: Q1; Número de citas: 14) 3. Rivera Pinto A., Aceta C., Kildal J., Fernández I., Lazkano E., Collaborative Robot Teleoperation in Mixed Reality Environment for Inspection Tasks. PRESENCE: Virtual and Augmented Reality, 2025; 1-31. (Factor de impacto: 0.4; Cuartil: Q4; Número de citas: 0) Cada publicación presentada corresponde a un hito de la sección anterior respectivamente. En relación al primer artículo, previo a su publicación se presentó un "Late Breaking Report": Rivera Pinto A., Kildal J., Visuo-Tactile Mixed Reality for Offline Cobot Programming. In: Companion of the 2020 ACM/IEEE International Conference on Human-Robot Interaction. HRI ’20. Association for Computing Machinery; 2020:403-405. Entre la primera y segunda publicación de revistas indexadas se publica un artículo describiendo los siguientes pasos dados tras haber analizado los resultados y las conclusiones de la primera. Rivera Pinto A., Kildal J., & Lazkano E. Análisis de una experiencia multimodal de realidad mixta para la programación de un cobot a través de su gemelo 1.4 Resultados de la investigación 7
digital. Revista de la Asociación Interacción Persona Ordenador (AIPO). 2021; 2(2), 19-33. Este artículo fue publicado a través de la invitación recibida tras la participación en el congreso de Interacción Persona-Ordenador (IPO), que fue parte del Congreso Español de Informática (CEDI 20/21), y tras haber ganado el premio a mejor trabajo de fin de máster en la temática de interacción otorgado por la Asociación Interacción Persona-Ordenador (AIPO). El tercer artículo fue escrito en el contexto del proyecto COGILE que es un proyecto del programa ELKARTEK, que tiene como objetivo promover proyectos de investigación con alto potencial industrial con el fin de mejorar la competitividad en los ámbitos de la industria inteligente, las energías más limpias y la salud personalizada2. 1.5 Entorno de investigación Esta investigación ha sido llevada a cabo en Tekniker, un centro de investigación y desarrollo que se encuentra en Eibar y que forma parte del Basque Research and Technology Alliance (BRTA). Cuenta con distintos campos de especialización entre los cuales se encuentra la robótica y la interacción persona-robot. Tekniker enfoca su actividad en la generación de conocimiento y transferencia a la industria en forma de soluciones tecnológicas. De hecho, como centro de investigación sirve a una amplia variedad de sectores como son los de máquina-herramienta, fabricación, energías renovables, aeronáutica y espacial, biomedicina o automoción entre otros. Sistemas Autónomos e Inteligentes: La unidad de Sistemas Autónomos e Inteligentes (SAI) cuenta con una amplia experiencia en el desarrollo de sistemas inteligentes con un alto nivel de autonomía. La unidad cuenta con tres líneas de trabajo principales: Robótica: La unidad lidera la línea de especialización en robótica de Tekniker, definiendo y coordinando las actividades de otros grupos de investigación relacionadas con este campo. La investigación en robótica se centra en dotar a los robots de inteligencia para la navegación de plataformas móviles y la ejecución de trayectorias en brazos robóticos así como proporcionar a estos últimos la capacidad de manipular objetos. 2https://www.spri.eus/es/ayudas/elkartek/ 8Capítulo 1 Introducción
Visión por computador: Su actividad se centra principalmente en la inspección y control de calidad de elementos, aunque también aporta un valor importante en la línea de robótica para múltiples tareas como pueden ser la calibración entre dispositivos o la detección de objetos. Interacción persona-máquina: Esta línea de trabajo tiene como objetivo mejorar la comunicación bidireccional entre las personas y los robots o máquinas haciendo uso de tecnología que facilite la transmisión, representación e interpretación de la información emitida tanto por la máquina como por el usuario. El principal propósito de Tekniker es aportar crecimiento y bienestar a través de la I+D+i al conjunto de la sociedad, contribuyendo de manera sostenible a la competitividad del tejido empresarial. El interés de Tekniker en el desarrollo de técnicas de interacción entre personas y robots en entornos industriales, ha propiciado este trabajo de investigación, además de proporcionar el entorno adecuado para ello. 1.5 Entorno de investigación 9
Parte II Actividad Investigadora y aportación científica
2 Estado del arte En el contexto de la cuarta revolución industrial, impulsada por tecnologías como la analítica de datos, la inteligencia artificial y el internet de las cosas, los procesos industriales deben adaptarse continuamente a las demandas específicas de cada tarea, aprovechando estas herramientas para optimizar su rendimiento y flexibilidad. Estas transformaciones también han introducido nuevas dinámicas en la interacción entre humanos y robots en el entorno laboral. Por ejemplo, en algunos casos, es necesario que los operarios colaboren directamente en ciertos procedimientos mientras el robot realiza una tarea. Sin embargo, este tipo de colaboración no es viable con la robótica tradicional, que depende de barreras físicas para garantizar la seguridad. 2.1 Robots colaborativos Los robots colaborativos (también conocidos como cobots) son robots capaces de trabajar en colaboración con el trabajador de una manera segura [47]. Para que un robot pueda ser denominado cobot debe cumplir unos requisitos: Seguridad: Los cobots deben tener características de seguridad integradas como son sensores de fuerza, sistemas de detección de colisiones y paradas automáticas en caso de entrar en contacto con un humano. Normalmente sus diseños cuentan con bordes redondeados, materiales suaves o que absorben impactos para reducir daños en caso de colisión. Interacción con humanos: Los cobots deben ser capaces de interactuar de forma intuitiva con los operadores, lo cual supone que puedan ser programados directamente por el usuario (guiado manual) o a través de interfaces de fácil comprensión. Limitación de fuerza y velocidad: Así se evita causar daños en caso de contacto con una persona. Estos parámetros se ajustan en función a la proximidad con los operarios. 13
Detección del entorno: Deben incorporar sensores para recoger información del entorno y reaccionar ante cambios en el mismo. Este factor es crucial en una tarea colaborativa donde robot y operario comparten el espacio de trabajo. Estos requisitos permiten que los robots y los trabajadores puedan coexistir en el mismo entorno de trabajo sin necesidad de contar con jaulas de seguridad o barreras físicas. La capacidad de los cobots para ser programados fácilmente por trabajadores sin conocimientos técnicos es un requisito importante para su fácil adopción por parte de las industrias manufactureras [33]. El guiado manual [44] es la técnica que permite programar un robot moviéndolo libremente con la mano alrededor de su campo de trabajo. La acción de programar un brazo robótico definiendo su comportamiento a través de un guiado manual también es conocido como programación por demostración [13]. Mover un brazo robótico de forma manual es intuitivo para los trabajadores que, de esta forma pueden transferir su conocimiento del trabajo a desempeñar sin la necesidad de contar con un conocimiento técnico de programación ni de la cinemática del robot. 2.2 Realidad Virtual, Mixta y Aumentada La programación por demostración es un método de programación de brazos robóticos muy interesante debido a que reduce considerablemente la dificultad que supone programar un brazo robótico. Esto se debe a que el operario transmite su conocimiento al robot o sistema. Además, se consigue reducir el tiempo empleado en la programación ya que no requiere de escribir código como se hace a través de la programación tradicional. Como se ha dicho anteriormente, es el método de programación adecuado para aplicaciones en entornos dinámicos. Sin embargo, este método de programación también tiene limitaciones. Por un lado, el brazo robótico requiere de su manipulación física. Por otro lado, las trayectorias generadas por demostración pueden no ser las más eficientes o seguras, ya que no se optimizan automáticamente. Para conocer la viabilidad de la programación por demostración a través de dispositivos de realidad virtual y mixta, se debe conocer a qué tipo de tecnología hacen referencia estos conceptos y la tecnología existente que permita al usuario trabajar con el gemelo digital [17] de un brazo robótico para poder programarlo. La realidad virtual es un concepto relativamente nuevo en la era tecnológica. Fue en 1994 cuando Paul Milgram y Fumio Kishino acuñaron el concepto de continuo de la 14 Capítulo 2 Estado del arte
virtualidad[48]. Este concepto describe un espectro continuo donde los extremos definen experiencias completamente reales y completamente virtuales. En la zona intermedia se encuentran aquellas experiencias que no son puramente reales ni puramente virtuales. Toda esa región del continuo se denomina realidad mixta. Para entonces ya habían sido creados distintos dispositivos y experiencias de realidad virtual. Una de las primeras experiencias es Sensorama [24], que fue creada por Morton Heiling en 1962. Este dispositivo ofrecía al usuario una experiencia visual estereoscópica, sonido estéreo, sensación de viento, efectos olfativos y movimiento del asiento. En cuanto al primer visor de realidad aumentada (Fig. 2.1), fue en 1968 cuando Ivan Sutherland creó un dispositivo que mostraba dibujos en forma de malla sobre el mundo real [71]. Fig. 2.1: Primer visor de realidad aumentada creado por Ivan Sutherland. En 2002, se introdujo la realidad virtual visual en la medicina [2, 21, 23, 67] con el fin de entrenar y mejorar las habilidades de los cirujanos. La tecnología de cascos de visión de realidad virtual y aumentada ha evolucionado lentamente hasta que en la década pasada resurgió con cierta fuerza. En 2012 Google presentó un prototipo avanzado de sus gafas de realidad aumentada. El propósito de Google con este dispositivo era que se integrase como una herramienta más en la sociedad. En los siguientes años le siguieron otros dispositivos similares pero de realidad virtual como son Oculus Rift [39], PlayStation VR o HTC Vive permitiendo la inmersión en simulaciones de entornos realistas. Estos dispositivos se desarrollaron principalmente con un propósito lúdico. Los dispositivos de realidad extendida no sólo están destinados a ofrecer una experiencia visual. También existen dispositivos que ofrecen información al usuario a través de otro sentidos. Los dispositivos hápticos [57] permiten una experiencia 2.2 Realidad Virtual, Mixta y Aumentada 15
3 Impacto de un sistema de realidad mixta multimodal para el guiado manual de un holograma de cobot En este capítulo se describe el primer paso de investigación realizado para estudiar si es posible programar un brazo robótico a través del guiado manual utilizando dispositivos de realidad mixta. Para ello, en este primer hito el objetivo propuesto consiste en la manipulación de un modelo simple de robot. Haciendo un análisis de los distintos dispositivos de realidad mixta existentes, consideramos que lo ideal sería poder ofrecer al usuario una experiencia inmersiva que le permita percibir una visión estereoscópica del robot que se pretende representar. Además, la manipulación directa del robot en lugar del uso de coordenadas y cuaterniones de rotación busca dar acceso al desempeño de esta tarea a personas que carecen de estos conocimientos técnicos. Por último, una de las mayores carencias de la realidad mixta visual es la ausencia de percepción de tangibilidad al interactuar con hologramas. El uso de dispositivos hápticos puede solventar esa falta de sensación de estar manipulando un holograma. Concretando el objetivo de esta primera fase, formulamos dos preguntas de investigación: ¿Es posible programar un brazo robótico por guiado manual usando dispositivos actuales de realidad mixta? ¿Qué impacto tiene el dispositivo háptico en el guiado manual? Para poder dar respuesta a estas dos preguntas se pretende desarrollar una aplicación prototipo que permita el guiado manual de un robot holográfico mientras que el usuario sea capaz de percibir la sensación de estar tocando físicamente esa representación holográfica del robot. 23
A través de un estudio de usuario es posible medir la viabilidad de este tipo de tareas mediante el uso de tecnologías de realidad mixta además de la opinión de los usuarios. 3.1 Dispositivos y tecnología utilizada Tras realizar un análisis de los distintos dispositivos de realidad mixta en el estado del arte, consideramos utilizar los siguientes dispositivos: HoloLens: Es un visor de realidad mixta estereoscópico que permite proyectar hologramas sobre el mundo real e interactuar con ellos haciendo uso de gestos que permiten trasladarlos, rotarlos y escalarlos. Por otro lado, el dispositivo cuenta con altavoces que permiten reproducir sonidos espaciales que, a través de la HRTF consigue que el oído perciba que el sonido procede de un punto concreto en el espacio (Fig. 3.1). Fig. 3.1: Dispositivo de realidad mixta HoloLens 1. Permite mostrar hologramas sobre el mundo real e interactuar con ellos. Ultrahaptics Stratos: Es un dispositivo háptico que utiliza ultrasonidos para generar sensaciones táctiles en el aire. La ventaja principal de esta tecnología es que no se requiere portar el dispositivo. Se coloca en una superficie y es capaz de emitir sensaciones táctiles en un espacio de hasta 70cm. Este dispositivo cuenta a su vez con Leap Motion, un dispositivo que es capaz de analizar la posición de la mano con gran precisión y a través del cual es posible virtualizar la pose de la mano de una manera más precisa que con HoloLens (Fig. 3.2). A diferencia de otros dispositivos hápticos, éste no ejerce fuerzas de resistencia al movimiento de la mano del usuario. 24 Capítulo 3 Impacto de un sistema de realidad mixta multimodal para el guiado manual de un holograma de cobot
Fig. 3.2: Dispositivo háptico que funciona a través de emisores de ultrasonidos y una cámara que permite virtualizar la pose de la mano. Entre el software que se ha necesitado para llevar a cabo el trabajo desarrollado se encuentran los siguientes programas y librerías: Unity 1 : Es un motor gráfico cuyo propósito principal es el de crear videojuegos. Además, es posible hacer uso de él para crear entornos gráficos 3D cuyo propósito no sea lúdico (como es el caso de esta aplicación o aplicaciones de simulación entre otros). Una de las ventajas de este motor gráfico es que permite compilar programas para múltiples plataformas entre las que se encuentra “Universal Windows Platform” que es compatible con HoloLens. Microsoft MRTK: Es una librería que permite configurar el entorno desarrollado en Unity para poder desplegarlo en HoloLens además de agilizar ciertos procesos del desarrollo como son el uso de comandos de voz o la interacción con los hologramas entre otros. Kits de desarrollo de Ultrahaptics y Leap Motion: Software específico que permite controlar y recibir información de estos dispositivos. 3.2 Configuración de los dispositivos Para poder hacer uso simultáneo de ambos dispositivos, se ha preparado la configuración que se muestra en la Fig. 3.3. Un PC es el encargado de ejecutar la aplicación de Unity. A dicho PC se encuentra conectado el dispositivo Ultrahaptics Stratos vía USB para poder recibir información de Leap Motion (dispositivo que permite captar la pose de las manos) y representarla en el entorno virtual, al igual que 1https://unity.com/ 3.2 Configuración de los dispositivos 25
enviar información a Ultrahaptics con el fin de generar zonas táctiles en su área de actuación. Fig. 3.3: Configuración de los distintos dispositivos utilizados en el desarrollo. Por otro lado, HoloLens se comunica con la aplicación al PC vía WiFi para poder mostrar distintos elementos en el entorno. HoloLens permite instalar las aplicaciones directamente en el dispositivo, lo cual hace que exista una menor latencia entre las acciones del usuario y la respuesta de la aplicación. Sin embargo, debido a que es necesario comunicar HoloLens con Ultrahaptics y Leap Motion, se mantuvo esta configuración. 3.3 Escena de entrenamiento: “Laboratorio virtual” Como trabajo previo al prototipo planteado, se desarrolló una escena con la que poder conocer las posibilidades y limitaciones de los dispositivos utilizados (Fig. 3.4). Esta aplicación cuenta con distintos elementos con los que el usuario puede interactuar a través de ambos dispositivos. El dispositivo háptico debe de alinearse con la escena holográfica para poder visualizar la representación de las manos en el entorno de realidad mixta superpuestas a las manos reales del usuario. Esta alineación se hacía de forma manual colocando las HoloLens en una posición fija predefinida relativa a Ultrahaptics Stratos. La escena consta de distintos elementos con distintas funcionalidades. Para poder interactuar con un elemento, el usuario debe centrar su mirada (que sirve como cursor del visor HoloLens) en el objeto y utilizar el comando de voz “mover” para desplazar automáticamente dicho objeto al centro de la escena, que es donde se encuentra la representación virtual de Ultrahaptics Stratos. Los distintos elementos interactuables son los siguientes: 26 Capítulo 3 Impacto de un sistema de realidad mixta multimodal para el guiado manual de un holograma de cobot
Fig. 3.4: Vista general de la escena de prueba para conocer las posibilidades de la tecnología utilizada. Esfera: Una bola con la que el usuario puede interactuar agarrándola y soltándola en la zona de detección de Leap Motion. Cuando el usuario se encuentra sujetando la esfera, el dispositivo háptico emite ultrasonidos generando al usuario la sensación de estar sujetando una esfera. Productos químicos: Consisten en dos vasos de precipitado los cuales contienen un líquido en su interior. La demostración consiste en indicar a través de sensaciones que el producto contenido en el vaso es peligroso. Además de generar zonas visuales indicando el peligro cuando el usuario aproxima su mano al vaso, se producen distintas sensaciones para advertir al usuario del riesgo. Mezcla química: La mezcla química se realiza a partir de los dos productos químicos descritos previamente. Tras su mezcla, el producto resultante es un líquido burbujeante cuyas pompas pueden ser explotadas con la palma de la mano del usuario. 3.4 Escena de guiado manual del robot holográfico Tras analizar las capacidades de los dispositivos utilizados mediante la escena del laboratorio virtual, se diseñó y desarrolló la escena que consta de un pequeño brazo robótico ficticio de 6 grados de libertad y que no cuenta con restricciones de giro en ninguno de sus ejes. Este aspecto no es un problema para responder a las preguntas de investigación planteadas, aunque sí es un detalle a tener en cuenta de cara a futuros diseños e implementaciones. Como se muestra en la Fig. 3.5, en el extremo 3.4 Escena de guiado manual del robot holográfico 27
del robot (en azul) se encuentra pegada una esfera (que está siendo sujetada por el usuario) igual que en la escena de entrenamiento, que permite mover el robot a una posición objetivo. En el centro de la escena se puede distinguir la representación del dispositivo háptico sobre la cual es posible mover el brazo robótico al igual que los objetos en la escena anterior. Por otro lado, en la escena también hay una caja de herramientas (objeto amarillo) que servirá como punto objetivo para realizar el guiado manual. Fig. 3.5: Punto de vista del usuario de la escena de realidad mixta durante la manipulación de un brazo robótico holográfico. En la esquina inferior derecha se muestra la perspectiva del entorno real. Cuando el usuario coloca su mano dentro de la zona de interacción de Ultrahaptics Stratos se muestra una versión holográfica de la mano superpuesta sobre su mano real. Esto indica que el dispositivo Leap Motion detecta correctamente la mano de la persona. En caso de no aparecer la mano virtual, significa que no es capaz de reconocer la pose de la mano. Una vez la mano virtual aparece en la escena, el usuario, al igual que en la escena anterior, puede sujetar la esfera que se encuentra en el extremo del robot y moverla por la zona de acción del dispositivo de seguimiento de la mano. Mientras se mueve dicha esfera, el robot calcula la cinemática inversa para representar los segmentos del robot de forma adecuada para que pueda llegar a su pose objetivo. Por otro lado, el dispositivo háptico transmite la sensación de tener una esfera sobre la palma de la mano. Por último, en la parte delantera de la representación del dispositivo háptico se muestra un control deslizante que permite rotar en el eje vertical la pinza del robot. La interacción del usuario con la escena no se limita al uso de gestos. También se ha implementado el uso de comandos de voz, facilitando la ejecución de órdenes frente al uso de menús que, además de hacer más complicada la interacción con ellos en realidad mixta, requiere que el usuario desvíe la mirada de la región de interés de la 28 Capítulo 3 Impacto de un sistema de realidad mixta multimodal para el guiado manual de un holograma de cobot
escena para seleccionar una opción. Mediante el comando de voz “mostrar ayuda” el sistema muestra la lista de comandos de voz disponibles en la aplicación que permiten detener el movimiento del robot (stop robot) o desbloquear la opción de mover el robot (resume robot). Se puede ver la aplicación de realidad mixta en funcionamiento en el correspondiente video2. 3.5 Estudio de usuario Para poder contestar las preguntas de investigación que han dado pie al desarrollo del prototipo presentado en este capítulo, se ha llevado a cabo un estudio de usuario. Las preguntas de investigación que se plantearon en un inicio son: ¿Es posible programar un brazo robótico por guiado manual usando dispositivos actuales de realidad mixta? ¿Qué impacto tiene el dispositivo háptico en el guiado manual? El estudio de usuario ha contado con la colaboración de 16 personas sin experiencia en realidad mixta ni en programación de brazos robóticos. La mayoría de participantes se encontraban en el rango de edad de entre 18 y 30 años, 2 participantes entre 31 y 50 años y otros 2 participantes entre 51 y 65 años de edad. Antes de comenzar el estudio de usuario, todos los participantes fueron informados de la recolección anónima de información durante el estudio de usuario y que eran libres de abandonar el estudio en cualquier momento y que, en tal caso, se descartaría toda la información recogida en su ejecución. No obstante, ningún participante decidió abandonar el estudio. La estructura del estudio de usuario es la siguiente: Cumplimentación del cuestionario demográfico: En el cuestionario se pregunta por distintos datos demográficos además de su experiencia previa con dispositivos de realidad virtual, mixta y aumentada. Entrenamiento en realidad mixta: Los participantes llevaron a cabo un entrenamiento de 15 minutos en el "laboratorio virtual". En él aprendieron cómo interactuar con los dispositivos y la escena para poder desenvolverse con mayor comodidad en la tarea experimental. 2https://youtu.be/caq8GvCCWtU 3.5 Estudio de usuario 29
Ejecución de la tarea experimental: La tarea experimental se llevó a cabo en la “escena de guiado manual del robot holográfico” donde los usuarios debían colocar el brazo robótico holográfico pinzando la caja de herramientas holográfica. Esta tarea debían realizarla seis veces, tres con el dispositivo háptico conectado y otras tres sin el dispositivo háptico. Cuestionario de evaluación: Entre ejecuciones los usuarios respondían al cuestionario SEQ que extendimos con la pregunta acerca de la satisfacción tras la ejecución de la tarea. El cuestionario SEQ permite al usuario expresar con simplicidad y rapidez la dificultad percibida en una tarea. Al ser una única pregunta, es ideal para medir la experiencia inmediata tras completar la ejecución de una tarea. Para concluir el estudio de usuario los usuarios rellenaron dos cuestionarios. El primero es el NASA-TLX que fue extendido con dos preguntas adicionales en relación a la sensación de control del robot y la “fisicalidad” del experimento. A través de este cuestionario se pretende evaluar la carga cognitiva percibida durante la realización de la tarea. Es ideal para realizar una comparación entre las diferentes condiciones de interacción presentadas en este trabajo. Por otro lado, también permite identificar los factores específicos que afectan a la carga cognitiva. El segundo cuestionario estaba formado por 4 afirmaciones. Los participantes debían responder cuánto de acuerdo o en desacuerdo estaban con ellas a través de una escala de tipo Likert de 7 puntos. La primera pregunta de investigación será respondida a partir de la ejecución de la tarea por parte de los participantes. Sin embargo, identificar el impacto del dispositivo Ultrahaptics Stratos en el desarrollo de la tarea, sí requiere un procedimiento especial para no comprometer el resultado del estudio. Durante el estudio de usuario, como se ha descrito previamente, los usuarios han realizado dos ejecuciones iguales con la diferencia de que en una de ellas no hubo sensación tangible durante el movimiento del robot. Haciendo que la mitad de los usuarios comiencen ejecutando la tarea en una de las condiciones y la otra mitad de los usuarios la otra, queda balanceada la posible influencia de haber experimentado antes una condición frente a la otra. 3.6 Resultados del estudio Los datos recogidos se pueden clasificar en cuantitativos (recogidos de las ejecuciones de los usuarios en la aplicación de realidad mixta) y cualitativos (recogidos a través de los cuestionarios). 30 Capítulo 3 Impacto de un sistema de realidad mixta multimodal para el guiado manual de un holograma de cobot
Para dar respuesta a la segunda pregunta de investigación, hemos analizado los resultados obtenidos con retroalimentación háptica y sin ella. Al contar con datos pareados, se aplicó el test de Wilcoxon [74] o el t-test [69] (en función si los datos recogidos seguían una distribución normal o no) para comprobar si existe una diferencia significativa. En este estudio hemos considerado que existe una diferencia significativa cuando el valor p resultante es inferior a 0.05. Dentro de los resultados cuantitativos se encuentran las siguientes variables: Distancia de error: Es la distancia entre el punto objetivo y el definido por el usuario. Los resultados ([(M V+H =2.706 mm, SD V+H =2.351 mm);( M V =2.259 mm, SD V =1.998 mm)]) no mostraron una diferencia significativa entre la ejecución con retroalimentación háptica (V+H) y sin ella (V). Tiempo total: Es el tiempo total dedicado a realizar una ejecución. Desde que el usuario comienza la tarea hasta que define la orientación de la pinza. En este caso, los resultados ([(M V+H =24.027 s, SD V+H =21.483 s);( M V =26.614 s, SDV=27.473 s)]) tampoco mostraron una diferencia significativa Tiempo moviendo el robot holográfico: Esta segunda variable temporal recoge el tiempo acumulado de todos los momentos en los que el usuario se encontraba sujetando la esfera del extremo del cobot holográfico. Al igual que en el caso anterior, los datos ([(M V+H =16.338 s, SD V+H =12.638 s);( M V =16.242 s, SDV=11.747 s)]) no muestran una diferencia significativa. Relación de tiempo moviendo el holograma: En tanto por uno, se ha recogido la relación del tiempo que ha estado el usuario moviendo el holograma con respecto al tiempo total dedicado a la tarea. Los datos ([(M V+H =0.748, SD V+H =0.165);( M V =0.715, SD V =0.197)]) sugieren una similitud entre ambas distribuciones. A través del t-test se puede observar que no existe una diferencia significativa entre ambas distribuciones. Pérdidas de la esfera: Por último se ha recogido cuántas veces el usuario ha soltado accidentalmente la esfera del extremo del holograma (dejando parado el holograma hasta que vuelve a sujetarlo de nuevo). Nuevamente, los datos ([(M V+H =6.681, SD V+H =7.553);( M V =7.319, SD V =8.577)]) no muestran una diferencia significativa. Por otro lado, las preguntas del SEQ sobre la opinión de los participantes respecto a la dificultad de la tarea y la satisfacción con el resultado, no revelaron una diferencia significativa, aunque los resultados fueron positivos (tarea sencilla y alta satisfacción con el resultado obtenido). Sin embargo, en el cuestionario NASA-TLX extendido, la 3.6 Resultados del estudio 31
4.3 Configuración y comunicación del sistema La configuración del sistema consta de dos partes comunicadas entre sí. Por un lado, se encuentra la aplicación desarrollada para HoloLens 2, mientras que en el otro lado se encuentra todo el sistema de ROS. Ambos lados se comunican a través de la librería ROS# (que ha sido integrada en la aplicación para HoloLens 2) y a través de websockets del módulo rosbridge de ROS. Fig. 4.2: Esquema del flujo de comunicación entre la aplicación de HoloLens 2 y ROS. Como se muestra en la Fig. 4.2, la comunicación comienza en el lado de Unity donde se publican las coordenadas de la pose definida para el extremo del robot. Posteriormente, un nodo en ROS es encargado de planificar la trayectoria para lograr mover el extremo del brazo robótico al objetivo y, desde la aplicación de Unity se lee un “topic” de ROS que muestra en tiempo real la rotación de cada eje del robot para actualizar el estado del gemelo digital holográfico de la aplicación de realidad mixta. Para este prototipo, se va a trabajar con el gemelo digital del robot colaborativo UR10 que cuenta con 7 grados de libertad y una longitud de 1300 mm. 4.4 Diseño de la interacción y la aplicación El diseño de la interacción y de la aplicación en su conjunto juega un papel fundamental en el desarrollo del sistema. De este diseño depende que el usuario pueda comprender de manera intuitiva cómo interactuar y cómo expresar sus intenciones 38 Capítulo 4 Un entorno de realidad mixta para programar un robot a través de su gemelo digital
para que el gemelo digital ejecute correctamente la operación deseada. Para ello se han tenido en cuenta algunos de los principios específicos de las interfaces 3D definidos en [5]. De entre los distintos principios que se definen, se ha tenido en cuenta principalmente el de la simplicidad en la interacción. De este modo, el usuario no requerirá de un largo proceso de aprendizaje que pueda provocar al usuario un rechazo a la aplicación. Es crucial encontrar el equilibrio entre poner disponibles las herramientas y elementos necesarios para desempeñar la tarea a través del dispositivo de realidad mixta y, a su vez, no saturar de elementos el entorno virtual. Si el escenario no está completo, no se puede realizar la tarea, y si se satura de elementos, puede confundir al usuario y dificulta la realización de la tarea. Con la ausencia de la retroalimentación háptica, surge la necesidad de utilizar otros canales de interacción, tanto para que el usuario reciba información como para que exprese su intención a la aplicación. En este caso, se ha hecho uso del canal auditivo para validar la ejecución de acciones como son agarrar y soltar el extremo del robot. Se han desarrollado dos enfoques principales para programar el brazo robótico: 1. Definición de la pose objetivo del extremo del brazo robótico. 2. Definición de la trayectoria del extremo del brazo robótico. El usuario puede seleccionar qué método de programación utilizar a través del menú principal. Por otro lado, a través del mismo menú, se puede acceder al módulo para el análisis del entorno con el fin de limitar los movimientos del brazo holográfico para evitar colisiones. Este módulo debe utilizarse antes de iniciar la programación. Los dos escenarios que permiten la programación del brazo robótico cuentan con elementos en común. Uno de ellos es el holograma que le sirve al usuario para definir la pose a la que quiere que llegue el extremo del brazo robótico. Este holograma ha pasado por varios diseños iterativos con el objetivo de facilitar su comprensión. El elemento central en todos los diseños es una esfera, elegida por su asociación intuitiva con la acción de ser agarrada. A continuación, se detallan las tres versiones principales del diseño que se muestran en la Fig. 4.5: Primera versión: Una esfera principal acompañada de un prisma rectangular alargado que indica la dirección deseada para orientar el extremo del robot. 4.4 Diseño de la interacción y la aplicación 39
Tras una evaluación interna, se concluyó que el prisma no aportaba claridad y se descartó. Segunda versión: Se sustituye el prisma por una esfera que orbita alrededor de la esfera principal, representando el sentido del extremo del robot. A través de este diseño, la definición de la pose se realizaría en dos pasos. Este diseño presentó dificultades en la selección precisa de la esfera orbitante debido a su tamaño. Versión final: Se regresa al diseño inicial, pero reemplazando el prisma por una flecha tridimensional que indica el sentido del extremo del robot. Esta geometría, por su forma intuitiva, facilita la comprensión del usuario sobre el propósito del holograma. Fig. 4.3: De izquierda a derecha, la evolución en el diseño del holograma que permite definir la pose del extremo del robot. Pintado en naranja el núcleo del objeto que define la posición objetivo y en color rojo la parte que define la orientación. La interacción con el holograma se basa en gestos naturales. El operario puede realizar un gesto de agarre simulando sujetar una esfera física. Mientras mantiene el agarre, cualquier movimiento de la mano se refleja en el desplazamiento del holograma. Además, la rotación de muñeca hace que el holograma realice una rotación equivalente, permitiendo una programación intuitiva. En cuanto a la programación, el usuario puede elegir entre dos métodos para definir el comportamiento del brazo robótico. Las siguientes secciones describen estas dos alternativas. 4.4.1 Definición de la trayectoria del extremo del robot Este método permite definir al usuario la trayectoria que quiere que el extremo del brazo robótico holográfico siga. Para ello, el usuario debe definir tanto la posición objetivo como los puntos intermedios a recorrer. La forma de registrar esos puntos 40 Capítulo 4 Un entorno de realidad mixta para programar un robot a través de su gemelo digital
es sujetando la esfera. Con una frecuencia de 3 Hz, el sistema graba la pose del holograma que representa el extremo del robot. Entonces, se dibuja una esfera amarilla semitransparente en esa posición para indicar que se ha grabado la pose como se puede apreciar en la Fig. 4.4. Cada pose está compuesta por una posición 3D y una rotación en forma de cuaternión. Fig. 4.4: Vista en primera persona del modo de programación describiendo la trayectoria. A través de comandos de voz, también es posible definir un punto en el espacio. Para ello, el usuario coloca el dedo índice de la mano previamente definida en el menú de configuración en la posición deseada y señalando en el sentido al que quiere que se oriente el extremo del brazo robótico holográfico y utiliza el comando de voz “posiciona bola”. En caso de que una pose objetivo se marque fuera del rango máximo que alcanza el brazo robótico, la aplicación lo pintará directamente en color rojo semitransparente indicando que no es posible alcanzar dicha pose. Si las poses objetivo se encuentran dentro del rango de alcance máximo del brazo robótico, se recogen en un nodo de ROS que se encarga de calcular a través de “MoveIt!” si es posible llegar a esos puntos objetivo o no. Durante la ejecución del movimiento del gemelo digital, el extremo del robot deja una estela (en forma de pequeños cubos naranjas) que representa la trayectoria seguida por el extremo del robot. Esta trayectoria puede ser limpiada a través del comando de voz “limpia trayectoria”. Los comandos de voz disponibles restantes son: Menú: Vuelve al menú principal. 4.4 Diseño de la interacción y la aplicación 41
Calibra: Desbloquea la posición del gemelo digital y moverlo a otra posición del espacio que rodea al usuario. Bloquea: Bloquea el holograma del gemelo digital en la posición en la que se encuentra ubicada. 4.4.2 Definición de la pose objetivo del robot La segunda forma de programar el gemelo digital del brazo robótico es definiendo únicamente la pose inicial y final del extremo del brazo robótico. Fig. 4.5: Diseños de las esferas del robot que permiten definir el punto inicial y final objetivo del extremo del gemelo digital del robot. En este caso, la escena holográfica cuenta con dos esferas (Fig. 4.5): la esfera naranja con la flecha roja que define la pose inicial y la esfera verde con flecha azul que representa la pose objetivo del brazo robótico. El usuario puede mover las esferas al igual que en el modo de programación anterior. Debido a que las dos esferas pueden estar a una distancia considerablemente larga, el usuario puede planificar la trayectoria y previsualizarla antes de ejecutarla. Por otro lado, el usuario puede ordenar directamente que se mueva al punto de destino en caso de que sea posible sin mostrar la previsualización. Si el usuario decide previsualizar la trayectoria, aparecerá un fantasma del robot en la misma posición en la que se encuentra el gemelo digital. En el caso de que “MoveIt!” alcance una solución, el fantasma del robot se pintará de color blanco (Fig. 4.6). En cambio, si no consigue calcular una solución, el robot se verá de color rojo. Este método de programación y el que consiste en la definición de la trayectoria del extremo del brazo robótico cuentan con comandos de voz en común (menú, calibrar, bloquear y limpiar trayectoria). Por otro lado, también cuenta con sus comandos de 42 Capítulo 4 Un entorno de realidad mixta para programar un robot a través de su gemelo digital
Fig. 4.6: Holograma del gemelo digital en su posición actual, junto al holograma que muestra la planificación del la trayectoria. voz específicos que son: “empieza” para mover el brazo robótico a la pose de inicio si es posible; “mover” para mover el brazo robótico a la pose objetivo si es posible; “planifica” para planificar la trayectoria a seguir entre los dos puntos definidos y previsualizar dicha planificación; y “ejecuta” para ejecutar el plan. 4.4.3 Escaneo del espacio circundante Con el fin de que la solución propuesta por MoveIt! sea compatible con el entorno real, el sistema necesita conocer qué obstáculos hay presentes. Para ello, se ha añadido una opción en la aplicación en la que el usuario debe posicionar el robot holográfico en la posición donde se ubicará el robot real. Después, el usuario da la orden de escanear a través del comando “escanea”. Durante el análisis, el usuario visualizará un mensaje indicando que debe esperar a que finalice el análisis. A través de la utilidad “spatial awareness” que ofrece HoloLens 2, que crea una malla que cubre todos los elementos del entorno cercanos al usuario como se aprecia en la Fig. 4.7, el programa envía la información de cada vértice de la malla al planificador de MoveIt! a través de un servicio de ROS y crea un cubo de 5 centímetros de lado como obstáculo para planificar la escena. Este método de escaneo se presenta como una alternativa a la definición manual de la zona de trabajo alrededor del robot, ofreciendo una primera aproximación hacia la automatización de esta tarea. Tras analizar su desempeño, hemos constatado que resulta adecuado para escenarios donde no es imprescindible considerar pequeños detalles del entorno. No obstante, a medida que aumenta la complejidad del área 4.4 Diseño de la interacción y la aplicación 43
Fig. 4.7: Imagen del mallado del entorno real que realiza HoloLens 2. de trabajo, se vuelve más desafiante lograr una representación suficientemente precisa. En el correspondiente enlace 3 se puede ver un vídeo con el sistema en funcionamiento. 4.5 Estudio de usuario Se llevó a cabo un estudio de usuario para validar el diseño e implementación de las dos técnicas de programación (definiendo la trayectoria o definiendo el punto inicial y final) desde la perspectiva de usuarios que no contaban con experiencia en la programación de brazos robóticos. Los 14 participantes que formaron parte de este estudio recibieron una explicación sobre el funcionamiento de HoloLens 2 independientemente de su experiencia con dispositivos de realidad mixta. Además, antes de comenzar con la parte principal del estudio, los participantes pudieron interactuar unos minutos con un escenario de entrenamiento con el fin de habituarse al uso de esta tecnología y la interacción con los hologramas. La interacción más relevante en esta escena fue con la representación holográfica del extremo del robot. Los usuarios podían probar de qué manera era la mejor para sujetar y mover los hologramas por toda la escena. Se decidió llevar a cabo el estudio en un entorno industrial con el fin de saber también el comportamiento del dispositivo bajo condiciones de ruido y donde el 3https://youtu.be/Ee8JChQE6zk 44 Capítulo 4 Un entorno de realidad mixta para programar un robot a través de su gemelo digital
reconocimiento de voz pudiese verse afectado. De esta manera, se podía evaluar el funcionamiento del dispositivo en un entorno de trabajo realista. El estudio estaba compuesto por las dos escenas descritas previamente. En la primera escena, los usuarios tenían que definir la pose deseada del extremo del robot haciendo uso de la esfera con la flecha. El objetivo para cada ejecución estaba marcado con un pequeño cubo rojo virtual. La tarea del participante consistía en desplazar el brazo robótico holográfico haciendo uso de la representación del extremo del robot a los cubos rojos virtuales. El participante debía realizar tres ejecuciones colocando el extremo del robot en los cubos que se encontraban en posiciones distintas dentro del alcance del robot. En la segunda escena, el usuario debía “dibujar” la trayectoria completa a realizar por el extremo del brazo robótico. Al igual que en la anterior escena, el usuario tenía como objetivo aproximar el brazo robótico al objetivo representado por un cubo rojo. Se recogieron dos tipos de información en el estudio de usuario. El primer tipo de información era cuantitativo y está relacionado con el rendimiento del usuario (obtenida a través de las acciones que realiza el usuario dentro de la aplicación). El segundo era cualitativo y el objetivo era conocer la opinión subjetiva del usuario. Para ello se pidió a los usuarios que respondieran 2 preguntas en una escala de 7 puntos de tipo Likert: SEQ: En general, ¿cómo de fácil te ha resultado la tarea? ¿Cómo de satisfecho has quedado con el resultado? En la primera pregunta, cuanto mayor sea el valor, más sencillo es realizar la tarea para el usuario. Del mismo modo, en la segunda pregunta, cuanto mayor sea el valor, mayor es la satisfacción del usuario con el resultado obtenido en la programación. Estas dos preguntas fueron formuladas tras cada ejecución obteniendo así un total de 6 respuestas por usuario. Tras finalizar el estudio, los usuarios respondían el cuestionario SUS que consiste en 10 preguntas que miden la usabilidad de un sistema o programa. Para concluir, los participantes fueron entrevistados con el fin de recoger su opinión. 4.6 Resultados del estudio Como se ha adelantado en la sección anterior, el estudio se llevó a cabo con 14 participantes de los cuales 8 eran hombres y 6 mujeres. 11 de ellos tenían entre 26 4.6 Resultados del estudio 45
y 30 años, 2 participantes tenían entre 31 y 35 años y una sola persona estaba en el rango de 41 a 45 años. Todos habían completado estudios universitarios en las siguientes áreas: informática, electrónica, automatización e ingeniería industrial. Aunque 4 de los participantes tuviesen experiencia con gafas de realidad virtual en videojuegos, ninguno de los participantes tenía experiencia con dispositivos de realidad mixta. El análisis comparativo entre las dos formas de programación tuvo como objeto de análisis las siguientes variables: tiempo moviendo el extremo del robot, tiempo sin mover el extremo del robot y distancia a la posición objetivo. La media de tiempo empleado para mover el extremo del robot fue menor definiendo la trayectoria (M=10.2 s, IC 95 % =[8.2 s, 12.2 s]) que únicamente definiendo la pose objetivo (M=12.41 s, IC95 %=[11.22 s, 13.6 s]). El promedio de tiempo computado de inactividad (tiempo en el que el participante reflexionaba sobre el movimiento a realizar, tiempo de cálculo requerido por MoveIt! o por el procesamiento de los comandos de voz, e intervalos en los que el robot está en movimiento) fue menor en la opción de programar el robot mediante la pose final (M=28.39 s, CI95 %=[24.84 s, 31.94 s]) frente a la programación mediante la definición de la trayectoria (M=30.87 s, CI95 %=[25.82 s, 35.92 s]). Respecto al error de aproximación cometido, la media obtenida en la definición del punto objetivo (M=10.988 mm, CI 95 % =[9.58 mm, 12.4 mm]) fue menor que cuando se definía la trayectoria (M=15.036 mm, CI 95 % =[13.13 mm,16.94 mm]). También se analizaron los cuestionarios respondidos por los usuarios. Los cuestionarios seleccionados ayudan a comprender la facilidad o dificultad que supone utilizar este tipo de aplicaciones desde el punto de vista subjetivo del usuario. Todas las respuestas recogidas en el SEQ para los dos casos evaluados fueron de 6 y 7, es decir, en ambos casos los participantes pensaron que la tarea era sencilla de completar. La media fue ligeramente superior en el caso de tener que definir el punto final (6.86) frente al de trazar la trayectoria (6.71). En la pregunta de qué de sencillo era realizar la tarea, la media fue la misma en ambos casos (6.52). Por último, esta aplicación obtuvo una puntuación de 87.68 en el cuestionario SUS, lo que significa que estaría en el rango ‘B’ de la escala de Bangor [7], 2.32 puntos por debajo de la marca más alta. Esta puntuación sitúa a la aplicación sobre el 98 % de las aplicaciones evaluadas a través de este cuestionario. Las preguntas 4 y 10 enfocan su interés en conocer la facilidad con la que el usuario es capaz de aprender 46 Capítulo 4 Un entorno de realidad mixta para programar un robot a través de su gemelo digital
el funcionamiento de la aplicación. El valor medio en las respuestas con una escala del 0 al 4 fueron un 3.36 y un 3.5, respectivamente. El estudio concluía con una entrevista semiestructurada. El objetivo era recoger información y comentarios que pudieran ayudar a interpretar las respuestas de los cuestionarios. Cualquier comentario sobre la aplicación y el estudio era bienvenido, aunque se les sugería que diesen su opinión acerca de los siguientes aspectos: El sistema y la aplicación como métodos de programación de robots. Su experiencia interactuando con un sistema de realidad mixta. Dificultad en la interacción con hologramas. Dificultad en aprender a utilizar el dispositivo de realidad mixta y la aplicación. Aspectos que se echan en falta o que pueden ser mejorados. La mayoría de los usuarios comentaron que la aplicación era sencilla de usar y la interacción intuitiva, lo que hacía que la curva de aprendizaje tuviera un crecimiento muy rápido. El corto periodo de práctica que tuvieron antes del estudio de usuario fue suficiente para poder interactuar con la escena durante el estudio. El principal problema para los participantes era la primera vez que intentaban agarrar un holograma ya que no les parecía sencillo hasta haberlo intentado unas cuantas veces. Por otro lado, los participantes dijeron que el uso de comandos de voz encajaba bien en este tipo de aplicaciones. También coincidieron en que era posible programar un brazo robótico a través de esta aplicación sin requerir conocimientos de robótica ni de programación. Además, era posible agilizar la manera de definir una pose objetivo frente a hacerlo a través de técnicas actualmente utilizadas en la industria (ej., usando el “teach pendant”). Sin embargo, los participantes cuestionaron la precisión que se puede llegar a alcanzar a través de la aplicación de realidad mixta. Dependiendo de la tarea a realizar, esta solución podría no ser lo suficientemente precisa como para satisfacer los requisitos de un escenario real. Como solución a este problema, algunos participantes sugirieron la posibilidad de aumentar la zona objetivo con el fin de mejorar la precisión de posicionamiento. En relación al método de programación por dibujado de la trayectoria, algunos participantes sugirieron habilitar la posibilidad de modificar poses intermedias tras definirlas, en lugar de tener que volver a dibujar la trayectoria entera. Por otro lado, pese a que los comandos de voz eran sencillos de recordar, bajo condiciones de un ruido ambiente elevado, algunos participantes tuvieron que repetir varias veces los comandos para que fueran detectados por el sistema. Además, ciertos comentarios apuntaron a la falta de información del estado de la aplicación durante el tiempo de cálculo de 4.6 Resultados del estudio 47
En el lado de ROS se encuentran los distintos nodos que ofrecen servicios relacionados al movimiento del robot y el proceso de calibración del alerón con respecto al robot, y tópicos que ofrecen información con respecto al estado del robot. Además, también está el nodo encargado de comunicarse con el controlador del robot. A través del controlador, es posible mandar comandos de posición al robot y leer el estado de cada uno de sus articulaciones, que servirá para actualizar el estado del gemelo digital en la escena de realidad mixta. 5.4 Diseño de la interacción y la aplicación La interacción con los distintos hologramas sigue siendo a través de los gestos que permite el dispositivo escogido (HoloLens 2) para la aplicación. El holograma de previsualización de la planificación de la trayectoria muestra distintos colores en función al estado en el que se encuentre. Durante el tiempo en el que el sistema calcula la trayectoria, el holograma se muestra pintado de color blanco. Una vez haya finalizado la planificación, en caso de que “MoveIt!” haya conseguido hallar una trayectoria válida, el holograma se verá de color verde mostrando la trayectoria calculada. En caso de no conseguir calcular una trayectoria válida, el color que muestre el holograma será el rojo y se quedará en estático. Por otro lado, el holograma de la herramienta que el usuario utiliza para indicar la pose objetivo para el brazo robótico fue modificado. Ahora, su forma es igual que la herramienta real. Pese a que la forma esférica de la base del diseño anterior (Fig. 4.3) pueda resultar más indicativa de objeto movible, haciendo uso de la geometría de la herramienta se puede ayudar al usuario a definir con una mayor precisión en qué posición quiere que esta se sitúe con respecto al objetivo a alcanzar. También se sustituyó el sistema de comandos de voz anterior por un sistema de diálogo con el que el usuario puede expresar de una manera más detallada y completa sus intenciones. El usuario debe utilizar el comando de voz “KUKA” para habilitar el reconocimiento de la frase hablada, que se dará por terminada con un silencio de 1 segundo. El audio grabado es enviado al servicio de AZURE Speech-to-Text que devuelve su transcripción. Tras recibir en forma de texto la transcripción del mensaje del usuario, se procesa su mensaje a través del sistema de diálogo KIDE4I. El sistema de diálogo cuenta con varios módulos: el encargado de extraer elementos clave y la polaridad de la frase, el encargado del conocimiento (ontología) y el encargado de gestionar el diálogo. Gracias a su modularidad es posible sustituir cualquiera de sus módulos por otra alternativa. Como ejemplo, el módulo de extracción de elementos clave y polaridad puede estar basado en reglas o en “Large Language Models” (LLMs). El procesamiento de la frase acaba cuando el sistema de diálogo devuelve como 54 Capítulo 5 Entorno de realidad mixta para teleoperar un robot físico en una tarea de inspección de alerones
respuesta una acción identificada, junto a información complementaria en caso de que fuese necesaria. Las distintas acciones que el diálogo puede procesar en este contexto son las siguientes: Calibra: Comienza el proceso de calibración para colocar el alerón correctamente con respecto al brazo robótico. Ejemplo: Calibra el alerón con respecto al robot. Recolocar: Desbloquea la posición de los hologramas de tal forma que el usuario pueda mover la escena a otra posición del mundo real. Ejemplo: Recoloca la escena. Bloquear: Bloquea la posición de los hologramas impidiendo que el usuario pueda moverlos de su posición. Ejemplo: Bloquea la escena. Planificar: Manda una señal al nodo de ROS encargado de planificar una trayectoria a la pose objetivo. Ejemplo: Planifica la trayectoria a la pose objetivo. Ejecutar: Ejecuta una trayectoria previamente planificada. Ejemplo: Ejecuta el plan calculado. Posicionar: Coloca el holograma de la herramienta en una posición predefinida y planifica la trayectoria desde la posición actual del robot. Ejemplo: Posiciona la herramienta en el nervio A. 5.5 Procedimiento de programación del robot para la tarea La tarea asociada a este desarrollo consiste en una inspección colaborativa de un alerón entre el robot y el usuario. Cada uno se encuentra en uno de los extremos del alerón e irán moviendo sus respectivas herramientas con el fin de inspeccionar los nervios en busca de alguna posible obstrucción en ellos. Antes de comenzar con la programación del brazo robótico, el usuario debe calibrar la pose del alerón con respecto al brazo robótico. Esta calibración se realiza de forma manual y para ello debe colocar la cámara situada a un lado del extremo del robot mirando hacia la placa de calibración situada en el soporte del alerón. La calibración se ha llevado a cabo usando el software de visión HALCON y la cámara IDS UI-5240CP encarando una placa de calibración de 100mm de ancho y alto (Fig. 5.2). 5.5 Procedimiento de programación del robot para la tarea 55
Fig. 5.2: Placa de calibración utilizada para obtener la pose del alerón con respecto a la cámara colocada junto al extremo del brazo robótico. Previo a esta calibración, también se ha tenido que realizar otra calibración entre la placa y el alerón. Para ello, se ha colocado la placa de calibración en la posición deseada y, a través de una cámara 3D se ha obtenido la nube de puntos del alerón y la placa. Teniendo los modelos 3D de cada uno de los elementos, a través de una función de matching, se puede conocer la relación existente entre ambos elementos de una manera precisa (error de posición del orden de 1 mm). Teniendo en cuenta el error que se puede cometer con el sistema de calibración y viendo las dimensiones de la herramienta y las dimensiones de los nervios del alerón, se puede observar que es viable el uso de este sistema. Por el lado de la herramienta, las dimensiones a tener en cuenta son 30 mm de ancho y 70 mm de alto. Por parte de los nervios del alerón, el nervio más pequeño cuenta con 80 mm de altura y 45 mm de anchura. Conociendo la relación de pose entre la placa y el alerón, la pose de la cámara con respecto a la base del robot y calculando la relación entre la cámara y la placa de calibración a través del nodo de calibración, el sistema calcula la relación que existe entre la base del robot y el alerón. De este modo, el sistema es capaz de colocar correctamente el alerón holográfico con respecto al gemelo digital del brazo robótico. Además, esta información también se cargará en el planificador de “MoveIt!” para poder tener en cuenta las restricciones de movimiento del robot. 56 Capítulo 5 Entorno de realidad mixta para teleoperar un robot físico en una tarea de inspección de alerones
Tras completar la configuración de la escena, el usuario lanza la aplicación de realidad mixta en HoloLens 2. En ella se pueden ver el gemelo digital del brazo robótico junto a la representación holográfica de la herramienta y el alerón. A través del sistema de diálogo, el usuario puede realizar las distintas acciones anteriormente descritas dentro de la escena. Para ello, previamente debe activar el sistema de diálogo a través del comando de voz “KUKA” y esperar a que se reproduzca un sonido que es el que indica que, a partir de ese momento, la frase que diga a través del micrófono será procesada por el sistema de diálogo. Cuando el usuario se queda en silencio durante 1 segundo, el sistema de diálogo entenderá que ha terminado la frase y procesará la información recogida. El primer paso de la programación consiste en la aproximación del brazo robótico a uno de los nervios del alerón. Para ello, el usuario debe sujetar la herramienta holográfica con la mano y desplazarla hasta un punto cercano del nervio objetivo. Una vez haya aproximado el holograma, el usuario deberá indicar que quiere planificar una trayectoria a través del sistema de diálogo. Fig. 5.3: Holograma semitransparente de color blanco que indica que se está calculando una trayectoria válida a la pose objetivo. Durante el tiempo de cálculo de la trayectoria, aparecerá un holograma semitransparente de color blanco como se puede observar en la Fig. 5.3 en la misma posición que el gemelo digital. Si la planificación de la trayectoria resulta exitosa y se obtiene una trayectoria válida, la aplicación cambia el color del holograma a verde (Fig. 5.4) y reproduce el movimiento calculado con el objetivo de previsualizar la trayectoria que es posible realizar, de tal forma que el usuario pueda analizarla antes de su ejecución. En caso contrario, es decir, si la planificación no obtiene una trayectoria válida, la aplicación cambia el color del holograma a rojo como se muestra en la Fig. 5.5 y se mantendrá estático indicando que no se ha logrado una trayectoria válida. 5.5 Procedimiento de programación del robot para la tarea 57
Fig. 5.4: Holograma semitransparente de color verde que indica que se ha podido calcular la trayectoria a la pose objetivo definida por el usuario. Una vez el usuario consigue planificar una trayectoria válida, puede ejecutarla a través del sistema de diálogo. El holograma de previsualización de la trayectoria desaparecerá y, tanto el gemelo digital como el robot real se moverán de manera sincronizada. Fig. 5.5: Holograma semitransparente de color rojo que indica que no ha sido posible calcular la trayectoria a la pose objetivo definida por el usuario. Finalmente, el robot debe introducir la herramienta en uno de los nervios. Para ello, se han definido unas poses fijas relativas a la herramienta (una pose por cada nervio disponible). A través del sistema de diálogo, el usuario posiciona la herramienta en uno de los nervios y se lanza una llamada a la planificación de la trayectoria hasta ese punto. Del mismo modo que antes, el holograma semitransparente indicará el resultado de la planificación. 58 Capítulo 5 Entorno de realidad mixta para teleoperar un robot físico en una tarea de inspección de alerones
A través del enlace1a pie de página se puede ver un video del sistema en funcionamiento. 5.6 Estudio de usuario Se realizó un estudio de usuario para observar si realmente es posible programar un brazo robótico a través de dispositivos de realidad mixta para realizar una tarea colaborativa de inspección de un segmento de alerón. 13 personas participaron en este estudio que tenía como objetivo el análisis de 2 de los 4 nervios del alerón y, para ello, debían introducir la herramienta que se encontraba en el extremo del robot en dichos nervios. Antes de iniciar el estudio de usuario, se recibió el consentimiento de los participantes a grabar datos de sus ejecuciones además de los cuestionarios para posteriormente utilizarlos de forma anónima en el análisis de uso de la aplicación desarrollada. La tarea asignada para cada usuario es la descrita en la sección 5.5. Durante la ejecución de la tarea se recogieron los siguientes datos: el tiempo de ejecución total de la tarea, el tiempo de interacción con la herramienta holográfica, el tiempo utilizado en la interacción por voz, el tiempo en el que el usuario permanece inactivo y el número de veces que el usuario ha movido la herramienta. Tras el estudio, al igual que en el trabajo de investigación anterior, se pidió a los usuarios rellenar el cuestionario SUS (System Usability Scale) y evaluar tres frases más relacionadas con el sistema de diálogo en una escala de tipo Likert de 5 puntos: La interpretación de las acciones comandadas por voz es precisa. La interacción con el sistema a través del diálogo por voz es eficiente. Considero que interactuar con el sistema a través de la interacción por voz es útil. A través del cuestionario SUS se puede realizar una medición estandarizada de la usabilidad de la aplicación desarrollada y permite realizar una comparación de usabilidad con diferentes sistemas. Es adecuado para evaluar interfaces gráficas de distinto tipo y da una visión general de cómo los usuarios perciben el sistema en su 1https://youtu.be/wpLqU5O21a4 5.6 Estudio de usuario 59
conjunto. Se decidió prescindir del NASA-TLX en este caso ya que no había intención de comparar distintas interfaces. Para concluir el estudio y conocer mejor la opinión y las sensaciones de los participantes durante la realización de la tarea, se les dio la oportunidad de responder cuáles eran para ellos los puntos más y menos favorables de la aplicación. Para reducir la diferencia de destreza entre ejecuciones, previo al comienzo del estudio, se permitió a los participantes practicar el movimiento de hologramas con las gafas y acostumbrarse a la visualización de los hologramas a través de ellas. 5.7 Análisis y discusión de los resultados El estudio de usuario fue llevado a cabo con 13 participantes con una edad media de 22.46 años (SD = 1.941). Para todos los participantes, esta fue su primera experiencia utilizando un dispositivo de realidad mixta. Los resultados cuantitativos fueron positivos aunque dejaban entrever lo que posteriormente reportarían los usuarios en la entrevista final como carencias o puntos de mejora. El tiempo medio empleado para realizar la tarea al completo fue de 337.4 segundos. De ellos, 46.7 segundos eran dedicados al movimiento de la herramienta y 125.2 a la interacción por voz. El tiempo restante (165.5 segundos) fue tiempo que el usuario estuvo inactivo, esperando la respuesta del sistema o visualizando el movimiento del robot. En cuanto al número de veces que los participantes movieron la herramienta, en promedio fue de 8.77 movimientos, siendo 3 movimientos el mínimo requerido por la tarea. Pese a haber obtenido un valor ligeramente alejado del valor mínimo, es comprensible que hayan requerido más movimientos para posicionar la herramienta. Entre los distintos motivos posibles se encuentra la comodidad a la hora de realizar grandes rotaciones del holograma de la herramienta y dependerá de la flexibilidad y capacidad de rotación de la muñeca de cada usuario. El análisis del cuestionario SUS reportó una puntuación de 79.23. De acuerdo con la escala de Bangor se puede considerar la aplicación como “aceptable” y en nota corresponde a una ‘C’ que dista 77 centésimas de la siguiente nota (‘B’). 60 Capítulo 5 Entorno de realidad mixta para teleoperar un robot físico en una tarea de inspección de alerones
Las preguntas relacionadas con el sistema de diálogo resultaron positivas en todos los casos. La eficacia de la interacción a través del sistema de diálogo fue señalada por el 72.2 % de los participantes mientras que el 27.8 % restante no la consideró ni eficaz ni un obstáculo para la interacción. Además, todos los participantes consideraron que interactuar con el sistema a través del sistema de diálogo era útil. 5.8 Conclusiones El tercer trabajo de investigación ha tenido como objetivo evaluar la viabilidad de programar un brazo robótico para realizar una tarea colaborativa de inspección de un alerón a través del gemelo digital del brazo robótico real. Además de realizar cambios en la interfaz de usuario de la aplicación, se ha integrado un sistema de diálogo orientado a tareas y basado en la semántica que permite una comunicación a través de voz el lugar de los hasta ahora utilizados comandos de voz básicos habilitando una interacción más natural con el sistema. Además de las dos principales ventajas de estas aplicaciones de realidad mixta (que son: no necesitar conocimientos de programación de robots y la teleoperación), el estudio de usuario muestra una recepción positiva por parte de los usuarios tras utilizar la aplicación. El uso de frases completas en lugar de comandos simples o palabras permite a los usuarios proporcionar detalles al sistema cuando sea necesario. Esta capacidad sumada a la manipulación directa de los hologramas usando las manos, permite una interacción completa con el escenario virtual. Por otro lado, proporcionar información no solo por el canal visual sino también utilizando el canal auditivo favorece la comprensión del estado de la aplicación y la teleoperación al usuario. Las 13 personas que llevaron a cabo el estudio no tenían experiencia con dispositivos de realidad mixta. Sin embargo, es importante señalar que el tamaño de la muestra puede ser limitado para extraer conclusiones generalizables, por lo que sería interesante ampliar el estudio más adelante con un número mayor de participantes. Las 13 personas fueron capaces de completar el estudio y colocar la herramienta anclada al robot real en los dos nervios a través del uso de la aplicación. El éxito de la tarea en todos los casos permite validar que la tecnología utilizada es compatible con la aplicación propuesta. Se identificaron puntos de mejora pendientes en la aplicación de cara a futuras versiones. Estos son la interacción con la herramienta utilizada para facilitar la acción de agarrar y posicionar la herramienta en la escena. 5.8 Conclusiones 61
6 Conclusiones y trabajo futuro Esta tesis ha tenido como objetivo principal estudiar la posibilidad de programar un brazo robótico a través de su gemelo digital haciendo uso de tecnologías de realidad mixta que permitan simplificar los requisitos cognitivos para realizar esta tarea. La complejidad incremental de los desarrollos realizados ha concluido con un prototipo que permite a usuarios sin conocimientos ni experiencia previa programar un brazo robótico para inspeccionar un alerón, a través de interacciones multimodales sencillas e intuitivas. La investigación realizada permite responder afirmativamente a la cuestión principal planteada: Sí, es viable utilizar la realidad mixta para programar un brazo robótico haciendo uso de una interfaz que ofrezca métodos de interacción alternativos al guiado manual de un robot físico. Sin embargo, esta investigación ha identificado también limitaciones en las técnicas de realidad extendida propuestas y estudiadas. En cuanto al hardware, por un lado, el rango de actuación de los dispositivos hápticos es limitado para muchas aplicaciones en tiempo real, aunque su uso tiene un impacto positivo. Además, la captura de sonido mediante los micrófonos integrados en los cascos de realidad mixta presentaron una calidad deficiente en entornos con ruido, obligando a algunos usuarios a repetir sus comandos. No se puede obviar que, en ciertas ocasiones, el sistema de reconocimiento de voz a texto interpreta incorrectamente las órdenes, ejecutando acciones no deseadas. En cuanto a las formas de interacción empleadas, dar al usuario la posibilidad de previsualizar la trayectoria del robot antes de ejecutarla mediante una vista 3D le permite colocarse en la perspectiva deseada para realizar un análisis adecuado. Este enfoque representa una mejora significativa en la comprensión del movimiento en comparación con las interfaces de pantalla tradicionales, donde es necesario interpretar el movimiento del brazo robótico a través de proyecciones bidimensionales. La manipulación del entorno virtual mediante gestos también ha tenido una buena acogida en los estudios realizados. Se ha comprobado que la definición de poses objetivo a través de la manipulación de hologramas resulta factible e intuitiva. Sin embargo, la precisión alcanzada por el usuario al posicionar el holograma sobre un objetivo es limitada. Es por ello, que la eficiencia de la aplicación estará condicionada por los requisitos de precisión y de tiempo de la tarea a realizar. 63
[40] M. B. Luebbers, C. Brooks, M. J. Kim, D. Szafir y B. Hayes. ‘Augmented Reality Interface for Constrained Learning from Demonstration’. En: Proceedings of the 2nd International Workshop on Virtual, Augmented, and Mixed Reality for HRI (VAM-HRI). 2019 (vid. pág. 19). [41] Z. Makhataeva y Huseyin A. V. ‘Augmented Reality for Robotics: A Review’. En: Robotics 9.2 (2020) (vid. pág. 16). [42] S. Makris, P. Karagiannis, S. Koukas y A. S. Matthaiakis. ‘Augmented reality system for operator support in human–robot collaborative assembly’. En: CIRP Annals 65.1 (2016), págs. 61-64 (vid. pág. 16). [43] I. Malý, D. Sedlᡠcek y P. Leitão. ‘Augmented reality experiments with industrial robot in industry 4.0 environment’. En: 2016 IEEE 14th International Conference on Industrial Informatics (INDIN). 2016, págs. 176-181 (vid. pág. 19). [44] D. Massa, M. Callegari y C. Cristalli. ‘Manual guidance for industrial robot programming’. En: Industrial Robot: An International Journal 42 (ago. de 2015), págs. 457 - 465 (vid. pág. 14). [45] A. Masurovsky, P. Chojecki, D. Runde et al. ‘Controller-Free Hand Tracking for Graband-Place Tasks in Immersive Virtual Reality: Design Elements and Their Empirical Study’. En: Multimodal Technologies and Interaction 4.4 (2020) (vid. pág. 20). [46] I. Maurtua, I. Fernández, A. Tellaeche et al. ‘Natural multimodal communication for human–robot collaboration’. En: International Journal of Advanced Robotic Systems 14.4 (2017), pág. 1729881417716043. eprint: https://doi.org/10.1177/ 1729881417716043 (vid. pág. 20). [47] G. Michalos, S. Makris, P. Tsarouchi et al. ‘Design considerations for safe human-robot collaborative workplaces’. En: Procedia CIrP 37 (2015), págs. 248-253 (vid. pág. 13). [48] P. Milgram y F. Kishino. ‘A Taxonomy of Mixed Reality Visual Displays’. En: IEICE Trans. Information Systems vol. E77-D, no. 12 (dic. de 1994), págs. 1321 - 1329 (vid. págs. 5, 15). [49] D. Mourtzis, J. Angelopoulos y N. Panopoulos. ‘Closed-Loop Robotic Arm Manipulation Based on Mixed Reality’. En: Applied Sciences 12.6 (2022) (vid. pág. 19). [50] D. Ni, A. YC. Nee, S. Ong et al. ‘Point cloud augmented virtual reality environment with haptic constraints for teleoperation’. En: Transactions of the Institute of Measurement and Control 40.15 (2018), págs. 4091-4104 (vid. pág. 18). [51] D. Ni, A. W. W. Yew, S. K. Ong y A. Nee. ‘Haptic and visual augmented reality interface for programming welding robots’. En: Advances in Manufacturing 5 (ago. de 2017) (vid. págs. 18, 20). [52] J. Norberto Pires. ‘Robot-by-voice: experiments on commanding an industrial robot using the human voice’. En: Industrial Robot: An International Journal 32.6 (ene. de 2005). Ed. por Guido Bugmann, págs. 505-511 (vid. pág. 20). [53] F. Obermair, J. Althaler, U. Seiler et al. ‘Maintenance with Augmented Reality Remote Support in Comparison to Paper-Based Instructions: Experiment and Analysis’. En: 2020 IEEE 7th International Conference on Industrial Engineering and Applications (ICIEA). 2020, págs. 942-947 (vid. pág. 64). 70 Bibliografía
[54] S.K. Ong, A.W.W. Yew, N.K. Thanigaivel y A.Y.C. Nee. ‘Augmented reality-assisted robot programming system for industrial applications’. En: Robotics and Computer-Integrated Manufacturing 61 (2020), pág. 101820 (vid. pág. 18). [55] SK. Ong, AYC. Nee, AWW. Yew y NK. Thanigaivel. ‘AR-assisted robot welding programming’. En: Advances in Manufacturing 8.1 (2020), págs. 40-48 (vid. pág. 18). [56] M. Ostanin y A. Klimchik. ‘Interactive Robot Programing Using Mixed Reality’. En: IFAC-PapersOnLine 51 (ene. de 2018), págs. 50-55 (vid. pág. 20). [57] J. Perret y E. Vander Poorten. ‘Touching Virtual Reality: A Review of Haptic Gloves’. En: ACTUATOR 2018; 16th International Conference on New Actuators. 2018, págs. 1 - 5 (vid. pág. 15). [58] T. Pettersen, J. Pretlove, C. Skourup, T. Engedal y T. Lokstad. ‘Augmented reality for programming industrial robots’. En: The Second IEEE and ACM International Symposium on Mixed and Augmented Reality, 2003. Proceedings. 2003, págs. 319 - 320 (vid. pág. 18). [59] D. Puljiz y B. Hein. Concepts for End-to-end Augmented Reality based Human-Robot Interaction Systems. 2019. arXiv: 1910.04494 [cs.RO] (vid. pág. 16). [60] D. Puljiz, E. Stöhr, K. S. Riesterer, B. Hein y T. Kröger. ‘Sensorless Hand Guidance using Microsoft Hololens’. En: 2019 14th ACM/IEEE International Conference on Human-Robot Interaction (HRI). IEEE. 2019, págs. 632-633 (vid. pág. 19). [61] D. Puljiz, B. Zhou, K. Ma y B. Hein. HAIR: Head-mounted AR Intention Recognition. 2021. arXiv: 2102.11162 [cs.RO] (vid. pág. 19). [62] C. P. Quintero, S. Li, M. KXJ. Pan et al. ‘Robot Programming Through Augmented Trajectories in Augmented Reality’. En: 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). 2018, págs. 1838-1844 (vid. pág. 18). [63] E. Rosen, D. Whitney, E. Phillips et al. ‘Communicating and controlling robot arm motion intent through mixed-reality head-mounted displays’. En: The International Journal of Robotics Research 38.12-13 (2019), págs. 1513-1526 (vid. pág. 20). [64] E. Rosen, D. Whitney, E. Phillips et al. ‘Communicating robot arm motion intent through mixed reality head-mounted displays’. En: Robotics Research. Springer, 2020, págs. 301-316 (vid. pág. 19). [65] P. Rückert, F. Meiners y K. Tracht. ‘Augmented Reality for teaching collaborative robots based on a physical simulation’. En: Tagungsband des 3. Kongresses Montage Handhabung Industrieroboter. Ed. por Thorsten Schüppstuhl, Kirsten Tracht y Jörg Franke. Berlin, Heidelberg: Springer Berlin Heidelberg, 2018, págs. 41-48 (vid. pág. 20). [66] A. San Martin y J. Kildal. ‘Audio-Visual Mixed Reality Representation of Hazard Zones for Safe Pedestrian Navigation of a Space’. En: Interacting with Computers 33.3 (oct. de 2021), págs. 311 - 329. eprint: https://academic.oup.com/iwc/article-pdf/33/3/ 311/41062380/iwab028.pdf (vid. pág. 19). [67] N. Seymour, A. Gallagher, R. Sanziana et al. ‘Virtual Reality Training Improves Operating Room Performance’. En: Ann Surg 236 (oct. de 2002), págs. 458-463 (vid. pág. 15). [68] R. Sodhi, I. Poupyrev, M. Glisson y A. Israr. ‘AIREAL: interactive tactile experiences in free air’. En: ACM Transactions on Graphics (TOG) 32.4 (2013), pág. 134 (vid. pág. 20). [69] Student. ‘The probable error of a mean’. En: Biometrika (1908), págs. 1 - 25 (vid. pág. 31). Bibliografía 71
[70] Y. Su, X. Chen, T. Zhou, C. Pretty y G. Chase. ‘Mixed reality-integrated 3D/2D vision mapping for intuitive teleoperation of mobile manipulator’. En: Robotics and ComputerIntegrated Manufacturing 77 (2022), pág. 102332 (vid. pág. 18). [71] I. E. Sutherland. ‘A head-mounted three dimensional display’. En: Seminal Graphics: Pioneering Efforts That Shaped the Field, Volume 1. New York, NY, USA: Association for Computing Machinery, 1998, págs. 295-302 (vid. pág. 15). [72] M. Walker, H. Hedayati, J. Lee y D. Szafir. ‘Communicating Robot Motion Intent with Augmented Reality’. En: Proceedings of the 2018 ACM/IEEE International Conference on Human-Robot Interaction. HRI ’18. Chicago, IL, USA: Association for Computing Machinery, 2018, págs. 316-324 (vid. pág. 19). [73] K. Wang y T. Tang. ‘Robot programming by demonstration with a monocular RGB camera’. En: Industrial Robot: the international journal of robotics research and application 50.2 (ene. de 2023), págs. 234-245 (vid. pág. 18). [74] F. Wilcoxon. ‘Individual comparisons by ranking methods’. En: Breakthroughs in statistics. Springer, 1992, págs. 196-202 (vid. pág. 31). [75] C. Xue, Y. Qiao y N. Murray. ‘Enabling Human-Robot-Interaction for Remote Robotic Operation via Augmented Reality’. En: 2020 IEEE 21st International Symposium on “A World of Wireless, Mobile and Multimedia Networks” (WoWMoM). 2020, págs. 194-196 (vid. pág. 17). [76] L. Yang, S. Xiao, T. Wang, Y. Shi y K. Driggs-Campbell. ‘Robotics via Experiential Active Learning With Immersive Technology: A Pedagogical Framework Designed for Engineering Curricula’. En: IEEE Access 12 (2024), págs. 65706-65715 (vid. pág. 19). 72 Bibliografía
Parte III Artículos publicados
7 Multimodal Mixed Reality Impact on a Hand Guiding Task with a Holographic Cobot Authors Andoni Rivera Pinto, Johan Kildal y Elena Lazkano Publisher Multidisciplinary Digital Publishing Institute (MDPI) Journal Multimodal Technologies and Interaction Year 2020 Quartile Q3 DOI https://doi.org/10.3390/mti4040078 75
Multimodal Technologies and Interaction Article Multimodal Mixed Reality Impact on a Hand Guiding Task with a Holographic Cobot Andoni Rivera Pinto 1,2,* , Johan Kildal 1and Elena Lazkano 2 1TEKNIKER, Basque Research & Technology Alliance (BRTA), C/ Iñaki Goenaga 5, 20600 Eibar, Spain; [email protected] 2Computer Sciences and Artificial Intelligence Department, University of the Basque Country/Euskal Herriko Unibertsitatea (EHU/UPV), 48940 Leioa, Spain; [email protected] *Correspondence: [email protected] Received: 17 September 2020; Accepted: 28 October 2020; Published: 31 October 2020 Abstract: In the context of industrial production, a worker that wants to program a robot using the hand-guidance technique needs that the robot is available to be programmed and not in operation. This means that production with that robot is stopped during that time. A way around this constraint is to perform the same manual guidance steps on a holographic representation of the digital twin of the robot, using augmented reality technologies. However, this presents the limitation of a lack of tangibility of the visual holograms that the user tries to grab. We present an interface in which some of the tangibility is provided through ultrasound-based mid-air haptics actuation. We report a user study that evaluates the impact that the presence of such haptic feedback may have on a pick-and-place task of the wrist of a holographic robot arm which we found to be beneficial. Keywords: human–robot interaction; cobot; augmented reality; mid-air haptics; hologram; multimodal 1. Introduction Industry 4.0 [ 1 ] is the term used for an ongoing fourth industrial revolution, according to which factories of the future will comprise machinery, warehousing and production facilities in the shape of cyber-physical systems. Collaborative robots, also known as ‘cobots’, are one such type of technological assets that interact and cooperate with workers within the same shared workspace. In order to tag a robot as collaborative, it needs to meet some key requirements [ 2 ]. The first requirement is trust, which is essential for collaboration and ensures the safety of the human worker during the task. A robot with this feature also needs to match or improve the efficiency of the task performance (e.g., the time spent doing the work). Related to trust, it is also necessary for the user to feel a sense of control. In case of feeling uncomfortable, the human worker must be able to intuitively stop the robot. Finally, the robot has to be predictable in order to avoid unexpected movements that could endanger the human worker. The ability of cobots to be programmed easily by technically naïve workers is an important requirement for their easy adoption by manufacturing industries [ 3 ]. ‘Manual guidance’ (the technique to program a robot by freely moving it by hand within the workspace [ 4 ]) is a very useful and widely used technique to program robots in an easy and technically non-demanding way. This is also known as programming by demonstration (PbD) [ 5 ]. Moving a robotic arm by hand is intuitive for workers, who are able to transfer their job expertise to the robot without the need of technical knowledge of robot kinematics or programming languages. Making use of hand guidance for programming imposes, however, the requirement is that the robot has to be available (and not engaged in any other automation-related task) for a worker to have Multimodal Technol. Interact. 2020,4, 78; doi:10.3390/mti4040078 www.mdpi.com/journal/mti
Multimodal Technol. Interact. 2020,4, 78 2 of 21 exclusive access to the robot while programming it for a new task. This non-productive time for robot re-training has an evident negative impact on productivity. As a way around this limitation, we propose performing manual guidance on the holographic representation of the digital twin [ 6 ] of the cobot that needs to be programmed, by using current off-the-shelf mixed-reality (MR) equipment. Creating a hologram of the robot that is only visual, however, has the limitation that is lacking the tangibility of grabbing and moving a physical robot. In order to replicate the experience of physical hand-guidance more closely, in this paper we propose that some aspect of the tangibility of the robot might be reproduced by providing tactile feedback to the hand that holds the robot. To do so in a seamless way (i.e., without attaching any equipment to the workers hand), we explore the possibilities offered by a mid-air haptics technology that makes use of ultrasound-based actuation. The goal is that the spatial and temporal combination of visual and tactile augmentations may convey a more realistic representation of the robot, supporting manipulation actions that resemble interacting with the physical robot. Relevant prior work aiming towards manipulating a hologram does exist in the literature, as it is reviewed in the next section. Such examples from the state of the art made use of different kinds of augmented reality and virtual reality devices in order to program a robotic arm movement. One of the outcomes from such work was the lack of naturalness of the manipulation, which further supports our approach of creating a visuo-tactile multimodal interface. In preliminary earlier work [ 7 ], we prototyped this concept by creating an application with which the user could perform a pick-and-place task of the wrist of a holographic robot. This task represents a part of the hand guidance programming process. We also reported an initial pilot evaluation with a small group of users, with the aim to obtain early indication of the impact that using haptic feedback in performing the task might have on task execution and user experience. Five inexperienced users collaborated in the preliminary study work where we obtained a first impression related to the user experience and distance error and total time spent in the task values. While results suggested that the interaction might be promising, that pilot study was non-conclusive, given the small user group and the limited amount of data collected. In this paper, we report a full study with 16 naïve users performing the same task and procedure, but collecting a larger set of results with the employment of additional research methodologies. With the goal to obtain more solid insight into performance and user experience aspects, we employed a broader range of both quantitative and qualitative methods and metrics. We monitored and analysed additional quantitative metrics, and we collected subjective experience data with the use of additional user research methods: extended versions of the validated SEQ questionnaire and the NASA-TLX index. The same four-question ad hoc questionnaire as in the previous paper was also still used for comparison (all questionnaires used in the studies are included in Appendix A) and we monitored some data related to the user performance to compare performance with visual only feedback and the multimodal experience. With the present study, we wanted to shed light on two aspects of the interaction proposed. First, we wanted to know how feasible it was for a user to perform hand guidance of a virtual robot using off-the-shelf mixed reality devices. The devices that we focused on were (i) a Hololens HMD device to create visual holograms and (ii) an Ulrahaptics Stratos Explore device to create the tactile feedback on the user’s hand. In addition, and more specifically, we wanted to learn about the role that adding ultrasound-based tactile stimulation might play in the user’s performance and in the UX obtained. Thus, we formulated the following two research questions: • Is it possible to program a robotic arm through hand guidance using current mixed reality devices? • What is the impact of the Ultrahaptics Stratos device in manual guidance through the application? To adapt the experimental task to naïve users (with no prior experience collaborating with robots), we simplified a programming-by-demonstration task to a task of pick-and-place of the wrist of the robot.
Multimodal Technol. Interact. 2020,4, 78 3 of 21 In this way, the user could grab the holographic robot end effector, displacing it into a new position. The simplification of the task comes from not taking into account the kinematics constraints of the physical robot, which is an important part in a PbD solution because it defines the rotation and movement limits of the robot. However, we put the research focus on the impact that the tangibility of the hologram has during the execution of the task. The rest of the paper is structured as follows. Section 2reviews the prior work. Section 3describes the design of the experimental user study reported in this paper. In the Section 4, we describe the experimental user study procedure and the results are reported in Section 5. These results are analysed in Section 6. Finally, in Section 7, we present the conclusions and future work. 2. State of the Art As we know it today, augmented reality almost exclusively addresses the visual channel, with some examples of auditory augmentation as well. However, as far back as 1962, the concept of multiple sensory modalities was proposed by Morton Heilig’s Sensorama [ 8 ]. This device is one of the first machines related to the multimodal virtual reality (VR), with stereoscopic 3D view, stereo sound, wind and smell effects and mobile seat. In 2002, visual virtual reality was introduced in medicine [ 9 ] in order to train and, thus, improve surgeon’s skills. During the following years, a number of studies continued to focus on this area [ 10 – 12 ]. Head-mounted display (HMD) technology has advanced in these recent years, allowing the representation of complex 3D scenes rendered in real time and with a good image quality and frame ratio per second. Oculus Rift [ 13 ], PlayStation VR or HTC Vive are examples of HMDs which allow a realistic immersive simulation of a virtual world. Regarding augmented reality HMDs, Google presented in 2012 the advanced prototype “Google Glasses” in 2012, a device which had the aim of coexisting in the society. In 2015 Google decided not to sell their device commercially to the public and, in 2017, they presented a new release aimed for business and industrial environment. Human–robot interaction (HRI) has made profilic use of augmented reality, as seen in recent review publications, such as in [ 14 ], which identifies a roadmap for the field. In the same way, in [ 15 ], a review of advanced robot programming approaches is presented where the authors explain the current augmented reality-based programming approaches. Augmented reality and virtual reality fit perfectly in the manufacturing industry as support for operators. A good example is found in Makris et al. [ 16 ] for the programming of welding robots, or in [ 17 ] for industrial robots supported by virtual reality. Using different methodologies, algorithms and tools like in [ 18 ], it is possible to improve training, programming, maintenance and process monitoring. Defining fiducial markers in the 3D world [ 19 ] that are directly linked to the real robot close spots working with augmented reality or defining the path that the robot has to follow [ 20 ] could also simplify the difficulty of the task. Related to the safety requirements for human–robot collaboration in industrial settings, in [ 21 ], authors present a depth-sensor based model for workspace monitoring and an interactive AR user interface for safe human–robot collaboration. Through these kinds of augmented reality devices, the virtual representation of a robot makes the safety factor easy to comply with. Moreover, due to the possibility of overlapping additional information over the real world, it is possible to improve time to complete the task using marks and labels [ 22 ]. In 2017 [ 23 ], Ni et al. developed an augmented reality system that mixes haptic and visual experiences for the programming of welding robots using a haptic PHANToM device. In this case, authors used a screen for visual augmented reality. Puljiz et al. [ 24 ] presented a work on programming a collaborative robot using Microsoft Hololens mixed reality HMD with a KUKA KR-5 collaborative robot and ROS as an operating system. Luebbers et al. [ 25 ] presented a novel augmented reality interface for the visualization and directed
Multimodal Technol. Interact. 2020,4, 78 4 of 21 control of robot skill learning from demonstration (LfD). Rosen et al. [ 26 ] proposed a mixed-reality HMD visualisation of the intended robot motion, allowing users also to adjust the goal pose of the end effector through hand gestures. In addiction, [ 27 – 29 ] are examples of systems in which systems that use mixed reality HMD were used for the interactive programming of industrial robots. A limitation common to all the prior work is the lack of haptic feedback that can increase the realism of the pick-and-place task. An exception is [ 23 ], where the autors did use a PHANToM haptic device, but the visual experience was limited to a screen. If the application does not contain predefined target points to perform the task, a naïve user would have difficulties to move the robot and have spatial perception. Moreover, we decided to use a mid-air haptic device as an alternative. With the currently available mixed reality (MR) head-mounted displays, the capability of representing a hologram of a robot and its workspace allows creating a better context to perform a task [ 30 ]. With hand tracking technologies, the worker could perform the pick-and-place task in a similar way as with a physical robot. In the examples from the literature mentioned above, the movement is done in the air without grabbing any physical object, and consequently with a lack of tangibility, which could be addressed by using a haptic device, making the execution of the task potentially more intuitive and natural [ 31 ]. In the work presented here, the haptic device selected is an Ultrahaptics Stratos device, which is able to present tactile sensation on the skin in mid-air, without the intrusion of a physical interface. The main goal of the present study was to combine the use of a haptic device with a MR virtual environment to measure the feasibility and the improvement introduced by the Ultrahaptics Stratos device. 3. Design of the Experimental User Study The goal was to investigate the research questions mentioned above, regarding the programming of the robots through manual guidance of the holographic representation of its digital twin, in the sub-task of performing a pick-and-place operation in the context of a virtual scene and robot. For this implementation, we used a Microsoft Hololens mixed reality device to provide visual augmented reality. This device consists of 2 transparent lenses in 16:9 aspect ratio, sensors including an inertial measurement unit, depth camera, environment understanding cameras and microphones among others, and input/output connections (Bluetooth, Wi-Fi, etc.) which allow communication between this head-mounted display with the haptic device. As mentioned, tactile feedback was generated using an Ultrahaptics Stratos Explore ultrasound actuator [ 32 ]. This option of mid-air rendering of tactile sensations was deemed more appropriate than alternatives such as haptic gloves, or an ‘AIREAL’ [ 33 ] air vortex emitter device. The Ultrahaptics Stratos device consists of 256 ultrasound transducers that are combined to create tangible sensations on specific points in space that the hand occupied at each moment in time. It presents the advantage that it does not require wearing any physical device on the hand, although it has the limitation that the interaction space is small (up to less than 70 cm above the actuating panel). This device integrates a Leap Motion hand tracker that permits the localization of the user’s hand within the interaction space. Through the position of the hand and the capacity to create local air disturbances, it is possible to create a range of sensations on the user’s skin. When the user places his/her hand in the interaction region, a holographic hand is superimposed with the real hand, which corresponds to the reconstruction of the user’s hand as it is tracked by the Leap Motion device. Using these devices and the Unity cross-platform game engine, we created a 3D environment which simulated a collaborative robot that moved within the interaction space of the Ultrahaptics device. The setup to run the application is shown in Figure 1. A PC managed the creation of the 3D environment and, with a script, it also collected all the data during the task execution. The Ultrahaptics Stratos Explore device was connected through a USB port to the same PC. The Microsoft Hololens
Multimodal Technol. Interact. 2020,4, 78 11 of 21 Like with the previous parameters, the distributions of the data obtained were very similar in both conditions ((M V+H = 6.681, SD V+H = 7.553); (M V = 7.319, SD V = 8.577)), and no statistically significant difference was obtained from a Wilcoxon test ( p = 0.698), suggesting that the presence of tactile feedback might not influence the number of times the hologram dropped off from the user’s hand. Summarising all the above analysis, none of the observed quantified variables showed significant differences between conditions. The following two sub-sections present subjective data captured and quantified with two standard questionnaire-based methods, extended SEQ questionnaires and the 210 extended NASA-TLX test. 5.2. Qualitative Results Now, we are going to analyse the users’ opinions extracted from the extended SEQ test, the extended RAW NASA-TLX test and the customized questionnaire. 5.2.1. SEQ—Task Difficulty The Single Ease Question is a test that consists of only one question about the difficulty of the task that the user has performed in the test. After each task execution, users register on a 7-point Likert scale how easy they found it to execute that task (three times for each test condition) where a low value means high difficulty and high value means a lower difficulty. Figure 8shows the distribution of the answers provided. The Wilcoxon test ( p = 0.147) found no perceived difference in difficulty that was statistically significant. The distribution of the obtained data shows the similarity ((M V+H = 6.291, SDV+H = 1.031); (MV= 6.104, SDV= 0.973)). xx 2 4 6 V + H SEQ Difficulty answers distribution V Figure 8. Distribution of users’ answers about the task difficulty after each execution, where a high value represents a difficult task and low value an easy task. 5.2.2. SEQ—Satisfaction with the Result We extended the SEQ questionnaire with a second question that participants responded to in the same way (same Likert scale) and immediately after the first one. In this scale, participants rated how satisfied they were with the result they had obtained in that task execution, meaning the higher the value, the better the performance. Results are plotted in Figure 9. In this case, the data distribution
Multimodal Technol. Interact. 2020,4, 78 12 of 21 obtained is [(M V+H = 6.354, SD V+H = 6.0); (M V = 0.887, SD V = 1.167)], and the Wilcoxon test returned a value (p= 0.033), corresponding to a statistically significant difference between series. 2 4 6 V SEQ Satisfaction with the result x x V + H Figure 9. Distribution of participants’ answers about the satisfaction with the obtained result. 5.2.3. NASA-TLX The index obtained from the “Raw NASA-TLX (Task Load Index)” questionnaire is a measure of the work load experienced by participants when executing the experimental task. The questionnaire consists of six dimensions, all of which are rated on a 21-point (0 to 20) Likert-like scale: mental demand, physical demand, temporal demand, performance level achieved, effort expended and frustration experienced. Keeping the same format, we added two dimensions to the questionnaire: sense of control and perceived physicality. The sense of control dimension has been used as part of an extended TLX questionnaire (in e.g., [ 40 ]), and is highly relevant in the task chosen in this study. We added also the perceived physicality [ 41 ] dimension to assess the subjective degree of realism that tactile feedback might be adding to the manipulation of the hologram. To calculate the index, we used the raw version of NASA-TLX [39], which does not make use of weighted pairwise comparisons. Figure 10 shows the distribution of the Raw NASA-TLX values obtained from the 16 users, suggesting a lower TLX with tactile feedback ((M V+H = 5.646, SD V+H = 3.999); (M V = 6.271, SDV= 4.484) ). However, after performing the t-test, we obtained a p -value greater than 0.05 (0.096), which means that no significant difference between the mean of both conditions was found from our data. Still, we went on to examine each dimension of TLX individually, to see which dimension(s) the difference in mean TLX value originated from. Figure 11 plots the data distribution of the six constituent dimensions of TLX. As a result, we found a highly significant statistical difference in the temporal demand ( p -value < 0.01). Participants reported a lower temporal demand with tactile feedback from the Ultrahaptics Stratos explore device. In addition, regarding the perception of control while carrying out the task, data from the experiment revealed ((M V+H = 14.625, SD V+H = 4.470); (M V = 12.25, SD V = 4.524)) a difference that was also statistically significant in their mean difference ( p -value = 0.04), with perception of control being superior when the user performed the task with tactile feedback from the Ultrahaptics Stratos Explore device (Figure 12).
Multimodal Technol. Interact. 2020,4, 78 13 of 21 0 5 10 15 20 V Task load indexes distribution XX V + H Figure 10. Distribution of the task load indexes of the task performance with and without haptic feedback respectively. 0 5 10 15 20 V + H V Mental demand 0 5 10 15 20 Physical demand 0 5 10 15 20 Temporary demand 0 5 10 15 20 Effort expended 0 5 10 15 20 Performance achieved 0 5 10 15 20 Frustartion experienced V + H V V + H V V + H V V + H V V + H V xx xxx x x x xx x x ** Figure 11. Distributions of responses in both conditions (with and without tactile feedback) in the six scales that form the NASA-TLX questionnaire. Notice that, in all cases, the polarity of the scales reflect a better outcome the lower the values, including the performance scale. 5 10 15 20 V + H x x Control sensation data distribution V Figure 12. Distribution of the answers to the control sensation while performing the task.
Multimodal Technol. Interact. 2020,4, 78 14 of 21 Similarly, perceived physicality increased when the users performed the task with haptic feedback as opposed to without it. This difference ( p -value = 4.751 × 10 4 ) can be seen in Figure 13 ((MV+H = 15.437, SDV+H = 6.375); (MV= 3.521, SDV= 4.843)). 0 5 10 15 20 V Physicality distribution data x x V + H Figure 13. Distribution of the answers to the physicality while performing the task. 5.2.4. Ad Hoc Questionnaire Finally, we administered the same ad hoc questionnaire as in the preliminary study [ 7 ]. This questionnaire consisted of four statements that the users had to answer to on a Likert 7-point scale, indicating whether they agreed with them (3) or not (−3): •Q1—With tactile feedback, I have perceived a certain advantage to carry out the task. •Q2—I haven’t done complete the task faster when I felt the ball. •Q3—I have achieved better accuracy with tactile feedback. • Q4—The perception that I was handling the robot was the same with and without tactile feedback. The distribution of the answers is shown in Figure 14. −3 −2 −1 0 1 2 3 Q1 Q2 Q3 Q4 Subjective perception of condition comparison Figure 14. Results of the subjective opinion from the participants about the multimodal robot manipulation task.
Multimodal Technol. Interact. 2020,4, 78 15 of 21 Some consensus was shown for Q1 and Q4, with opinions regarding Q3 and Q4 distributed around the neutral region of the scale (with a rather even split of opinions). Thus, for Q1, 13 out of the 16 users perceived quite a clear advantage in carrying out the task with tactile feedback. As for Q4, participants largely supported the opinion that handling the hologram with or without tactile feedback felt different. 6. Discussion of the Results With the results obtained, we can evaluate the impact of the presence of tactile feedback from an Ultrahaptics Stratos device on the task performance and on the user experience. Regarding the objective quantitative metrics measured from the participants’ task executions (distance error, time spent on the task, net time grabbing the ball, ratio between both time values, number of times the participant lost the ball), we observed that the effect of introducing tactile feedback was small and, in fact, we found no statistically significant differences between conditions. An immediate conclusion that can be drawn from this is that the presence of the tactile feedback from an Ultrahaptics did not affect performance on the task of the experiment, either positively or negatively. In contrast, with the analysis of quantified subjective data obtained through the various questionnaires administered, some results emerged that showed positive effects from the presence of tactile feedback, with no negative effects detected. Starting with the extended SEQ questions (ease of task execution and satisfaction with result obtained), mean values of the distributions suggested positive effects from the presence of feedback, although those differences were not found to be statistically significant. A similar trend was found for the TLX index and for its six constituent dimensions, where, in all cases, mean values of response distributions were numerically lower with tactile feedback, suggesting positive effects from its presence. Statistical comparison of distributions showed that the difference was significant in the case of temporal demand, where it was reported to be lower (with a high level of statistical significance) when tactile feedback was present. This same trend showing a positive effect was also present in the categories that extended the NASA-TLX questionnaire (not included in the calculation of the TLX index): sense of control and perceived physicality. In both cases, not only were distribution means higher in the condition with tactile feedback, but the differences were statistically significant in both cases, with a particularly high level of significance in the case of the perceived physicality. Further confirmation of the positive impact from the presence of tactile feedback was found in the data collected from the ad hoc questionnaire (the four statements that participants could agree or disagree with). Regarding the two questions gathering most consensus in the responses, Q4 showed that a clear difference was perceived when handing the hologram with or without tactile feedback and, according to the broad consensus around Q1, the presence of tactile feedback offered an advantage for the execution of the experimental task. We hypothesize that the positive impact of having tactile feedback (reflected in the data from several of the metrics) might have been due to the effectiveness of feedback as a mechanism to confirm to the participant that he/she was grasping the hologram, as long as the feedback was felt. If such grasp got lost and the hologram was left behind in the process of dragging it, participants could notice immediately the change in sensation on the palm of their hand and react quickly to fetch the hologram again and resume the task. This could account for the significant differences found in lowering temporal demand, improving the sense of control, and providing an enhanced perception of physicality of the hologram. This interpretation of the results was reinforced by discussions with the participants about their experience during the study, who provided comments stating that it was helpful to receive tactile feedback, mostly because they could know better if they were holding the robot at each moment or not. Based on the data presented above, our interpretation is that the presence of tactile feedback was fulfilling the expectation of participants to be feeling in their hands the (virtual) object that they were holding. Tactile feedback increased the naturalness of that experience, and the sensation felt was reassuringly familiar, paving the way for an interaction that did not need to be learned.
Multimodal Technol. Interact. 2020,4, 78 16 of 21 7. Conclusions and Future Work In response to the first research question in the introduction, this paper presents the implementation of a functional demonstrator based on a Hololens head mounted display, with which a user can manipulate (move) a holographic robot, for the execution of a pick and place task, achieving positioning precision that remains under 3 mm. This functionality is a building block of interactions in more complex programming-by-demonstration scenario. The second research question (impact of tactile feedback from mid-air haptics actuators on a pick and place task) motivated the main contribution of this paper. The paper reports a user study in which one such task was performed by participants in two conditions, with and without tactile feedback. As discussed above, the results obtained suggest that feedback does not affect observed performance (no significant numerical differences were recorded, although the trend was for better scores obtained with tactile feedback). As for the qualitative results, subjective scores from participants on a set of questionnaire-based methods employed supported also the trend that tactile feedback was noticeable and with an impact that was positive. The strongest evidence of that was found on the reduction of temporal demand, improved sense of control, enhanced perceived physicality of the hologram, plus the consensus among participants that the noticeable effect introduced by tactile feedback gave an advantage in the successful execution of a pick-and-place task of a holographic robot arm. Building on the results obtained in this paper, our next steps are focused on implementing a programming-by-demonstration method and scenario, based on the handing of a holographic robot arm. We will reproduce a realistic context in which a real robot can reproduce the procedure demonstrated to the hologram. Alongside this process, we will investigate further haptic actuation techniques and multimodal interaction that can lead to improved performance and user experience, as well as individual the impact of each device used in the interaction task. Author Contributions: Conceptualization, A.R.P., J.K. and E.L.; methodology, A.R.P. and J.K.; software, A.R.P.; validation, A.R.P., J.K. and E.L.; formal analysis, A.R.P. and J.K; investigation, A.R.P. and J.K.; resources, J.K.; data curation, A.R.P.; writing—original draft preparation, A.R.P.; writing—review and editing, J.K. and E.L.; supervision, J.K. and E.L; project administration, J.K. and E.L. All authors have read and agreed to the published version of the manuscript. Funding: This research received no external funding. Acknowledgments: We would like to thank the people who participated in this study for the time they dedicated. Conflicts of Interest: The authors declare no conflict of interest. Appendix A. Evaluation Questionnaire All the data related to this paper can be found in the following link: https://drive.google.com/ drive/folders/12HaSSw9czindIjUhUptYFsSgGqphuofn?usp=sharing.
Multimodal Technol. Interact. 2020,4, 78 17 of 21 Figure A1. Demographic questionnaire some data about the participant. Figure A2. Extended SEQ. The user had to fill out one of these papers for each execution.
Multimodal Technol. Interact. 2020,4, 78 18 of 21 Figure A3. Extended RAW NASA TLX questionnaire. The users had to complete one row after the three executions with haptic feedback. The same procedure after the three executions without haptic feedback.
Multimodal Technol. Interact. 2020,4, 78 19 of 21 Figure A4. Customized questionnaire. The users had to complete it at the end of the evaluation. References 1. Henning, K. Recommendations for Implementing the Strategic Initiative Industrie 4.0; Forschungsunion: Berlin, Germany, 2013. 2. Michalos, G.; Makris, S.; Tsarouchi, P.; Guasch, T.; Kontovrakis, D.; Chryssolouris, G. Design considerations for safe human-robot collaborative workplaces. Procedia CIRP 2015,37, 248–253. [CrossRef] 3. Kildal, J.; Tellaeche, A.; Fernández, I.; Maurtua, I. Potential users’ key concerns and expectations for the adoption of cobots. Procedia CIRP 2018,72, 21–26. [CrossRef] 4. Massa, D.; Callegari, M.; Cristalli, C. Manual guidance for industrial robot programming. Ind. Robot Int. J. 2015,42, 457–465, [CrossRef] 5. Billard, A.; Calinon, S.; Dillmann, R.; Schaal, S. Survey: Robot programming by demonstration. In Handbook of Robotics; Springer: Berlin/Heidelberg, Germany, 2008; Volume 59. 6. El Saddik, A. Digital Twins: The Convergence of Multimedia Technologies. IEEE Multimed. 2018 ,25, 87–92. [CrossRef] 7. Rivera-Pinto, A.; Kildal, J. Visuo-Tactile Mixed Reality for Offline Cobot Programming. In HRI ’20: Companion of the 2020 ACM/IEEE International Conference on Human-Robot Interaction; Association for Computing Machinery: New York, NY, USA, 2020; pp. 403–405, [CrossRef] 8. Heilig, M.L. Sensorama Simulator. U.S. Patent 3,050,870, 28 August 1962. 9. Seymour, N.; Gallagher, A.; Sanziana, R.; O’Brien, M.; Vipin, B.; Andersen, D. Virtual Reality Training Improves Operating Room Performance. Ann. Surg. 2002,236, 458–463. [CrossRef] [PubMed] 10. Gallagher, A.G.; Ritter, E.M.; Champion, H.; Higgins, G.; Fried, M.P.; Moses, G.; Smith, C.D.; Satava, R.M. Virtual reality simulation for the operating room: Proficiency-based training as a paradigm shift in surgical skills training. Ann. Surg. 2005,241, 364. [CrossRef] [PubMed] 11. Grantcharov, T.P.; Kristiansen, V.; Bendix, J.; Bardram, L.; Rosenberg, J.; Funch-Jensen, P. Randomized clinical trial of virtual reality simulation for laparoscopic skills training. Br. J. Surg. 2004 ,91, 146–150. [CrossRef] [PubMed] 12. Aggarwal, R.; Ward, J.; Balasundaram, I.; Sains, P.; Athanasiou, T.; Darzi, A. Proving the effectiveness of virtual reality simulation for training in laparoscopic surgery. Ann. Surg. 2007 ,246, 771–779. [CrossRef] [PubMed] 13. Luckey, P. Oculus Rift. 2012. Available online: https://en.wikipedia.org/wiki/Oculus_Rift (accessed on 30 October 2020). 14. Makhataeva, Z.; Varol, H.A. Augmented Reality for Robotics: A Review. Robotics 2020,9, 21, [CrossRef] 15. Zhou, Z.; Xiong, R.; Wang, Y.; Zhang, J. Advanced Robot Programming: A Review. Curr. Robot. Rep. 2020 , [CrossRef] 16. Makris, S.; Karagiannis, P.; Koukas, S.; Matthaiakis, A.S. Augmented reality system for operator support in human–robot collaborative assembly. CIRP Ann. 2016,65, 61–64, [CrossRef] 17. Burghardt, A.; Szybicki, D.; Gierlak, P.; Kurc, K.; Pietru´s, P.; Cygan, R. Programming of Industrial Robots Using Virtual Reality and Digital Twins. Appl. Sci. 2020,10, 486, [CrossRef]
Multimodal Technol. Interact. 2020,4, 78 20 of 21 18. Andersson, N.; Argyrou, A.; Nägele, F.; Ubis, F.; Campos, U.E.; de Zarate, M.O.; Wilterdink, R. AR-Enhanced Human-Robot-Interaction—Methodologies, Algorithms, Tools. Procedia CIRP 2016,44, 193–198, [CrossRef] 19. Pettersen, T.; Pretlove, J.; Skourup, C.; Engedal, T.; Lokstad, T. Augmented reality for programming industrial robots. In Proceedings of the Second IEEE and ACM International Symposium on Mixed and Augmented Reality, Tokyo, Japan, 10 October 2003; pp. 319–320. 20. Ong, S.; Yew, A.; Thanigaivel, N.; Nee, A. Augmented reality-assisted robot programming system for industrial applications. Robot. Comput. Integr. Manuf. 2020,61, 101820, [CrossRef] 21. Hietanen, A.; Pieters, R.; Lanz, M.; Latokartano, J.; Kämäräinen, J.K. AR-based interaction for human-robot collaborative manufacturing. Robot. Comput. Integr. Manuf. 2020,63, 101891, [CrossRef] 22. Akan, B.; Çürüklü, B. Augmented reality meets industry: Interactive robot programming. In Proceedings of the SIGRAD 2010: Content Aggregation and Visualization, Västerås, Sweden, 25–26 November 2010; Linköping University Electronic Press: Linköping, Sweden, 2010; pp. 55–58. 23. Ni, D.; Yew, A.W.W.;; Ong, S.K.; Nee, A. Haptic and visual augmented reality interface for programming welding robots. Adv. Manuf. 2017,5, [CrossRef] 24. Puljiz, D.; Stöhr, E.; Riesterer, K.S.; Hein, B.; Kröger, T. Sensorless Hand Guidance using Microsoft Hololens. In Proceedings of the 2019 14th ACM/IEEE International Conference on Human-Robot Interaction (HRI), Daegu, Korea, 11–14 March 2019; pp. 632–633. 25. Luebbers, M.B.; Brooks, C.; Kim, M.J.; Szafir, D.; Hayes, B. Augmented Reality Interface for Constrained Learning from Demonstration. In Proceedings of the 2nd International Workshop on Virtual, Augmented, and Mixed Reality for HRI (VAM-HRI), Daegu, Korea, 11–14 March 2019. 26. Rosen, E.; Whitney, D.; Phillips, E.; Chien, G.; Tompkin, J.; Konidaris, G.; Tellex, S. Communicating and controlling robot arm motion intent through mixed-reality head-mounted displays. Int. J. Robot. Res. 2019,38, 1513–1526. [CrossRef] 27. Ostanin, M.; Klimchik, A. Interactive Robot Programing Using Mixed Reality. IFAC-PapersOnLine 2018,51, 50–55, [CrossRef] 28. Rückert, P.; Meiners, F.; Tracht, K. Augmented Reality for teaching collaborative robots based on a physical simulation. In Tagungsband des 3. Kongresses Montage Handhabung Industrieroboter; Schüppstuhl, T., Tracht, K., Franke, J., Eds.; Springer: Berlin/Heidelberg, Germany, 2018; pp. 41–48. 29. Rosen, E.; Whitney, D.; Phillips, E.; Chien, G.; Tompkin, J.; Konidaris, G.; Tellex, S. Communicating robot arm motion intent through mixed reality head-mounted displays. In Robotics Research; Springer: Cham, Switzerland, 2020; pp. 301–316. 30. Brooks, F.P. Grasping Reality Through Illusion&Mdash;Interactive Graphics Serving Science. In CHI ’88: Proceedings of the SIGCHI Conference on Human Factors in Computing Systems; ACM: New York, NY, USA, 1988; pp. 1–11, [CrossRef] 31. Ikits, M.; Brederson, J.D. The Visual Haptic Workbench. In Visualization Handbook; Hansen, C.D., Johnson, C.R., Eds.; Butterworth-Heinemann: Burlington, MA, USA, 2005; pp. 431–447, [CrossRef] 32. Carter, T.; Seah, S.A.; Long, B.; Drinkwater, B.; Subramanian, S. UltraHaptics: Multi-point mid-air haptic feedback for touch surfaces. In Proceedings of the 26th Annual ACM Symposium on User Interface Software and Technology, St Andrews, UK, 8–11 October 2013; pp. 505–514. 33. Sodhi, R.; Poupyrev, I.; Glisson, M.; Israr, A. AIREAL: Interactive tactile experiences in free air. ACM Trans. Graphics 2013,32, 134. [CrossRef] 34. Student. The probable error of a mean. Biometrika 1908,6, 1–25. [CrossRef] 35. Wilcoxon, F. Individual comparisons by ranking methods. In Breakthroughs in Statistics; Springer: Berlin/Heidelberg, Germany, 1992; pp. 196–202. 36. R Core Team. R: A Language and Environment for Statistical Computing; R Foundation for Statistical Computing; R Core Team: Vienna, Austria, 2018. 37. Byers, J.; Bittner, A.; Hill, S. Advances in Industrial Ergonomics and Safety; Taylor & Amp: London, UK, 1989; pp. 481–485. 38. MeasuringU: 10 Things To Know About The Single Ease Question (SEQ). Available online: https:// measuringu.com/seq10/ (accessed on 19 September 2020). 39. Hart, S.G. NASA-task load index (NASA-TLX); 20 years later. In Proceedings of the Human Factors and Ergonomics Society Annual Meeting; Sage Publications: Los Angeles, CA, USA, 2006; Volume 50, pp. 904–908.
requirements and more which we describe below, the devices used that permit us to carry out this design, the software and hardware setup required, and the descriptions of both proposed programming methods in this paper (by defining the last pose and by defining the trajectory). 3.1. Design requirements To achieve that, we identified the following requirements: Simplify the required knowledge related to spatial coordinates and about how to use the robotic arm model the user works with. The holographic digital twin of the robot. An accurate virtual model of the robot (holographic digital twin), in terms of physical dimensions, kinematics, and dynamics. The planning and the government of movements of the virtual robot have to be done with the same operating system and software that will be used by the physical robot. This guarantees that both the teaching experience of the user of the holographic robot and the program generated during the teaching session are directly transferable from the virtual to physical HRC setup. As the robotic arm could be surrounded by obstacles (e.g., tools or boxes) and over a table, it is needed to define the rest of the objects which compose the scene to avoid collisions while planning its movement. A visual AR device from the current state-of-the-art ecosystem renders the different holograms which compose the virtual scene. Easy-to-learn user interaction methods with the application by direct hologram manipulation using the hands and by voice commands to execute different actions. Avoid information overload from an excessive number of holographic representations, which might interfere with the perception of objects in the real scene. Make use of the auditory feedback to obtain the benefits of multimodal information representation. For example, auditory confirmation of actions. A representative hologram of the end effector that conveys the affordance of interacting with it using a natural interaction, using a gesture recognition module. In this way, the user will not spend a long time learning how to interact with the system. With all these requirements, after analyzing the current state-of-the-art about AR devices and available software, we decided to use HoloLens 2 HMD as it was the MR device that best suited our requirements, such as being able to use natural interaction through voice and gestures. In addition, Unity was chosen as a development platform, and Mixed Reality Toolkit to accelerate the programming process for MR applications. In order to compute all what concerns to the robot behavior, we decided to use ROS. Thanks to ‘MoveIt!’we will be able to compute the inverse kinematics of the robotic arm taking into account its movement restrictions and limitations. To communicate the MR application with ROS we use the adapted version of ROS# Bischoff & Vollenweider (2020) library (base project developed by Siemens Bischoff (2019)) in Unity side, through which it is possible to establish a connection, read and write in topics, and also make use of services and actions. 3.2. System configuration and communication In this section we describe the setup used to allow the communication between the HoloLens application with ROS which is the responsible of computing the robotic arm movements. As it is described in Figure 1, the flow of the application starts with the worker defining one target pose through HoloLens 2 application. The pose is defined using a hologram that represents the 6D pose of the robotic arms end effector. Thanks to the communication created between ROS and the HoloLens application with ROS#, the target pose (geometry_msgs/Pose) is sent to a ROS service that is responsible for planning with ‘MoveIt!’, whether it exists or not, a valid trajectory for that request. In the same call, the action to be performed (std_msgs/String) is sent: “add”to stack the different points of a trajectory to be planned, “plan”to plan the trajectory, “execute”to execute the previously planned trajectory and “reset”to remove the possible stored points or planned trajectories. If so, the application will receive a positive answer that will allow the user to execute the planned path. Otherwise, the answer will not permit the execution of a plan as it is not successfully computed. On the other hand, on the ROS side, a topic is publishing the state of the robot (sensor_msgs/JointState) all the time. This information is read in the HoloLens application to represent the current status of the robot each time. When a message is published with the current joint states in ROS, the application in HoloLens retrieves them and updates the digital twin’s pose. 3.3. AR application and interaction design Related to the application design requirements, we took them into consideration for the design described below. We mainly focused our efforts on reducing the visual Figure 1. Setup scheme to run the application with the connected devices. The unity side represents the application running in HoloLens 2. INTERNATIONAL JOURNAL OF HUMAN–COMPUTER INTERACTION 4749
information shown by the MR glasses and representing it using the auditory channel. We also took into account the auditory feedback to represent the action of grabbing and releasing the end effector to replace the haptic feedback used in the previous work. With a simple user interface, we can achieve the less skilled and naïve users learn how to program a robotic arm while forgetting about the technical part of the task. For this purpose, we also worked on direct hologram manipulation using the hands. We developed an application that addressed all the previously defined requirements. When the user launches the application, a menu panel (Figure 2) is displayed in front of them with some selectable buttons, two of which are related to the developed programming methods. Another button opens the settings menu, and there is also one button with which the user could move the digital twin and scan the surrounding area in order to take into account the possible obstacles that the real robotic arm should avoid. The available settings allow the user to select his/her dominant hand, keep or automatically remove the programmed trajectory visualization (explained in this section below) after a few seconds. The user can define the pose of the robotic arm end effector using an orange sphere (the center of which is indicated by the position of the end effector) with a red arrow (to indicate the orientation of the end effector) (Figure 3). The end effector representation had two different designs before the current one. The first one consisted of a sphere with a cylinder indicating the direction the user wanted the end effector to look at. In the second alternative, the end effector was composed of a sphere and a smaller sphere orbiting around the main one. The disadvantage of the first design was that the meaning of the cylinder was confusing for the user. In the second one, the problem was the difficulty of moving the smaller sphere, even if the design was visually cleaner than the first one. With the selected design, to define a new target position and publish new coordinates, the user moves the sphere by interacting with it using his/her hand (grab-move-release). The hand tracking module integrated in HoloLens 2 allows the user to intuitively use his/her hand and make the grab gesture in the sphere, move it to another position, and leave it in a new one. Similarly, to define a new rotation, the user only needs to twist the wrist. This rotation is easily recognizable for the user, thanks to the arrow. When the hand tracking module detects the picking gesture on the sphere, as there is no tactile feedback, the system plays a sound to confirm to the user that the sphere is being grabbed. Similarly, when the user releases the sphere or drops it accidentally, the system plays another sound to confirm the action of releasing voluntarily or to warn that the end effector has been dropped accidentally. This sound is very useful while programming the robotic arm because it indicates rapidly that the user lost the ball if this action was not made on purpose. The study in Pinto et al. (2020) showed that highlighting sensorially the event of accidentally dropping the robot from the hand was key to preserving a good UX. Currently, this is not a common problem now due to the improvement of the hand-tracking module. Regarding the programming task, the user can choose between two ways to define the behavior of the robotic arm. 3.3.1. Method 1: Defining the trajectory In this method, while the end effector’s sphere is being moved by the user, it defines the trajectory to be performed. The rate of the system for defining poses is 3 Hz (i.e., 3 poses per second). The objective is that the robot follows the trajectory defined by the user, by passing along the poses defined by the user on the trajectory. Each pose is composed of a 3D position and a quaternion for the rotation. Through a voice command (locate ball), it is also possible to define a point in space using the hand. In this case, the sphere is placed in the position of the index finger of the user and with the orientation the finger is pointing towards. A new sphere is created in the same location, and it is colored in yellow until ‘MoveIt!’calculates if the robot can reach the defined position. If the sphere is out of the robot bounds, this is directly painted in red and the new pose will not be computed by ‘MoveIt!’. After calculating the new pose, if that pose is feasible, the defined sphere is painted in green and in red otherwise (Figure 4). The movement of the robotic arm draws a trail between contiguous defined poses indicating the trajectory that the Figure 2. Main menu of the application where the user can select the programming method or go into the settings menu. Figure 3. Representation of the end effector location. Through these spheres, the user could define where he/she wants to start and end the robot’s end effector movement. 4750 A. RIVERA-PINTO ET AL.
end effector of the robotic arm makes during the movement from the first position to the second. The representation of the trail is composed of small orange boxes (15 mm side length). The goal with this feedback was to add understandability for the user since it shows both, the planned trajectory together with the checkpoints. In order to execute or enable/disable the different options available, the user must use the voice commands. The voice commands are the following: Menu: Returns from programming space to the main menu. Calibrate: Allows the user to move the digital twin picking and placing it in the surrounding space. Lock: Locks the digital twin hologram in the space so that it cannot be moved. Clean Trajectory: Removes all the defined spheres and trajectory boxes that had been drawn due to the movement of the robot. Locate Ball: Moves the sphere to the index finger of the user and looking at the direction of the index finger. In case of detecting both hands, the sphere will be located in the main hand (previously defined in the settings menu). 3.3.2. Method 2: Defining the starting and ending poses The second way to define the robotic arm’s behavior is by defining exclusively the starting and ending points (Figure 5). In this case, the user has to move two spheres. The orange sphere with the red arrow defines the starting pose (position and rotation), while the green sphere with the blue arrow defines the ending pose (Figure 3). In the same way, the user can move both spheres as they can in the other programming mode. Regarding the holographic digital twin location, the system allows the user to move it around and locate it, for instance, on a table instead of leaving it floating in the air. Figure 4. User point of view. The different goal spheres are colored in yellow (poses that are not calculated yet), green (reached poses) and red (unreachable poses). Figure 5. User and third-person perspective of the application. Starting and ending poses definition mode running a movement and drawing the planned trajectory. INTERNATIONAL JOURNAL OF HUMAN–COMPUTER INTERACTION 4751
Using voice commands, it is possible to allow the movement or the lock of the digital twin in the current position. In this programming method, the user only defines two points that can be quite far from each other. The user can plan the trajectory and preview it before executing it, or they can move the end effector and plan a new trajectory. There is also the possibility to execute the movement without previewing the robot’s trajectory. In case the user decides to preview the trajectory, a faded hologram of the robotic arm will appear in the same location as the digital twin representation. If MoveIt! reaches a solution, it will be colored in white and will represent the planned movement. Otherwise, the hologram will be colored in red, indicating that it was not possible to obtain a solution. For this method, the first four voice commands of the previous method (menu, calibrate, lock and clean trajectory) are available again, as well as some specific voice commands, which are the following: Start: Moves the robotic arm to the starting pose (orange sphere) if it is possible. Move: Moves the robotic arm to the ending pose (green sphere) if it is possible. Plan: A trajectory is planned and displayed through HoloLens 2 starting from the current position to the pose of the green sphere. Execute: After planning a path, this command executes the planned path. 3.4. Surrounding space scanning For the trajectory returned by MoveIt! to be compatible with the physical environment, the system needs to know what obstacles are present. The user locates the hologram of the digital twin in the position in which the real robot would be placed later and scans it after recognizing the voice command “scan”. The text message “please wait”is shown in the user’s field of view while the task is being performed. Using the spatial awareness utility that HoloLens 2 provides, which creates a mesh covering all the physical objects in the surroundings, we send the position of each vertex that the mesh is composed of to add a 5 centimeter size box as an obstacle to the ‘MoveIt!’planning scene (Figure 6). To perform this task, we created a service in ROS, the input of which was an array of 3-dimensional vectors representing the position of each vertex. The service itself was responsible for publishing each obstacle in the ‘MoveIt!’ scene. Once it finishes the task, the HoloLens 2 application receives a feedback message to notify the user that the task is done, and the text message is then removed. This scan method intends to be an alternative for manually defining the workspace of the robot. The module we developed is a first approach to automate this task. Validating this method, we could check that scanning simple areas (e.g., a working table or a conveyor belt) can be considered good enough to define the movement restrictions. Nevertheless, more complex workspaces (e.g., a table with differently shaped objects) will not be represented with enough accuracy, it can be seen in Figure 6. 4. User study We conducted a user study to validate the design and implementation of these two versions of the interactive robot programming technique, from the perspective of naïve users in robot programming. We first describe the design and procedure of the user study we conducted with 14 participants, and then we report the results obtained. 4.1. Study procedure Before starting the study, we obtained informed consent from each participant. We let each participant know that all the data we would obtain related to their participation would be anonymously used and collected. They were free to leave the study at any moment during their participation in case they disagreed with anything. The design of the study was approved by the ethics committee of the organization hosting the study. At first, whether they had used in the past an extended reality (XR) device or not, the participants received an explanation about how the HoloLens 2 device works. Then, they had some minutes to interact in a training scenario with some holograms in order to learn the different interactions. The most relevant interaction shown in this scene was with the end effector holographic representation. The users could try the best way for them to pick and move holograms around the interactive scene and, thus, reduce the impact of being naïve using this technology. We decided to carry out this user study in an industrial environment in order to know also the behavior of the device under noisy conditions where the voice detection could be affected, understanding the participant’s voice. After testing the MR device, the users started the main part of the study. In the first scene, the users had to define the desired pose of the end effector with the sphere with an arrow. The target goal was marked with a small virtual red cube. This red cube emitted a spatial sound in order to help the Figure 6. View of the RViz application showing the current robotic arm pose and the workspace. 4752 A. RIVERA-PINTO ET AL.
participant to locate it easier in the virtual scene. As explained before, the user grabbed the spherical hologram by using the hand and moved and rotated it to perform the pose definition part of the task. Once the final pose was defined, the user planned the trajectory to be performed and had the robot execute it by saying the voice command. The participant completed the execution of the task by saying the voice command “finish”. After that, the red cube moved to a new target point for the next repetition of the task. During the execution, we saved the time spent performing the task, which part of that time was moving the sphere, and the distance from the user-defined pose to the target. The participants performed this task three times and, after each execution, they answered to two questions. Both made use of the same 7-point Likert scale. The questions asked were the following: SEQ: Overall, how difficult or easy you find this task? How satisfied are you with the result? In the first question, the higher the value, the easier the task is for them. In the second question, the higher the value, the more satisfied they are with the result. The first question was the ’Single Ease Question’(SEQ) which asks about the difficulty of the task. With the same format as the SEQ, questionnaire, we added a second question about the user’s satisfaction with the result obtained. Later, in the second scene, they had to define the digital twin’s trajectory from the initial pose to the target moving the sphere and thus, setting different waypoints in the space. As in the previous scene, the same values were recorded. The users also answered the same questions after each execution. Once they finished the part they had to interact with the application, they completed the system usability scale (SUS) questionnaire, consisting of 10 statements for which participants had to express their degree of agreement with in a 5point scale. Finally, we conducted a semi-structured interview in which we asked participants to elaborate about their experience interacting with the system during the study. The following video (https://youtu.be/Ee8JChQE6zk) shows a walkthrough of what the user study consisted in. 4.2. Analysis of the results We analyzed the collected data from the participants and the results obtained from the study. From the 14 participants that took part in the study, 8 were male and 6 female. Most of them were in the age range of 26–30 (11 participants), 2 were in the age range of 31–35 and 1 was the range of 41–45. All the participants had completed their bachelor’s degree in the following expertise areas: computer science, electronics, automation and industrial engineering. None of them had tried a MR device before, but 4 had tested VR glasses for gaming. 4.2.1. Task performance We monitored the time (in seconds) the participants spent moving the sphere as well as the time dedicated to planning and executing the proposed trajectory and the distance gap between the target position and final pose of the robot hologram produced during the study. The data about time spent by the users to move the end effector representation (data from three executions by each participant for each design version) is reported in Figure 7. The mean time to move the sphere in case of defining the whole trajectory (M ¼10.2s, CI 95% ¼[8.2, 12.2]) was lower than when they just defined the end point (M ¼12.41s, CI 95% ¼[11.22, 13.6]). Related to the time concerning to the period spent thinking on which movement is necessary, the time required by ‘MoveIt!’to compute the path to perform the movement, the execution time and the voice commands required to perform the task, we obtained the results in Figure 8. In this case, the mean is lower when the user defines the goal only (M ¼28.39s, CI 95% ¼[24.84, 31.94]) by less than 3 seconds compared to the time registered from the trajectory executions (M ¼30.87s, CI 95% ¼[25.82, 35.92]). Analyzing the gap (in millimeters) between the target position and the ending location of the robot’s end effector (Figure 9), we got that the mean was lower defining the target point (M ¼10.988 mm, CI 95% ¼[9.58, 12.4]) than defining the trajectory (M ¼15.036 mm, CI 95% ¼[13.13,16.94]). Figure 7. Mean and 95% confidence interval of the time (in seconds) spent by users moving the sphere to the target position. Figure 8. Part of the execution time mean and 95% confidence interval of time (in seconds) that participants spent without manipulating the hologram. INTERNATIONAL JOURNAL OF HUMAN–COMPUTER INTERACTION 4753
This difference might have been caused by the time difference performing the task in both programming methods. 4.2.2. Questionnaire data With regard to the SEQ questionnaire assessing task difficulty in a 1 to 7 Likert scale, among all the answers from the users, all recorded responses for either version of the interface were in the levels 6 and 7 of the scale, indicating that in both cases participants thought that the task was very easy to complete. Noticeably, the 14 participants considered slightly easier to execute the task defining the ending point (6.86) than drawing the full trajectory (6.71). The answers to the satisfaction with the result of the task question showed high values that indicates that they felt satisfied with the results obtained after performing the task. The participants reported that both tasks were equally easy (6.52). The last questionnaire the users filled is System Usability Scale (SUS). SUS is a tool used to measure the usability of an object, device or application. The resultant score obtained from the users was 87.68. This score in a acceptability scale is considered as acceptable and in a Bangor’s grade scale Bangoret al. (2009) as ‘B’, 2.32 points below the highest mark, which places the application over the 98% percentile. The questions 4 and 10 are related to the application learnability. The mean obtained from these questions in a scale from 0 (worse) to 4 (better) were 3.36 and 3.5 respectively. 4.2.3. Semi-structured interview As the final stage in the user study, a semi-structured interview was conducted with each participant. The goal of this interview was to gather subjective information and comments that could help to interpret their responses. An experimenter held the interview as a conversation in which participants were asked to provide their opinion about various aspects regarding their interaction with the system and about the system itself. Specifically, participants were encouraged to comment on and elaborate on the following aspects: The system and the application as a method to program robots. Their subjective experience of interacting with the AR system. Ease of interaction with the holograms. Ease of learning to use HMD device and AR application. Aspects that were missing or could be improved. Any additional comments. A general comment from the participants was that the application was easy to use and the interaction was intuitive. The learning curve was very fast. The brief tutorial received in the introduction to the study was reported to be sufficient to then interact with the scene during the study. The main problem for the participants was the first contact with the holograms when grabbing them was not done easily before a few trials of practicing. The voice commands and the way to move the holograms fit well to the task in participants’ opinion, helping to accelerate the robot programming process. Participants also agreed that it was possible to perform this task for users that did not have the skills nor the knowledge to program a robotic arm with current state-of-the-art techniques (e.g., using a teach pendant). As using the application is possible to set faster the target pose than doing it physically, it is feasible to test different goal poses to plan and choose the most appropriate plan that satisfies the needs of the user and improving the time performance. However, some of the users questioned the accuracy obtained through these programming methods. Depending on the task, some thought that this solution in this current form might not be sufficiently accurate to satisfy the requirements in a real scenario. One participant suggested testing the accuracy by performing the programming task with the application in a real scene to measure the resulting error. Other participants proposed a functionality to allow zooming into the target to increase the pose accuracy. Related to the programming method by drawing the trajectory, some participants suggested allowing modification of the pose of the intermediate points instead of needing to reset the entire path. In the same vein, the concatenation of actions could be interesting, for instance, in order to interact with the tool handled by the robotic arm. Another alternative proposed was to interpolate all the defined points so as to smooth out the defined trajectory. The voice commands were easy to remember. However, under noisy conditions (as was the case in the study, which was conducted in an industrial laboratory environment), some participants had to repeat the same voice command more than once before it was understood by the system. Some participants also complained that the system did not produce enough feedback informing about its status, such as while the system was computing a trajectory or executing the planned path. Regarding the AR HMD device, some participants pointed out that, even if HoloLens 2 suits the task well, its field of view is not yet big enough to provide full user comfort. Linked to this limitation, some participants had difficulties picking the hologram when it was located at a low height. Figure 9. Mean and 95% confidence interval of distance error positioning the end effector’s hologram in the goal position. 4754 A. RIVERA-PINTO ET AL.
5. Discussion This section provides a discussion of the user study data analysis presented in the previous section. The data related to the user performance in the task of teaching an action to the robot, suggests that both designs provided similar average performance in a range between 10.2 and 12.4 seconds on average. While no direct comparison was performed with teaching a physical robot by demonstration, the expected time to complete such task would be expected to be comparable. Such a direct comparison should be performed in the future, with a fully implemented MR robot programming interface. Regarding the positioning accuracy obtained, we observed that positioning error was in the approximate range between 11 and 15 mm. To put this in perspective, we compare it with data from another study, published in Arevalo Arboleda et al (2021). They reportedly achieved very similar accuracy values, in the approximate range between 11 and 13.5 mm, with a system designed for a different mixed reality application. Whether such accuracy is sufficiently high will depend on the nature of the robotic application being programmed. Strategies to improve it could include the use of computer vision techniques to reach the objective. As a separate strategy (and as the participants in the study suggested), the current end effector representation could be replaced with a different hologram design that helped the user achieve more precise positioning. As for the usability and user experience obtained while programming the robot with this interface, all the participants found the application easy to use, intuitive, and useful for the task of robot programming. They felt comfortable using the application after the training session with the device, as well as satisfied with the result obtained, as reflected in the SEQ questionnaire data. Many participants reported that the possibility of having control over changing the pose of individual points of the trajectory that they were defining could improve the programming experience. In this way, to make any modifications to a trajectory, they would not need to draw the whole trajectory again. Instead, they could just make changes in specific intermediate points to tweak and improve the proposed trajectory. Some participants also suggested granting the user control over when each intermediate dot of a trajectory is defined in space. In the current implementation, as the user moves the hologram of the end effector, a trail of the trajectory followed by the end effector is drawn behind by adding one point in space at a regular 3 Hz rate, at the precise position where the end effector was at each interval. In the proposed new design, no trail would be automatically drawn as the end effector moved. Instead, the user would decide in which positions of the trajectory that she was following should a passing point be defined. In relation to the HoloLens 2 device, some users complained about its narrow field of view. Although this specification of the device has improved since its first version, it seems not to be sufficiently broad to work with holograms as big as the ones rendered in the study. As an exception to this general opinion, one of the users (who had experience using VR devices) declared that they preferred having less field of view than in VR glasses to avoid getting dizzy after using the device continuously for an extended period of time. According to the results of the SUS questionnaire, the usability of the interactive system was suitable for the task in the study. From among the 10 questions in SUS, 2 of them (questions 4 and 10) indicate learnability of the system. Such learnability level was high for all participants, suggesting that the requirement for users to easily learn and work with the application was satisfied. Regarding other requirements stated in section 3, results obtained from the study and our own assessment of the system suggest that this first design and implementation address them all. 6. Conclusions and future work In this paper, we describe the design and implementation of a mixed-reality multimodal interface with which a user can teach new trajectories to a robot arm by holding and moving the hologram of its digital twin by hand, in a mixed-reality scenario. This interactive codeless technique to program the robot is not only similar to teaching the physical version of the robot arm by demonstration, but it also uses the same software architecture for the calculation of trajectories and subsequent generation of the program. It is thus a technique intended to program the robot by interacting with its physical or holographic digital twin version alike. Our purpose was to develop a user-friendly HRI interface to ease the robot programming experience, allowing an inexperienced worker to execute a task that requires specific skills and experience. The advantages of this design compared to the state-ofthe-art remain in the mix of direct manipulation of the digital twin by an interaction module based on hand tracking and programming a robotic arm by defining the end effector’s final pose instead of defining individually the end pose’s position and rotation (the interface proposed supports the integrality of the interaction). We have described the application design process and its implementation. We have then reported a user study (n ¼14) to validate the fundamentals of the present design approach, as well as the implementation of this interface. We have assessed the outcome of the study with quantitative and qualitative metrics of efficiency and precision, as well as by observing and enquiring about its usability and obtained user experience in a task of programming by demonstration new robot trajectories, solely interacting with the holographic digital twin of the robot in a mixed reality scenario. Findings from the study suggest that it may be easy and feasible for a naïve user to program new trajectories for a robotic arm through the use of this interface and by solely interacting with its holographic digital twin. Quantitative results suggest that the time required to program a new INTERNATIONAL JOURNAL OF HUMAN–COMPUTER INTERACTION 4755
trajectory (from dragging the hologram along it to validating the trajectory calculated by the robot’s operating system) may be acceptably short. The average distance error to the target position may also be small according to results from the study, although the sufficiency of the precision observed will depend on the requirements of each specific robotic application. The paper discusses possible approaches to improve precision in future versions of the interface. Some such approaches present the potential to grant the user with additional control over the trajectory teaching process, specifically a technique in which the user can choose explicitly which points in space to include in the trajectory. Modifying the interaction in this way should not only result in quicker programming (due to less repetitions of the whole trajectory stroke), but most importantly in an improved user experience due to the enhanced control over the teaching process. Future work will not only investigate these alternative versions of the interface, but will also validate the whole robot programming approach by executing in the physical twin the program created by interacting with its holographic digital twin. Disclosure statement No potential conflict of interest was reported by the author(s). Funding This research received funding from The Centre for the Development of Industrial Technology (CDTI). 5RRed Cervera de Tecnolog ıas rob oticas en fabricaci on inteligente (CER-20211007). ORCID Andoni Rivera-Pinto http://orcid.org/0000-0001-8550-5312 Johan Kildal http://orcid.org/0000-0002-0630-7260 Elena Lazkano http://orcid.org/0000-0002-7653-6210 References Arevalo Arboleda, S., R€ ucker, F., Dierks, T., & Gerken, J. (2021). Assisting manipulation and grasping in robot teleoperation with augmented reality visual cues [Paper presentation]. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. https://doi.org/10.1145/3411764.3445398 Bambu^ sek, D., Materna, Z., Kapinus, M., Beran, V., & Smr z, P. (2019). Combining interactive spatial augmented reality with head-mounted display for end-user collaborative robot programming [Paper presentation]. 2019 28th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN) (pp. 1–8). https:// doi.org/10.1109/RO-MAN46459.2019.8956315 Bangor, A., Kortum, P., & Miller, J. (2009). Determining what individual SUS scores mean: Adding an adjective rating scale. Journal of Usability Studies,4(3), 114–123. https://dl.acm.org/doi/10.5555/ 2835587.2835589 Barentine, C., McNay, A., Pfaffenbichler, R., Smith, A., Rosen, E., & Phillips, E. (2021). A VR Teleoperation Suite with Manipulation Assist [Paper presentation]. Companion of the 2021 ACM/IEEE International Conference on Human-Robot Interaction (pp. 442–446), Boulder, CO, USA. https://doi.org/10.1145/3434074.3447210 Bischoff, M. (2019). ROS#. https://github.com/siemens/ros-sharp/. https://github.com/siemens/ros-sharp/ Bischoff, M., Vollenweider, E. (2020). ROS#for HoloLens.https:// github.com/EricVoll/ros-sharp/. Bolano, G., Roennau, A., Dillmann, R., & Groz, A. (2020). Virtual reality for offline programming of robotic applications with online teaching methods [Paper presentation]. 2020 17th International Conference on Ubiquitous Robots (UR) (pp. 625–630). https://doi. org/10.1109/UR49135.2020.9144806 Burghardt, A., Szybicki, D., Gierlak, P., Kurc, K., Pietru s, P., & Cygan, R. (2020). Programming of industrial robots using virtual reality and digital twins. Applied Sciences,10(2), 486. https://doi.org/10.3390/ app10020486 Chan, W. P., Sakr, M., Quintero, C. P., Croft, E., & Van der Loos, H. F. M. (2020). Towards a multimodal system combining augmented reality and electromyography for robot trajectory programming and execution [Paper presentation]. 2020 29th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN) (pp. 419–424). https://doi.org/10.1109/ RO-MAN47096.2020.9223526 Cruz Ulloa, C., Dom ınguez, D., Del Cerro, J., & Barrientos, A. (2022). A mixed-reality tele-operation method for high-level control of a legged-manipulator robot. Sensors,22(21), 8146. https://doi.org/10. 3390/s22218146 De Pace, F., Gorjup, G., Bai, H., Sanna, A., Liarokapis, M., & Billinghurst, M. (2020). Assessing the suitability and effectiveness of mixed reality interfaces for accurate robot teleoperation [Paper presentation]. 26th ACM Symposium on Virtual Reality Software and Technology.https://doi.org/10.1145/3385956.3422092 Delmerico, J., Poranne, R., Bogo, F., Oleynikova, H., Vollenweider, E., Coros, S., Nieto, J., & Pollefeys, M. (2022). Spatial computing and intuitive interaction: bringing mixed reality and robotics together. IEEE Robotics & Automation Magazine,29(1), 45–57. https://doi. org/10.1109/MRA.2021.3138384 Fang, H., Ong, S.-K., & Nee, A. Y. (2014). Novel AR-based interface for human-robot interaction and visualization. Advances in Manufacturing,2(4), 275–288. https://doi.org/10.1007/s40436-0140087-9 Gadre, S. Y., Rosen, E., Chien, G., Phillips, E., Tellex, S., & Konidaris, G. (2019). End-user robot programming using mixed reality [Paper presentation]. 2019 International Conference on Robotics and Automation (ICRA) (pp. 2707–2713). https://doi.org/10.1109/ICRA. 2019.8793988 Hietanen, A., Pieters, R., Lanz, M., Latokartano, J., & K€ am€ ar€ ainen, J.-K. (2020). AR-based interaction for human-robot collaborative manufacturing. Robotics and Computer-Integrated Manufacturing, 63(June), 101891. https://doi.org/10.1016/j.rcim.2019.101891 Jacob, R. J., Sibert, L. E., McFarlane, D. C., & Mullen, M. P. Jr, (1994). Integrality and separability of input devices. ACM Transactions on Computer-Human Interaction,1(1), 3–26. https://doi.org/10.1145/ 174630.174631 Kildal, J., Tellaeche, A., Fern andez, I., & Maurtua, I. (2018). Potential users’key concerns and expectations for the adoption of cobots. Procedia CIRP,72(2018), 21–26. https://doi.org/10.1016/j.procir.2018. 03.104 LeMasurier, G., Allspaw, J., & Yanco, H. A. (2021). Semi-autonomous planning and visualization in virtual reality. LeMasurier, G., Allspaw, J., Wonsick, M., Tukpah, J., Padir, T., Yanco, H., Phillips, E. (2022). Designing a user study for comparing 2D and VR human-in-the-loop robot planning interfaces.5th International Workshop on Virtual, Augmented, and Mixed Reality for HRI.https://openreview.net/forum?id=HNeeXUiQjkc Lotsaris, K., Fousekis, N., Koukas, S., Aivaliotis, S., Kousi, N., Michalos, G., & Makris, S. (2021). Augmented reality (AR) based framework for supporting human workers in flexible manufacturing. Procedia CIRP, 96(2021), 301–306. https://doi.org/10.1016/j.procir.2021.01.091 Mal y, I., Sedl a cek, D., & Leit~ ao, P. (2016). Augmented reality experiments with industrial robot in industry 4.0 environment [Paper presentation]. 2016 IEEE 14th International Conference on Industrial Informatics (INDIN) (pp. 176–181). https://doi.org/10.1109/INDIN. 2016.7819154 4756 A. RIVERA-PINTO ET AL.
Masurovsky, A., Chojecki, P., Runde, D., Lafci, M., Przewozny, D., & Gaebler, M. (2020). Controller-free hand tracking for grab-and-place tasks in immersive virtual reality: Design elements and their empirical study. Multimodal Technologies and Interaction,4(4), 91. https:// doi.org/10.3390/mti4040091 Neves, J., Serrario, D., & Pires, J. N. (2018). Application of mixed reality in robot manipulator programming. Industrial Robot: The International Journal of Robotics Research and Application,45(6), 784–793. https://doi.org/10.1108/IR-06-2018-0120 Ong, S. K., Yew, A., Thanigaivel, N., & Nee, A. Y. (2020). Augmented reality-assisted robot programming system for industrial applications. Robotics and Computer-Integrated Manufacturing, 61(February), 101820. https://doi.org/10.1016/j.rcim.2019.101820 Ong, S., Nee, A., Yew, A., & Thanigaivel, N. (2020). AR-assisted robot welding programming.Advances in Manufacturing,8(1), 40–48. https://doi.org/10.1007/s40436-019-00283-0 Ostanin, M., Mikhel, S., Evlampiev, A., Skvortsova, V., & Klimchik, A. (2020). Human-robot interaction for robotic manipulator programming in mixed reality [Paper presentation]. 2020 IEEE International Conference on Robotics and Automation (ICRA) (pp. 2805–2811). https://doi.org/10.1109/ICRA40945.2020.9196965 Pinto, A. R., Kildal, J., & Lazkano, E. (2020). Multimodal mixed reality impact on a hand guiding task with a holographic cobot. Multimodal Technologies and Interaction,4(4), 78. https://doi.org/10. 3390/mti4040078 Puljiz, D., & Hein, B. (2019). Concepts for end-to-end augmented reality based human-robot interaction systems. https://arxiv.org/abs/ 1910.04494 Puljiz, D., St€ ohr, E., Riesterer, K. S., Hein, B., & Kr€ oger, T. (2019). General hand guidance framework using Microsoft HoloLens. 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (pp. 5185–5190). https://doi.org/10.1109/IROS40897.2019.8967649 Puljiz, D., Zhou, B., Ma, K., & Hein, B. (2021). HAIR: Head-mounted AR Intention Recognition. https://arxiv.org/abs/2102.11162 Quintero, C. P., Li, S., Pan, M. K., Chan, W. P., Machiel Van der Loos, H. F., & Croft, E. (2018). Robot programming through augmented trajectories in augmented reality [Paper presentation]. 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (pp. 1838–1844). https://doi.org/10.1109/IROS.2018.8593700 San Martin, A., & Kildal, J. (2021). Audio-visual mixed reality representation of hazard zones for safe pedestrian navigation of a space. Interacting with Computers,33(3), 311–329. https://doi.org/10.1093/ iwc/iwab028 Su, Y., Chen, X., Zhou, T., Pretty, C., & Chase, G. (2022). Mixed realityintegrated 3D/2D vision mapping for intuitive teleoperation of mobile manipulator.Robotics and Computer-Integrated Manufacturing, 77(October), 102332. https://doi.org/10.1016/j.rcim.2022.102332 Walker, M., Hedayati, H., Lee, J., & Szafir, D. (2018). Communicating robot motion intent with augmented reality. Proceedings of the 2018 ACM/IEEE International Conference on Human-Robot Interaction (pp. 316–324). https://doi.org/10.1145/3171221.3171253 Xue, C., Qiao, Y., & Murray, N. (2020). Enabling human-robot-interaction for remote robotic operation via augmented reality [Paper presentation]. 2020 IEEE 21st International Symposium on “A World of Wireless, Mobile and Multimedia Networks”(WoWMoM) (pp. 194–196). https://doi.org/10.1109/WoWMoM49955.2020.00046 Zhou, Z., Xiong, R., Wang, Y., & Zhang, J. (2020). Advanced robot programming: A review. Current Robotics Reports,1(4), 251–258. https://doi.org/10.1007/s43154-020-00023-4 About the authors Andoni Rivera Pinto is a researcher at Tekniker and PhD candidate at University of the Basque Country (UPV/EHU) since 2019. He has a MSc in Computational Engineering and Smart Systems (UPV/EHU) and researches human-computer interaction in the context of robotics by using VR and AR devices. Johan Kildal specializes in interaction design and research for accessibility, mobility and Human-Robot Collaboration. He was Principal Researcher for eight years at Nokia Research (Finland), filing eight patents for innovative interactions. Currently, he researches at TEKNIKER about collaboration with robots in industry. Elena Lazkano is Associate Professor in the Computer Science and AI Department, University of the Basque Country (UPV/EHU). MSc in AI (KUL, Belgium) and PhD in Computer Sciences (UPV/EHU). She leads the robotics branch of the RSAIT research group, as researcher in the field of intelligent robotics. INTERNATIONAL JOURNAL OF HUMAN–COMPUTER INTERACTION 4757
9 Collaborative Robot Teleoperation in Mixed Reality Environment for Inspection Tasks Authors Andoni Rivera Pinto, Cristina Aceta, Johan Kildal, Izaskun Fernández y Elena Lazkano Publisher MIT Press Journal PRESENCE-VIRTUAL AND AUGMENTED REALITY Year 2024 Quartile Q4 DOI https://doi.org/10.1162/PRES_a_00439 111
Rivera et al. 49 the domain knowledge consists of the actions to be executed by the target system (in this case, the robot) in a system-readable format and the words that the user can use to trigger those actions. Dialogue-related knowledge includes the interactions that the system may have with the user in a parametric form, and the implications of different situations that may arise in the interaction (e.g., what to do when the user does not provide all the necessary information to execute a specific action). One of the main advantages of the use of semantic technologies in KIDE4I is that it allows an easy implementation and adaptation to different industrial use cases, as one of its main premises is reuse. In this case, most of the necessary instances for dialogue management are reused from an initial KIDE4I use case (Aceta et al., 2022) and, as for the domain knowledge, KIDE4I also provides resources for the exploitation of existing lexical databases for instantiation. As noted above, KIDE4I is a modular framework. That is, in order to obtain an action that is executable by the target system, which is triggered by a user command, different components are involved. Once a transcribed user command is received by the dialogue system, the first step for its interpretation is to extract its key components (i.e., the verb—which usually triggers the target action—and its complementary information—which usually determines the specific details of the action, such as the target robot component for that action), which is performed by the key element extraction module. Once the key elements are extracted, this information is consulted against the semantic repository, which will determine which is the action to be executed by the robot and any additional information. All this information flow is orchestrated by the dialogue manager, which is also in charge of following the modelling in the dialogue knowledge in the semantic repository, determining the next step in the dialogue process. Finally, the polarity interpreter is responsible for determining the polarity of user commands (i.e., whether they mean a confirmation or a negation) when the dialogue manager determines that a specific system request requires a polar answer (i.e., yes or no). In this particular case, KIDE4I’s implementation has been named KIDE4Inspection. Similarly to KIDE4I adaptations, KIDE4Inspection (Aceta et al., 2022) follows a simple flow where the first step is to determine the action to be executed according to the user command. If that task is not successful, KIDE4Inspection will trigger a system response informing the user that the action could not be interpreted, followed by a request to reformulate their command. However, if the system is able to determine the action, it will proceed, if required by the interpreted action, to assign the corresponding arguments of that action (e.g., the target towards which the action will be executed). Again, if the system is not able to obtain the required action arguments, it will trigger a message informing the user and requesting to provide such information. If the information provided by the user is correct (according to what is modelled in the ontology), the flow will continue. Otherwise, the system will give a second chance to the user to provide the arguments by asking the information again. Finally, the output will be the action and the relevant information, in a format that can be executable by the HoloLens application, i.e., the intermediary between the user and the robot. 3.4 Robotic Arm Programming Procedure As mentioned before, the task presented in this paper involves the inspection of the nerves of an aileron of 2.44 m in length. Each end of the aileron was sectioned and the longitudinal nerves were visible, going from end to end inside the aileron. On one end of the aileron, the robotic arm must position a tool with a mirror over one of the nerves. On the other end, the user must hold a laser measuring device over the same nerve. If the laser beam reflects off the mirror, it indicates that the nerve is not blocked (see Figure 3). The flowchart in Figure 4 outlines the procedures followed for the programming task. The gray nodes represent the beginning and the ending of the program, the yellow nodes represent actions the user performs in the application, and the blue diamond boxes are decisions to be taken by the user. The first step of the task involves determining the relative pose between the robot and the aileron. The
50 PRESENCE: VOLUME 34 Figure 3. On the left, an image depicting the layout for the collaborative task where the user is on the opposite side of the aileron from where the robotic arm is located. On the right, an image showing the same layout with a person interacting with the holographic tool. Figure 4. Flowchart of user’s interactions with the MR application. aileron is at the same time the target and an obstacle to be avoided. A calibration plate is placed on one of the support bars of the aileron in a known position relative to the aileron. Additionally, an RGB camera is positioned on board next to the robot’s end effector. The user moves the real robotic arm manually to align it with the calibration plate while the camera captures the scene. Using the HALCON machine vision software, a ROS node is in charge of capturing an image from the
Rivera et al. 51 onboard camera (IDS UI-5240CP) facing the calibration table (100 mm ×100 mm size) and computing the pose of the calibration plate relative to the camera. With this information, along with the camera’s known location relative to the base of the robotic arm and the calibration plate’s known location relative to the aileron (established through an external calibration process), the system can determine the position of the aileron relative to the robotic arm’s base. According to the results reported by the calibration of the camera, the committed error in the acquisition of the calibration plate’s position is less than 1 mm. It is important that this value is as low as possible because the nerves of the aileron are quite narrow for the end effector tool we mounted on the robotic arm’s end effector, and the tool could hit the nerves. For the calibration of the aileron with respect to the calibration plate, we used a Photoneo 3D camera which provides a 2D image and a pointcloud. Through the 2D image, we obtained the pose of the calibration plate with respect to the camera. Besides, by using the point cloud, we matched the 3D model of the aileron to obtain the pose of the aileron with respect to the camera. With this information we got the pose of the aileron with respect to the calibration plate. The expected error in this calibration was below 1.5 mm. The calibration of the onboard camera and the robotic arm was made by capturing multiple images of a known calibration pattern from different angles and distances. We used an IDS UI-5240CP camera with which we obtained an error below a millimeter. The information of the pose of the aileron with respect to the onboard camera, alongside the 3D model of the aileron, are loaded into “MoveIt!” to compute the trajectories while avoiding collisions. After completing the scene setup, the user launches the MR application on the HMD. The holographic digital twin of the robotic arm, along with the virtual aileron and a representation of the tool attached to the end effector, is displayed for the user in the MR scene. Using the dialogue system, the users can order a variety of actions to the system. By issuing the voice command “KUKA,” the dialogue system is activated. A validation sound is played to confirm that the voice command has been detected, and a text appears in the user’s field of view, prompting them with the question “Which action do you want to perform?” The user answers by speaking aloud and a silence of 1 second is interpreted by the system as the end of the order. The captured voice is processed, and the transcription is displayed in front of the user as well as conveyed through audio. Once the phrase is transcribed, it is sent to the semantics-based task-oriented dialogue system that returns the action through a REST API communication. These are the different actions the dialogue can process from users’ input: Calibrate: Starts the calibration process in order to place the aileron in the correct position with respect to the robotic arm. (Example phrase: “Calibrate the relation between the robot and the aileron.”) Relocate: Unlocks the holographic scene allowing the user to move it in the surrounding space as well as rotating it. (Example phrase: “Relocate the scene.”) Lock: Locks the holographic scene preventing its movement from the user. (Example phrase: “Lock the scene.”) Plan: Plans a trajectory for the robotic arm from the current configuration to the target pose. (Example phrase: “Plan the trajectory to the target position.”) Execute: In case it exists and it is valid, executes the planned robot trajectory. (Example phrase: “Execute the planned path.”) Place: Sets the pose of the end effector in one of the nerves (that has to be defined by the user). (Example phrase: “Place the end effector on the nerve A.”) At first instance before starting the programming procedure, the user can move and lock the position of the whole holographic scene keeping the calibrated distances between the different holograms. This action will allow the user to place the scene in a more comfortable zone to interact with it. In addition, through voice
52 PRESENCE: VOLUME 34 Figure 5. User’s point of view while the system is planning the trajectory. interaction, the user can retrieve the previously calculated pose of the aileron with respect to the base of the robot and display that relative position in the virtual environment. The first step of the programming procedure consists in setting the target pose of the robotic arm’s end effector. This step involves placing the tool representation in the desired real-world pose. To achieve it, users grab the hologram with their hand and move it to the target position and orientation. They then release the hologram to finalize the positioning. In order to define the orientation, the users must rotate their wrist so the hologram rotates by pivoting around the grabbed point of the object. The first action in robot trajectory programming is planning. The system attempts to find a path from the current pose of the robot to the target pose defined by the holographic tool. When this action is detected, a semi-transparent white hologram of the robotic arm appears over the digital twin until the trajectory plan is resolved by “MoveIt!” (see Figure 5). Once the system obtains a result, if it returns a valid trajectory, the white hologram will be colored green and will display a preview of the planned trajectory (see Figure 6). Otherwise, the hologram will be colored red and remain stationary (see Figure 7). If the planned trajectory is satisfactory, the user can command the robot to execute the plan by means of the dialogue system. The preview of the trajectory will then disappear, and the hologram of the digital twin will start moving according to the plan, simultaneously with the real robotic arm. The planning and execution steps can be repeated as many times as necessary to reach the desired target pose. The final step of the task involves refined positioning. This step is specific to the proposed task, where the target poses are predefined with respect to the worked object (i.e., the user who uses the application does not define this pose). In the context of aileron inspection, we have predetermined one pose for each nerve, aligning the tool to enable sensor readings from the other side. The user does not have to move the holographic tool because this was already done in the previous step. Interaction with the system is limited to the dialogue and the tool will be placed automatically in the requested position and a trajectory plan will also be requested. As for the collaborative task, after the user moves the real robotic arm to one of the aileron nerves by interacting with the application, the user places the real laser on the opposite side of the aileron and checks the status of the inspected nerve. The laser has two lights that provide feedback to the user. The first indicates whether the
Rivera et al. 53 Figure 6. User’s point of view while the system was able to compute a valid trajectory (green hologram). Figure 7. User’s point of view while the system was not able to compute a valid trajectory (red hologram). laser is on or off. The second indicates whether or not the laser beam has returned to the beam emitter. 4 User Study We conducted a user study to validate and assess how the proposed MR design aligns with the robot programming task from the perspective of users who have no prior experience with MR. In this section, we will describe the procedure that was followed in the user study, which involved 13 participants. We will also discuss the data that was collected during and after the user study and present an analysis of the results obtained. 4.1 Study Procedure All participants provided their consent before starting the user study, allowing for the collection and
54 PRESENCE: VOLUME 34 anonymous analysis of data. They were also informed that they could withdraw from the study at any time without providing an explanation. The ethics committee of the hosting organization approved the study design and provided oversight throughout the process. The objective of this user study was to assess the feasibility of programming a real robotic arm in an actual task using its holographic representation, as well as to evaluate participants’ perceptions of the MR application in terms of its interest, helpfulness, and user-friendliness. During the study, participants were tasked with programming the robotic arm, following the procedure described in the previous section, for two of the four available nerves. The task involved aligning the tool with the nerve, planning a valid trajectory, and executing it. Participants started with the first nerve at the top and then progressed to the third nerve at the bottom. Additionally, participants had to plan a new trajectory to position the tool inside the aileron according to predefined poses. Users were asked to place the tool first aligned with the aileron so that the second movement would be in a straight line within the nerve. The second movement was exclusively made by using the voice input. Once the target was reached by the robotic arm, participants were required to place the laser on the other side of the same nerve and validate the successful return of the laser beam. Regarding the recorded data during the execution, we collected (as dependent variables) the time spent on each execution, the time spent manipulating the holographic tool, the number of times the holographic tool was moved per execution, and the sentences used with the dialogue system to perform various actions. Following the completion of the study, participants were asked to complete the System Usability Scale (SUS) questionnaire (Brooke, 1995), extended with three additional statements related to the dialogue system. •The interpretation of the voice commands is accurate. •The interaction with the system through voice commands is efficient. •I consider that interacting with the system through voice commands is useful. Furthermore, participants were given the opportunity to answer two open-ended questions about what they liked the most and what they disliked the most about the application. Prior to the aforementioned procedure, participants were provided with approximately seven to ten minutes to familiarize themselves with a similar environment. This training session aimed to teach them how to interact with a hologram and reduce any significant skill gap between the two executions they performed during the study. 4.2 Analysis of the Results The user study was carried out with 13 people who gave their consent to collect data anonymously from their executions while they were performing the task. The average age of the participants was 22.46 (SD = 1.941). All participants were studying for a bachelor’s degree in mechatronics. This was the first experience with a MR device for all the participants. We collected time related to the time spent performing the whole task, the time spent moving the holographic tool, the time spent using the dialogue system, the time while the user was doing nothing, the number of times the user moved the holographic tool, and the dialogue interaction success ratio. 4.2.1 Total Task Time. This metric measures the whole time spent planning the movement to both nerves and also executing the trajectories. The data reports the following results [M =337.4 s; CI95% =[314.0, 360.8]]. 4.2.2 Time Moving the Tool. We recorded the time users spent moving the tool to get it into the pose they wanted. The count started when the users grabbed the tool and stopped when they dropped it. The results for the whole task are the following: [M =46.7 s; CI95% = [41, 7, 51.7]]. 4.2.3 Voice Interaction Time. This metric reports the time that the user spent producing voice commands and having them understood by the dialogue system (which,
Rivera et al. 55 in some cases, involved repeating the same voice command more than once): [M =125.2 s; CI95% =[110.9, 139.3]]. 4.2.4 Time Stopped. This metric computes the total time users spent not performing any active action. They were typically waiting for a step in the process to be completed, or progress was halted because the action they requested could not be executed: [M =165.5 s; CI95% =[148.1, 182.9]]. 4.2.5 Number of Tool Movements. The data reported above represents the number of times the user grabbed and dropped the sphere used to indicate a target position in space: [M =8.77; CI95% =[7.8, 9.74]]. 4.2.6 Dialogue Interaction Success Ratio. The dialogue interaction success ratio was calculated by dividing the total sentences the sematics-based dialogue system could process by the total processed sentences. The value measured for this metric was 88.37%. The sentences used by the participants were similar to the sentences shown in Section 3.4. Those phrases were given as examples to the participants, although they were free to use different sentences with equivalent meaning. Failures in the speech to text service were due to the background noise of the industrial environment. In addition, one reason participants were unsuccessful in their attempt to give a command through the dialog system was that they did not complete the sentence or got stuck mid-sentence, so information was missing and the system could not perform an action. 4.3 Analysis of Questionnaire Results The mean SUS score obtained was 79.23. According to Bangor’s grade scale (Bangor et al., 2009), usability of the application is acceptable, and in the grade scale it corresponds to a “C,” although only 77 hundredths below a rating that would classify it as grade “B.” The distribution of the answers given to each question is reported in Figure 8. For the figure, as well as for computing the score, the polarity of the negative questions was inverted. Thus, higher scores indicate more positive answers. According to the answers for the additional three questions, reported in Figure 9, the results show that, from the users’ perspective, the interpretation of the voice commands was accurate in relation to the action they intended to request the system to perform (Q11). The effectiveness of voice commands was pointed out by 72.2% of the participants, while the remaining 27.8% considered it neither effective nor an obstacle to interaction (Q12). In addition, all participants considered that interacting with the system by using voice commands was useful (Q13). The participants’ answers to the open questions about what they liked and disliked the most about the system help interpret the results obtained from the questionnaire. The most relevant positive aspect of the application for the participants was the voice system (46% of the answers highlighted the pros of using a dialogue system for interacting with the application). The responses also highlighted the teleoperability choice of the system, preventing risks from exposure to the robot, that is, being able to operate a robot from a remote location, which also helps reduce risk from exposure to the robot. On the other hand, the main drawback participants found in the application was the difficulty to interact with the tool in its holographic representation (38% of the participants remarked on this aspect). In addition, participants noted that voice detection was not sufficiently accurate. It is likely that this lack of accuracy was caused by the background noise of the industrial setting (representative of the environmental conditions this system would be in a real production line). This was probably aggravated by the quality of the microphone in the HMD device, which should ideally be superior. 5 Discussion The results from the user study show that every participant was able to complete the task of programming a robotic arm by interacting with its holographic digital twin using the system described in this paper.
56 PRESENCE: VOLUME 34 Figure 8. Distribution of the answers from the users to the SUS questionnaire. Figure 9. Distribution of the answers from the questions related to the dialogue system. This result should be put in perspective by the fact that every participant was naive in the use of MR devices as well as in the programming of robotic arms. The results from the user study help characterize the performance obtained. Users finished the whole task in an average of 337.4 seconds (CI95% =[314.0, 360.8]), that is, 5–6 minutes in total. In all the cases, the task was successfully completed, and the user could finish the whole task. This gross total time included the programming of at least five poses for the robot, in each one with the interaction with the dialogue system, computation of trajectories, inspection and validation by the participants, plus execution.
Rivera et al. 57 From the total average time spent performing the task, the time users spent moving the tool to define a new pose accounts for only 13.85% of the total time (46.7 seconds on average, CI95% =[41.8, 51.7]). With an average of nearly nine interactions performed to define each new pose, each interaction took an average of around 5 seconds to perform. Considering the time values reported in somewhat similar studies in the literature (such as in Le et al., 2020), the time required by users to define a target pose is comparable to the time we observed in our study. They reported an average of 39.33 seconds to complete their task, which consisted of a drag-and-drop task. In our case, we recorded an average of 46.7 seconds as a sum of every instance the user moved the holographic tool representation during its positioning on the two nerves used in the user study. Regarding the time spent with voice interaction, this activity represented more than a third of the total task execution time (125.2 seconds). The voice interaction time started with the activation of the dialogue system by pronouncing the prompt voice command, and it lasted until an action was recognized. It is relevant to emphasize that in this part of the task execution, some users encountered difficulties to be understood by the system. This might be due to the presence of background noise in the laboratory (located in an industrial building with robots and machines operating inside cells), in which background noise was similar to the noise typically found in an industrial environment. The remaining time spent in task execution corresponded to periods in which the user was idle (e.g., while the system was computing or executing a part of the sequence), or time with non-productive actions (e.g., with the user trying to grab the holographic tool). The average time that the user was idle or nonproductive was 165.5 seconds (CI95% =[148.1, 182.9]). These 2.5–3 minutes accounted for almost half of the total execution time. Participants (who, as said, had no prior experience manipulating holograms) reported in some cases in the post study open question that they had found it difficult to grab the hologram of the tool at first, but that the little amount of practice from participating in the study already diminished such difficulties. Training and practice received prior to executing the experiment task lasted only 10 minutes per participant. It is likely to expect that devoting more time to the training part prior to the study would have reduced the time and number of attempts required to grab the hologram. As a result, it is likely that it would have also reduced the frustration experienced by some of the participants, and it might have improved their UX. Conversely, participant reports in the open question about the best aspects of the application suggest that the interaction with the system through a dialogue system is a good option, as it simplifies the process of performing the task. Altogether, participants found that the way to teleoperate the robot with this XR system was easy. They felt that more training will be beneficial beforehand to better control the input interface of the device as well as to know what options are available for dialog interaction. However, with the short time they were provided to learn about the use of the application, they felt they were sufficiently prepared to perform well. 6 Conclusions and Future Work In this paper, we described the design of an application with which it is possible to command a robotic arm through its holographic digital twin and we evaluated it in an aileron inspection task in an industrial environment. This design solves the challenge of teleoperating a robot over a distance by replicating next to the operator the remote robot and its relevant surrounding scenario. The solution proposed is independent of how far the robot is located from the operator. It allows the operator not to need to be next to the robotic arm to program its behavior. Because the operator does not need to displace repeatedly all the way to the location of the robot, the time required and the physical effort exerted by the operator to program the robot are drastically reduced. These reductions will be bigger the further the robot is in different production scenarios. With the purpose of improving the interaction with the scene, we developed and integrated in the XR application a semantic-based task-oriented dialogue system
58 PRESENCE: VOLUME 34 that allows a voice communication with the application by using natural language rather than basic commands. Apart from the two main advantages of MR applications for robot programming, which are not needed for programming skills and teleoperability, the user study results report a positive reception by users in using multimodal technology to program a robotic arm through its holographic digital twin. The use of more extended sentences instead of simple commands allowed the users to provide details of the action to the system when it was required. This capability in addition to the direct interaction with the holograms permitted a complete interaction with the virtual scenario. On the other hand, the feedback provided not only by the visual channel but also using the auditory channel helped the users to understand every time the status of the application and the teleoperation. A user study was carried out with 13 people who had never used an XR HMD. They all could complete the task of positioning the robotic arm’s end effector in the target nerves. Findings from the study suggest that the application enables inexperienced users to program the robot by using XR technologies. The success of the task performance in all cases validates that the technology used is compatible with the proposed application. Regarding future work, also linked to the feedback received in the user study, we identified that the interaction with the holographic tool needs to be improved to enable a smoother and more consistent performance of the users. Besides the design and implementation of the module, two aspects that could be key for this improvement are the performance of hand tracking implemented in future HMD devices, and the habituation of users for grabbing and moving a hologram. In the context of a future user study, such habituation should be eased by providing a more significant training stage before data collection for system evaluation. Limitations Thirteen participants took part in the user study. While this number may not be sufficient to consider the results fully representative of all potential users, future studies should involve more participants to achieve more robust conclusions. Declarations Funding The research for this paper has been financially supported by the Elkartek Programme, Basque Government (Spain), project COGILE (KK-2021 / 00016). Ethical Approval An internal board has reviewed the ethical design of the user study, following guidance from the European Commission on Ethics in Social Science and Humanities.1 Conflict of Interest The authors have no relevant financial or nonfinancial interests to disclose. The authors have no competing interests to declare that are relevant to the content of this paper. REFERENCES Aceta, C., Fernández, I., & Soroa, A. (2021). Todo: A core ontology for task-oriented dialogue systems in industry 4.0. In Further with Knowledge Graphs (pp. 1–15). IOS Press. 10.3233/SSW210031 Aceta, C., Fernández, I., & Soroa, A. (2022). Kide4i: A generic semantics-based task-oriented dialogue system for human-machine interaction in industry 5.0. Applied Sciences,12(3). 10.3390/app12031192 Bambuˆ sek, D., Materna, Z., Kapinus, M., Beran, V., & Smrž, P. (2019). Combining interactive spatial augmented reality with head-mounted display for end-user collaborative robot programming. 2019 28th IEEE International Conference on Robot and Human Interactive Communication,1–8. 10.1109/RO-MAN46459.2019.8956315 1https://ec.europa.eu/info/funding-tenders/opportunities/ docs/2021-2027/horizon/guidance/ethics-in-social-science-and -humanities_he_en.pdf