External Robot Localization in 6DoF
Full text
FACULDADE DE ENGENHARIA DA UNIVERSIDADE DO PORTO External Robot Localization in 6DoF Nuno Filipe Leonor Nebeker Mestrado Integrado em Engenharia Eletrotécnica e de Computadores Supervisor: Armando Jorge Miranda de Sousa (Ph.D.) October 27, 2015
© Nuno Filipe Leonor Nebeker, 2015
ii
Abstract Mobile robots have shown the potential to change the way humans work, by taking on strenuous or repetitive tasks such as transporting heavy loads and by starting to interact with human workers. These robotic systems require robust and precise localization. External localization is relevant to complement self-localization and to provide valuable information when Automated Guided Vehicles (AGVs) move in irregular terrain. Recent developments in Red, Green and Blue, and Depth (RGB-D) sensors have made valuable visual information available to researchers at affordable costs prompted new research into video tracking. This masters thesis proposes a versatile and cost effective architecture for external 6 Degrees of Freedom (DoF) localization. The system uses RGB-D video tracking with multiple sensors and leverages Robot Operating System (ROS) and Point Cloud Library (PCL). The proposal includes sensor fusion strategies to improve overall localization performance. A suggested set-up would cost 2570C (2015 prices), using 3 Kinect for Xbox One sensors and computers . The proposed system takes advantage of the target’s natural geometry. Simulation analysis, on a Xubuntu 14.04 system with a 2006 Intel Pentium 4 processor and 3 GB of RAM, shows the proposed tracker, based on Iterative Closest Point (ICP) point cloud registration, performs with maximum translation error under 5 cm in either axis and maximum rotation error under 5°, with an update frequency above 2 Hz and is robust to noise and partial occlusion. In conjunction with the Cooperative Robot for Large Spaces Manufacturing (CARLoS) project, several datasets were collect and made publicly available. These include RGB-D data from two Kinect for Xbox 360 devices with different views of a moving AGV as well as the robot’s sensor data and 6 DoF ground truth provided by a commercial motion capture system. iii
iv
Resumo Os robôs móveis têm revelado o seu potencial para mudar o modo como os humanos trabalham: assumem tarefas fisicamente exigentes e repetitivas como transportar cargas pesadas e começam a interagir com trabalhadores humanos. Estes sistemas robóticos exigem localização robusta e precisa. A localização externa é relevante para complementar auto-localização e para oferecer informação valiosa quando Automated Guided Vehicles (AGVs) (Veículos Autoguiados) se deslocam em terreno irregular. Desenvolvimentos recentes em sensores Red, Green and Blue, and Depth (RGB-D) (cor e profundidade) vieram disponibilizar informação visual valiosa a investigadores a baixo custo levando a um novo interesse em tracking por vídeo. Esta dissertação de mestrado propõe uma arquitetura versátil e com boa relação qualidadepreço para localização externa em 6 Degrees of Freedom (DoF) (Gruas de Liberdade). O sistema usa tracking por vídeo RGB-D com vários sensores e tira proveito de Robot Operating System (ROS) e Point Cloud Library (PCL). A proposta inclui estratégias de fusão de sensores para aumentar o desempenho de localização. Uma montagem sugerida custaria 2570C (preços de 2015), usando 3 sensores Kinect for Xbox One e 3 computadores . O sistema proposto tira partido da geometria natural do alvo. Revela-se por análise de simulações, num sistema Xubuntu 14.04 com processador Intel Pentium 4 de 2006 e 3 GB de RAM, que o tracker proposto, baseado em registo de nuvens de pontos por Iterative Closest Point (ICP), funciona com erro máximo em translação inferior a 5 cm em cada eixo e erro máximo em rotação inferior a 5° com uma frequência de atualização acima de 2 Hz e é robusto em relação a ruído e oclusão parcial. Em conjunto com o projeto Cooperative Robot for Large Spaces Manufacturing (CARLoS), foram recolhidos e disponibilizados publicamente vários conjuntos de dados. Incluem-se dados RGB-D de dois dispositivos Kinect for Xbox 360 com vistas diferentes de um AGV em movimento bem como dados de sensores do robô e ground truth em6DoF conseguida com um sistema comercial de captura de movimento. v
vi
Acknowledgments First of all, I would like to thank Armando Jorge Miranda de Sousa for his orientation and for cementing my fascination with automation in his lectures. I also thank Carlos Miguel Correia da Costa for his patience and availability and for collaborating in capturing datasets for this dissertation. I would like to acknowledge the researchers and staff at ROBIS and LABIOMEP for providing the great working environment I enjoyed. I will also drop a line for my colleagues and friends, in particular José Pedro Fortuna Araújo and João Pedro Cruz Silva, for their valuable editorial notes, friendship and collaboration. Finally, I am always grateful to my mother Maria da Luz and my brother Pedro for their constant support in my personal and academic life, and my girlfriend Soraia for the past few years of companionship, which I hope will extend long into the future. Nuno Filipe Leonor Nebeker vii
xiv CONTENTS
List of Figures 2.1 ICPdemonstration ................................. 7 2.2 Kalman filter example — Black: truth, green: filtered process, red: observations . 8 2.3 Particle filter example — particles for a robot’s location . . . . . . . . . . . . . . 9 2.4 PCAinuse(c).................................... 10 2.5 SIFT features used in image recognition . . . . . . . . . . . . . . . . . . . . . . 11 2.6 DetectedSURFfeatures .............................. 11 2.7 ORB features used in image matching . . . . . . . . . . . . . . . . . . . . . . . 12 2.8 FPFH region of influence for point Pq....................... 13 2.9 SAC-IA registration — two partial view on left; alignment result on right . . . . 13 2.10 Exemplary pose estimation and tracking results . . . . . . . . . . . . . . . . . . 14 2.11Exemplaryresults.................................. 15 2.12Occlusionhandling ................................. 16 3.1 TypicalROSgraph ................................. 18 3.2 PCLlibraries .................................... 19 3.3 Kinect for Windows components . . . . . . . . . . . . . . . . . . . . . . . . . . 20 3.4 Kinect for Windows v2 components . . . . . . . . . . . . . . . . . . . . . . . . 21 3.5 Qualisys Oqus Motion Capture System in use at LABOMEP (Qualisys Oqus cameras highlighted in red and Kinect for Xbox 360 devices in blue) . . . . . . . . . 22 4.1 Proposed physical setup alternatives — Multiple views of target (left) and greater trackingarea(right)................................. 25 4.2 Captured point cloud (white) and preprocessed point cloud (blue) . . . . . . . . . 27 4.3 ROI from frame t−1 (left) and target within ROI on frame t(right) . . . . . . . 28 4.4 Data fusion sequence diagram . . . . . . . . . . . . . . . . . . . . . . . . . . . 30 5.1 Map (left) and captured point cloud and outliers (right) in background subtraction 34 5.2 Map (left) and captured point cloud and out-lier clusters (right) in background subtraction...................................... 34 5.3 Point clouds from two Kinect sensors simultaneously pointing at chair . . . . . . 35 5.4 Tracking test — point cloud and transforms . . . . . . . . . . . . . . . . . . . . 36 5.5 Tracking error during translation . . . . . . . . . . . . . . . . . . . . . . . . . . 36 5.6 Tracking error during rotation . . . . . . . . . . . . . . . . . . . . . . . . . . . 37 5.7 Tracking error during translation and rotation . . . . . . . . . . . . . . . . . . . 37 5.8 Simulation target mesh (left) and point cloud (right) . . . . . . . . . . . . . . . . 38 5.9 Simulation 1 — Simulated target (white) and cropped point cloud (blue) while stationary ...................................... 39 5.10 Simulation 1 — Tracking error with stationary target . . . . . . . . . . . . . . . 39 xv
xvi LIST OF FIGURES 5.11 Simulation 1 — Simulated target (white) and cropped point cloud (blue) during translation...................................... 40 5.12 Simulation 1 — Tracking error during translation . . . . . . . . . . . . . . . . . 40 5.13 Simulation 1 — Simulated target (white) and cropped point cloud (blue) during rotation ....................................... 41 5.14 Simulation 1 — Tracking error during rotation . . . . . . . . . . . . . . . . . . . 41 5.15 Simulation 1 — Simulated target (white) and cropped point cloud (blue) during rotation ....................................... 42 5.16 Simulation 1 — Tracking error during rotation (with smaller voxel grid) . . . . . 42 5.17 Simulation 1 — Simulated target (white) and cropped point cloud (blue) during translationandrotation ............................... 43 5.18 Simulation 1 — Tracking error during translation and rotation . . . . . . . . . . 43 5.19 Simulation 2 — Simulated target with noise (white) and cropped point cloud (blue) whilestationary................................... 45 5.20 Simulation 2 — Tracking error with stationary target . . . . . . . . . . . . . . . 45 5.21 Simulation 2 — Simulated target with noise (white) and cropped point cloud (blue) duringtranslation.................................. 46 5.22 Simulation 2 — Tracking error during translation . . . . . . . . . . . . . . . . . 46 5.23 Simulation 2 — Simulated target with noise (white) and cropped point cloud (blue) duringrotation ................................... 47 5.24 Simulation 2 — Tracking error during rotation . . . . . . . . . . . . . . . . . . . 47 5.25 Simulation 2 — Simulated target with noise (white) and cropped point cloud (blue) during translation and rotation . . . . . . . . . . . . . . . . . . . . . . . . . . . 47 5.26 Simulation 2 — Tracking error during translation and rotation . . . . . . . . . . 48 5.27 Simulation 3 — Simulated target with noise and occlusion (white) and cropped point cloud (blue) while stationary . . . . . . . . . . . . . . . . . . . . . . . . . 50 5.28 Simulation 3 — Tracking error with stationary target . . . . . . . . . . . . . . . 50 5.29 Simulation 3 — Simulated target with noise and occlusion (white) and cropped point cloud (blue) during translation . . . . . . . . . . . . . . . . . . . . . . . . 51 5.30 Simulation 3 — Tracking error during translation . . . . . . . . . . . . . . . . . 51 5.31 Simulation 3 — Simulated target with noise and occlusion (white) and cropped point cloud (blue) during rotation . . . . . . . . . . . . . . . . . . . . . . . . . . 52 5.32 Simulation 3 — Tracking error during rotation . . . . . . . . . . . . . . . . . . . 52 5.33 Simulation 3 — Simulated target with noise and occlusion (white) and cropped point cloud (blue) during translation and rotation . . . . . . . . . . . . . . . . . 52 5.34 Simulation 3 — Tracking error during translation and rotation . . . . . . . . . . 53 5.35 Simulation 4 — Simulated target with noise and occlusion (white) and cropped point cloud (blue) while stationary . . . . . . . . . . . . . . . . . . . . . . . . . 55 5.36 Simulation 4 — Tracking error with stationary target . . . . . . . . . . . . . . . 55 5.37 Simulation 4 — Simulated target with noise and occlusion (white) and cropped point cloud (blue) during translation . . . . . . . . . . . . . . . . . . . . . . . . 56 5.38 Simulation 4 — Tracking error during translation . . . . . . . . . . . . . . . . . 56 5.39 Simulation 4 — Simulated target with noise and occlusion (white) and cropped point cloud (blue) during rotation . . . . . . . . . . . . . . . . . . . . . . . . . . 57 5.40 Simulation 4 — Tracking error during rotation . . . . . . . . . . . . . . . . . . . 57 5.41 Simulation 4 — Simulated target with noise and occlusion (white) and cropped point cloud (blue) during translation and rotation . . . . . . . . . . . . . . . . . 57 5.42 Simulation 4 — Tracking error during translation and rotation . . . . . . . . . . 58
LIST OF FIGURES xvii 5.43 Simulation 4 — Tracking error during translation and rotation (before target is lost) 58 6.1 CARLoSAGV ................................... 61 6.2 Data collection environment at LABIOMEP . . . . . . . . . . . . . . . . . . . . 62 6.3 Guardian platform mesh (left) and point cloud (right) . . . . . . . . . . . . . . . 63
xviii LIST OF FIGURES
List of Tables 2.1 Comparison between SIFT, SURF and ORB (matches) . . . . . . . . . . . . . . 12 2.2 Comparison between SIFT, SURF and ORB (speed) . . . . . . . . . . . . . . . . 12 3.1 Kinect for Xbox 360 characteristics . . . . . . . . . . . . . . . . . . . . . . . . 20 3.2 Kinect for Xbox One characteristics . . . . . . . . . . . . . . . . . . . . . . . . 21 4.1 Proposed tracker cost estimate . . . . . . . . . . . . . . . . . . . . . . . . . . . 31 4.2 Summary of tracker design considerations . . . . . . . . . . . . . . . . . . . . . 32 5.1 Simulation 1 — Errors while stationary . . . . . . . . . . . . . . . . . . . . . . 44 5.2 Simulation 1 — Errors during translation . . . . . . . . . . . . . . . . . . . . . 44 5.3 Simulation 1 — Errors during rotation . . . . . . . . . . . . . . . . . . . . . . . 44 5.4 Simulation 1 — Errors during rotation . . . . . . . . . . . . . . . . . . . . . . . 44 5.5 Simulation 1 — Errors during translation and rotation . . . . . . . . . . . . . . . 44 5.6 Simulation 2 — Errors while stationary (with noise) . . . . . . . . . . . . . . . . 49 5.7 Simulation 2 — Errors during translation (with noise) . . . . . . . . . . . . . . . 49 5.8 Simulation 2 — Errors during rotation (with noise) . . . . . . . . . . . . . . . . 49 5.9 Simulation 2 — Errors during translation and rotation (with noise) . . . . . . . . 49 5.10 Simulation 3 — Errors while stationary (with noise and occlusion) . . . . . . . . 54 5.11 Simulation 3 — Errors during translation (with noise and occlusion) . . . . . . . 54 5.12 Simulation 3 — Errors during rotation (with noise and occlusion) . . . . . . . . . 54 5.13 Simulation 3 — Errors during translation and rotation (with noise and occlusion) 54 5.14 Simulation 4 — Errors while stationary (with noise and occlusion) . . . . . . . . 59 5.15 Simulation 4 — Errors during translation (with noise and occlusion) . . . . . . . 59 5.16 Simulation 4 — Errors during rotation (with noise and occlusion) . . . . . . . . . 59 5.17 Simulation 4 — Errors during translation and rotation (with noise and occlusion) 59 5.18 Average tracker update rates (Hz) . . . . . . . . . . . . . . . . . . . . . . . . . 59 xix
xx LIST OF TABLES
Acronyms and Abbreviations API Application Programming Interface AGV Automated Guided Vehicle BRIEF Binary Robust Independent Elementary Features CARLoS Cooperative Robot for Large Spaces Manufacturing DoF Degrees of Freedom EKF Extended Kalman Filter FAST Features From Accelerated Segment Test FPFH Fast Point Feature Histograms GPU Graphics Processing Unit GUI Graphical User Interface HSV Hue, Value, Saturation IR Infrared ICP Iterative Closest Point ORB Oriented FAST Rotated BRIEF RANSAC Random Sample Consensus ROI Region of Interest RGB Red, Green and Blue color RGB-D Red, Green and Blue, and Depth RMS Root Mean Square ROS Robot Operating System PCA Principal Components Analysis PDF Probability Density Function PCL Point Cloud Library PFH Point Feature Histograms xxi
xxii Acronyms and Abbreviations SDK Software Development Kit SAC-IA Sample Consensus Initial Alignment SIFT Scale-Invariant Feature Transform SIS Sequential Importance Sampling SURF Speeded Up Robust Features
Chapter 1 Introduction 1.1 Context This document is the final report for the Dissertation Thesis (Major in Automation) of the Master in Electrical and Computers Engineering at the Faculty of Engineering of the University of Porto. The subject of the dissertation is “External Robot Localization in 6DoF” and it is supervised by Armando Jorge Sousa. 1.2 Defining the Problem The current problem is tracking a rigid object in a dynamic environment, such that the object’s poses are known for each frame of video, in 6 Degrees of Freedom (DoF). Pose in 6 DoF is defined as translation and rotation from a reference and can be expressed by x,yand zCartesian coordinates and roll, pitch and yaw angles about fixed axes. Object tracking is still a technologically complex problem in environments where service robots, such as robotic manipulators, operate, due to the need for quality results and for safety, as well as an interest in inexpensive systems. In these scenarios, the need arises to use exterior localization techniques so that the real time control loop can interact with the objects of interest, identifying and tracking them in 3D with all available sensors. Tracking the Cooperative Robot for Large Spaces Manufacturing (CARLoS) robot is an example scenario. 1.3 Objectives The objective of this project is to develop a robotic system that will accomplish the tracking of mobile robots, mobile manipulators and other goods in semi-structured environments with realtime restrictions, exploring the advantages of the most recent Red, Green and Blue, and Depth (RGB-D) sensors. 1
8State of the Art are based on a system model, estimates of noise covariance matrices and initial estimates for the system’s state and error covariance. The filter algorithm has two steps: a time-update or prediction step and a measurement-update or correction step. The prediction step uses information up to t−1 to estimate system variables in t. The correction step begins by calculating the Kalman gain, based on system and noise parameters and on the current error covariance. The corrected estimate for tis then calculated by adding the prediction from t−1 and the measurement from t, compared to the expected output, considering the predicted state and weighed by the Kalman gain. Finally, the error covariance is updated based on its previous prediction, system model and Kalman gain. Due to its recursive nature, the Kalman filter algorithm only needs to store the system and noise models, the latest predictions and corrected estimates, their respective covariances and the latest input and output. For nonlinear systems, Extended Kalman Filter (EKF) linearizes the system around the current mean and covariance. This filter is no longer optimal, but is still the most used estimation algorithm for nonlinear systems. This is noted in [10], as well as the fact that it is only effective if the sampling and update time is small enough for the dynamics to be properly approximated by linearization at each step. Figure 2.2: Kalman filter example — Black: truth, green: filtered process, red: observations [11] The Kalman filter is widely used in tracking, as noted in [2]. Using a Kalman filter to predict the new pose of a tracked object, thereby saving time on new pose estimation, is proposed in [5]. 2.3.3 Particle Filter Sequential Monte Carlo estimation methods or particle filters, introduced in [12] (as the bootstrap filter) and explained in [13], are based on state-space models like the Kalman filter, but do not make assumptions on linearity. In Baysian filters, the objective it to calculates a posterior Probability Density Function (PDF), based on available information including most recent measurements, that contains all the statistical information and is a complete solution to the problem.
2.3 Algorithms of Note 9 A recursive algorithm, such as the Kalman filter, does this in two steps: prediction, where the prior PDF is calculated and correction, where the posterior PDF updates it with new measured data, acquiring better information on the system’s state — based on all available information. Analytically calculating this PDF is not possible, unless certain restrictions are considered. In the Kalman filter, the system is restricted to being linear and this function is restricted to a Gaussian distribution parametrized by a mean and a covariance, which are both calculated. If the state-space is discrete and there is a finite number of states, the posterior density function can be calculated by summing the conditional probability of each state. Without such restrictions, the PDF cannot be calculated, but it can be approximated. Knowing there are several particle filter algorithms, this section will continue by explaining only the Sequential Importance Sampling (SIS) algorithm as this is sufficient for understanding the concept of particle filtering and it is the basis for most such methods. This algorithm calculates the PDF in a number of samples, or particles. The final estimate results from each particle’s contribution, affected by its weight. The weight is the result of evaluating an importance density function, the design of which is the center of designing a particle filter for a given application. Figure 2.3: Particle filter example — particles for a robot’s location [7] As the algorithm is iterated, one particle tends to represent more of the PDF, while the others’ weights approach zero - the degeneracy problem. This can require a re-sampling of particles to promote diversity. Particle filters are widely used in localization, such as in [7] and in tracking such as in [14]. 2.3.4 Principal Component Analysis Principal Components Analysis (PCA) was first described in [15], named in [16] and explained in [17] is a method to fit lines and planes to measured data points, thereby finding the underlying relationships between variables, while reducing the problem’s dimensionality.
10 State of the Art The results of this analysis are a new orthonormal basis for the data, such that the each axis accounts, in sequence, for most of the data’s variance. These axes constitute the principal components. Data points can then be projected into this base. An algorithm for applying PCA to a dataset starts by centering the data on the origin, by subtracting its centroid, then calculates the principal components, either through eigenvalue decomposition of the covariance matrix or through singular value decomposition (SVD) of a form of the data matrix [17]. Figure 2.4: PCA in use (c) [18] PCA is used in [18] to find a tracked object’s orientation and as noted in section 2.2 to combine and compare appearance features. 2.3.5 Local Feature Descriptors As mentioned in section 2.2,SIFT and other local feature descriptors are popular in model-based tracking. These aim to quantify the properties of points in images, so that they can be matched and objects can be recognized. SIFT and Speeded Up Robust Features (SURF) also propose a key-point detector, which first selects points of interest. !(!) relies on an established detector. 2.3.5.1 SIFT SIFT was introduced in [19] and patented in the United States under [20]. It produces feature vectors from an image that are invariant to scale, translation, and rotation and nearly invariant to illumination changes and projection. SIFT generates in the order of 1000 keys (features) for an image and requires as little as 3 to match to identify an object, so it can tolerate significant occlusion. Furthermore, enough translation is allowed so that 2D objects can be rotated up to 60° from the camera plane and 3D objects can be rotated up to 220° and still be identified. 2.3.5.2 SURF SURF, introduced in [21] aims to be faster and as robust as SIFT and other previous descriptors. In the interest of speed, SURF is not explicitly invariant to anisotropic scaling, and perspective effects.
2.3 Algorithms of Note 11 Figure 2.5: SIFT features used in image recognition [19] In the authors’ tests, considering viewpoint changes, rotations, zooming and illumination variations, SURF outperformed SIFT in speed by 2 times and in recall for the same levels of precision. Figure 2.6: Detected SURF features [21] 2.3.5.3 ORB Oriented FAST Rotated BRIEF (ORB) was proposed in [22] with the intention of providing performance comparable to SIFT and greater speed (claimed to be two orders of magnitude higher than SURF). It adapts the Features From Accelerated Segment Test (FAST) detector [23] with an added orientation component and scale invariance and the Binary Robust Independent Elementary Features (BRIEF) descriptor [24] with orientation invariance. Both detector and descriptor were designed with efficiency and speed in mind. In particular, BRIEF descriptors are represented by
12 State of the Art Table 2.1: Comparison between SIFT, SURF and ORB (matches) [22] Inliers (%) Npoints Magazines data set ORB 36.180 548.50 SURF 38.305 513.55 SIFT 34.010 584.15 Boat data set ORB 45.8 789 SURF 25.6 795 SIFT 30.2 714 Table 2.2: Comparison between SIFT, SURF and ORB (speed) [22] Detector ORB SURF SIFT Time per frame (ms) 15.3 217.3 5228.7 binary strings, so calculating the difference between two features is reduced to calculating the Hamming distance between their descriptors, a very efficient operation. Due to its high speed it is well suited for real-time operations. Figure 2.7: ORB features used in image matching [22] The evaluation of ORB offers a comparison between this descriptor and SIFT and SURF. Comparison reults for matched points for two data set are shown in Table 2.1 and average detection times per frame from a set of 2686 are shown in Table 2.2. 2.3.6 FPFH feature descriptor and SAC-IA registration Feature descriptors can also be determined for 3D points, such as Fast Point Feature Histograms (FPFH), introduced in [25] (based on Point Feature Histograms (PFH) in [26,27]). This descriptor relates geometrical relationships between a point and its neighbors and is calculated based upon point coordinates and surface normals. FPFH improves over PFH by caching previous results and streamlining some formulations. Six DoF registration algorithms such as ICP risk falling into local minima in the optimization process returning incorrect results [3,25]. Sample Consensus Initial Alignment (SAC-IA), also
2.4 Recent Work in 6DoF RGBD Video Object Tracking 13 Figure 2.8: FPFH region of influence for point Pq[25] introduced in [25], based on Random Sample Consensus (RANSAC) [28] attempts to correct this by considering and ranking several sets of correspondences. SAC-IA selects many samples from one point cloud, but guarantees their pairwise distances are larger than a user-set distance. The corresponding point in the other point cloud is randomly selected from a set of points whose feature histograms are similar. These correspondences are used to calculate the 6 DoF transform between the point clouds and an error metric. Repeating these steps and using an optimization algorithm on the best transform yields the final result. Figure 2.9: SAC-IA registration — two partial view on left; alignment result on right [25] 2.4 Recent Work in 6DoF RGBD Video Object Tracking 2.4.1 Shape and Model-Based Tracking in Stereo Video An algorithm is proposed in [29] for shape-based and then iterative model-based 6 DoF pose estimation of single colored objects in stereo Red, Green and Blue color (RGB) images.
14 State of the Art This algorithms first segments the target object, using Hue, Value, Saturation (HSV) color segmentation, separating a single color in the image. Initial translation is estimated through centroids of object in both captured images. Rotation is obtained from matching the left image with a multi-view model. This is accurate if the object is in the same position as in training. To correct errors resulting from a difference in translation, a corrective rotation is calculated such that if the object were affected by this new rotation, its appearance would be the same as if it were in fact in the training position (center of the left image). This correction affects the translation, so that is recalculated, using simulated views from a 3D model of the object and the process is repeated iteratively. The algorithm generally converges in two iterations. Final orientation error depends on the angle resolution of training images. This method was tested in simulation. Based on this estimate, the algorithm proposed in [14] tracks arbitrarily shaped objects using a full 3D model in stereo video. The algorithm consists of pose update, probability function, normalization and annealed particle filter. The update consists in adding Gaussian noise to the current estimate, a translation vector and a rotation vector. To calculate the probability function for each updated pose, object views are rendered from that pose and compared to the captured images via a bitwise AND operation on binary detected edges on both images. The number of matching pixels is divided by the number of pixels in the rendered view edge image and this result is used in calculating the a-posteriori probabilities for all particles. This technique allows for matching even with partial occlusions. Because both normalizing over the number of edge pixels and normalizing over the total number of pixels in the image both present issues with unexpected edge filter responses and varying distance respectively, and additional weighted normalization step is performed. The annealed particle filter runs several filters with fewer particles in layers. This aims to reduce the overall umber of particle evaluations. Figure 2.10: Exemplary pose estimation [29] and tracking [14] results
2.4 Recent Work in 6DoF RGBD Video Object Tracking 15 2.4.2 Pose Estimation of Rigid Objects Using RGB-D Imagery The algorithm proposed by [18] uses 480 ×640 RGB-D imagery from a Kinect sensor. The object model consists of ORB features and their associated 3D positions on the objects, collected from different angles and distances. ORB features in the captured image are matched with model features. A homography matrix is iteratively computed and then used to remove outliers. A region of interest (ROI) is determined by fitting a 3D bounding box to the matching keypoints in the captured point cloud. The oject points are segmented through clustering. The object’s 6D pose is determined in two steps. First, translation is calculated by taking the centroid of the object points and finding its depth in the camera space. Rotation is obtained through a bounding box along the principal axes of the point distribution. Object detection is subsequent frames is accelerated by considering only a quadrangular ROI in the image space, corresponding to the previous 3D bounding box. Objects are tacked at 15 Hz with a 92% detection rate. Future work includes using GPU optimization and using contours to track untextured objects. Figure 2.11: Exemplary results [18] 2.4.3 Princeton Tracking Benchmark and RGBD Tracking Algorithms A dataset consisting of 100 RGB-D video samples and a 6 DoF tracking method are introduced in [6]. The samples were recorded with a Kinect sensor and used to evaluate the proposed baseline algorithms. The authors present a set of algorithms based on the state of the art in 2D and 3D tracking. The occlusion detection strategy is of particular note, due to its use of depth information: the depth data within the object’s 2D bounding box is divided into an object region (depth range of target points, which are the closest to the camera) and a background region (further away). When other occluding objects enter the bounding box, their depth is lower.
16 State of the Art Figure 2.12: Occlusion handling [6] 2.5 Conclusion This review of the state of the art in 6 DoF video tracking shows that this problem is non-trivial and that challenges remain. There are a number of significant design decisions that impact performance and accuracy, as well as the range of application of a tracker. Current efforts attempt to leverage new technology and are often tailored to specific targets and environments.
Chapter 3 Tools and Equipment 3.1 Introduction Tackling any programming project involves finding the appropriate libraries, frameworks and tools. For a problem that is focused on robotics and computer vision, this means finding robotics middleware [30] that will provide an adequate hardware abstraction and common services including package management and communications (between a robot’s services and between different robots), a computer vision library that provides real-time processing functions of 2D and 3D images and a complete sensor system that provides a quality imaging solution. Finally, validation requires a solution for obtaining a Ground Truth trajectory. 3.2 ROS Robotics MiddleWare Among current solutions for robotics middleware [30], Robot Operating System (ROS) stands out. It was designed to be: [31] • "Peer-to-peer" • "Tools-based" • "Multi-lingual" • "Thin" • "Free and Open-Source" ROS allows researchers to develop robotics software from hardware drivers to high-level decision making, by integrating different components by different developers. 3.2.1 Nodes and Packages The peer-to-peer nature of ROS allows different processes to operate separately and communicate. These can be visualized in graph form and are called "nodes." 17
24 Proposed Solution 4.2 Prerequisites Some necessary information and services are assumed to be available to the tracking system. At the same time, some decisions are left to the final user. 4.2.1 Sensor Localization The system’s sensors’ poses must be known in reference to the global frame world, as the tracking target’s pose is referenced to this frame. The method by which this information is made obtained is outside the scope of this system, but it should be made available to tf (see: Section 3.2.3). Amapping and a localization node are provided for testing purposes and to allow for the correction of a sensor’s pose when moved from a known location by a small rotation or translations (single digits in degrees of rotation or centimeters of translation). The user can go as far as to manually publish a transformation between or can use one of many widely available ROS mapping and self-localization packages. 4.2.2 Communication Communication between the sensors, tracking nodes and data-fusion node is handled transparently through ROS services and may include transitions between different computers. This will depend on a local network whose bandwidth can support the necessary messages. For reference, a Microsoft Kinect for Windows point cloud with resolution 640×480 and color for each point encoded in 8 floats takes up 9.3Mb and is captured at 30 Hertz, resulting in a bandwidth of 279MB/s. These considerations must be taken for other types of sensors. 4.2.3 Physical Setup The advantages of having a multiple sensor tracking system are twofold: For one, occlusions effectively disappear when different sensors have their own views of the target, so that even if the target is occluded for one sensor it will be in full view for the global system. On the other hand, the effective tracking area can be transparently increased by adding new sensors. It is up to the end-user to balance these advantages. Fig. 4.1 demonstrates this duality. 4.3 Tracking The tracking node uses Fast Point Feature Histograms (FPFH) descriptors and Sample Consensus Initial Alignment (SAC-IA) registration to determine an initial pose and Iterative Closest Point (ICP) registration to track the target. Algorithm 1summarizes this node’s work flow.
4.3 Tracking 25 PC2Sensor2 PC1 Sensor1 Tracking area PC2 Sensor2 PC1 Sensor1 Tracking area Occlusion handling Large working area Figure 4.1: Proposed physical setup alternatives — Multiple views of target (left) and greater tracking area (right) 4.3.1 Point Cloud Preprocessing Real world sensors are subject to noise to varying degrees and return large datasets that must be filtered to become manageable in a real-time system. For each new point cloud, a pass-through filter removes points whose depth is beyond the sensor’s effective working range. In the case of the Microsoft Kinect for Windows the working range is 0.4 m to 4 m. A statistical out-lier removal filter removes points that are further from their neighbors than a set multiple of the standard deviation of pairwise distances. A voxel grid filter collapses points within small voxels whose size is compatible with the senor’s actual resolution. Finally, the point cloud is transformed into the global frame so that further processing uses global frame coordinates. Examples of this preprocessing routine are shown in Chapter 5. The example of Fig. 4.2 is of particular note, as it shows the result of filtering and under-sampling by voxel grid a noisy point cloud. 4.3.2 Initial Pose Estimation Upon starting or having lost track of the target, the tracking node will initiate initial pose estimation. The objective of this process is to obtain an approximate transform for the target in the scene that will lie within convergence range for the continuous tracking algorithm. This step uses the whole scene point cloud, as the target’s position is completely unknown and no assumptions are made about it, other that it should be visible at this point. This process relies on FPFH and SAC-IA, as described in Section 2.3.6.
26 Proposed Solution Algorithm 1 Single-sensor tracker Input: M(target model point cloud), C0(captured point cloud), Tcw (camera to world transform), t(maximum error threshold) Output: T(new target pose transform), e(registration error) Cc(cropped point cloud) 1: function PROCESS(M,C0,Tcw) 2: C←PREPROCESS(C0,Tcw).Algorithm 2 3: if trackerinitialized =false then 4: T,e←FIND(M,C).Algorithm 3— get initial estimate 5: else 6: T,e←TRACK(M,C,T0).Algorithm 4— continuous tracking 7: end if 8: if e>tthen 9: trackerinitialized ←false .Lost target and should reinitialize 10: return 11: else 12: Cc←CROP(M,C,T).Algorithm 5 13: end if 14: return T,e,Cc 15: end function FPFH descriptors depend on surface normals, so these are estimated for the whole scene and for the target model (target model normals are estimated on node startup upon loading the respective point cloud). The target model is aligned to the scene using SAC-IA which returns the approximate initial pose and the alignment error. If the alignment error is greater than a threshold, the target has not been found and the tracker is reinitialized. 4.3.3 Continuous Tracking The tracker assumes proximity as a constraint on motion. This means the displacement from frame i−1 to frame iis assumed to be smaller than a predetermined distance: |xi−xi−1|<d. As long as this assumption holds, using the latest pose as an initial estimate for ICP registration will yield good results in few iterations. Algorithm 2 Preprocess point cloud Input: C0(unprocessed point cloud), Tcw (camera to world transform) Output: C(preprocessed point cloud) 1: function PREPROCESS(C) 2: Ctemp ←PASSTHROUGHFILTER(C0,min,max,z).Applied over the zaxis 3: Ctemp ←STATISTICALOUTLIERFILTER(Ctemp,k,a).Number of neighbors and multiple of std 4: Ctemp ←VOXELGRIDFILTER(Ctemp,lea f size) 5: C←TRANSFORM(Ctemp,Tcq).Transforms to global frame 6: return C0 7: end function
4.3 Tracking 27 Figure 4.2: Captured point cloud (white) and preprocessed point cloud (blue) This assumption has a useful corollary: points that are very far from the latest pose will not be part of the target. With this consideration in mind, each new point cloud is cropped around the target’s location. The crop box is translated and rotated to match the latest pose and is enlarged, compared to the target model, to account for the object’s movement between frames. This Region of Interest (ROI) is demonstrated in Fig. 4.3. This cropped point cloud is used in ICP registration, instead of the full scene, to improve performance, by reducing the number of points in the cloud. It is also published for use by a data fusion node. At any point during tracking, if the registration score is worse than a threshold, the target is considered lost and a new initialization is required. This can be avoided with the use of multiple sensors, whereupon an individual tracker will be able to use the global tracking result for the next iteration. Algorithm 3 Estimate initial pose Input: M(target model point cloud), C(captured point cloud) Output: T(target pose transform), e(registration error) 1: function FIND(M,C) 2: C←PREPROCESS(C).Algorithm 2 3: NT←COMPUTENORMALS(T) 4: NC←COMPUTENORMALS(C) 5: FT←COMPUTEFPFH(T,NT) 6: F C←COMPUTEFPFH(C,NC) 7: T,e←SAC-IA(T,C,FT,F C) 8: return T,e 9: end function
28 Proposed Solution Figure 4.3: ROI from frame t−1 (left) and target within ROI on frame t(right) 4.4 Data Fusion The work flow of the fusing data fusion node is described in Algorithm 6. This node takes the local results from single-sensor tracking nodes, merges their partial views and returns a higher quality tracking result. 4.4.1 Data Flow Taking advantage of the publisher/subscriber paradigm offered by ROS, any number of singlesensor local trackers can broadcast their pose estimates and have access to the latest global estimate. The obvious advantage of data fusion is having a higher quality global result, as the fusion tracker has access to more information, be that different views of the target or a wider operating range. The less obvious, but equally important advantage is that local trackers can use the latest global pose as an estimate for the next iteration. Not only does this allow for faster loss recovery, as a local node will not have to reinitialize, but will instead use global information, it also means non-deterministic tracking methods can be used, without burdening individual single-sensor trackers. The global tracker can, for example, use an Extended Kalman Filter (EKF) and provide the filter’s prediction as an estimate. This multi-directional information flow is displayed by the sequence diagram in Fig. 4.4. Note that all messages are published, but only subscribed messages are represented, for simplicity and readability. The first node, Tracker 1, estimates the target’s initial pose and publishes it, along with the registration error and cropped point cloud. The global tracker subscribes to these and performs its Algorithm 4 Track target Input: M(target model point cloud), C(captured point cloud), T0(latest target pose transform) Output: T(new target pose transform), e(registration error) 1: function TRACK(M,C,T) 2: Cc←CROP(M,C,T).Algorithm 5 3: T,e←ICP(M,Cc,T0).Uses last pose as initial estimate 4: return T,e 5: end function
4.5 Output 29 Algorithm 5 Crop point cloud Input: M(target model point cloud), C(captured point cloud), T(target pose transform) Output: Cc(cropped point cloud) 1: function CROP(M,C,T) 2: minpoint,maxpoint ←MINMAX3D(M).Get bounding box corners 3: Cc←CROPBOXFILTER(C,minpoint +d,maxpoint +d,T).Crop around target with orientation leaving extra space 4: return Cc 5: end function own tracking iteration, after which it publishes a global pose and error, as well the merged point cloud, to which other ROS nodes would subscribe. The single-sensor tracker also subscribes to the global pose and error, but these will only truly come into play once Tracker 2 is started. Instead of initializing by estimating an initial pose, it will subscribe to the global data and work from there. From this time, the tracking system continuously shares its results, so that the global results are refined by having access to extra views and quantitative data and the local trackers don’t lose the target and perform better, based on a quality previous pose. 4.4.2 Off-line Fusion As mentioned in Section 3.2.4,ROS offers data logging tools that can, upon playback, emulate nodes’ communication over their published topics. When operating independently, the output from single-sensor trackers can be recorded and later played back for off-line data-fusion. 4.5 Output The tracking problem’s output can be expressed in the form of a homogeneous transformation matrix that includes information on rotation and translation: Algorithm 6 Data fusion tracker Input: M(target model point cloud), Cc(cropped point clouds), t(maximum error threshold), T0 (target pose transforms), e0(registration errors) Output: T(new target pose transform), e(registration error) Cm(merged point cloud) 1: function PROCESS(M,Cc,T0) 2: for n←1,nsensors do 3: Cm←Cm+Cc[n].Merge point clouds 4: end for 5: Tbest ←T0e0 best 6: T,e←TRACK(M,Cm,T0e0 best).Algorithm 4— continuous tracking 7: return T,e,Cm 8: end function
30 Proposed Solution Initial pose estimate Tracker 1 Sensor 1 Global tracker Initial pose estimates; covariances; cropped cloud Global pose; covariances Other nodes Global pose; covariance; merged cloud Point cloud Tracked pose;covariance; cropped cloud Global pose; covariance; merged cloud Sensor 2 Tracker 2 Point cloud Point cloud Continuous tracking loop Tracker n Sensor n Point cloud ... Initial pose estimatePoint cloud Tracked pose; covariance; cropped cloud Global pose; covariance Global pose; covariance Figure 4.4: Data fusion sequence diagram Translation vector: d= x y z .(4.1) Rotation matrix: R= r11 r12 r13 r21 r22 r23 r31 r32 r33 (4.2) Combined into the homogeneous transformation matrix: T="R d 0 1 #(4.3) This is the output of Point Cloud Library (PCL) registration methods. Rotation can be represented by matrix Rand by quaternion q: q=Dx y z w E(4.4) ROS provides transparent conversions between Rand q. This information is transmitted in ROS encapsulated into a geometry_msgs/TransformStamped message which includes: • A time stamp; • The name of the parent frame; • The name of the child frame;
4.6 Cost Analysis 31 Table 4.1: Proposed tracker cost estimate Component Cost (C) General purpose computer 650 Microsoft Kinect for Xbox One 200 Network switch 20 Total (2 sensors; 2 computers) 1720 Total (3 sensors; 3 computers) 2570 • The translation vector; • The rotation quaternion. This is particularly useful, because it includes the tracked object in the system’s tf2 transform tree (see Section 3.2.3). Along with this information, a confidence rating is also useful. Registration algorithms also return this in the form of the internal error metric for the last iteration. This is transmitted along with the target’s frame name as a float64. This should be further refined into a covariance matrix. 4.6 Cost Analysis The proposed tracker is financially competitive, as it uses off the shelf hardware. Commercial RGB-D sensor solutions, such as the proposed Kinect for Xbox 360, are more and more cost effective and have adequate characteristics. Other hardware requirements are limited to general purpose computers and basic networking hardware. A Microsoft Kinect for Xbox One with a PC adapter has a total estimated retail price of C200, in Portugal (or 200 USD, in the United States) [44]. The hardware specifications for this device are: [56] • "64-bit (x64) processor" • "Physical dual-core 3.1 GHz (2 logical cores per physical) or faster processor" • "USB 3.0 controller dedicated to the Kinect for Windows v2 sensor or the Kinect Adapter for Windows for use with the Kinect for Xbox One sensor" • "4 GB of RAM" • "Graphics card that supports DirectX 11" A market survey of 2015 prices leads to an estimate of C650 for a desktop computer with these characteristics. A networking switch can cost as little as C20. Cost estimates for different configurations are listed in Table 4.1.
32 Proposed Solution Table 4.2: Summary of tracker design considerations Design consideration Decision Object representation Shape-based: point cloud Feature selection FPFH; point cloud Object detection FPFH,SAC-IA Joint detection and tracking Yes Tracking method Deterministic; ICP Tracking assumptions Maximum velocity Loss recovery; occlusion handling Reinitialization; data fusion For comparison, the tracking system used in data collection, discussed in Chapter 6has an approximate cost of 150000C. This system offers better precision than the proposed solution, but its price is two orders of magnitude greater than what would be expected for the proposed system and it depends on Infrared (IR) markers. 4.7 Conclusion Revisiting the tracker design decisions outlined in Section 2.2, the proposed tracker is summarized by Table 4.2.
Chapter 5 Development and Evaluation 5.1 Introduction This chapter chronologically details development efforts in various areas, examining avenues of development that were eventually not followed and providing greater insight into design decisions. Development logs are available on the project’s website (http://paginas.fe.up.pt/ ~ee10192/). 5.2 Familiarization with Kinect, PCL and ROS The first weeks of development brought a greater practical understanding of Robot Operating System (ROS), Point Cloud Library (PCL) and the Microsoft Kinect for Windows sensor. There is a great wealth of information and resources for each ranging from software and hardware documentation [34,41,57] to research papers [37,38,39,40] and hobbyist tutorials. 5.3 Mapping and Self-Localization As stated in Section 4.2, tracking a target in the global frame requires sensors’ locations. These, in turn, can be acquired through map-based self-localization. Initial mapping efforts consisted of simply taking one point cloud, filtering it (as described in Section 4.3.1) and saving it as a map. Self-localization was performed with Iterative Closest Point (ICP) registration. This approach is computationally inexpensive and produces good results for small translations and rotations around the initial pose (few centimeters and degrees). It should be noted that large translations and rotations (on the order of 0.5 m and 45°) result in ICP registration falling into local minima and unreasonable output poses. In an effort to improve the performance of the self-localization node, ICP was replaced by Generalized ICP [58]. This extension to the original algorithm combines ICP with point-to-point variants into a probabilistic framework and promises more robustness. In testing, this proved 33
40 Development and Evaluation Figure 5.11: Simulation 1 — Simulated target (white) and cropped point cloud (blue) during translation 5.9.3 Rotation For this simulation, the target rotates around the zaxis (yaw) at an angular velocity of 1.5×10−2rads−1. This scenario focuses more directly on the system’s interest in 6 Degrees of Freedom (DoF) pose.A simulation visualization is shown on Fig. 5.13, where the published cropped cloud almost completely matches the simulated target. As shown in the graphs of Fig. 5.14, error on yaw is well bounded, but have a greater variance than pitch and roll. Translation error on the zaxis is small and constant, but on xand yit is sinusoidal. Causes of this behavior include numerical errors, incorrect interpolations in the tf2 query and point cloud preprocessing (see Section 4.3.1). The voxel grid leaf size is 2 ×10−2m. Figure 5.12: Simulation 1 — Tracking error during translation
5.9 Simulation 1 — No Noise and No Occlusion 41 Figure 5.13: Simulation 1 — Simulated target (white) and cropped point cloud (blue) during rotation 5.9.4 Rotation With Smaller Voxel Grid Using a voxel grid leaf size of 10−2m as shown in Fig. 5.15 with error in the graphs of Fig. 5.16 and in Table 5.4. This shows that the error in orientation is reduced, but the error in translation shows only a slight reduction. This leads to the conclusion that the major causes of this error are numerical errors and lag between tf2 data. 5.9.5 Translation and Rotation For this simulation, the target moves along the yaxis at a velocity of 0.01 ms−1and rotates around the zaxis (yaw) at an angular velocity of 1.5×10−2rads−1. Figure 5.14: Simulation 1 — Tracking error during rotation
42 Development and Evaluation Figure 5.15: Simulation 1 — Simulated target (white) and cropped point cloud (blue) during rotation In this case the difference in pose from one frame to the next is greater and requires more ICP iterations. This lower the update rate, which can eventually result in a lost target. While the pose in a frame is computed, the object moves away from the previous position, causing a greater computational cost for the next frame and further lowering the update rate. As the lag increases, the target moves out of the neighborhood of its last know position, leaving the cropped point cloud, which results in unsuccessful registration. This process of failure is demonstrated in Fig. 5.17, where, on the left, a significant error in localization is visible and, on the right, the target has been lost and is leaving the area of the cropped point cloud. As is evident in the graphs of Fig. 5.18, the target was lost. In fact, at t=17, the tracker is reinitialized with the initial pose, due to a very high registration error, and there is a sharp increase in error, as the object is no longer being tracked. 0 50 100 150 200 250 300 350 400 450 −0.02 −0.015 −0.01 −0.005 0 0.005 0.01 0.015 0.02 Time (s) Error (m) Error on translation over time x y z 0 50 100 150 200 250 300 350 400 450 −2 0 2 4 6 8 10 x 10−3 Time (s) Error (rad) Error on rotation over time Roll Pitch Yaw Figure 5.16: Simulation 1 — Tracking error during rotation (with smaller voxel grid)
5.9 Simulation 1 — No Noise and No Occlusion 43 Figure 5.17: Simulation 1 — Simulated target (white) and cropped point cloud (blue) during translation and rotation Figure 5.18: Simulation 1 — Tracking error during translation and rotation
44 Development and Evaluation Table 5.1: Simulation 1 — Errors while stationary Error Translation (m) Rotation (rad) x y z Euclidean Roll Pitch Yaw Max 0.0050 0.0180 0.0200 0.0087 0.0020 0 0.0120 RMS 0.0049 0.0177 0.0198 0.0085 0.0011 0 0.0071 Table 5.2: Simulation 1 — Errors during translation Error Translation (m) Rotation (rad) x y z Euclidean Roll Pitch Yaw Max 0.0070 0.0320 0.0200 0.0121 0.0070 0.0030 0.0170 RMS 0.0059 0.0267 0.0190 0.0102 0.0048 0.0018 0.0083 Table 5.3: Simulation 1 — Errors during rotation Error Translation (m) Rotation (rad) x y z Euclidean Roll Pitch Yaw Max 0.0220 0.0230 0.0200 0.0381 0.0090 0.0040 0.0210 RMS 0.0140 0.0144 0.0194 0.0243 0.0046 0.0016 0.0103 Table 5.4: Simulation 1 — Errors during rotation Error Translation (m) Rotation (rad) x y z Euclidean Roll Pitch Yaw Max 0.0200 0.0200 0.0200 0.0346 0.0020 1.00 ×10−30.0090 RMS 0.0140 0.0143 0.0200 0.0243 4.9×10−41.4×10−40.0019 Table 5.5: Simulation 1 — Errors during translation and rotation Error Translation (m) Rotation (rad) x y z Euclidean Roll Pitch Yaw Max 0.0150 0.0890 0.0200 0.0260 0.0390 0.0020 0.2160 RMS 0.0101 0.0101 0.0184 0.0175 0.0202 0.0013 0.1241 5.10 Simulation 2 — Noise and No Occlusion Identical simulations were carried out, after introducing random noise in simulated points. Each point was affected by random noise in its x,yand zcoordinates, up to 1 ×10−2m. The aim of these simulations is to evaluate the tracker’s robustness to noise levels that would be expected of current Red, Green and Blue, and Depth (RGB-D) sensors. 5.10.1 Stationary Target Tracking performance with a stationary target is comparable to the previous case. The visual effect of noise is shown in Fig. 5.19 and error results are shown in the graphs of Fig. 5.20 and in
5.10 Simulation 2 — Noise and No Occlusion 45 Table 5.6. Figure 5.19: Simulation 2 — Simulated target with noise (white) and cropped point cloud (blue) while stationary 0 5 10 15 20 25 30 35 40 0 0.005 0.01 0.015 0.02 0.025 0.03 Time (s) Error (m) Error on translation over time x y z 0 5 10 15 20 25 30 35 40 −12 −10 −8 −6 −4 −2 0 2x 10−3 Time (s) Error (rad) Error on rotation over time Roll Pitch Yaw Figure 5.20: Simulation 2 — Tracking error with stationary target 5.10.2 Translation During translation the tracker lags further behind the target, as is notable in Fig. 5.21. Error values increase, approaching the set limits, as shown in Fig. 5.22 and Table 5.7. 5.10.3 Rotation Once again, performance in rotation is comparable to the original simulation without incorporating noise. The results are visualized in Fig. 5.23. Error values are well within defined limits as shown in the graphs of Fig. 5.24 and in Table 5.8.
46 Development and Evaluation Figure 5.21: Simulation 2 — Simulated target with noise (white) and cropped point cloud (blue) during translation 5.10.4 Translation and Rotation As was the case without noise, the target is quickly lost, this time after 14 s. the phenomenon of the target leaving the previous frame’s neighborhood is clearly visible in Fig. 5.25. Error values in the graphs of Fig. 5.26 and Table 5.9 are larger than in the previous simulation and show that the target was lost 3 s sooner. This is a result of compounding the increased tracking error as seen in the translation with noise simulation with the adverse results of lag in this dynamic situation. 0 10 20 30 40 50 60 0 0.005 0.01 0.015 0.02 0.025 0.03 0.035 0.04 0.045 Time (s) Error (m) Error on translation over time x y z 0 10 20 30 40 50 60 −0.005 0 0.005 0.01 0.015 0.02 0.025 Time (s) Error (rad) Error on rotation over time Roll Pitch Yaw Figure 5.22: Simulation 2 — Tracking error during translation
5.10 Simulation 2 — Noise and No Occlusion 47 Figure 5.23: Simulation 2 — Simulated target with noise (white) and cropped point cloud (blue) during rotation 0 50 100 150 200 250 300 350 400 450 −0.03 −0.02 −0.01 0 0.01 0.02 0.03 Time (s) Error (m) Error on translation over time x y z 0 50 100 150 200 250 300 350 400 450 −0.005 0 0.005 0.01 0.015 0.02 0.025 Time (s) Error (rad) Error on rotation over time Roll Pitch Yaw Figure 5.24: Simulation 2 — Tracking error during rotation Figure 5.25: Simulation 2 — Simulated target with noise (white) and cropped point cloud (blue) during translation and rotation
48 Development and Evaluation 0 5 10 15 20 25 30 35 0 0.05 0.1 0.15 0.2 0.25 0.3 0.35 Time (s) Error (m) Error on translation over time x y z 0 5 10 15 20 25 30 35 −0.1 0 0.1 0.2 0.3 0.4 0.5 0.6 Time (s) Error (rad) Error on rotation over time Roll Pitch Yaw Figure 5.26: Simulation 2 — Tracking error during translation and rotation
5.11 Simulation 3 — Noise and 25% Occlusion 49 Table 5.6: Simulation 2 — Errors while stationary (with noise) Error Translation (m) Rotation (rad) x y z Euclidean Roll Pitch Yaw Max 0.0020 0.0260 0.0250 0.0035 0.0030 0.0020 0.0110 RMS 9.5×10−40.0247 0.0237 0.0016 9.4×10−48.7×10−40.0026 Table 5.7: Simulation 2 — Errors during translation (with noise) Error Translation (m) Rotation (rad) x y z Euclidean Roll Pitch Yaw Max 0.0030 0.0420 0.0250 0.0052 0.0050 0.0020 0.0250 RMS 0.0017 0.0365 0.0238 0.0029 0.0027 6.7×10−40.0175 Table 5.8: Simulation 2 — Errors during rotation (with noise) Error Translation (m) Rotation (rad) x y z Euclidean Roll Pitch Yaw Max 0.0270 0.0270 0.0250 0.0468 0.0060 0.0030 0.0230 RMS 0.0178 0.0179 0.0242 0.0309 0.0020 9.1×10−40.0104 Table 5.9: Simulation 2 — Errors during translation and rotation (with noise) Error Translation (m) Rotation (rad) x y z Euclidean Roll Pitch Yaw Max 0.0100 0.0880 0.0240 0.0173 0.0300 0.0030 0.1910 RMS 0.0064 0.0602 0.0231 0.0111 0.0180 0.0011 0.1080 5.11 Simulation 3 — Noise and 25% Occlusion To evaluate the tracker’s robustness when faced with occlusions, the simulated sensor was moved away from the target’s initial pose so that the pass-through filter would discard a portion of its points. The sensor was placed 3.75 m away from the target’s center, while the pass-through filter was set to discard points further than 3.9 m. This leaves up to ¼ of the object’s volume occluded. During these simulations, it quickly became apparent that the t(maximum error threshold) parameter (used in Algorithm 1) would have to be increased, compared to previous simulations. This parameter is used to internally determine whether continuous ICP tracking is successful or the tracker has lost its target. Increasing it allows the system to maintain confidence in its performance and continue to follow the object in this challenging situation. This comes with the trade-off of not reinitializing, when a succession of incorrect estimations occurs. 5.11.1 Stationary Target Tracking performance with a stationary target is comparable to the previous case, despite a greater variance in rotation error. The visual effect of noise combined with occlusion is shown in Fig. 5.27
56 Development and Evaluation Figure 5.37: Simulation 4 — Simulated target with noise and occlusion (white) and cropped point cloud (blue) during translation 5.12.4 Translation and Rotation Tracking performance during translation and rotation is similar to the case of rotation and fails in the same way: too few points are visible for tracking, while the update rate is not high enough. Results for this test are shown in Fig. 5.41. In this case, relevant error values in the graphs of Fig. 5.42 and Fig. 5.43 increase over time, as the tracker maintains confidence in its pose estimates, but does not accomplish effective tracking. This results in the higher error values of Table 5.17. 0 20 40 60 80 100 120 0 0.005 0.01 0.015 0.02 0.025 0.03 0.035 Time (s) Error (m) Error on translation over time x y z 0 20 40 60 80 100 120 −0.015 −0.01 −0.005 0 0.005 0.01 0.015 Time (s) Error (rad) Error on rotation over time Roll Pitch Yaw Figure 5.38: Simulation 4 — Tracking error during translation
5.12 Simulation 4 — Noise and 75% Occlusion 57 Figure 5.39: Simulation 4 — Simulated target with noise and occlusion (white) and cropped point cloud (blue) during rotation 0 50 100 150 −0.03 −0.02 −0.01 0 0.01 0.02 0.03 Time (s) Error (m) Error on translation over time x y z 0 50 100 150 −0.02 0 0.02 0.04 0.06 0.08 Time (s) Error (rad) Error on rotation over time Roll Pitch Yaw Figure 5.40: Simulation 4 — Tracking error during rotation Figure 5.41: Simulation 4 — Simulated target with noise and occlusion (white) and cropped point cloud (blue) during translation and rotation
58 Development and Evaluation 0 50 100 150 −0.2 0 0.2 0.4 0.6 0.8 1 1.2 1.4 1.6 Time (s) Error (m) Error on translation over time x y z 0 50 100 150 −0.5 0 0.5 1 1.5 2 2.5 Time (s) Error (rad) Error on rotation over time Roll Pitch Yaw Figure 5.42: Simulation 4 — Tracking error during translation and rotation 0 20 40 60 80 100 120 140 −0.02 −0.01 0 0.01 0.02 0.03 0.04 0.05 Time (s) Error (m) Error on translation over time x y z 0 20 40 60 80 100 120 140 −0.03 −0.02 −0.01 0 0.01 0.02 Time (s) Error (rad) Error on rotation over time Roll Pitch Yaw Figure 5.43: Simulation 4 — Tracking error during translation and rotation (before target is lost)
5.13 Conclusion 59 Table 5.14: Simulation 4 — Errors while stationary (with noise and occlusion) Error Translation (m) Rotation (rad) x y z Euclidean Roll Pitch Yaw Max 0.0020 0.0280 0.0260 0.0035 0.0040 0.0070 0.0090 RMS 0.0010 0.0258 0.0243 0.0018 0.0017 0.0029 0.0046 Table 5.15: Simulation 4 — Errors during translation (with noise and occlusion) Error Translation (m) Rotation (rad) x y z Euclidean Roll Pitch Yaw Max 0.0020 0.0330 0.0280 0.0035 0.0050 0.0110 0.0120 RMS 0.0014 0.0297 0.0247 0.0025 0.0019 0.0037 0.0047 Table 5.16: Simulation 4 — Errors during rotation (with noise and occlusion) Error Translation (m) Rotation (rad) x y z Euclidean Roll Pitch Yaw Max 0.0270 0.0260 0.0270 0.0468 0.0160 0.0080 0.0270 RMS 0.0198 0.0155 0.0240 0.0344 0.0048 0.0020 0.0080 Table 5.17: Simulation 4 — Errors during translation and rotation (with noise and occlusion) Error Translation (m) Rotation (rad) x y z Euclidean Roll Pitch Yaw Max 0.0320 0.0300 0.0260 0.0554 0.0150 0.0080 0.0160 RMS 0.0223 0.0170 0.0238 0.0386 0.0049 0.0023 0.0052 Table 5.18: Average tracker update rates (Hz) Simulation Stationary Translation Rotation Translation and Rotation 1 — No noise 3.7456 2.6143 2.7370 3.1314 1 — Smaller voxel grid — — 3.6500 — 2 — With noise 3.3353 2.6965 2.6100 2.9979 3 — Noise, 25% Occlusion 3.6977 2.7049 3.0167 2.9076 4 — Noise, 75% Occlusion 3.5073 3.0408 3.1109 2.9501 5.13 Conclusion The project brought many challenges including familiarization with different technologies and demanding hardware requirements that could not fully be met. During the development of this project, several approaches were explored, bringing a better understanding of their advantages and limitations and an appropriate path was established. Simulation results are satisfactory with error in translation under 5 cm and rotation under 5°. Tolerable inaccuracies originate in necessary preprocessing and finite numerical precision. This operation adapts the system to sensor noise and increases the update rate, by reducing the number
60 Development and Evaluation of points that are processed in ICP registration. The system is also shown to be robust to random sensor noise and partial occlusion. The perceived point of failure is a low update rate. Average rates for the different simulations are collected in Table 5.18. These are consistently higher than 2 Hz, but should still be increased. Further reducing the number of points in registration would have a positive impact in this, but it would reduce accuracy. Other parameters that affect update rate are ICP internal parameters such as the maximum number of iterations and minimum estimation error. This problem can be better approached by using better dedicated hardware, as suggested in Section 4.6 that would take advantage of GPU optimized algorithms in PCL.
Chapter 6 Data Collection 6.1 Introduction The Cooperative Robot for Large Spaces Manufacturing (CARLoS) project aims to apply modern robotics solutions to the automation of stud welding in ship structures, while cooperating with human operators [64,65]. Several organizations are involved in this project, including INESC TEC, making the Automated Guided Vehicle (AGV) robot platform available for data collection for external robot localization. Figure 6.1: CARLoS AGV [64] 6.2 Localization Requirements As shown in Fig. 6.1, the AGV is tracked and can traverse irregular terrain, such as ramps or steps. This possibility requires a localization system in 6 Degrees of Freedom (DoF). It is, therefore, an example of a system that could gain from the proposed tracker. 61
62 Data Collection 6.3 Collection Environment Data was collected at LABIOMEP [53] with two Kinect for Xbox 360 devices mounted on tripods collecting Red, Green and Blue, and Depth (RGB-D) video and a Qualisys Oqus motion capture system tracking passive Infrared (IR) markers on the robot and on the sensors to provide their poses. This set-up is seen on Fig. 6.2. Note the IR markers on both the AGV and the sensors. The robot moves in the room, including climbing a ramp. Figure 6.2: Data collection environment at LABIOMEP The LABIOMEP tracking set-up uses 12 Qualisys Oqus cameras and cost around 150000C. Some of these cameras are highlighted in Fig. 3.5. The placement of Kinect devices followed the recommendation of Fig. 4.1 for improved occlusion handling and constructive data fusion. 6.4 Collected Data The results of the session were: • Data collected by the CARLoS robot from: –Structure IO 3D sensor –Hokuyo URG-04LX front Laser Range Finder –Hokuyo URG-04LX_UG01 back Laser Range Finder –SparkFun 9DOF MPU-915 IMU –CH Robotics UM7 IMU •RGB-D datasets from each Kinect for Xbox 360 device, while robot moves; • One dataset from a hand-held Kinect for Xbox 360 for mapping; • Ground truth using Qualisys motion tracking system with 12 Oqus cameras –4 high speed Oqus 310+ cameras with 640 ×512 resolution @ 1764fps –8 high resolution Oqus 400 cameras with 1696 ×1710 resolution @ 480fps
6.5 Conclusion 63 6.4.1 Model Point Clouds A point cloud can be sampled from 3D meshes of the AGV platform, to be used as a target model. Examples are shown in Fig. 6.3. These meshes are available with the rest of the dataset. Figure 6.3: Guardian platform mesh (left) and point cloud (right) 6.5 Conclusion These datasets are available for the CARLoS team for further development and validation of the self-localization and mapping systems and have been made available for future development in 6 DoF tracking at https://github.com/carlosmccosta/dynamic_robot_localization_ tests. This dataset shows dynamic movements of a robot in irregular terrain from two different perspectives with RGB-D data and ground truth, as well as the sensors’ poses in the global frame, making it very interesting for future work.
64 Data Collection
Chapter 7 Conclusion and Future Work Mobile robots and Automated Guided Vehicles (AGVs) operate in diverse environments and require adequate localization. The burden of this task can be taken from the robot by using video tracking for external localization. Six Degrees of Freedom (DoF) video tracking is a non-trivial problem. There are a number of significant design decisions that impact performance and accuracy, as well as the range of application of a tracker. Current efforts attempt to leverage new technology and are often tailored to specific targets and environments. A set of relevant software and hardware has been identified, for both development and deployment use, including a flexible robotics middleware platform and Red, Green and Blue, and Depth (RGB-D) library and devices. The system was developed with and will use Robot Operating System (ROS) and Point Cloud Library (PCL). The hardware used in development was the Microsoft Kinect for Xbox 360 sensor and Microsoft Kinect for Xbox One is recommended for deployment. The project brought many challenges including familiarization with different technologies and demanding hardware requirements that could not fully be met. During the development of this project, several approaches were explored, bringing a better understanding of their advantages and limitations and an appropriate path was established. Simulation results are satisfactory and promising for real-word application. The system is also shown to be robust to random sensor noise and partial occlusion. The perceived point of failure is a low update rate. This problem can be approached by using better dedicated hardware, as suggested in Section 4.6 that would take advantage of Graphics Processing Unit (GPU) optimized algorithms in PCL. The collected RGB-D datasets with 6 DoF ground truth are available for the Cooperative Robot for Large Spaces Manufacturing (CARLoS) team for further development and validation of the self-localization and mapping systems and has been made available for future development in 6DoF tracking. This dataset shows dynamic movements of a robot in irregular terrain from two different perspectives with RGB-D data and ground truth, as well as the sensors’ poses in the global frame, 65
72 REFERENCES [63] Open Perception Foundation. Documentation - Point Cloud Library (PCL), 2015. URL http://pointclouds.org/documentation/tutorials/gpu{_}install.php. [Accessed 2015-09-03]. Cited on page 37. [64] The CARLoS Project. The CARLoS Project, 2013. URL http://carlosproject.eu/. [Accessed 2015-08-19]. Cited on page 61. [65] Carlos M Costa, Héber M Sobreira, Armando J Sousa, and Germano M Veiga. Robust and Accurate Localization System for Mobile Manipulators in Cluttered Environments. 2015 IEEE International Conference on Industrial Technology, At Sevilla, (October):3308–3313, 2015. doi: 10.1109/ICIT.2015.7125588. Cited on page 61.