Benchmarking_Open
Abstract
The HARTU (Handling with AI-enhanced Robotic Technologies for flexible manufactUring) project, funded by the European Union, is committed to accelerating advancements in robust and flexible robotic handling. To achieve this, a comprehensive research and benchmarking dataset has been compiled and will be published openly.This dataset is created as a key result of the system validation efforts conducted through iterative test sprints, ensuring its relevance to real-world industrial application. It is designed to facilitate further research and technological development in AI-based handling and manipulation systems.
Full text
Test Sprint Tracking Evaluation of Primitive Pose Estimation KER15 WP2 ULMA – Use Case 7: Box to box order preparation Date: 30/03/2025
Handling with AI-enhanced robotic technologies for flexible manufacturing Test Sprint Tracking 1 This project has received funding from the European Union’s Horizon Europe Research and Innovation Programme under grant agreement Nº. 101092100 (HARTU).
Handling with AI-enhanced robotic technologies for flexible manufacturing Test Sprint Tracking 2 Index 1 TEST SPRINT IDENTIFICATION ....................................................................................................................................... 3 2 EXPERIMENT DESCRIPTION ........................................................................................................................................... 4 3 FUTURE PLAN/ACTIONS ................................................................................................................................................ 7 4 HOW THIS TECH CAN AFFECT TO END-USER (OPTIONAL) .............................................................................................. 7
Handling with AI-enhanced robotic technologies for flexible manufacturing Test Sprint Tracking 3 1 TEST SPRINT IDENTIFICATION Title Evaluation of Primitive Pose Estimation Iteration number 1 Description Testing of developed perception module after integration in the HARTU App Manager to analyse the performance achieved in this setup. Objectives Determine a measure of repeatability the module can achieve in this use case. Partner Name ASOCIACIÓN DE INVESTIGACIÓN METALÚRGICA DEL NOROESTE (AIMEN) Date and Location 30/03/2025 - Polígono Industrial de Cataboi SUR-PPI-2 (Sector 2) Parcela 3 36418, O Porriño (Pontevedra) KPIs T2.2 – Object and environment monitoring. UR-05, NR-29
Handling with AI-enhanced robotic technologies for flexible manufacturing Test Sprint Tracking 4 2 EXPERIMENT DESCRIPTION 2.1 Experiment setup The following materials were used for this experiment: a Photoneo MotionCam-3D Color Camera, a standard input container, several objects of the use case, and a computer. The following software modules were used for this experiment: • App Manager. • Image Acquisition Module. • Manual Segmentation Module. (for in-house testing only) • Primitive Pose Estimation Module. 2.2 Experiment execution The goal of this experiment is to provide a measure of repeatability of the pose estimation for relevant objects under the conditions defined for the setup of this use case. To make sure the module works well for the wide range of objects available, the following scenarios were selected to represent this variability: • Scenario 1 – Three objects were selected from each of three cardboard boxes variations. This selection is aimed to represent some variability in the pharmaceutical objects present in this use case. • Scenario 2 – Two objects were selected from each of three cylindrical object variations. This selection includes metallic cans and plastic pots. In this scenario, objects are placed in upright position. • Scenario 3 – Two objects were selected from each of three cylindrical object variations. This selection includes metallic tubes and plastic bottles. In this scenario, objects are placed sideways in the container. For each of these scenarios, the objects were placed inside the container in randomly selected positions and orientations, as shown in the following images. These positions were kept for the three iterations of this experiment to compare the estimated poses in each repetition. Data was collected to measure the distribution of cartesian centroids of the primitives, in X, Y, and Z coordinates, and the angular distribution of their orientation axis. In the case of the cardboard boxes, the orientation axis is the normal vector of the top surface. For cylinders, it is the vector along the main axis of the cylinder. These tests were executed by using a simple behaviour tree built with the App Manager. This tree contains three perception nodes: image acquisition, object segmentation, and pose estimation. 2.3 Experiment results The visual results of the scenarios tested during this experiment are shown in the following figures:
Handling with AI-enhanced robotic technologies for flexible manufacturing Test Sprint Tracking 5 For box primitives and cylinders viewed from the top, the centroid of the surface is represented as a red sphere for each estimated pose. Since each surface appears to have only one red sphere, this means that the precision on the estimated cartesian coordinates for the centroid is very good. Additionally, the normal vector to the estimated primitive is also represented as a line coming out of the sphere. For sideways cylinders, the estimated centroid is inside of the primitive shape, but the position and orientation of the primitive shape itself is a measure of the quality of the result. For scenario 1, the algorithm achieves a 0.11 mm of error on X, 0.14 mm of error on Y, and 0.02 mm on Z. Regarding the angular error, the results show that there is an average error of 0.52°. In scenario 2, results show that that there is an error of 0.69 mm for X, 0.70 mm for Y, and 0.07 mm for Z. As for the angular error, an average of 0.41° was observed for all objects across the three repetitions. Finally, in scenario 3, the algorithm shows an average position error of 0.43 mm for X, 0.83 mm on Y, and 0.65 mm on Z. Regarding the angular error, the algorithm shows an average error of 0.40° across all objects. The detailed results for all three scenarios can be found in the following tables. Translation Error (Scenario 1) Repetition 1 Repetition 2 Repetition 3 X (mm) Y (mm) Z (mm) X (mm) Y (mm) Z (mm) X (mm) Y (mm) Z (mm) Object A 0.1923 0.0012 0.0262 0.1893 0.0003 0.0191 0.3816 0.0008 0.0071 Object B 0.3731 0.0012 0.0143 0.1862 0.0003 0.0108 0.1869 0.0009 0.003 Object C 0.1873 0.1999 0.0317 0.1874 0.1998 0.0309 0.3747 0.3998 0.0627 Object D 0.0017 0.0042 0.0261 0.0005 0.0013 0.0084 0.0011 0.0028 0.0177 Object E 0.0003 0.3951 0.0269 0.0007 0.198 0.0084 0.0009 0.1971 0.0185 Object F 0.0054 0.364 0.0216 0.0017 0.1814 0.0032 0.0037 0.1826 0.0184 Object G 0.013 0.0027 0.0422 0.0067 0.0014 0.0218 0.0063 0.0013 0.0204 Object H 0.3902 0.3682 0.0105 0.1996 0.1821 0.0208 0.1906 0.1861 0.0103 Object I 0.0009 0.2098 0.0285 0.0014 0.2109 0.0285 0.0005 0.4207 0.0504 Orientation Error (Scenario 1) Repetition 1 Repetition 2 Repetition 3 Normal Error (°) Normal Error (°) Normal Error (°) Object A 0.3379 0.1297 0.4022 Object B 1.4893 0.8423 0.6472 Object C 0.1286 0.2059 0.1383 Object D 0.9071 1.0072 1.9145 Object E 0.0633 0.0758 0.0773 Object F 0.3243 1.1729 1.4906 Object G 0.5396 0.1818 0.3661
Handling with AI-enhanced robotic technologies for flexible manufacturing Test Sprint Tracking 6 Object H 0.2808 0.122 0.1589 Object I 0.3415 0.2343 0.5757 Translation Error (Scenario 2) Repetition 1 Repetition 2 Repetition 3 X (mm) Y (mm) Z (mm) X (mm) Y (mm) Z (mm) X (mm) Y (mm) Z (mm) Object A 0.1862 0.3725 0.0321 0.9308 1.3071 0.1078 0.7446 0.9346 0.0757 Object B 0.7282 0.7282 0.0088 1.274 1.2745 0.0388 2.022 2.0027 0.0299 Object C 0.3798 1.1307 0.0361 0.7814 1.1457 0.0981 0.4016 2.2764 0.0619 Object D 0.2135 0.1936 0.0258 0.2192 0.3989 0.0206 0.4357 0.2052 0.0464 Object E 0.6423 0.2204 0.0643 1.2506 0.9322 0.1417 1.8929 0.7118 0.2059 Object F 0.1947 0.0038 0.0232 0.1956 0.0047 0.0328 0.3907 0.0086 0.0561 Rotation Error (Scenario 2) Repetition 1 Repetition 2 Repetition 3 Normal Error (°) Normal Error (°) Normal Error (°) Object A 0.1348 0.1922 0.0573 Object B 0.3178 0.7329 1.0507 Object C 0.1472 0.1625 0.3097 Object D 0.2171 0.2601 0.4772 Object E 0.2515 0.4569 0.2055 Object F 0.2841 0.1612 0.1228 Translation Error (Scenario 3) Repetition 1 Repetition 2 Repetition 3 X (mm) Y (mm) Z (mm) X (mm) Y (mm) Z (mm) X (mm) Y (mm) Z (mm) Object A 0.0785 1.7575 1.0372 0.6896 0.7185 0.1501 0.7681 1.0391 0.8872 Object B 0.1226 0.1201 0.1931 0.2741 0.1001 0.2336 0.1515 0.0201 0.4268 Object C 0.0021 2.4248 1.2022 0.4604 0.2566 1.9912 0.4625 2.6814 0.789 Object D 0.0511 0.0034 0.2813 0.0477 0.0568 0.4792 0.0034 0.0564 0.1979 Object E 0.1364 1.5668 0.5425 0.2346 2.1676 0.7929 0.0982 0.6008 0.2503 Object F 2.0501 0.1512 0.9724 0.7891 0.3483 0.4436 1.2611 0.1971 0.5288 Rotation Error (Scenario 3 Repetition 1 Repetition 2 Repetition 3 Main Axis Orientation Error (°) Main Axis Orientation Error (°) Main Axis Orientation Error (°) Object A 0.1733 0.4765 0.4706 Object B 0.9317 0.5175 0.4761 Object C 0.2858 0.8484 0.5625 Object D 0.2054 0.2204 0.1711 Object E 0.0445 0.0594 0.0471 Object F 0.1901 0.0856 0.1287
Handling with AI-enhanced robotic technologies for flexible manufacturing Test Sprint Tracking 7 During the tests for scenario 3, it was noticed that for objects with a “step” in the body, our algorithm may differ slightly from one test to another regarding the start of this smaller section. Nonetheless, in every repetition the main body of the object was accurately represented by the primitive shape. In the end, this larger part is what is most important to represent since the grasping of the object will be done in this section. Overall, the primitive pose estimation module achieves good results for the wide range of objects and geometric primitives found in this use case. 3 FUTURE PLAN/ACTIONS Future activities will involve the testing of pose estimation in this use case for grasping in a more complex behaviour tree. This will allow the validation of these results with the robot. 4 HOW THIS TECH CAN AFFECT TO END-USER (OPTIONAL) This technology provides a flexible solution to perform pose estimation using geometric primitives as a base. Our module first performs a classification stage to identify what is the geometric primitive that most resembles the segmented object, which provides the system the flexibility to work on objects of varying shapes with minor reconfiguration. Afterwards, according to the identified primitive, the algorithm performs a best fit approach that defines the size, position, and orientation of the geometric primitive. The output of this module provides valuable data that may be used downstream for robotic grasping and manipulation. The performance of this module can be the difference between a correct and an incorrect grasp, so it is critical to test and validate these results under several conditions.