scieee AI-readable full text Open interactive document viewer

HAWK4Label: Efficient Weak-Supervised Method for Object Segmentation Labeling

Carlon, Nicola; Bastianello, Riccardo; Gottardi, Alberto; Tonello, Stefano; Menegatti, Emanuele

Abstract

Data labeling poses significant challenges in data quality, domain discrepancies, overfitting, and the intensive effort required for accurate annotation. This is significantly pronounced in industrial contexts, where existing datasets often fail to represent real-world complexities. To address these challenges, we introduce an innovative image data labeling solution. HAWK4Label exploits a weak-supervised approach for the auto-labeling process mixed with the human's feedback to improve the label quality. Experiments demonstrate that the instance segmentation model trained on datasets generated with HAWK4Label performs better than a manually labeled dataset with an IoU of 97\% (w.r.t. 90\%).

Full text

HAWK4Label: Efficient Weak-Supervised Method for Object Segmentation Labeling Nicola Carlon1,∗, Riccardo Bastianello1,∗, Alberto Gottardi1,†, Stefano Tonello1, Emanuele Menegatti2 1: IT+Robotics srl, 36100 Vicenza, Italy 2: Dept. of Information Engineering, University of Padova, Italy {name.surname}@it-robotics.it,[email protected] Abstract—Data labeling poses significant challenges in data quality, domain discrepancies, overfitting, and the intensive effort required for accurate annotation. This is significantly pronounced in industrial contexts, where existing datasets often fail to represent real-world complexities. To address these challenges, we introduce an innovative image data labeling solution. HAWK4Label exploits a weak-supervised approach for the auto-labeling process mixed with the human’s feedback to improve the label quality. Experiments demonstrate that the instance segmentation model trained on datasets generated with HAWK4Label performs better than a manually labeled dataset with an IoU of 97% (w.r.t. 90%). I. INTRODUCTION The rapid advancement of deep learning (DL) has significantly broadened the research landscape, demonstrating that supervised learning techniques can substantially improve recognition performance and support the execution of complex tasks in industrial scenarios [1], [2]. Data labeling is a major bottleneck in developing DL pipelines since annotating largescale datasets is costly, time-consuming, and prone to human error and inconsistency. Research has focused on reducing these burdens through transfer learning [3], which transfers knowledge from a source to a target domain with limited annotated samples. However, achieving high predictive accuracy requires large amounts of labeled data. Early self-supervised robotic perception exploited manipulation or environmental consistency to generate labels without manual annotation, often relying on temporal scene changes [4], [5]. These methods are not able to handle occlusions and cluttered scenes. Other works [6], [7] introduced grasp-based segmentation by subtracting known backgrounds and isolating the objects. Beyond physical interaction-based labeling, a complementary line of research has focused on simulation-driven and synthetic approaches to automatically annotate training data [8], [9]. However, these approaches required detailed object models and calibration, and the results can suffer from visual artifacts when used in the real world. This paper proposes HAWK4Label, an innovative method for efficient object segmentation labeling. HAWK4Label addresses the limitations of manual and fully supervised labeling by mixing 2D and 3D data during the image collection phase ∗Authors equally contributed to the work, †Corresponding Author This work has been partially funded by the European Union under the project AI REDGIO 5.0, grant agreement No 101092069 Fig. 1: HAWK4Label system schema. and integrating a weakly supervised approach with human-inthe-loop feedback mechanisms. The fusion of 2D-3D allows for managing occlusions, object variability, and the industrial environment, overcoming the current SOTA methods. II. METHODS During the depalletizing process, images are acquired and automatically annotated in real-time by a self-supervised labeling module. The generated annotations are subsequently reviewed through a dedicated user interface, where the operator can provide qualitative feedback to refine label accuracy. If there are artefacts in the annotations due to subtraction or reprojection errors, the operator can modify the annotation via the graphical interface. This closed-loop process incrementally improves the labeling performance with minimal manual intervention. As shown in Fig. 1, HAWK4Label is composed of 4 main phases: (i) Data Acquisition and labels generation, where the segmented masks are computed; (ii) Self-supervised annotation refinement that leverage Segment Anything (SAM) v2 [10] to refine the segmented masks; (iii) Human Feedback to improve label quality, when the worker can modify the annotations and provide qualitative feedback; (iv) Final robotics picking system, when the dataset is used to train instance segmentation models. A. Data Acquisition and Labels Generation The RGB-D vision system monitors manual depalletization cycles by capturing detailed 3D point clouds to generate annotations aligned with 2D images automatically. The method subtracts point clouds from consecutive acquisitions—before and after an item is removed—to isolate the object’s 3D 2025 I-RIM Conference October 17-19, Rome, Italy ISBN: 9788894580570 10.5281/zenodo.17629650 67 Fig. 2: Acquisition and annotation generation phase. The images in the blu box represent the consecutive data acquisition. The resulting labels are depicted in the bottom image. (a) (b) Fig. 3: Results of the subtraction techniques only in 2D. 3a is the original scenario, while 3b is the corresponding generated masks. The masks are incomplete because the missing parts have the same colors as the background, while the shadow is recognized as a part to be segmented. representation, which is then projected into the initial 2D image to form a precise segmentation mask. Starting from T0(full scene), each removal triggers a new acquisition Tj; subtracting Tjfrom Tj−1yields the removed item’s cloud, which is projected onto the T0image. This continues until all items are removed, with the final step using the empty scene Tn. Leveraging 3D data avoids artifacts from colored backgrounds and shadows that would occur with 2D-only subtraction, ensuring accurate object labeling (Fig. 2-3). III. EXPERIMENTS The proposed HAWK4Label system was validated in a realworld depalletization task with water bundles, can packs, and cookie boxes. HAWK4Label automatically generated segmentation masks and trained an instance segmentation model, which was later deployed on an ABB manipulator for pickand-place operations. The same datasets were manually annotated to establish a baseline. Experiments were run on a Lenovo ThinkPad (Intel i7, 16 GB RAM, NVIDIA T500), and a Zivid2 MR60 camera was used. Using Intersection over Union (IoU ≥0.75) as the identification criterion, the model trained with HAWK4Label annotations achieved 97.4% accuracy for water bundles and 96.7% for both can packs and cookie boxes, outperforming manual labels (84.6% for water and cans, 90.3% for cookies). This improvement shows the pipeline’s robustness across different object textures, including challenging reflective shrink-wrapped surfaces. The small accuracy gap between object classes likely stems from differences in visual complexity and occlusion. Moreover, HAWK4Label can label 8 pallets/day versus 2 pallets/day manually, offering a fourfold productivity boost without quality loss. When deployed for real depalletization, the HAWK4Label-trained model yielded a 90.4% dexterous grasp success rate compared to 81.3% with manual labels, confirming better grasp point identification and execution. The deadlock rate dropped to 0.21% (vs. 0.37%), indicating fewer perception or planning failures. These results highlight that accurate automated labeling accelerates dataset creation and translates directly into more reliable, continuous robotic operation in industrial settings. IV. CONCLUSIONS This paper introduced HAWK4Label, a weakly supervised method that combines 3D differencing, SAM v2 refinement, and minimal human validation to automate dataset annotation for robotic manipulation. In depalletization tests, it achieved up to 97.4% accuracy, surpassing manual labeling. The system is robust, scalable, and suitable for industrial deployment. Future work targets advanced data augmentation and multiview dynamic scene understanding to reduce operator effort. REFERENCES [1] K. Makantasis, K. Karantzalos, A. Doulamis, and N. Doulamis, “Deep supervised learning for hyperspectral data classification through convolutional neural networks” In 2015 IEEE international geoscience and remote sensing symposium. [2] L. Barcellona, A. Bacchin, A. Gottardi, E. Menegatti, and S. Ghidoni, ”Fsg-net: a deep learning model for semantic robot grasping through few-shot learning,” in 2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2023, pp. 1793–1799. [3] N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, ”Parameter-efficient transfer learning for nlp,” in International conference on machine learning. PMLR, 2019, pp. 2790–2799 [4] T. Faulhammer, R. Ambrus¸, C. Burbridge, M. Zillich, J. Folkesson, N. Hawes, P. Jensfelt, and M. Vincze, ”Autonomous learning of object models on a mobile robot,” IEEE Robotics and Automation Letters, vol. 2, no. 1, pp. 26–33, 2017. [5] A. Zeng, K.-T. Yu, S. Song, D. Suo, E. Walker, A. Rodriguez, and J. Xiao, ”Multi-view self-supervised deep learning for 6d pose estimation in the amazon picking challenge,” in 2017 IEEE international conference on robotics and automation (ICRA). IEEE, 2017, pp.1386–1383. [6] V. Florence, J. J. Corso, and B. Griffin, ”Robot-supervised learning for object segmentation,” in 2020 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2020, pp. 1343–1349. [7] Y. Liu, X. Chen, and P. Abbeel, ”Self-supervised instance segmentation by grasping,” in 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2023, pp. 1162–1169 [8] T. Muller, A. Evans, C. Schied, and A. Keller, ”Instant neural graphics primitives with a multiresolution hash encoding,” ACM transactions on graphics (TOG), vol. 41, no. 4, pp. 1–15, 2022 [9] J. Dirr, J. C. Bauer, D. Gebauer, and R. Daub, ”Cut-paste image generation for instance segmentation for robotic picking of industrial parts,” The International Journal of Advanced Manufacturing Technology, vol. 130, no. 1, pp. 191–201, 2024 [10] N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Radle, C. Rolland, L. Gustafson, et al., ”Sam 2: Segment anything in images and videos,” arXiv preprint arXiv:2408.00714, 2024 68