Robot Programming by Demonstration: Segmentation and Via-Set Optimization
Abstract
Programming by Demonstration (PbD) offers an intuitive way to program robots, but turning hand-guided demonstrations into precise and adaptable skills remains challenging, especially for complex interactions. In this work, we present a novel PbD framework that improves skill programming and adaptive execution for robots operating with physical interfaces, such as the one available on aircraft cockpits.
Full text
Robot Programming by Demonstration: Segmentation and Via-Set Optimization Isacco Zappa Politecnico di Milano [email protected] Mohammad H. Saeidi Mostaghim Politecnico di Milano [email protected] Andrea Maria Zanchettin Politecnico di Milano [email protected] Paolo Rocco Politecnico di Milano [email protected] Abstract—Programming by Demonstration (PbD) offers an intuitive way to program robots, but turning hand-guided demonstrations into precise and adaptable skills remains challenging, especially for complex interactions. In this work, we present a novel PbD framework that improves skill programming and adaptive execution for robots operating with physical interfaces, such as the one available on aircraft cockpits. Index Terms—Programming by Demonstration, Trajectory Alignment and Segmentation, Via-Set Optimization, Robot Skills Programming I. INTRODUCTION Programming by Demonstration (PbD) provides an accessible solution for programming robot skills [1]. Recent advancements have shown that hand-guided demonstrations can be used to learn both the action models required for task planning and their execution policies in terms of an objectcentered waypoint, or via-set, list [2]. However, applying PbD effectively to tasks involving complex interactions inside of cluttered environments remains a significant challenge. Developing PbD frameworks for programming robust skills with a few demonstrations, that can be executed efficiently in real time and adapted to variable environments remains a key challenge. A major difficulty is translating demonstrations into compact, structured execution policies that generalize well and support runtime adaptation. Existing methods for trajectory segmentation or feature extraction often fail to consistently capture task-relevant elements (e.g., keyframes) from different demonstrations [3]. Indeed, identifying critical movement primitives, sensitive to task constraints, remains difficult for PbD approaches. Further beyond the representation problem, translating learned skills into reliable, adaptive robot motions under strict runtime constraints adds complexity. The challenge extends from achieving accurate target poses to designing computationally tractable optimization and motion planning strategies that produce feasible, collision-free trajectories in near real time. This paper is therefore motivated by the need to overcome these limitations in PbD for robust skill programming and efficient, adaptive execution. This work makes three main contributions. First, we develop a custom event-informed Dynamic Time Warping (DTW) This work was partly supported by Ministero dell’Industria e del Made in Italy under Progetto Accordo per l’Innovazione “ARTO - Automatic Robotic for Testing Optimisation”, code F/350238/01-03/X60, CUP: B49J25000480005. Fig. 1. Visualization of learned and refined skill (SC: State Change, WP: Waypoint, GC: Gripper Change). method that aligns multiple variable demonstrations by incorporating both kinematic data and skill-relevant events. Second, we introduce Segmentation via Dynamic Programming (SEGDP), a multi-modal, multi-dimensional algorithm for extracting skill primitives. Third, we propose a via-set optimization module that leverages statistical models at detected keyframes to refine execution at runtime. A representation of the pipeline’s outcome can be seen in Figure 1. The framework is implemented on a UR5e cobot for an input device testing task within an aerospace cockpit mock-up. Experiments with kinesthetic teaching show that the system can learn adaptable skills even from one-shot demonstrations, execute testing tasks precisely, and adjust to changes and cluttered environments. II. METHODOLOGY The proposed PbD framework comprises dedicated teaching and execution modules that convert kinesthetic demonstrations into precise, adaptive robot skills. It can operate from a single demonstration but achieves best performance with multiple demonstrations. 2025 I-RIM Conference October 17-19, Rome, Italy ISBN: 9788894580570 10.5281/zenodo.17629840 205
Fig. 2. a) Cost matrix of custom DTW algorithm for two trajectories. b) Results of SEGDP and fastSEGDP A. Trajectory Alignment and Segmentation To address temporal variability in human demonstrations, we introduce a custom DTW algorithm. Unlike standard approaches based solely on kinematic similarity, our method augments the cost function with penalties for mismatching critical task events such as input device signals, gripper actions, or waypoint saving triggers by the user. This eventweighting produces semantically meaningful alignments, crucial for learning from event-rich interaction tasks. Figure 2-a illustrates the Dynamic Programming (DP) cost matrix of the event-informed DTW (Equation 1). C(i, j) = d(p1,i, p2,j ) + min{C(i−1, j), C(i, j −1), C(i−1, j −1)} (1) The standard cost d(p1,i, p2,j)is augmented with the event penalty term defined in Equation 2 to penalize misalignments of same-class events. Ce(p1,i, p2,j) = X k∈Ecommon Sk(p1,i)=Sk(p2,j ) we,k (2) After alignment, the trajectory tube from multiple demonstrations is segmented to extract keyframes. The goal is to approximate it with an optimal number of linear hypertubes anchored at the boundaries, reducing decision variables and easing optimization. To prevent overfitting, the objective penalizes excessive segmentation, rewarding only segments with substantial contribution. Additional penalties discourage covering free space or neglecting the trajectory tube. Although the cost manifold is non-convex, DP provides an effective solution, hence the name SEGDP. After the initial sweep, further recursion follows Equation 3, with the geometric segment cost refined to exclude foreign events: Cseg(sa, sb) = Cg(sa, sb) + Cevent(sa, sb), where Cevent(sa, sb)would be a non-negative penalty applied if events co-occur undesirably within the segment [sa, sb]. DP [t, m] = min m−2≤j<t{ DP [j, m −1] + Cg(j+ 1, t)} (3) While flexible and intuitive, SEGDP has a complexity of O(N3). To address this, we introduce fastSEGDP, which replaces latched linear hyper-tubes with Ordinary Least Squares (OLS) fits of the upper and lower tube trajectories. This Fig. 3. Forward pass of sliding window optimization. modification ensures super-additivity in the cost function, enabling the use of PELT with prefix sums for efficient cost evaluation. The resulting complexity is reduced to O(N). Figure 2-b shows the segmentation results. B. Via-Set Optimization After segmentation, the method extracts the multidimensional mean and covariance (µk,Σk) at each keyframe. At runtime, an optimal inverse kinematic branch b∗is chosen by minimizing the cost to the first keyframe, considering manipulability, obstacle avoidance, and C-space trajectory length. This branch guides the Via-Set Optimization, which refines trajectories through hyper-ellipsoidal sets defined by scaled covariances. The objective combines C-space trajectory length minimization with task-space Mahalanobis distance to the µk penalties, constrained to remain within the safety hyperellipsoids defined by the scaled covariance c2 k. Since solving this optimization online is expensive, we employ a sliding-window strategy over sub-trajectories to keep it feasible (Figure 3). The refined trajectory is then passed to a motion planner, with fallback to alternative branches b∗ (2) if needed. For tasks requiring sub-millimeter accuracy, the covariance bounds are expanded proportionally to enforce stricter refinement. III. EXPERIMENTAL VALIDATION The framework was validated through experiments on six cockpit input-device skills, including button pressing, lever pushing, knob twisting, and switch actuation. Each skill was taught via singleand multi-shot demonstrations under varying conditions, achieving a 91% success rate, confirming robustness and precision. Performance can improve further by refining keyframes and adding critical waypoints. Finally, a usability study with 10 non-experts yielded a System Usability Scale score of 79.2 and a 70% execution success rate, showing that even first-time users can provide effective demonstrations. REFERENCES [1] Lucci, N., Montini, E., Zappa, I., Zanchettin, A. & Rocco, P. Intuitive Cobot Programming for Small-Medium Enterprises. European Robotics Forum. pp. 241-246 (2024) [2] Zanchettin, A. Symbolic representation of what robots are taught in one demonstration. Robotics And Autonomous Systems.166 pp. 104452 (2023) [3] Sen, B., Elfring, J., Torta, E. & Molengraft, R. Semantic learning from keyframe demonstration using object attribute constraints. Frontiers In Robotics And AI.11 pp. 1340334 (2024) 206