Teaching contact-rich tasks from visual demonstrations by constraint extraction
Full text
Teaching contact-rich tasks from visual demonstrations by constraint extraction Christian Hegeler1, Filippo Rozzi2, Loris Roveda3, and Kevin Haninger1 Abstract—Contact-rich manipulation involves kinematic constraints on the manipulated object’s motion, typically with discrete transitions between constraints. Allowing the robot to detect and reason about these contact constraints can support robust and dynamic manipulation, but how can these contact models be efficiently learned? Purely visual observations are an attractive data source, allowing passive task demonstrations with unmodified objects. To use visual demonstrations for contact-rich tasks, we propose a method where object pose trajectories are clustered by contact modes, and then parameterized constraint types are selected and fit to minimize constraint violation. The fit constraints are then used to (i) detect contact online with force/torque measurements and (ii) change the robot policy with respect to the active constraint. We demonstrate the approach with real experiments, on cabling and rake tasks, showing the approach gives robust manipulation through contact transitions. Index Terms—Learning from Demonstration, Force Control, Contact-rich Manipulation I. INTRODUCTION For robots to dynamically interact with their environments, robust control policies for contact are needed [1]. Articulated environments such as drawers and doors, as well as manipulation tasks such as cabling and assembly tasks, involve kinematic constraints which typically change during task execution, have geometric uncertainty, but still must be respected to avoid excessive force. A further challenge is that these kinematic constraints are application-specific, depending on the manipulated object and environment [2]. Model-based contact planning methods are capable of providing dynamic interaction that respects collision forces, Coulomb friction constraints, and actuation limits [3], [4]. However, they need contact models and a reference trajectory, which are application-specific and often derived from simulation models [5]. However, they can also be learned from demonstrations, e.g. with an instrumented tool that measures the pose and forces of a human [2]. Contact monitoring methods can use force and position data to identify contact mode and respond [6], [7], [8], but 1Department of Automation at Fraunhofer IPK, Berlin, Germany 2Politecnico di Milano, Department of Mechanical Engineering, Milano, Italy 3Istituto Dalle Molle di Studi sull’Intelligenza Artificiale (IDSIA), Scuola Universitaria Professionale della Svizzera Italiana (SUPSI), Universit` a della Svizzera Italiana (USI) IDSIA-SUPSI, Lugano, Switzerland [email protected] Corresponding author: [email protected]. The authors have no relevant financial or non-financial interests to disclose. This project has received funding from the European Union’s Horizon 2020 research and innovation program under grant agreement 101058521 — CONVERGING. Fig. 1: Setup of the approach, where the keypoint trajectory of an object is tracked by an RGBD camera (left), the data are clustered, and kinematic constraints are fit. The constraint model is then exploited by the robot online (right), where measured force is compared with the constraint Jacobians to identify constraint online and the robot’s trajectory is adapted to enter and maintain constraints. also require contact models. The constraints in a task provide high-level task information which can be used for longerhorizon planning and monitoring [9]. Learning from demonstration (LfD) for contact-rich tasks can reproduce position and force trajectories, but require an additional force/torque sensor to separate human and environment forces [10], [11], [12]. Learning from demonstration (LfD) can specify robot tasks in an efficient and natural way [13], but also has challenges teaching contact-rich tasks. Either a 2nd force/torque sensor is needed [10], [11], [12], or the robot must be teleoperated [14], [15], limiting ease of demonstration and dexterity in fine tasks, respectively. Furthermore, this approach must balance whether the force or position trajectory should be tracked, a question of the robot’s impedance [16]. This can be optimized when the robot reproduces the trajectories [17], [11], however, introducing additional complexity in the definition of the reward function, data collection, and implementation. The need for a contact geometry model or an additional force/torque sensor limits the scalability of model-based or LfD approaches. We consider a contact-rich task where the manipulated object is subject to a discrete collection of kinematic constraints and propose a method that extracts these contact geometries from visual demonstrations 2025 IEEE 21st International Conference on Automation Science and Engineering (CASE) August 17-21, 2025. Los Angeles, CA, US 979-8-3315-2246-9/25/$31.00 ©2025 IEEE 3500 2025 IEEE 21st International Conference on Automation Science and Engineering (CASE) | 979-8-3315-2246-9/25/$31.00 ©2025 IEEE | DOI: 10.1109/CASE58245.2025.11164000 Authorized licensed use limited to: University of Patras. Downloaded on November 04,2025 at 20:17:43 UTC from IEEE Xplore. Restrictions apply.
Demonstrate Online Learn constraints trajectory mode [1, ..., M] Cluster hm(x,θm) ID constraint type Detect current constraint Force Jm Fit constraint params Relative pose (TCP and object) TCP Object Approach next constraint Maintain current constraint Pose reference Mode-dependent trajectory object trajectory Constraint Impedance Control Force reference Jm+1 Fig. 2: The components of the proposed approach, where demonstrations and constraint learning are covered in Section III and online control in Section IV. of an object pose trajectory. We make three motivating observations: (i) due to motor control noise and low stiffness, repeatable aspects of human demonstrations are induced by constraints, (ii) constraints are one of a few categories (e.g. point contact, line on plane, plane on plane, hinge) where only certain parameters must be identified, and (iii) exact tracking of a force trajectory is often not needed - simply a lower bound to keep contact and upper bound for safety is often sufficient. We propose a method that records demonstrations with object keypoint trajectories, clusters them by discrete contact mode, identifies a suitable constraint type, and fits the constraint parameters to minimize the constraint violation. These constraints are then used for online control in two ways: (i) to derive the direction of expected force by the Jacobian of the constraint, allowing online classification of the contact condition from F/T measurements, and (ii) to set a desired force to maintain the current constraint and approach the next one. We validate this approach with real experiments to demonstrate a cabling and raking task. We find that clustering and fitting parameterized models is effective even on noisy RGBD data and that online detection with constraint Jacobians is effective – provided the Jacobians are distinguishable. Compared to other constraint-learning work [2], this method doesn’t need additional motion markers or F/T sensors to collect demonstrations and the learned models are applied online. Compared with deep-learning approaches for LfD [18], [19], this method addresses high-stiffness contact-rich manipulation and uses human-readable visual information; keypoints. Compared with variance-aware [20] or stiffness-aware [21] LfD, our method explicitly models constraints, and changes in contact, and can detect changes in contact online. II. PROBLEM STATEMENT AND APPROACH This section formalizes the type of considered manipulation problem and then overviews the proposed approach. The contact-rich manipulation problem is expressed with the pose of a manipulated object, x∈RD, where Dis either R3for a point observation or SE(3) for an object pose. During a task, the object has a trajectory of Tpoints, X= [x1, . . . , xT]. There are discrete contact conditions m∈[1, . . . , M], where which mode active at timestep n, mn, is unknown. Each mode has a holonomic constraint of the form hm(x)=0, where h:RD→Rnmand nmis the degree of the constraint in m[22]. A constraint hm(x)is enforced with constraint forces of the form F=λTJm(x), where F∈RDare the forces expressed in the coordinate system of x,λ∈Rnmthe Lagrange multipliers representing the magnitude of the force and Jm=∂hm/∂x the constraint Jacobian. Given only pose trajectories X, where a sequence of contact modes [1, . . . , M]implicitly occur, it is desired that the robot can safely reproduce this sequence of contacts over small geometric variations in the constraints. The proposed approach can be seen in Figure 2. First, the data is clustered to produce Mdatasets, each corresponding to a constraint mode. Then, a constraint model is fit for each data cluster, where it is assumed that the constraint is one of a library of parameterized models hm(x, θm), where the parameters θmare parameters fit to the clustered data. The fit models then give Jacobians Jm, which give the direction or subspace in which the mode’s constraint forces are expected. Force measurements allow the identification of the constraint, and can inform the robot policy. III. VISUAL DEMONSTRATIONS This section describes how the visual demonstrations are collected, the data segmented, and the constraints fit. A. Pose estimation Poses of objects, such as the plug and rake seen in Figure 1, are estimated with keypoints. Keypoints have a fixed relation to local visual object features and are ideally invariant to rigid transformations of object poses. The keypoint location and object pose is estimated from RGB images in 2D pixel space based on the work of [23] and are subsequently transformed into 3D world coordinates with the associated depth information and known camera pose calibrated with [24]. As shown in [23], the used framework is able to detect keypoints on unseen objects within trained object classes, providing some generalizability. For learning keypoint detection for new object classes, image data is collected autonomously by the robot and labeled semiautomatic as keypoints are labeled in one image, then projected into all other collected images. When using three or more distinguishable keypoints, related object poses can be unambiguously described in SE(3). For an initial keypoint configuration, an initial object pose is assigned with the center at the mean of all related keypoints and identity as rotation. When detecting keypoints for any new object pose, the pose is described by a transformation of the initial object pose that aligns both sets of keypoints. 3501 Authorized licensed use limited to: University of Patras. Downloaded on November 04,2025 at 20:17:43 UTC from IEEE Xplore. Restrictions apply.
0.50 0.25 0.00 Position (m) 0.0 0.1 Likelihood 0 1 2 0 50 100 150 200 250 300 350 400 Time step 0.00 0.05 0.10 Violation free_space plane hinge Fig. 3: Segmentation results for Zen Garden Raking. The upper plot shows the mean positions, the middle the projected likelihood, and the bottom the violation function ∥hm(xt)∥of the fit constraints to indicate when the different constraints are active. The estimated transitions (marked with vertical dotted lines) are within 10 time steps of sudden transitions in the violation function, indicating sufficient accuracy in segmentation. B. Segmentation The full demonstrated trajectory Xhas unknown transitions between the discrete constraints. To partition into a temporally ordered sequence of clusters, in which just one constraint per cluster is present, we model the overall trajectory including the time component as a Gaussian Mixture Model (GMM) [25]. By including time in the regressor, the linear effects of time on trajectory are modeled with a covariance term. Thus, the changes in velocity and local motion between contact modes are better fit by different covariances for the mixture models. The main advantage of such an estimation framework is that there is no need to manually predefine any internal parameters according to the types of given tasks and/or trajectories. The trajectory is augmented with time, X=X × T ∈ R(D+1)×T, where Xis the D-dimensional pose observations, T= [τ1, . . . , τT]the one-dimensional temporal variable (i.e., time stamp), and Tis the length of the trajectory. Including time allows temporal correlations to be fit to capture motion characteristics. The GMM is fit with the standard Expectation Maximization approach, giving a Gaussian distribution for the mth cluster of µm= [µm,τ , µm,x],Σm=Σm,τ Σm,τx Σm,xτ Σm,x ,(1) where τand xrefer to the one-dimensional temporal variable and the D-dimensional spatial variable. Additionally, the weight vector w1, . . . , wMis returned from the fitting process, giving the prior probability of each mode. The posterior likelihood of the mode mat time step tis calculated by γm,t =wmN(τt;µm,τ ,Σm,τ ) PM i=1 wiN(τt;µi,τ ,Σi,τ ),(2) where γm,t is the posterior likelihood of the mode mbeing active at time t. The results of this process can be seen in Figure 3. To assign modes, each data point is assigned to a set based on its maximum likelihood, and the overall trajectory is thus split into Mmodes, with corresponding data Xm. D free D radius D line D hinge free 0.00 0.00 0.00 0.00 radius 1.77 0.18 0.91 0.05 line 3.07 4.61 1.69 20.50 hinge 9.17 9.11 12.27 0.39 TABLE I: Constraint selection, showing the optimized loss from (3) for each pair of the model (row) and hand-labeled data segments (column). C. Constraint Fitting We now have the segmented data, where we assume each trajectory segment contains one constraint and next identify which type of constraint and fit the parameters. We assume that the kinematic constraints are holonomic constraints in the form hm(x, θm)=0, where θmare parameters describing the geometry and are to be fit, and hm is a violation function which characterizes the constraint. The constraint is assumed to belong to a set of possible contact primitives [2], such as point-on-plane, or line constraint. The constraint parameters are fit by min θmX xt∈X m ∥hm(xt, θm)∥2 2+βm(θm) + Ns (3) s.t.∥hm(xt, θm)∥2 2≤s∀xt∈ Xm(4) where βmis a mode dependent regularization term and sa slack variable describing the maximum of hmwhich adds an H∞norm penalty on hm(xt, θm). Identifying the proper constraint from the library can be done by identifying which constraint has the best fit (i.e. lowest loss from (3)) for the data segment. We show the results of this for hand-labeled data segments in Table I. Note that free-space must be otherwise identified, as it has a loss of 0. In this case, free-space is typically the first mode, so can be automatically labeled. This framework provides a general way to cluster and fit constraints, where each type of constraint must only specify hm(x, θm). Next, we describe several constraints used in the demonstrations, along with their parameterizations. 1) Radius Constraint: The radius constraint describes a point x∈R3with a fixed radius r∈Rto a point xc∈R3 in base coordinates. All possible values for xare therefore distributed on the surface of a sphere centered at xc. All points xtmust therefore fulfill the equality constraint ∥xt− xc∥2=r. The radius constraint is thus fit by using hm(xt, θm) = r− ∥xt−xc∥2, βm(θm) = kr, (5) where kis a regularization constant. 2) Line on Plane Constraint: The line-on-plane constraint describes an object in line contact with a fixed plane. The object can be translated and rotated along the plane such that the line contact between the object and the plane is never broken. The pose of the line is described by two introduced points p0, p1∈R3on the line with a relative position in object coordinates and a plane in world coordinates parameterized with its normal vector np∈R3and offset d∈R. The described constraint is formulated such that for all object poses in a data set Xmthat correspond to the described 3502 Authorized licensed use limited to: University of Patras. Downloaded on November 04,2025 at 20:17:43 UTC from IEEE Xplore. Restrictions apply.
constraint mode the two introduced points are located on the plane to fulfill the equality constraint with hm(xt, p0, p1, np, d) = |∥nT p(T(xt)p0)∥2−d| |∥nT p(T(xt)p1)∥2−d|+∥nT p(p0−p1)∥2 ∥nT p(p0−p1)∥2,(6) βm(θm) = k(|∥np∥2-1|+|∥p0-p1∥2-kD|+∥p0∥2+∥p1∥2),(7) where T(x)∈R4×4is the transformation matrix built from the pose x,kis a regularization constant and kDis an arbitrary but fixed distance of p0and p1in meter. 3) Hinge constraint: A single fixed line in the object coordinate system has a fixed position in world coordinates. The line contact is parameterized with two introduced points on this line xc,0, xc,1∈R3in world coordinates and p0, p1∈R3 in object coordinates. Using transformation matrices from the object to the world coordinate system T(xt)one can express the described relationship as xc,i =T(xt)xp,i,i= 1,2. The constraint is thus fit with hm(xt, θm) = ∥xc,0−T(xt)xp,0∥2 ∥xc,1−T(xt)xp,1∥2,(8) βm(θm) = k(|∥p0−p1∥2−kD|+∥p0∥2+∥p1∥2),(9) where kand kDfollow the description from III-C2. The upper bound sfor this constraint by the H∞norm is omitted. IV. ONLINE CONTROL This section describes how the contact models can be employed in the online control of the robot. A. Contact Detection Given a constraint hm(x, θm)∈Rnm, the forces which would enforce that constraint can be found as F=JT m(x)λ+ϵ, (10) where force F∈RDis in the coordinate system of x, Jacobian Jm(x) = ∂hm(x,θm) ∂x ∈Rnm×D,λ∈Rnmare Lagrange multipliers which enforce the constraint and are not measured, and ϵerror forces. The error forces ϵcan be from gravitational or inertial forces of the payload, friction in contact, or other unmodelled effects. In order to efficiently detect the active constraint in the online framework we need a robust metric that could scale properly for constraints with higher dimensions (nm>1). For a given m,xt, and Ft, we can find ∥ϵ∥2, the magnitude of force not explained by the constraint; i.e. the residual. This least-squares problem is solved by the pseudoinverse Jm(x)† of the constraint Jacobian with ϵ(F, x, m) = min λ∥F−JT m(x)λ∥(11) =∥F−Jm(x)†Jm(x)F∥2.(12) This residual tells us the amount of measured force that is not explained by the constraint and gives a metric for online detection: the constraint with a lower residual is selected as the current constraint. As the measured wrench vector Ftis measured at the TCP, it needs to be transformed to match the orientation of the coordinate system of the pose x, expressed with XY Z extrinsic Euler angles. We transform the wrench into a frame at the origin of the object frame and with the same orientation as the base frame [26]. Free space is also detected online but has no force residual. To robustly detect free space in the presence of acceleration forces we classify the mode as free space when ∥Fn∥< F∀l∈[t−L, t], where F∈Ris a threshold and La time window. B. Constraint-aware Control The constraint information can also be directly used in the control strategy. We use a lower-level admittance controller, such that M¨xr t+D˙xr t+Kxd t−xr t=Fd t−Ft,(13) where Fd∈R6is the force reference, Ftthe measured force in robot TCP, xd tthe desired pose of the robot, xr tthe robot pose and M,D,K∈R6×6the mass, damping, and stiffness admittance parameters. All quantities in (13) are expressed in the TCP frame, and the matrices M,D, and Kare diagonal. The following sections propose how the reference signals Fd tand xd tcan be adjusted online according to the mode information. 1) Contact-triggered Control: At each time step t, we infer the current mode mtas shown in Section IV-A. We then suppose that each mode has its own motion strategy as xd t=πmt(xr t),(14) where πmtis a mode-specific controller which produces a desired pose xd t. 2) Making and Maintaining Contact: Many constraints are unilateral constraints more accurately considered as h(x)≥0, for example sliding on the contact of a surface. This means the constraint can be lost by moving away from the constraint, which may be undesired for the task and limits the contact detection. Additionally, the pose where contact is entered may vary. Additional motion in the direction of the next constraint, as given by the Jacobian of the next constraint, can help ensure that contact is made. This is done by adding to the desired contact wrench. We add a small desired force to the robot force controller as Fd t=TwJT mt(xt)α+JT mt+1 (xt)α+,(15) where Twtransforms the wrench from object coordinates to TCP, mtis the current mode, mt+1 the next contact, α∈Rnmt, α > 0specifies the force to apply to maintain the current constraint and α+the force to approach the next constraint. 3503 Authorized licensed use limited to: University of Patras. Downloaded on November 04,2025 at 20:17:43 UTC from IEEE Xplore. Restrictions apply.
Fig. 4: Segmentation of cable pulling data with fit parameters for the cable constraint. The distance between the fit center for constraint 2 and ground truth is 3.2cm. V. VALIDATION The approach is validated on a UR16e robot, which has an integrated F/T sensor at the flange and a Realsense D435 camera. The CasADi framework is used for parameter fitting and calculating Jacobians [27]. The admittance control is implemented via ROS [15]. The code and experiment data is available at https://github.com/khaninger/contact monitoring. A. Cable pulling The data collected for the radius constraint parameter estimation consists of 3D positions of a plug that is constrained with a cable that is fixed via two different pivots to a workbench. The demonstration data is shown in Fig. 4, where the observed plug positions are segmented into: free space (yellow), constraint by cable fixed at constraint 1 (cyan), and constraint 2 (magenta). The segmentation process has some noise in detecting the transitions but provides sufficient accuracy for the constraint fitting. The location of the fit center 2 and its ground truth are shown in blue respectively red. Additionally, the sphere for the estimated radius around center 2 is shown. The fit constraint results in an error of 3.2cm with respect to the measured pivot. Using the fit constraints, the controller is applied as seen in the attached video. The parameters used for constraint detection are F= 6, detection window L= 8, and constraint forces are applied of α= 12,α+= 6. The stiffness parameters of the impedance controller are K= [250,250,550,20,20,20]. The motion strategy πmt(xt)finds the closest point in the demonstration trajectory, generates xd t by transforming from the object frame to TCP, and advances along the trajectory when the distance to the current xd tis less than 1cm. Results of the online detection can be seen in the attached video and Fig. 5, where the residual per constraint and forces are shown. Near the starting pose, the residual between pivot 0 and pivot 1 can be clearly distinguished - at that point these constraints are in different directions. However, as the robot moves near when the two constraints are in a similar direction (around 20 sec), the residuals become more similar, and a few centimeters of error in the mode detection are noticed. This is attributed to the constraints being in a similar direction - (a) Disturbed cabling process Fig. 5: Forces and contact residual in the cabling task Fig. 6: Poses from rake data with fit parameters for the plane contact, with corresponding contact points p0, p1 the vectors J0and J1for pivots 0 and 1 are almost identical, and distinguishing the forces is difficult. Similar difficulty in distinguishing certain constraints has been noted by other authors [2], [28]. To validate the robustness of detecting changes in contact, the cable is unhooked by hand from pivot 1 multiple times. Additionally, at 60 seconds the robot is manually perturbed 20 cm from the path, and at 80 seconds the cable slips on its own. The robot is able to recover from all these perturbations, as seen in the attached video. The force residuals can be seen in Fig. 5(b), where it can be seen that this online change is robustly detected, and the robot can recover with the modedependent control. B. Zen Garden Raking This task involves a rake over a planar surface, with the additional constraint of reaching the edge of the raking area. These two constraints are extracted from demonstrations, where the pose of the rake is tracked as the rake moves in free space, line contact with the workbench, and then a hinge constraint with the border of the raking area is met. The rake pose is tracked with four keypoints, and the residual error of the correspondence to the reference keypoints is less than 3 mm, as seen in Fig. 6, which also shows the points of the identified contact point on the line constraint (black points) and the fit plane location (purple). The results of the live control experiments on the rake can be seen in Fig. 7 and the attached video. The parameters 3504 Authorized licensed use limited to: University of Patras. Downloaded on November 04,2025 at 20:17:43 UTC from IEEE Xplore. Restrictions apply.
Fig. 7: Measured forces in TCP, constraint residual, and desired force Fd t in the rake task used are F= 1.4,w= 6,α= 4 and α+= 6. The contact with the second constraint (hinge) results in an impulse in the forces, but again this impulse does not appear in the residual for the hinge constraint, indicating the impulse is occurring primarily in the subspace expected by the constraint. VI. CONCLUSION The paper has proposed a pipeline for learning constraints from the passive observation of markerless human demonstrations. A general constraint fitting process is introduced, and several example constraints are validated from real data. The accuracy of keypoint-based pose estimation, GMMbased clustering, and constraint model identification is shown to have sufficient accuracy for some contact-rich tasks. These constraint models are shown to result in effective online detection of constraint from F/T information, provided the Jacobian between them is sufficiently distinct. The usefulness of this information was demonstrated by a simple modedependent trajectory followed by a desired force that maintains the current constraint. REFERENCES [1] M. Suomalainen, Y. Karayiannidis, and V. Kyrki, “A survey of robot manipulation in contact,” Robotics and Autonomous Systems, vol. 156, p. 104224, 2022. [2] G. Subramani, M. Gleicher, and M. Zinn, “Recognizing geometric constraints in human demonstrations using force and position signals,” IEEE Robotics and Automation Letters, vol. 3, no. 2, pp. 1252–1259, 2018. [3] A. Aydinoglu, A. Wei, and M. Posa, “Consensus Complementarity Control for Multi-Contact MPC,” Apr. 2023. [4] F. R. Hogan and A. Rodriguez, “Reactive planar non-prehensile manipulation with hybrid model predictive control,” The International Journal of Robotics Research, vol. 39, no. 7, pp. 755–773, Jun. 2020. [5] E. Huang, X. Cheng, and M. T. Mason, “Efficient Contact Mode Enumeration in 3D,” in Algorithmic Foundations of Robotics XIV, S. M. LaValle, M. Lin, T. Ojala, D. Shell, and J. Yu, Eds. Cham: Springer International Publishing, 2021, vol. 17, pp. 485–501. [6] H. Kato, D. Hirano, and J. Ota, “Contact-Event-Triggered Mode Estimation for Dynamic Rigid Body Impedance-Controlled Capture,” in 2019 International Conference on Robotics and Automation (ICRA), pp. 3600–3606. [7] G. Bledt, P. M. Wensing, S. Ingersoll, and S. Kim, “Contact Model Fusion for Event-Based Locomotion in Unstructured Terrains,” in 2018 IEEE International Conference on Robotics and Automation (ICRA), pp. 4399–4406. [8] K. Haninger and D. Surdilovic, “Multimodal Environment Dynamics for Interactive Robots: Towards Fault Detection and Task Representation,” in Proc. IEEE/RSJ Intl Conf on Intelligent Robots and Systems (IROS), pp. 6932–6937. [9] S. Cheng, C. Garrett, A. Mandlekar, and D. Xu, “Nod-tamp: Multi-step manipulation planning with neural object descriptors,” arXiv preprint arXiv:2311.01530, 2023. [10] T. Tang, H.-C. Lin, and M. Tomizuka, “A learning-based framework for robot peg-hole-insertion,” in Dynamic Systems and Control Conference, vol. 57250. American Society of Mechanical Engineers, 2015, p. V002T27A002. [11] C. Chang, K. Haninger, Y. Shi, C. Yuan, Z. Chen, and J. Zhang, “Impedance Adaptation by Reinforcement Learning with Contact Dynamic Movement Primitives,” in 2022 IEEE Intl. Conf. on Advanced Intelligent Mechatronics (AIM), 2022. [Online]. Available: arXiv:2203.07191 [12] A. T. Le, M. Guo, N. van Duijkeren, L. Rozo, R. Krug, A. G. Kupcsik, and M. B¨ urger, “Learning forceful manipulation skills from multi-modal human demonstrations,” in 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2021, pp. 7770–7777. [13] H. Ravichandar, A. S. Polydoros, S. Chernova, and A. Billard, “Recent advances in robot learning from demonstration,” Annual Review of Control, Robotics, and Autonomous Systems, vol. 3, pp. 297–330, 2020. [14] M. Vecerik, O. Sushkov, D. Barker, T. Roth¨ orl, T. Hester, and J. Scholz, “A Practical Approach to Insertion with Variable Socket Position Using Deep Reinforcement Learning.” [Online]. Available: http://arxiv.org/abs/1810.01531 [15] S. Scherzinger, A. Roennau, and R. Dillmann, “Contact Skill Imitation Learning for Robot-Independent Assembly Programming,” in 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Nov. 2019, pp. 4309–4316. [16] F. Stulp, J. Buchli, A. Ellmer, M. Mistry, E. A. Theodorou, and S. Schaal, “Model-free reinforcement learning of impedance control in stochastic environments,” IEEE Transactions on Autonomous Mental Development, vol. 4, no. 4, pp. 330–341, 2012. [17] K. Chatzilygeroudis, V. Vassiliades, F. Stulp, S. Calinon, and J.-B. Mouret, “A survey on policy search algorithms for learning robot controllers in a handful of trials.” [Online]. Available: http://arxiv.org/abs/1807.02303 [18] T. Z. Zhao, V. Kumar, S. Levine, and C. Finn, “Learning fine-grained bimanual manipulation with low-cost hardware,” arXiv preprint arXiv:2304.13705, 2023. [19] C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” arXiv preprint arXiv:2303.04137, 2023. [20] G. Franzese, A. M´ esz´ aros, L. Peternel, and J. Kober, “ILoSA: Interactive learning of stiffness and attractors,” in 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2021, pp. 7778–7785. [21] F. J. Abu-Dakka, L. Rozo, and D. G. Caldwell, “Force-based variable impedance learning for robotic manipulation,” Robotics and Autonomous Systems, vol. 109, pp. 156–167, 2018-11. [Online]. Available: https://linkinghub.elsevier.com/retrieve/pii/S0921889018300125 [22] V. Acary and B. Brogliato, Numerical Methods for Nonsmooth Dynamical Systems: Applications in Mechanics and Electronics. Springer Science & Business Media, 2008. [23] K. Blomqvist, J. J. Chung, L. Ott, and R. Siegwart, “Semi-automatic 3d object keypoint annotation and detection for the masses,” in 2022 26th International Conference on Pattern Recognition (ICPR). IEEE, 2022, pp. 3908–3914. [24] K. Daniilidis, “Hand-eye calibration using dual quaternions,” The International Journal of Robotics Research, vol. 18, no. 3, pp. 286– 298, 1999. [25] S. Calinon, F. Guenter, and A. Billard, “On learning, representing, and generalizing a task in a humanoid robot,” IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), vol. 37, no. 2, pp. 286–298, 2007. [26] H. BO, “Six-Degree-Of-Freedom Active Real-Time Force Control of Manipulator,” 2001. [27] J. A. Andersson, J. Gillis, G. Horn, J. B. Rawlings, and M. Diehl, “CasADi: A software framework for nonlinear optimization and optimal control,” Mathematical Programming Computation, vol. 11, no. 1, pp. 1–36, 2019. [28] T. J. Debus, P. E. Dupont, and R. D. Howe, “Distinguishability and identifiability testing of contact state models,” Advanced Robotics, vol. 19, no. 5, pp. 545–566, 2005. 3505 Authorized licensed use limited to: University of Patras. Downloaded on November 04,2025 at 20:17:43 UTC from IEEE Xplore. Restrictions apply.