scieee AI-readable full text Open interactive document viewer

Robotic Cloth Manipulation: Real Implementation using Model Predictive Control and Reinforcement Learning

Luque Acera, Adrià

Full text

Double Master’s Thesis Master’s Degree in Industrial Engineering (MUEI) Master’s Degree in Automatics and Robotics (MUAR) Robotic Cloth Manipulation: Real Implementation using Model Predictive Control and Reinforcement Learning MEMORY Author: Adrià Luque Acera Directors: Adrià Colomé Figueras Carlos Ocampo Martínez Date: September 2021 Escola Tècnica Superior d’Enginyeria Industrial de Barcelona Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 3 Abstract In recent years, robots have expanded past factories and industrial environments, being used for urban and assistive applications. Given many of our daily life tasks involve handling non-rigid textile objects, robotic cloth manipulation is a relevant field of study nowadays. For example, a robot can assist a person with reduced mobility in tasks like dressing up, folding clothes or extending a tablecloth. However, highly deformable objects present complex behaviors, adopting multiple configurations during manipulation, making robotic implementations more challenging. The present Double Master’s Thesis explains the process of designing a Model Predictive Controller for cloth manipulation tasks, improving it with Reinforcement Learning techniques to find the optimal model parameters and controller tuning, and finally implementing the full closed-loop control scheme in a real setup, to test its performance. The developed controller includes an already existing linear model of the cloth, to consider not only the present configuration, but also predict the future behavior of the cloth depending on the chosen control inputs and states. Concretely, it is assumed that some key points of the manipulated cloth piece (corners) must follow a defined trajectory in space without controlling them directly with a robot. This controller was tested in simulation using a more complex nonlinear cloth model, resulting in successful reference tracking almost in real-time. Both the linear cloth model and the resulting controller presented multiple parameters to be tuned for each specific case. To obtain their optimal values, Reinforcement Learning techniques were applied with successful results, enabling applications in more varied conditions with better tracking performance. The developed control scheme was then implemented in a real setup, using a WAM robot, an existing Cartesian controller, and adapting previously developed Computer Vision algorithms for cloth identification and segmentation to obtain real feedback data. All these systems were connected using ROS, and adapted to work together in real-time. Several experiments were conducted to add a filtering process and find the optimal working conditions empirically, and then the implementation was tested in more adverse conditions (rotations, fast movements and disturbances), all with positive results and low tracking errors, validating the developed implementation as a suitable one for robotic cloth manipulation tasks with trajectory tracking. 4 ABSTRACT Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 5 Contents Abstract 3 List of Figures 8 List of Tables 9 Notation 11 1 Introduction 15 1.1 Objectives.......................................... 15 1.2 Scope ............................................ 16 1.3 ThesisOutline........................................ 17 2 Model Predictive Control for Cloth Manipulation 19 2.1 TheoreticalBackground................................... 19 2.2 ProblemDefinition ..................................... 25 2.3 Starting Point: Models and Baseline Controller . . . . . . . . . . . . . . . . . . . . . . 27 2.4 ControllerRedesign..................................... 32 2.5 Generalizations and Improvements . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38 2.5.1 Accepting Models of any Size . . . . . . . . . . . . . . . . . . . . . . . . . . . 38 2.5.2 Using a Local Cloth Base . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41 2.5.3 Outputting a Single TCP Pose . . . . . . . . . . . . . . . . . . . . . . . . . . . 46 2.5.4 UsingotherModels................................. 48 2.5.5 Minimizing the Slew Rate . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53 2.5.6 Simulating Closer to Reality . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57 2.5.7 Flexibility Options and Summary . . . . . . . . . . . . . . . . . . . . . . . . . 63 3 Reinforcement Learning Enhancements 65 3.1 TheoreticalBackground................................... 65 3.2 Learning the Parameters of the Linear Model . . . . . . . . . . . . . . . . . . . . . . . 70 3.2.1 Gathering and Processing Real Data . . . . . . . . . . . . . . . . . . . . . . . . 71 3.2.2 Learning with Real Cloth Data . . . . . . . . . . . . . . . . . . . . . . . . . . . 74 3.2.3 AnalysisofResults................................. 76 3.3 Learning the Parameters of the Controller . . . . . . . . . . . . . . . . . . . . . . . . . 82 3.3.1 Initial Considerations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83 3.3.2 Obtaining the Most Suitable Structure . . . . . . . . . . . . . . . . . . . . . . . 86 3.3.3 Tuning the Resulting Controller . . . . . . . . . . . . . . . . . . . . . . . . . . 89 6 CONTENTS 4 Real Implementation 95 4.1 Overall Structure and Considerations . . . . . . . . . . . . . . . . . . . . . . . . . . . . 95 4.2 Adapting the Model Predictive Controller . . . . . . . . . . . . . . . . . . . . . . . . . 98 4.2.1 Code Translation to C++ . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 98 4.2.2 Integration in the ROS Structure . . . . . . . . . . . . . . . . . . . . . . . . . . 100 4.3 The Cartesian Controller Nodes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102 4.4 The Vision Feedback Nodes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 104 4.4.1 Obtaining the Cloth Mesh . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105 4.4.2 Robot-Camera Calibration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106 4.4.3 ClosingtheLoop..................................109 4.5 Implementation in Real Time . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 113 5 Experimental Results 117 5.1 Effects of Working in Real Time . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 117 5.2 Vision Feedback Filter Selection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 120 5.3 Analysis of Control Parameters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 123 5.4 Tracking in Adverse Conditions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127 6 Budget and Impact 135 6.1 ProjectBudget........................................135 6.2 EnvironmentalImpact....................................137 Conclusions 139 Acknowledgements 143 Bibliography 145 Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 7 List of Figures 2.1 Basic MPC block diagram in a simulation scenario (full-state feedback) . . . . . . . . . 24 2.2 MPC block diagram. General case with disturbances and state estimation (output feedback) 24 2.3 Two UR10 arms manipulating a cloth piece, picking up two corners . . . . . . . . . . . 25 2.4 Visualization of the tracking problem . . . . . . . . . . . . . . . . . . . . . . . . . . . 26 2.5 Linear mass-spring-damper model, with 𝑁=4×4=16 nodes.............. 27 2.6 Closed-loop block diagram of MPC applied to cloth manipulation in simulation . . . . . 31 2.7 Results obtained executing the baseline MPC implementation . . . . . . . . . . . . . . . 31 2.8 Results obtained executing the Redesigned MPC implementation . . . . . . . . . . . . . 38 2.9 Mesh node numbering and connections (4×5example).................. 39 2.10 Corner trajectories rotating the cloth trajectory and/or parameters . . . . . . . . . . . . . 41 2.11 Evolutions obtained rotating the trajectory and swapping the parameters accordingly . . 42 2.12 Top view of a cloth piece and representation of local parameters and projections . . . . . 42 2.13 Process to obtain the local cloth base . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43 2.14 Worst cases for the two considered methods of computing the cloth base . . . . . . . . . 44 2.15 Results for a new trajectory with rotations, using a local cloth base . . . . . . . . . . . . 45 2.16 New setup with a rigid piece between the TCP and upper corners of the cloth . . . . . . 46 2.17 Comparison between the new nonlinear model and the linear one (1 s movement) . . . . 49 2.18 Comparison between the new nonlinear model and the linear one (0.2 s movement) . . . 49 2.19 Comparison between three scenarios with the same maximum speed . . . . . . . . . . . 50 2.20 Large 13 ×13 mesh with nodes taken on reduced 7×7and 4×4meshes highlighted . . 51 2.21 Unstable evolutions when changing model size or sampling time but not the parameters . 52 2.22 Analysis of the effects of 𝐻𝑝minimizing 𝑢with 𝑄𝑘=0.05, 𝑅 =1............ 55 2.23 Analysis of the effects of 𝐻𝑝minimizing Δ𝑢with 𝑄𝑘=0.05, 𝑅 =1........... 56 2.24 Analysis of the effects of 𝐻𝑝minimizing Δ𝑢with 𝑄𝑘=0.2, 𝑅 =1............ 56 2.25 Resulting distributions for the results of the DMPC and SMPC simulations . . . . . . . . 59 2.26 Results for a simulation using SMPC . . . . . . . . . . . . . . . . . . . . . . . . . . . . 60 2.27 Closed-loop block diagram of MPC applied to cloth manipulation in simulation . . . . . 61 2.28 3D plot with the results of a simulation, including the WAM Robot . . . . . . . . . . . . 62 3.1 Basic diagram of Reinforcement Learning . . . . . . . . . . . . . . . . . . . . . . . . . 65 3.2 3D plots of all the used input TCP trajectories . . . . . . . . . . . . . . . . . . . . . . . 71 3.3 Raw data gathered through Computer Vision . . . . . . . . . . . . . . . . . . . . . . . . 72 3.4 Comparison between raw and filtered data for the lower left corner . . . . . . . . . . . . 73 3.5 Zoomed in comparison between original data points and regularized ones . . . . . . . . 73 3.6 Evolution of means and deviations of all parameters in one learning experiment . . . . . 76 3.7 Evolution of the rewards for the same learning experiment . . . . . . . . . . . . . . . . 77 8 LIST OF FIGURES 3.8 Comparison between learnt parameters and their rewards depending on input trajectory . 78 3.9 Rewards for all the considered experiments, grouped by 𝑛and 𝑇𝑠............. 79 3.10 Example of computing the final model parameters for 𝑛=4, 𝑇𝑠=10 ms......... 80 3.11 Simulation results with a longer trajectory in more demanding conditions . . . . . . . . 82 3.12 Different possibilities to represent two weights with the same proportion . . . . . . . . . 84 3.13 Results of the first round of learning experiments (RMSE) . . . . . . . . . . . . . . . . 87 3.14 Results of the first round of learning experiments (time) . . . . . . . . . . . . . . . . . . 87 3.15 Results of the learning experiments using the triple model scheme . . . . . . . . . . . . 89 3.16 Results obtained by simulating across all possible 𝑄, 𝑅 values .............. 90 3.17 Results obtained by executing 2 epochs of REPS with 100 samples each . . . . . . . . . 91 3.18 Results for all the analyzed trajectories . . . . . . . . . . . . . . . . . . . . . . . . . . . 91 3.19 Situation where a strict upper bound would leave the optimum out (a), and solution (b) . 92 3.20 Comparison of results using the triple model scheme with disturbances . . . . . . . . . . 94 4.1 Full diagram of the final implementation in ROS . . . . . . . . . . . . . . . . . . . . . . 96 4.2 Picture of the WAM used in the real setup, and piece that connects to the cloth . . . . . . 97 4.3 Picture of the full real setup during an experiment . . . . . . . . . . . . . . . . . . . . . 97 4.4 Results obtained with the translated closed-loop simulation in C++ . . . . . . . . . . . . 99 4.5 Diagram of an open-loop implementation in ROS using the Read Node . . . . . . . . . . 104 4.6 A Kinect camera, used to capture RGB-D images and obtain current mesh positions . . . 105 4.7 Diagram of the setup with World, End-Effector and Camera coordinate frames shown . . 106 4.8 Comparison between the different considered filters in a scenario without noise . . . . . 110 4.9 Diagram of the ROS implementation in real time . . . . . . . . . . . . . . . . . . . . . 113 5.1 Experimental results without Vision feedback nor running in real time . . . . . . . . . . 117 5.2 Experimental results with Vision feedback but not running in real time . . . . . . . . . . 118 5.3 Experimental results with Vision feedback and running in real time . . . . . . . . . . . . 119 5.4 Evolution of the 𝑋-axis position of the lower right corner for all considered filters . . . . 120 5.5 Detail of the found situation where the corners of the Vision mesh are not placed accurately121 5.6 Results obtained on the filter selection experiments . . . . . . . . . . . . . . . . . . . . 122 5.7 Obtained RMSE using different linear model sizes (5 experiments each) . . . . . . . . . 123 5.8 Results of all the executed experiments depending on 𝑇𝑠, 𝐻𝑝, 𝑊𝑉............124 5.9 Cloth corners and TCP evolutions for 𝑇𝑠=20 ms, 𝐻𝑝=25,𝑊𝑉=0.2..........126 5.10 Cloth corners and TCP evolutions tracking a trajectory with rotations . . . . . . . . . . . 127 5.11 Comparison between evolutions of the same cloth corner executing at different speeds . . 128 5.12 Cloth corners and TCP evolutions blocking the camera in two instants . . . . . . . . . . 130 5.13 Human walking in the setup, creating a barrier between camera and cloth . . . . . . . . 131 5.14 Human agent pulling and pushing the elbow joint during an execution . . . . . . . . . . 132 5.15 Disturbance created by a person poking the cloth . . . . . . . . . . . . . . . . . . . . . 132 5.16 Cloth corners and TCP evolutions blocking the camera and with human interaction . . . 133 Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 9 List of Tables 2.1 Comparison of KPIs between the original and the redesigned controller . . . . . . . . . 37 2.2 Original linear model parameters, tuned using the nonlinear one . . . . . . . . . . . . . 41 2.3 Quantitative results of DMPC and SMPC simulations . . . . . . . . . . . . . . . . . . . 59 3.1 Final linear model parameters, depending on size and 𝑇𝑠................. 81 3.2 Optimal 𝑄, 𝑅 weights, using Δ𝑢and no 𝑄𝑎, for all other combinations . . . . . . . . . . 88 3.3 Learnt optimal 𝑅/𝑄ratios for different 𝑇𝑠,𝐻𝑝....................... 93 4.1 DH parameters of the Barrett WAM . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102 5.1 Conditions of the filter selection experiments . . . . . . . . . . . . . . . . . . . . . . . 122 6.1 Completeprojectbudget ..................................136 16 CHAPTER 1. INTRODUCTION These shortcomings gave reason to continue this line of work, and can be reinterpreted as the main objectives of the next phase of the research project, i.e., the work described in this Thesis document. Therefore, these objectives are: •Improving the existing baseline MPC in simulation, generalizing it to more cases and models, and preparing it for real scenarios. •Implementing a Model Predictive Controller into a real robot, translating and adapting the code as needed. This includes the design of a full closed loop, adding a Cartesian controller and Vision feedback, and their execution in real time. •Adding Learning techniques to automatically tune all the parameters present in the controller so the tasks are executed optimally, improving the results of the predictive controller. This can be separated in two categories: learning the parameters of the cloth model, and learning the parameters of the controller itself. 1.2 Scope The original project proposed by the IRI had some progress and was later adapted into “Reinforcement Learning and Visual Servoing for Model Predictive Control”, proposed on the IRI-UPC Internship Program 2021-1 under the “María de Maeztu Unit of Excellence” Seal, tied to a Research Initiation (INIREC, “INIciació a la RECerca” in Catalan) grant at the UPC. This Thesis starts at this point, as a 4-month granted internship at the IRI to continue working on the project, and before starting with the main body of the document, it is worth mentioning its limits and general scope. First of all, the starting point is not a blank slate. The work carried out at the IRI before starting this Thesis resulted in the creation of nonlinear and a linear cloth models, and also a baseline MPC in a closed-loop simulation. This conforms the foundation of the work done in this Thesis, as will be detailed in Section 2.3. The fact that the project was proposed from the IRI and had some work done beforehand meant that, from the start, there was not absolute freedom, so the type of control was already decided and set as MPC, and the cloth models to be used were the ones developed, but also meant there was no need to search for other options in the vast research publications. This Thesis covers a project with clear guidelines and builds upon previous work, without considering radically different options, which are to be considered out of scope. This Thesis does not try to create or come up with the best possible control technique for a cloth manipulation application, nor does it involve modeling cloth pieces. The main scope can be derived from the objectives and the context given in the previous paragraphs. It includes improving and contributing to the previous work, done in simulation, transforming the result of that process so it can be executed in a real robot, connecting all the necessary systems, gathering data and analyzing the results, and using Reinforcement Learning techniques to improve the results even further. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 17 Of course, to create a working control scheme with a real robot, there is a lot of assembly work putting pieces together, and some of them had been developed for other projects at the IRI. This is the case of the code responsible of obtaining a cloth mesh from the data of a camera, detailed in Section 4.4, and the used Cartesian Controller, adapted from an available code as explained in Section 4.3. Once again, the scope of the Thesis does not include developing these codes, but it does include their understanding, adaptation, improvement, usage and connections with the rest of the control scheme. Finally, it is also worth mentioning that the results of this project (this Thesis plus previous work on the topic done at the IRI, that is out of the scope of this Thesis but forms its basis), will be published on a journal paper. Even though the scope of the paper is a bit larger, this proves the contribution of this Thesis in the general Robotics research. 1.3 Thesis Outline This document is organized revolving around the three axes described by the three main objectives mentioned in Section 1.1: MPC, RL and the implementation of the complete control scheme into a real robot. Although the table of Contents shows every subsection and links to it, here a small summary of every chapter is presented to help the reader know what to expect and decide what to read. •Chapter 2: Model Predictive Control begins with a theoretical background on control systems and this type of controllers, to then describe the problem at hand, what was done before this Thesis, and the work done involving MPC, starting with a controller redesign, and then a series of minor improvements organized in subsections. •Chapter 3: Reinforcement Learning also has a brief theoretical base at the start, followed by the application of RL techniques in two different ways: first, to improve the modeling of a real cloth piece, learning the parameters of its linear model. Second, to tune the controller automatically, considering all the different options available. •Chapter 4: Real Implementation is dedicated to all the steps taken to go from simulation to a real robot, starting by translating the code. The final system uses ROS (Robot Operating System), so all the nodes and overall structure will be described in this section. Finally, the real system must also work in real time, which implies a series of changes and considerations explained at the end of this chapter. •Chapter 5: Experimental Results shows the outcome of the real implementation, and contains some analyses of the obtained results discussing, for example, types of filters for the Vision feedback or combinations of control parameters. •Chapter 6: Budget and Impact discusses the budget of this project and its effects on the environment, both positive and negative, derived from its applications and resources used. •Conclusions closes the main body of the document reflecting on the work done in relation with the objectives, and describes some future work that can be done in upcoming projects. 18 CHAPTER 1. INTRODUCTION Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 19 2. Model Predictive Control for Cloth Manipulation This chapter revolves around Model Predictive Control (MPC) and its use in this Thesis. Section 2.1 introduces the necessary theoretical concepts and properties of MPC, while Section 2.2 defines the specific control problem in hand. Section 2.3 specifies the controller developed prior to this Thesis. Section 2.4 introduces its redesign, and then Section 2.5 delves into all the modifications introduced to create a generalized and improved version, ready to be applied to a real scenario. 2.1 Theoretical Background Model Predictive Control stems from Optimal Control [3]. Its main idea is to use a dynamic model to forecast, or “predict”, the behavior of the system, formulating an optimization problem to obtain the control signals that will produce the best outcome. At any given time, one can predict until a determined moment in the future, or horizon, optimize to obtain a sequence of states and control signals, and apply the first one to the system so it evolves into a new state, from which the process can start again. Dynamic models are usually described by the differential equations on (2.1), where 𝑥∈R𝑛is the state, 𝑢∈R𝑚is the control signal or input, 𝑦∈R𝑝is the output and 𝑡∈Ris time. Equation (2.1c) is the initial condition, which specifies the value of the state at the starting instant (usually defined as 𝑡0=0). ¤𝑥(𝑡)=𝑑𝑥(𝑡) 𝑑𝑡 =𝑓(𝑥, 𝑢, 𝑡)(2.1a) 𝑦(𝑡)=ℎ(𝑥, 𝑢, 𝑡)(2.1b) 𝑥(𝑡0)=𝑥0(2.1c) The previous formulation is the most general one, valid for a generic system. A specific but common case is to have the dynamic model as a Linear and Time-Invariant (LTI) one. In that case, the formulation becomes the one shown in (2.2). Even though the time invariance means the matrices do not depend on time, the other three variables are still time functions (𝑥(𝑡), 𝑢(𝑡), 𝑦(𝑡)). For simplicity, this dependency is commonly not explicitly written. ¤𝑥=𝐴𝑥 +𝐵𝑢 𝑦=𝐶𝑥 +𝐷𝑢 𝑥(0)=𝑥0 (2.2) This formulation is also known as the State-Space (SS) Representation, and can also be used for discrete-time models, when the systems are sampled at a given rate. This rate (time step, sampling time 20 CHAPTER 2. MODEL PREDICTIVE CONTROL FOR CLOTH MANIPULATION or sampling period, 𝑇𝑠) must be chosen carefully according to the dynamics of the system, and if done properly, the resulting system equations can be seen in (2.3), where instead of a derivative, we obtain the next state sample with data of the current sample, and instead of a differential equation, we have a linear difference equation. Here the corresponding time step must be specified to avoid confusion, either between brackets or as a subscript, depending on the used notation. 𝑥(𝑘+1)=𝐴𝑥(𝑘) + 𝐵𝑢(𝑘) 𝑦(𝑘)=𝐶𝑥(𝑘) + 𝐷𝑢(𝑘) 𝑥(0)=𝑥0 (2.3) The previous expressions describe deterministic models, where all the dynamics are described with a State-Space realization, but when studying a real system, we use instruments with a certain precision and uncertainty, and sensors that can have some noise, resulting in fluctuations in the obtained data. Stochastic models try to consider these uncertainties, and their research is an active and important topic nowadays, as improving the techniques to model real complex systems helps designing and testing controllers, and is essential in those that depend directly on the model of the system, such as MPC. The simplest stochastic model is a variation of the previously presented one, but adding a random sensor noise 𝑛(𝑘) ∈ R𝑝to the output, and some unmodeled disturbance 𝑑(𝑘) ∈ R𝑔that affects the state of the system, as seen in (2.4). Matrix 𝐵𝑑∈R𝑛×𝑔is used to apply the effects of the disturbances to the states when the relations are known. Sometimes, each disturbance affects one state (thus 𝑔=𝑛), so this matrix becomes the identity. 𝑥(𝑘+1)=𝐴𝑥(𝑘) + 𝐵𝑢(𝑘) + 𝐵𝑑𝑑(𝑘) 𝑦(𝑘)=𝐶𝑥(𝑘) + 𝐷𝑢(𝑘) + 𝑛(𝑘) 𝑥(0)=𝑥0 (2.4) Now that we know how to represent the model of a system, the main objective is to obtain a way to control it so that the system evolves in a desired way, e.g., as close as possible to a given reference, with the minimum possible error, or optimizing a series of other performance indicators. This means obtaining a sequence of control signals 𝑢such that the resulting evolution of the system is as close to the desired one as possible. Depending on the approach we choose, we obtain different types of controllers. One of the most common ones is using State Feedback, where the control signal directly depends on the current state vector through a gain matrix (𝑢=𝐾𝑥), which is obtained with the given specifications. While the equations presented until now can be defined as open-loop, because the states and outputs of the system depend on themselves and arbitrary independent inputs, feeding the states back to obtain the control signals like this creates a closed-loop system, as we can think of the dependencies between variables as a complete circle or back and forth between states and controls. Although not as simple as a constant matrix, most controllers are based on feeding back output data from the system through different processes. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 21 For this project, a Model Predictive Controller was chosen as the specific option to be tested. Previous research at the IRI had already tested simpler control techniques, which were too sensitive to the initial conditions and possible disturbances of the environment [4], and a first approach to MPC had been tested with satisfactory results, so the choice of MPC itself and the comparative analysis with other control techniques had been done prior to this Thesis and thus are out of its scope. Regardless of that, before continuing, we can list some of the properties of MPC, and its advantages and disadvantages [5], [6], [7], to give a brief background on the basis of the choice. They are: Advantages: •Usage of all system dynamics described by the model. •Multiple and varied control goals can be considered. •Simple and optimal control policies even on complex systems. •Straightforward consideration of physical constraints. •Use of feed-forward features that reject disturbances. Disadvantages: •An accurate model of the real system is required. •Complex models can lead to high computational loads. •Need of numerical optimization solvers. •High number of parameters, which must be tuned. The fact of using a model already comes with its pros and cons, as it gives a lot of information that can be used to predict how the system will evolve and what is the best course of action, but the model must be accurate to reality to have a direct correspondence with the evolution of the real system, and complex dynamics come at the cost of computational load and time. The most important advantage nowadays, however, is probably being able to ask for multiple control objectives, and weigh them as needed (e.g., tracking a reference trajectory while also keeping the energy consumption as low as possible) to set priorities among them. For the drawbacks, nowadays we can use a wide array of optimization solvers that keep getting more common and less resource intensive as time goes on, and there are techniques to reduce complexity without losing important dynamics on the model, so the main point that always stands is the large number of parameters that need to be tuned. This is the principal reason for the addition of a Learning algorithm that tries to obtain the optimal combination of parameters automatically once the control structure is set and a large number of simulations can be executed. 22 CHAPTER 2. MODEL PREDICTIVE CONTROL FOR CLOTH MANIPULATION With the controller type chosen and its overall characteristics described, we can proceed into a formal formulation of a general MPC. The control signals, which are what we are looking to obtain in any controller, are found via an optimization problem. In it, we define an objective function to optimize (usually minimize) and a set of constraints. The basic objective function of MPC measures the deviation of both states and control signals from zero, and takes the form of weighted squares as seen in (2.5), where matrices 𝑄and 𝑅weigh the state and control input deviations, respectively, and 𝐻𝑝is the prediction horizon, the number of steps that we look into the future to predict the behavior of the system. 𝐽= 𝐻𝑝 Õ 𝑘=1h𝑥(𝑘)𝑇𝑄 𝑥(𝑘)i+ 𝐻𝑝−1 Õ 𝑘=0h𝑢(𝑘)𝑇𝑅 𝑢(𝑘)i(2.5) It is worthy to point out how from the two distinct parts, the state sum goes from 1 to the horizon 𝐻𝑝itself, while the control signals are shifted a step before that. Of course, this is because the states at step 𝑘=𝐻𝑝, which is the furthest in the future we want to predict, are obtained with states and control signals at step 𝑘=𝐻𝑝−1, as per (2.3), so the next control signal is not needed to reach the horizon. Additionally, 𝑥(0)is the initial condition, so there is no need to penalize it. The last predicted state, at the horizon, is sometimes separated from the rest with a unique weighting matrix to penalize the final deviation even further. The objective function in (2.5) assumes that the states need to tend to the origin and penalizes (the function will be minimized, so a higher value is worse) their deviation from zero, and the same happens with the control signals. In a general scenario, this is not what we want, as the states will probably need to reach a certain value or to follow a series of values. This is why the most common version of the MPC objective function includes the error, or difference between the states and their reference or intended value, at each time step, as shown in (2.6). 𝐽= 𝐻𝑝 Õ 𝑘=1h𝑥(𝑘) − 𝑟(𝑘)𝑇𝑄𝑥(𝑘) − 𝑟(𝑘)i+ 𝐻𝑝−1 Õ 𝑘=0h𝑢(𝑘)𝑇𝑅 𝑢(𝑘)i(2.6) There are, more alternatives, including ones for the control signals. For example, instead of trying to reduce their absolute value to decrease the energy consumption, one might want to reduce the variation between consecutive steps, to reduce how nervous the input is and decrease potential damage in real mechanisms. This can be done by using the slew rate (Δ𝑢(𝑘)=𝑢(𝑘) − 𝑢(𝑘−1)) in the objective function instead of the control signal (𝑢) itself. This can be especially useful when disturbances and noise increase nervousness of the resulting control signals. Besides the objective function, the set of constraints can also be varied. However, it should always include the dynamic equations of the model, as each state and control signal on every time step is a different optimization variable, but they are related in a known way from one step to the next along the horizon. The initial condition must be present too. Other constraints usually include physical limitations Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 23 of variables (upper and lower bounds, or rate limits), as in the real world actuators, sensors and systems cannot take any real value or change instantaneously, as well as other hard or soft limits that must be considered [8]. In the end, the general MPC-related optimization problem results in what can be seen on (2.7), where the model equations are expressed as general discrete functions which may be nonlinear, and the new Xand Uare the subsets of their respective spaces that satisfy all bounds affecting states and control signals, respectively. All constraints must hold for the entirety of the prediction horizon. min 𝑢(𝑘)𝐽(𝑥, 𝑢) s.t. 𝑥(𝑘+1)=𝑓(𝑥, 𝑢, 𝑘) 𝑦(𝑘)=ℎ(𝑥, 𝑢, 𝑘) 𝑥(0)=𝑥0 𝑥(𝑘) ∈ X⊆R𝑛 𝑢(𝑘) ∈ U⊆R𝑚                         ∀𝑘∈ [0, 𝐻𝑝−1] (2.7) To implement a Model Predictive Controller, the optimization problem in (2.7) must be solved at every time step, updating the initial condition to the current state, predicting an horizon, and then applying only the first control input of the obtained optimal sequence. This means that the index 𝑘is a local variable, which always starts at 0 (meaning “current time”) and ends at the horizon 𝐻𝑝, which globally would be the moment equal to current time plus this many steps. In a real application, the controller has a model of the system inside for the optimization problem as described, and its output, the control signal, is fed into the real system, from where we measure some outputs to apply as feedback and close the loop. However, before implementing a controller in the real world, it is common to test it in simulation. Simulating a real system requires an additional model, called Simulation-Oriented Model (SOM), separate from the one used inside the controller, which is the Control-Oriented Model (COM). In the most basic case, the used SOM can be the same model as the COM, so they will evolve equally and the prediction will be perfect, which is not very realistic. An easy way to fix this is adding random disturbances and noise, which will be present in a real case and produce some variation between models. Other ways involve having different models for each case. It is not unusual to simplify the COM as much as possible, without losing important dynamics, with the objective of making the optimization problem easier and faster to solve, so the SOM can be a more complex model of the system, e.g., making the COM a linear model (Linear MPC is much faster and simpler) and the SOM a nonlinear one. Another issue we can encounter when applying an MPC into a real system is how to obtain the outputs of the system to feed back into the controller. In other words, at every time step the optimization problem needs an initial state to predict from, but in real systems we usually do not have access to all the states, or measuring them could be expensive, slow, or hard to the point of altering the behavior of the system to obtain their value. In these cases, we need to analyze which are the outputs of the system that can be 24 CHAPTER 2. MODEL PREDICTIVE CONTROL FOR CLOTH MANIPULATION measured with reasonable ease, and try to estimate the necessary states from them. There exist several options to choose from [9], [10]. To list some of the widely used ones, we have Luenberger Observers, Kalman Filters and their Extended variants (KF, EKF), and Moving Horizon Estimators (MHE). The last ones are based on the same principles as MPC [11], [12], applying an optimization problem on a sliding window with a model of the system, but instead of looking into the future to obtain the optimal control signal to apply, we look into the past known data to estimate the current states. All in all, the minimal Model Predictive Control scheme in a simulation scenario can be represented as the block diagram seen in Fig. 2.1, where the output of the SOM is the full state vector that can get fed back directly into the MPC block, together with a reference. Inside the MPC, we find the Optimizer and the COM used to obtain initial states and constraints. We can also represent a more realistic scenario, with all the details described in this section, as done in Fig. 2.2, where we have a “Plant” block, which could be the real system or a SOM, which is affected by some disturbance 𝑑and outputs a vector 𝑦that is not the state vector, so it must go through an estimator before closing the loop with the estimated states ˆ𝑥. The output also receives some sensor noise 𝑛in the measurement process. Figure 2.1: Basic MPC block diagram in a simulation scenario (full-state feedback) Figure 2.2: MPC block diagram. General case with disturbances and state estimation (output feedback) Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 25 2.2 Problem Definition With the concepts explained in the theoretical introduction to MPC, we can now proceed to explain the specific problem at hand. The overarching objective is to manipulate cloth pieces as accurately as possible with robots, considering they are not rigid objects and they are picked from specific spots, and not taken fully inside a gripper, folded, wrapped or otherwise, as seen in Fig. 2.3. The resulting control objective is simply described as tracking: following a trajectory with the lowest possible error, but the nature of cloth pieces and the environments where they must be manipulated (usually with humans in the very near proximity) increase the difficulty of this task significantly. Figure 2.3: Two UR10 arms manipulating a cloth piece, picking up two corners We could consider both robot and cloth into a big, complex model to use in our predictive controller, where we would have to describe all the dynamics of the robot, consider what needs to be simplified to reduce computation time, decide between a dynamic or a purely kinematic model, and end up with a high order model with joint torques as inputs and cloth positions as outputs, paralleling the real system. The connection would then be direct: the outputs of the controller would be these torques sent to the robot, the feedback signal (sensed, for example, with a camera) would be the cloth positions, and the reference is a trajectory to be followed. This apparent simplicity in the connections, however, does not compensate for the clear complexity inside the controller and its model, having to work with two very distinct subsystems, robot and cloth, with equations describing all the important dynamics well enough to have useful predictions, while being simple enough to compute them fast. This is one of the main reasons not to follow this approach, and, conversely, separate the robot and the cloth as two systems, focusing only on the cloth piece as the one to be modeled and controlled using MPC. Another reason is that considering the robot as its own system also comes with clear advantages. Now the problem is simply to track a reference trajectory of the End Effector (EE) or Tool Center Point (TCP) of the manipulator, given in Cartesian space (position and orientation in the regular 3-dimensional space). This task is done by Cartesian controllers, a well-known and researched area of robotics that extends beyond cloth manipulation, as they can be used in all kinds of applications. As this controller will not be designed using MPC, it will only be discussed when used for the real implementation of the scheme, in Section 4.3. 32 CHAPTER 2. MODEL PREDICTIVE CONTROL FOR CLOTH MANIPULATION 2.4 Controller Redesign The provided baseline implementation proves to be valuable in many regards: it uses both the nonlinear and the linear cloth models, so the code for those is already present and available, is a first breakthrough implementing MPC with these models obtaining low errors and times, and the optimization uses CasADi in its Matlab version. This open-source tool is great to code rapid and efficient optimizations, and specifically to code MPC implementations. To top it off, it has Python and C++ versions, with quite direct equivalencies resulting in easy to translate codes between programming languages. This is a key feature needed for this Thesis, knowing the code made for simulation must be translated to implement it in a real scenario. It is worth to point out that CasADi is not a solver by itself, but can parse and connect to multiple ones, some of them built-in within the base CasADi package in all of its versions. With that being said, the specific implementation also had severe shortcomings, the most notable one being the simulation time. Of course, in simulation, each iteration includes the optimization problem inside the controller and also the computations related to the SOM, which would not be necessary in a real case with the real system instead of a SOM. The results are impressive comparing them to the cited literature ([17] [18], with the cloth model actually included in the MPC, as mentioned in Section 2.3), reaching average iteration times in the order of tens of milliseconds, but with a time step of 10 ms, that is not enough. Even when only counting the optimization time, a trajectory of 7 s takes more than 30 s (adding all steps together) to be computed, averaging at a bit under 50 ms per iteration, which is almost 5 times the limit fixed by 𝑇𝑠. There are two approaches to improve this: change the optimization problem inside the controller to make it faster, or change the sampling time to have more room to compute. We will tackle both, but knowing the second one implies changes also in the parameters of the model (they all depend on 𝑇𝑠), this section will be dedicated to redesigning the controller as a first change. Before continuing, another point to be addressed is the computation of the error. The provided code computed the MAE (Mean Absolute Error) for each coordinate to obtain 6 error values (2 lower corners times 3 coordinates), and then averaged them to obtain a final mean error. While this is completely valid, in this Thesis we will use the RMSE (Root Mean Squared Error), as it penalizes large errors more. The previously reported error of 3.15 mm was an averaged MAE, and the average RMSE of all coordinates goes a bit up, to 3.51 mm. Furthermore, as they are errors in the three spatial coordinates, we can compute the 2-norm to get the total position error on each time step, and then compute the RMSE along time, averaging the 2 corners at the end to have a singular value. With this method, the error of the previous simulation goes up to 6.18 mm, which is a considerable difference. Whenever an error is reported, we will always note which method has been used to obtain it, either “Average” or “Norm”. This section will now explain the new controller design, with all the changes introduced to the optimization problem (objective function and constraints), comparing the results with the baseline case. No other changes to the code will be made for a fair comparison, resulting in the same structure and functionality. For more specific changes and improvements made progressively to bring the simulation to a point that can be implemented in a real scenario, refer to Section 2.5. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 33 Adding an artificial reference as a step between the real state and the desired reference can be helpful, as it can differ from the input reference to guarantee feasibility and finally converge to it to enforce stability (reasoning behind its inclusion in the baseline case), but it increases complexity, with new variables for each coordinate and step within the horizon, a term in the objective function, and a terminal constraint dedicated to it. The first change to the optimization problem is to remove it completely. This removes the mentioned theoretical guarantees, but in practice the references are feasible by construction and it was also seen how the system could have unstable evolutions even with the artificial reference, because the solver has a limited time to find the optimal solution and depending on the initial conditions, some constraints would not be satisfied when giving an output. The new formulation is closer to the typical one, as seen in (2.15), where 𝑓𝑐𝑡 =𝑇𝑠𝑓0. min 𝑢(𝑘)𝐽= 𝐻𝑝−1 Õ 𝑘=0h𝑦(𝑘+1) − 𝑟(𝑘+1)2 𝑄+𝑢(𝑘)2 𝑅i s.t. 𝑥(𝑘+1)=𝐴𝑥(𝑘) + 𝐵𝑢(𝑘) + 𝑓𝑐𝑡 𝑦(𝑘)=𝐶𝑥(𝑘) 𝑥(0)=𝑥0 𝑥(𝑘) ∈ X⊆R𝑛 𝑢(𝑘) ∈ U⊆R𝑚                         ∀𝑘∈ [0, 𝐻𝑝−1] (2.15) This way of writing the new objective function 𝐽is a more compact version of (2.6), both because we have used the norm symbols instead of the equivalent expression of vector transposed times the weighting matrix times the vector again, and because we have joined both terms in a single sum, shifting the indices of the error part accordingly. This has the same reasoning as before, we have no control over 𝑦(0)as the initial state is given, so there is no need to penalize its error (it will always be present), and we do not compute 𝑢(𝐻𝑝), as we only need until a step before to compute 𝑥(𝐻𝑝). The main idea of the implementation is to have a “solver” or “controller” object defined, that wraps all the equations and relations between variables in a way that makes the computations as fast as possible, and such that it behaves like a function, to be able to call this object with the updated data during simulation (reference and initial state) and get the control signal to apply as its output. This can be done thanks to the symbolic tools from, for example, CasADi, following a defined syntax and making all variables depend on input parameters and optimization variables, the ones that we actually want to know the value of. The next changes have to do with this specific implementation of the MPC into Matlab using CasADi, and therefore will be accompanied by code snippets showing the differences and details of how the theoretical optimization problem is actually materialized inside this “solver” object. These snippets are not full codes, and they cannot be executed independently, as they rely on variables defined on previous lines, and also can have lines omitted in between. Attaching the whole codes here would hide the important details that have been changed, so this is the best option. The resulting final version of the used codes can be found in the Appendices document found along this memory. 34 CHAPTER 2. MODEL PREDICTIVE CONTROL FOR CLOTH MANIPULATION The first change has to do with the definition and handling of symbolic variables, and what is used as an actual optimization variable. Theoretically, both states and control signals are optimization variables, depending on each other along the horizon with the equations defined by the model. This means we would have, for 𝑁nodes, (6𝑁+6)𝐻𝑝variables (the outputs are just some states). Knowing both 𝑁and 𝐻𝑝can be quite large, this is a big amount of variables to be handled, more if we consider that, in practice, all the states depend purely on the initial value, the matrices of the system that do not change along time, and the control signals, which are the real values we are looking for. With this reasoning, we can create a symbolic variable that represents all the state vectors along the horizon, containing explicit relations only to optimization variables and input parameters, and substitute the model constraint with it. The provided code did something similar to this, but instead defined a custom mapping function 𝑓:(𝑥𝑘, 𝑢𝑘) → 𝑥𝑘+1 with the system equations, ready to be used later, as seen in Lst. 2.1. Listing 2.1: Original definition of the state and control variables and their relation through a function 1% Declare model variables 2phi = SX.sym('phi',3*nr*nc); 3dphi = SX.sym('dphi',3*nr*nc); 4x = [phi; dphi]; 5u = SX.sym('u',6); 6 7% Mapping function f(x,u)−>(x_next) 8f = Function('f',{x,u},{A*x+B*u + COM.dt*f_ext}); 9 10 % Parameters of the optimization problem: initial state and reference 11 P = SX.sym('P',6*nr*nc,Hp+1); 12 13 Xk = P(:,1); 14 for k = 1:Hp 15 % Control variables 16 Uk = SX.sym(['U_'num2str(k)],6); 17 18 % Obtain the states for the next step 19 Xk_next = f(Xk,Uk); 20 Xk = Xk_next; 21 22 % [Create symbolic artificial reference ra. Add it and Uk to Opt.v Vector] 23 % [Update objective function with Xk_next, Uk, ra, P] 24 end But here we can also see how, to create this function, we need to define some temporary symbolic variables to represent 𝑥and 𝑢, that are only used internally by the function to create a link between them, and later define the actual variables used to represent the states and controls, Xk and Uk, respectively. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 35 This definition can prove a bit confusing and overcomplicate the translation process for the real implementation. By transforming it into a matrix, we have a much simpler code, and no need for so many intermediate and local symbolic variables. The resulting code is shown in Lst. 2.2, where we also have the symbolic control signal as a matrix for convenience. There are a couple more changes done to make things easier later, such as the definition of the initial state and reference as their own variables, x0 and Rp. This is not detrimental to the performance of the code, as inside we still have a reference to the original symbolic variable, P, which will be the input to our “solver” object. This means that every time we call the solver, Pwill have known values inside, and all the instances of this variable will get substituted automatically. Listing 2.2: New definition of the state variables with a matrix 1% Declare model variables 2x = [SX.sym('pos',3*nr*nc,Hp+1); 3SX.sym('vel',3*nr*nc,Hp+1)]; 4u = SX.sym('u',6,Hp); 5 6% Initial parameters of the optimization problem 7P = SX.sym('P', 1+6, max(6*nc*nr, Hp+1)); 8x0 = P(1, :)'; 9Rp = P(1+(1:6), 1:Hp+1); 10 11 % Fill the state matrix so it depends only on x0 and u 12 x(:,1) = x0; 13 for k = 1:Hp 14 x(:,k+1) = A*x(:,k) + B*u(:,k) + COM.dt*f_ext; 15 16 % [Update objective function with x, u, Rp] 17 end This modification is useful for humans reading the code and for later, but it does not change the optimization variables or their number by itself, if we compare codes. The change comes, once again, from removing the artificial reference, which was an added optimization variable to the problem too. With the states computed with the previously explained method instead of being proper optimization variables, this leaves us with just the control inputs as the only unknown variables to be optimized in the problem. This is the minimum achievable number, which in our case corresponds to 6𝐻𝑝, half of what we had, and not depending on the mesh size. Having all the control inputs in a single matrix, it can be converted into the vector of optimization variables anywhere outside of the shown loop. With only one type of variable in this vector, the bounds are easier to apply too, as there is no need to consider which indices correspond to which variable. This is done with the few lines shown in Lst. 2.3, which in the full code are actually after line 9 of Lst. 2.2. 36 CHAPTER 2. MODEL PREDICTIVE CONTROL FOR CLOTH MANIPULATION Listing 2.3: New definition of the vector of optimization variables and their bounds 1% Optimization variables 2w = u(:); 3lbw = −ubound*ones(6*Hp,1); 4ubw = +ubound*ones(6*Hp,1); Right after this definition, there is also the initialization of the objective function to 0, and the definition of the constraints vector and bounds as empty variables. Removing the artificial reference also has an effect here, as the terminal constraint, which ensured the final output was exactly equal to the final value of the artificial reference, now also disappears. We can substitute it by a terminal constraint making the final state the same as the real reference at the end of the horizon, but taking a practical approach, we will not add it unless we need it, as it would increase the computational time. This means that, with the model equations implemented as explained, there are no constraints other than bounds. In fact, only bounds on the control signals remain, as together with the initial state, they will define the bounds on the whole state matrix by construction. The final important change inside the solver is the definition of the objective function itself. As shown in Lst. 2.4, it is directly the implementation of function 𝐽from (2.15), and, of course, these three lines substitute the comment on line 16 of Lst. 2.2. It is worth pointing out that, given Matlab is a 1-based indexing language (all arrays and matrices start with index 1, as opposed to many other programming languages such as Python or C++, which are 0-based and indices start with 0), the index kmust go from 1 to 𝐻𝑝in the iterative loop, meaning everything is shifted one step from the theoretical expressions, which also start at 𝑘=0. Listing 2.4: New objective function definition (for each step from 1 to 𝐻𝑝) 1% Objective function 2x_err = x(COM.coord_lc,k+1) − Rp(:,k+1); 3objfun = objfun + x_err'*Q*x_err + u(:,k)'*R*u(:,k); Besides a slight code optimization to the adaptive weight computation, and moving some hard-coded values to have them as parameters, these are all the major changes done in the controller. Lst. 2.5 shows the code to create the structures necessary to finally define the solver object, which actually uses the IPOPT solver [22], can be called during simulations and will return the corresponding control signals. Listing 2.5: Creating the solver object and its settings 1opt_prob = struct('f', objfun, 'x',w,'g',g,'p', P); 2config = struct; 3config.print_time = 0; 4config.ipopt.print_level = 0; %0−3 min−max print 5config.ipopt.warm_start_init_point = 'yes'; 6solver = nlpsol('ctrl_solver','ipopt', opt_prob, config); Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 37 There are, however, some changes outside of the construction of the solver. First of all, the variable of input parameters, containing the reference and initial state, has been reorganized to have less zero elements, which can be important in the real implementation. Secondly, the initial guess has been changed not only to remove the part corresponding to the artificial reference, but also to give a better one to the control signals: instead of 0, the initial guess is the displacement between consecutive steps of the reference. This assumes the same movement for the upper (controlled) and lower (reference) corners, which is valid in slower trajectories, and usually a better start than 0. We can now compare the two codes fairly under the same conditions. Three different experiments have been executed for each controller, with the resulting values for our Key Performance Indicators (KPIs), times and errors, shown in Table 2.1. Experiment 1 corresponds to the base simulation, executing the original controller as it was, and the redesigned one with the same parameters: control signals 𝑢 (displacements) bounded to 0.8 mm per step, matrix 𝑄is left as the adaptive value, and 𝑅=10. Seeing that the new controller also follows the reference correctly, the bound for 𝑢was increased progressively until 5 cm per step, which is the value used for Experiment 2. Finally, 𝑅was tuned to find the best results in both cases for Experiment 3, resulting in 𝑅=20 for both controllers. Table 2.1: Comparison of KPIs between the original and the redesigned controller Exp. Controller Total time Avg. time per MAE (Avg.) RMSE (Norm) [s] iteration [ms] [mm] [mm] 1Original 34.3011 49.0717 3.1479 6.1751 Redesigned 23.5757 33.7277 2.9531 6.4159 2Original 20.5052 29.3350 3.4077 6.8595 Redesigned 11.6055 16.6030 3.2317 7.2620 3Original 17.6295 25.2210 3.2210 6.6777 Redesigned 10.9973 15.7328 2.3946 4.9087 First of all, we can see how the original controller took around 5 times the time step (𝑇𝑠=10 ms) in the first conditions, but it can be improved to take half that time with the tuning of Exp.3. Another important result is that the redesigned controller is consistently faster, and even if exact timing values depend on execution (background tasks running), in the optimal conditions of Exp.3 it is still around 10 ms (a whole 𝑇𝑠) faster than the original. This result is promising, but we must still remember 15 ms is a time step and a half, so we are still not within the range of real time feasibility. But with such an improvement, now increasing 𝑇𝑠and/or reducing 𝐻𝑝seem reasonable strategies. A new study of times and errors depending on 𝐻𝑝can be seen on Subsection 2.5.5, where new changes are introduced. Of course, being faster would mean nothing if the error increased significantly, losing the remarkable tracking of the original. We can see this is not the case, as the MAEs (averaged across coordinates) are even lower in the new controller, and although the RMSEs (using the “Norm” version, explained at the beginning of this section) are a bit higher on the first two experiments, under optimal conditions it is also reduced in the redesigned controller, reaching the absolute best result under all metrics. 38 CHAPTER 2. MODEL PREDICTIVE CONTROL FOR CLOTH MANIPULATION To finish this section, we show the results of Exp.3 in the redesigned controller in Fig. 2.8. This is the graphical demonstration that the tracking capabilities have been maintained, if not improved, with the changes introduced in this section. Figure 2.8: Results obtained executing the Redesigned MPC implementation 2.5 Generalizations and Improvements After redesigning the core of the controller to make it simpler and faster, there are still some other changes needed to be done in order to get a more generalized code that can be applied to a real robot performing actual movements. This section will tackle this changes one by one in different subsections, as they do not depend on each other, even if they were implemented in the shown order, and thus sometimes benefit from changes introduced before, the order could have been different leading to the same result. 2.5.1 Accepting Models of any Size The first change involves reducing the amount of hard-coded values in the code, specifically regarding the size of the linear model. The function responsible for creating the system matrices (𝐴,𝐵,𝑓𝑐𝑡 ) starts with the computation of a connectivity matrix, which is a square matrix with size 𝑁×𝑁that indicates connections between nodes with a 1, and leaves unconnected pairs with a 0. In other words, in this matrix, each row and each column represent a node, so we can check, for example, all the connections to node in row 𝑖by finding all the columns that are 1 in this row. In principle, this matrix is symmetric, as if a node 𝑖is connected to another node 𝑗, the connection goes both ways, and both elements (𝑖, 𝑗)and (𝑗, 𝑖)of the matrix must be 1. While this is true for a generic connectivity matrix, in our case, we will remove some connections with the controlled nodes, to ensure their movement is governed only by the control signals (a robot moving them) and not affected by their neighboring nodes. It is clear how this matrix depends on the total number of nodes 𝑁, as well as the proportions of the mesh, as the neighboring nodes are not the same for a 4×3and a 3×4mesh. For our application, we will always use square meshes, but it can be beneficial to future works based on this one to expose and code the general case, and leave the sizes as parameters. For this reason, we will talk about a mesh of 𝑛𝑟 rows and 𝑛𝑐columns, with 𝑁=𝑛𝑟·𝑛𝑐, instead of 𝑁=𝑛2like in the square case. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 39 Fig. 2.9 shows an example of a mesh of arbitrary size, chosen to be 𝑛𝑟=4, 𝑛𝑐=5to avoid having a square for the general case. A generic node is connected to its four neighbors, but edges and corners are special cases with less connections. With the nodes numbered left to right and bottom to top, the connections of an interior node 𝑖are to the nodes on both sides (𝑖±1), and vertically connected nodes (𝑖±𝑛𝑐). Edge nodes, marked in a yellowish orange shade, only have three of these connections, and corner nodes (red and violet) only have two. We have colored the upper corners differently because they are the controller nodes in our application, and their connection will be only one-way, as mentioned before. Figure 2.9: Mesh node numbering and connections (4×5example) With the help of this example, it is easy to deduce how to identify edge nodes, while also knowing which edge are they part of, with corners simply being part of two edges. The bottom edge is composed of nodes the indices of which are lower or equal to 𝑛𝑐. Analogously, the top edge has all the nodes with indices greater than (𝑛𝑟−1)𝑛𝑐. The right edge has all the nodes with indices divisible by 𝑛𝑐, and if we divide the indices of the nodes in the left edge by 𝑛𝑐, the remainder is always 1. This can be deduced from a more formal identification of the right edge, as those all have remainder 0. This means that, instead of hard-coding a matrix of zeroes and ones that is only valid for a fixed mesh size (in the original case, 4×4), we can build a connectivity matrix for any mesh with a few lines of code, as shown in Lst. 2.6. Listing 2.6: Creation of the connectivity matrix for a mesh of any size 1conn = zeros(nr*nc); 2for i=1:nr*nc 3if(mod(i,nc) ~= 1), conn(i, i−1) = 1; end 4if(mod(i,nc) ~= 0), conn(i, i+1) = 1; end 5if(i > nc), conn(i, i−nc) = 1; end 6if(i <= (nr−1)*nc), conn(i, i+nc) = 1; end 7end 8conn(nctrl,:) = 0; % Remove connections on controlled nodes 40 CHAPTER 2. MODEL PREDICTIVE CONTROL FOR CLOTH MANIPULATION As shown in the last line, once the actual connectivity is built, we set all the rows corresponding to controlled nodes (set N𝐶) to 0, indicating their neighbors have no effect in them. However, it is very important to remark that the controlled nodes do affect their neighbors when moving, so while, for example, element (16,11)is zero in the mesh of Fig. 2.9, element (11,16)will be 1. On the construction of the model, this connectivity matrix (𝑀𝐶) is then used to apply the spring and damper forces to neighbor nodes. In steady state, the sum of forces on all nodes must be zero, so we can create a new “unitary force” matrix (𝐹), where the diagonal element of each row is the result of adding all the columns from the connectivity matrix together, and then we subtract the connectivity matrix itself, so the sum of all the elements of any row is 0, as shown in (2.16). This matrix is then used to define the previously mentioned 𝐹𝑝and 𝐹𝑣, multiplying all the elements by the stiffness or damping in the corresponding coordinate. In the construction shown in (2.17), each block is a square matrix of size 𝑁, and the code had to be changed from the specific 𝑁=16 to a general case. 𝐹→            𝐹𝑖,𝑖 = 𝑁 Õ 𝑘=1 𝑀𝐶(𝑖, 𝑘) 𝐹𝑖, 𝑗≠𝑖=−𝑀𝐶(𝑖, 𝑗) (2.16) 𝐹𝑝= 𝑘𝑥𝐹0 0 0𝑘𝑦𝐹0 0 0 𝑘𝑧𝐹 , 𝐹𝑣= 𝑏𝑥𝐹0 0 0𝑏𝑦𝐹0 0 0 𝑏𝑧𝐹 (2.17) The same reasoning applies to the creation of system matrix 𝐴, which apart from using 𝐹𝑝and 𝐹𝑣, as shown in (2.12), it also has identity matrices of size 3𝑁, and this dependency must be explicit, not a hard-coded number. Additionally, for both 𝐴and 𝐵, the set of controlled coordinates is now introduced as a parameter of the linear model, and is not just defined as the upper corners with the indices depending on size, but is left open in case this code is used with an application where other nodes are the ones being controlled. Lst. 2.7 shows this new code, which also materializes (2.13) and the considerations explained in the paragraphs immediately after. Listing 2.7: New code to create matrices A and B 1A = [eye(3*N) Ts*eye(3*N); 2(Ts/m)*Fp eye(3*N)+(Ts/m)*Fv]; 3A(cctrl, 3*N+1:end) = 0; 4A(cctrl+3*N, :) = 0; 5 6B = zeros(2*3*N,6); 7for i=1:length(cctrl) 8B(cctrl(i), i) = 1; 9end Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 41 Finally, the computation of 𝑓𝑐𝑡 is also done depending on mesh sizes: gravity affects the rows corresponding to 𝑣𝑧for any size, the force corresponding to the initial length is set to the correct coordinates, and the rows of controlled coordinates (matching an index of C𝐶) are set to zero. This completes the changes to create a linear model of any size. The only other change from hard-coded indices to parametric ones is in the function to take a reduced mesh, which now can take the original and the final sizes as parameters, and selects the correct nodes to keep. 2.5.2 Using a Local Cloth Base Another issue of the simulation comes from the definition of the parameters of the linear model and the possible movements to be executed in consequence. As mentioned in Section 2.3, the linear model has seven parameters in total, with the first six being stiffness and damping in all three space coordinates. As one would expect by the properties of a thin textile material such as the one being considered, the stiffness in the direction perpendicular to the cloth plane (when it is extended) is much lower than in the two axes of that plane, and the same can be said about the damping, as seen in Table 2.2, where the cloth is in the 𝑥𝑧 plane, and the 𝑦component is much lower in both stiffness and damping. Table 2.2: Original linear model parameters, tuned using the nonlinear one Var. x y z 𝑘-305.6028 -13.4221 -225.8987 𝑏-4.0042 -2.5735 -3.9090 Δ𝑙0- - 0.0312 This difference has serious implications, as the cloth orientation has a direct effect on the parameters. As fast proof, we can rotate the original trajectory 90 degrees, along with the initial mesh position, while keeping the parameters in the same order (or vice versa). The results can be seen in Fig. 2.10, and it is clear how, with this implementation, rotations are not feasible, as they result in an unpredictable evolution. Figure 2.10: Corner trajectories rotating the cloth trajectory and/or parameters 48 CHAPTER 2. MODEL PREDICTIVE CONTROL FOR CLOTH MANIPULATION TCP base as 𝑅𝑡𝑐 𝑝 =[𝑌𝑐𝑙𝑜𝑡ℎ 𝑋𝑐𝑙𝑜𝑡ℎ −𝑍𝑐𝑙𝑜𝑡ℎ](this matches the axes shown in Fig. 2.16). Finally, we convert this rotation matrix into a quaternion (with an off-the-shelf function) to obtain the full TCP Pose. In summary, with just the control signals on each step, we can compute the absolute positions of the upper corners, and with those we can get both the position and orientation of the TCP. 2.5.4 Using other Models During the course of this Thesis, and parallel to it, there were some updates in the implementation of the nonlinear model, meaning there was a need to update its use in the simulations, and also compare the new version with the linear mass-spring-damper model, to check the linearization still holds and they have similar dynamics. Furthermore, while starting the real implementation and checking the rates of the different parts of the control scheme, the vision feedback proved to be the slowest part, creating a need for a “backup” fast model to fill in the gaps between real feedback data. To be the fastest possible, the model must be linear, so the simulation code must be adapted to also work with a linear SOM. This subsection is dedicated to these two changes. The new nonlinear model is quite similar to the old one, with a few parameter adjustments. The main difference is found in the functions used to simulate a step. The old codes had first and second order simulators, needing one or two previous states, respectively, and the simulations with this model as SOM used the second order step simulator. With the new update, the first order simulator is improved, and the second order one is removed. Apart from that, there was no independent function to initialize the model on its own, and in demanding or chaotic scenarios, the simulation got stuck trying to minimize an error than never went down. To enable the use of this new model, each one of this issues had to be addressed. As a first change, a timeout was introduced on the loop that simulates a step. Analyzing how long a regular simulation took, and then trying different values, a final threshold was set to 0.1 s. This updated simulation function used new variables and a different nomenclature, so a “parser” function to translate between names was created. Then, to have an independent function that initializes the model as its own object and nothing more, parts of the code were separated and adapted to reach that point. The calls to these functions can be seen in the codes shown in the Appendices document, and even if the functions themselves are not included, a link to the entire source code is provided there. Changing the nonlinear model, even if slightly, means the comparison with its linear counterpart must be done again, to check it still stands, and to what degree. We can initialize both models on the same position and enter the same control inputs, to check how they evolve. Starting with a movement of 10 cm in only one axis and completed over 1 s (a speed of 10 cm/s with 𝑇𝑠=0.01 s means a displacement of 𝑢=1mm each step), the two models behave almost identically for all axes and directions. On constant velocities, including with no movement, the linearization seems to perform the closest to the nonlinear model, and the notable differences happen when accelerating, decelerating or changing direction. This means that there is not a limit for the prediction time (𝑇𝑠·𝐻𝑝) fixed by the similarity between models, but there is a limit on how fast these movements can be, or more accurately, change. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 49 To demonstrate this, a combined step input was created, with movement in all three axes. The upper corners must move 10 cm in 𝑥, 30 cm in (negative) 𝑦, and 20 cm in 𝑧. In Fig. 2.17, this distance is set to be travelled within 1 s. We can see how the inputs change instantly to the corresponding displacement per step values, and back to 0 when a second has passed. The evolutions of the lower corners are almost identical between models (solid lines for the nonlinear, NL, and dashed for the linear, L). Changing the movement time to 0.2 s results in the evolutions shown in Fig. 2.18, where the inputs are five times larger (maximum of 15 mm per step), and the lower corners of the linear model present oscillations more pronounced than those of the nonlinear one. Even if they stabilize to the same steady state after a few seconds, the differences between both models can be significant. The results are worsened by the start and stop being so close together, as keeping the same speed for longer and then stopping, or even fast speeds with progressive accelerations, give better results, as shown in Fig. 2.19. Figure 2.17: Comparison between the new nonlinear model and the linear one (1 s movement) Figure 2.18: Comparison between the new nonlinear model and the linear one (0.2 s movement) 50 CHAPTER 2. MODEL PREDICTIVE CONTROL FOR CLOTH MANIPULATION Figure 2.19: Comparison between three scenarios with the same maximum speed These results prove how the limitations to consider are not the maximum prediction time nor the maximum displacement (pure bound of 𝑢), but how the displacement values change from one step to the next. The change introduced in Subsection 2.5.5 ties with this idea. While changing the function that initializes the nonlinear model, some extra changes were introduced. Before, the mesh was always initialized in a fixed position of space, with the cloth plane being 𝑥𝑧. It made sense, given how rotations were not allowed, but now, if we want to try new trajectories starting on other points, this limitation must be removed. The new process to initialize the model is to first load the desired reference trajectory, and then deduce the implied parameters from it, which are: cloth side length, initial cloth center and angle. The cloth side length is defined by the distance between the first points of the reference trajectories for each of the lower corners (even if their distance is not fixed like on the upper corners, the range of possible movements is now limited, and at the initial position, we will always have the cloth extended), its initial center is halfway between the lower corners and half a side length up, because all trajectories will start with a vertical cloth, and thus the only allowed rotation will be along the global 𝑍axis, with an angle that can be obtained with an arctangent. These three values are then given to the nonlinear model initializer function, which creates a mesh with size and position according to them. A function to create a mesh for the linear model was created too, and if the initial angle is different from zero, the local base is updated accordingly, defining the COM in local axes so the stiffness and damping parameters are applied in the correct directions. Setting this new nonlinear model aside, looking at a possible real implementation, there was a need for a way to keep the execution going even without constant feedback, given how it is the slowest part of the control scheme, returning data after several samples. The easiest solution is to have a backup model, a SOM even in a real scenario and not in simulation, to update the initial states between feedback updates. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 51 This backup model needs to be as fast as possible, because it will have to be updated every step, adding computational time to a controller that already takes longer than a time step. Thus, the model must be a linear one, of the minimum size possible (but not smaller than the COM), and it will be updated on every step in the real case as if it was a simulation. To prepare this case, we can alter our closed-loop simulation to work with a linear SOM, as an intermediate step that will simplify the real implementation. The definition of the controller is exactly the same, and the first half of every iteration during simulation is also unchanged: get a sliding window of the reference, rotate it and the initial state of the COM, call the solver, and obtain an increment in local coordinates, which we can convert back to the global base and add with previous data to obtain the absolute positions of the corners and the TCP pose, as seen in Subsections 2.5.2 and 2.5.3. However, now the SOM is also linear, so instead of using 𝑢or 𝑢𝑎𝑏𝑠 to update it, we must simulate a step using local coordinates too. The process involves taking the previous state in global coordinates, rotating it (positions and velocities), then applying the linear system equation (with new dedicated matrices 𝐴𝑆𝑂𝑀 , 𝐵𝑆𝑂𝑀 , 𝑓𝑆𝑂𝑀 ) to obtain the next states in local coordinates, which we then rotate back to the global base to save the evolution and update the initial COM state. It is clear how the intention is to still have slightly different models for COM and SOM, even if they are both linear. For example, if the feedback gives a mesh of side size 13 (𝑁=132), we can have a SOM of side size seven and a COM of side size four, as they can both be obtained by taking a reduced version of the largest one, as seen in Fig. 2.20. In fact, for square meshes, we can use 𝑛𝑏𝑖𝑔 =𝑘(𝑛𝑠𝑚𝑎𝑙𝑙 −1) + 1to find the bigger sizes that can be reduced into a given small size. However, changing the size of a linear model will modify its parameters, as, for example, spring forces depend on length and the nodes will be at different distances, and there is no given expression to relate size and parameters. Even worse, the method used to match the behavior of the linear model with the nonlinear one is based on an optimization, which reaches the maximum number of variables or times out just with models of side size 7, even using shorter trajectories to compare evolutions. Figure 2.20: Large 13 ×13 mesh with nodes taken on reduced 7×7and 4×4meshes highlighted 52 CHAPTER 2. MODEL PREDICTIVE CONTROL FOR CLOTH MANIPULATION And that is not all, parameters of the linear model also depend on the sampling time (𝑇𝑠). This means we cannot change this value at all without having to find the corresponding parameters, which can only be done for a model of side size 4 and even then the optimization can fail to converge and the behavior of the linear models ends up not being similar to the nonlinear one. To prove how much this limits us, two scenarios where the model parameters change have been simulated in open-loop, keeping the original values, and are shown in Fig. 2.21. The first one corresponds to changing the COM side size from 4 to 7, and input a very slight movement in 𝑧. Even before starting the movement, the cloth falls under its own weight, causing a reaction that unstabilizes the system. The second case, on the right, is changing 𝑇𝑠from 10 to 12 ms, and even just that slight change produces an oscillation in 𝑧that instead of disappearing, suddenly increases again. Figure 2.21: Unstable evolutions when changing model size or sampling time but not the parameters These results show how there is no way of using the original model parameters in a general case, and together with the limitations of the original method to tune them, justify the search for a new technique to obtain them for multiple sizes and time steps, for example, one involving learning algorithms. This is the motivation behind what will be explained in Section 3.2. Once the parameters for different conditions have been found, we will be able to apply them to our new codes. As a last note regarding the implementation using a linear SOM, if the model is exactly the same as the COM, and everything is done in simulation so there is no disturbances or noise, it is logical that both models will evolve in exactly the same way. In simulation, this can mean getting to the point of computing the whole control sequence from the start, and then applying the corresponding pre-computed control signal on every step. This is of course going back and against the basis of MPC and its receding horizon, which is implemented as explained because real systems do not behave perfectly and exactly as their models and predictions, so it does not make much sense. However, in a real implementation, we could explore this special case, and reduce the computational load not calling the optimizer unless we receive a feedback signal. Once we do and predict enough steps, the “backup” SOM will evolve exactly as predicted, so we can keep applying the computed control inputs until a new feedback signal arrives. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 53 This approach will be left as a last resort in case the optimization takes longer than a time step in the real implementation and cannot be called on every iteration. Theoretically, it should have the same behavior as increasing the sample time to values multiple of the considered one, which cannot be done in our system without changing the cloth model parameters, but it is not commonly found in literature. For the simulation code, the general case accepting a linear SOM different than the COM will be kept, even if before applying learning algorithms we can only use the exact same model for both, which of course produces almost perfect results, with close to no tracking error, and low times even adding the time taken to simulate the SOM evolution (still over 10 ms per iteration, but under 20 ms). 2.5.5 Minimizing the Slew Rate Formulating the optimization problem to be solved in the controller, the two terms of the objective function (minimizing errors, 𝑦−𝑟and control signals, 𝑢) were justified to improve tracking while not having disproportionate energy consumption, as mentioned in Section 2.4. However, in a real implementation, a robot will consume energy also while staying in place, as the motors must produce the necessary torque for gravity compensation, i.e., making the robot not fall under its own weight, so the consumption is not directly related to the absolute value of the control signals, which are displacements. This section explores a suitable alternative for the term related to 𝑢in the objective function, and analyzes the implications of this change. The control inputs 𝑢of the considered scheme are the upper corner displacements in a time step, as seen in (2.13). They are closely related to the speed of the upper corners, as with a fixed 𝑇𝑠, they tell the distance to be traveled in this amount of time. While reducing speed can have its benefits, especially having an upper bound to avoid really fast movements in an environment that will probably have humans nearby, this does not justify their inclusion in the objective function to minimize energy consumption, as a fast movement at constant speed and in a straight line can consume less than one constantly changing speed and direction. This hints at acceleration being a better variable to be minimized, or more precisely the rate of change of the displacements represented by the control signals, rather than the displacements themselves or the absolute positions of the corners. For a generic control input 𝑢, the rate of change can be written as Δ𝑢, following (2.20), and this new variable is commonly known as slew rate [25]. Δ𝑢(𝑘)=𝑢(𝑘) − 𝑢(𝑘−1)(2.20) Penalizing slew rates instead of control signals has clear advantages in the studied problem. Faster speeds are not considered worse by themselves anymore, but sudden changes in speed are. Starting or stopping a movement is now penalized, but the entire cruise time has no effect if done at constant speed, and, given the quadratic nature of the cost function, it is worse to have an instant change of velocity than to progressively alter it (e.g., Δ𝑢2>2(Δ𝑢/2)2). This also includes changes of direction, given how they are accelerations in one direction and/or decelerations in another. These cases with high speed changes are, coincidentally, exactly the ones where the linear model differs the most from the nonlinear one, as seen in Subsection 2.5.4 and Fig. 2.19. This means that minimizing the slew rate not only makes 54 CHAPTER 2. MODEL PREDICTIVE CONTROL FOR CLOTH MANIPULATION more theoretical sense given the nature of the variables at hand, but also helps improving the similarity between the linear model and the nonlinear one (and thus the real cloth too). This is of course because less changes in 𝑢result in a smoother movement, which is preferred, and if the reference trajectory does not demand it, sudden and constant changes will be avoided, which is especially useful in situations with noise, where the initial state might vary slightly on every iteration, and the resulting control signal with it if these constant changes were not penalized. The new optimization problem formulation is shown on (2.21). As explained, in the objective function we now minimize the slew rates instead of the pure control signals, maintaining the weighting matrix 𝑅 and the error part unchanged. The constraints are shown to make the changes introduced up until now explicit in one place. We can see the definition of the slew rates as a new constraint, which must hold also for 𝑘=0, where we have Δ𝑢(0)=𝑢(0) − 𝑢(−1), using the local time indices where 0 is the current time step of the overall execution. This 𝑢(−1)must be defined externally as an initial condition 𝑢0, which is the applied control signal in the previous step. Before the initial state, we find a new constraint, introduced after the changes done in Subsection 2.5.3, forcing the distance between the upper corners (here compacted to 𝑑𝑢𝑐, a 3-component vector, result of subtracting the corresponding states) to always remain constant and equal to the side length of the cloth, 𝐿. In the implementation, this equality is squared to remove a square root and keep a QCQP. The time instants are shifted like on the error computation, as we have no control over the current position at 𝑘=0, but the constraint applies to the end of the horizon. The original control signals are still needed and bounded as before (maximum speeds are still relevant), so there is no need for a model or bounds change. min 𝑢(𝑘)𝐽= 𝐻𝑝−1 Õ 𝑘=0h𝑦(𝑘+1) − 𝑟(𝑘+1)2 𝑄+Δ𝑢(𝑘)2 𝑅i s.t. 𝑥(𝑘+1)=𝐴𝑥(𝑘) + 𝐵𝑢(𝑘) + 𝑓𝑐𝑡 𝑦(𝑘)=𝐶𝑥(𝑘) Δ𝑢(𝑘)=𝑢(𝑘) − 𝑢(𝑘−1) 𝑑𝑢𝑐 (𝑘+1)2=𝐿2 𝑥(0)=𝑥0 𝑢(−1)=𝑢0 𝑥(𝑘) ∈ X⊆R𝑛 𝑢(𝑘) ∈ U⊆R𝑚                                           ∀𝑘∈ [0, 𝐻𝑝−1] (2.21) It is worth noting that in the code implementation, the slew rates are coded in a similar way to the state matrix. The optimization variables are still the absolute control signals, so a vector is created subtracting two consecutive ones (and the initial one in the first step), to be then used in the objective function, but internally, CasADi and IPOPT still have the same number of optimization variables. Of course, the previous control signal (in local coordinates) is also added to the input parameter matrix. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 55 Changing the optimization problem like this is a considerable difference, which affects the tuning of all control parameters. Consequently, it also has an effect on computation times and tracking errors, which will be different for the new optimally tuned configuration. We must remember that after the MPC redesign introduced in Section 2.4, the computational time went down considerably, but was still around the 15 ms per iteration when 𝑇𝑠=10 ms. As mentioned in that section, with the new times, altering 𝑇𝑠 or 𝐻𝑝seem like reasonable strategies, and we have seen in Subsection 2.5.4 how altering 𝑇𝑠or the mesh size changes the parameters of the model, and how finding new ones is not such an easy task without Learning techniques, so we are left with an analysis of the effects of the prediction horizon. With the baseline controller used with the linear model, a study of computational time and mean errors was carried out prior to this Thesis, resulting in the selection of 𝐻𝑝=30, used also in the simulations shown in this document. For the redesign, this analysis will be shown here, for both versions of the objective function, minimizing pure control signals and slew rates, to check the effects of the change introduced in this subsection, and new suitable values of 𝐻𝑝for both versions. The first step then is to obtain some well-tuned values for weighting matrices 𝑄and 𝑅for both cases. In fact, 𝑄is computed in an adaptive way as described previously, but we can call this matrix 𝑄𝑎and add a constant weight 𝑄𝑘. Now we can alter these two weights at will, knowing a higher 𝑄𝑘will prioritize low errors, as they will be penalized more in the objective function, and a higher 𝑅will make minimizing the control signal term a priority. Beforehand, we found that 𝑅=20 had the best results minimizing the control signals themselves (with 𝑄𝑘=1), but we can express these values normalized so the largest one is 1, as what is important is the proportion between them, not the value itself, and this way all weights will be between 0 and 1. With the same proportion, now 𝑄𝑘=0.05, 𝑅 =1. Another alternative is to make the weights add 1, leaving only one free value, which makes sense because the real degree of freedom is the proportion between two weights. However, with this method neither of the two weights is 1 (unless the other is 0), so the proportion itself is not as clear. It is also common to normalize the ranges of the variables being weighted, because sometimes the errors and the controls can be orders of magnitude apart. This is not the case in the considered optimization, though, because both errors and control signals are in the order of millimeters. With all these considerations, the resulting analysis for the new controller penalizing 𝑢can be seen in Fig. 2.22. Figure 2.22: Analysis of the effects of 𝐻𝑝minimizing 𝑢with 𝑄𝑘=0.05, 𝑅 =1 56 CHAPTER 2. MODEL PREDICTIVE CONTROL FOR CLOTH MANIPULATION In blue, corresponding to the left axis, the RMSE (“Norm” method) are shown, starting with large values for low 𝐻𝑝, which in fact represent simulations where the found optimal course of action is to stay in place. With such a short horizon, the tracking error is relatively low, and the minimal cost comes from reducing the energy consumption term, which is also weighted 20 times more. The errors decrease quickly between 𝐻𝑝=5and 15, and then they improve much more slowly. In fact, for 𝐻𝑝=17, the RMSE is just a 10% higher than with 𝐻𝑝=50. This percentage is an adequate threshold to divide low and high errors, and is marked with a blue dashed line. Looking at the right axis, in red, we find the mean computational time per iteration, counting only the optimization and base changes, but not the SOM simulation. The evolution is the opposite of the error, as expected, taking considerably more time the larger the horizon is, as the optimizer must compute a prediction more steps into the future. The threshold value here is fixed by 𝑇𝑠=10 ms, as optimizations must be completed within a time step. This value is crossed at 𝐻𝑝=22, marked with a dashed red line. Thankfully, a region to the right of the dashed blue line and to the left of the red one exists, giving us a range of values that result in fast and accurate enough simulations. This range is of course 𝐻𝑝=[17,22], and for example, choosing 20, we get a mean time around 8.9 ms with an RMSE of 11.5 mm. We can now proceed to do the same analysis for the controller minimizing slew rates. Two different tunings were analyzed, with 𝑄𝑘=0.05 (Fig. 2.23) and 𝑄𝑘=0.2(Fig. 2.24). Figure 2.23: Analysis of the effects of 𝐻𝑝minimizing Δ𝑢with 𝑄𝑘=0.05, 𝑅 =1 Figure 2.24: Analysis of the effects of 𝐻𝑝minimizing Δ𝑢with 𝑄𝑘=0.2, 𝑅 =1 Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 57 Figures 2.23 and 2.24 are left close together for a better visual comparison. The first tuning is the same one as in the case minimizing 𝑢. The overall behavior is the same, as one would expect but the errors decrease faster, reaching the 10% threshold at 𝐻𝑝=14. This improvement is countered by the fact that mean times increase with lower horizons, crossing 10 ms at just 𝐻𝑝=17, leaving a very thin range of possibilities. Simulating with the new controller minimizing Δ𝑢to find better results using other weights, an optimal tuning was found at 𝑄𝑘=0.2, 𝑅 =1. However, even with a horizon of 20 steps, this resulted in an average computational time of 13.6 ms per step and a RMSE around 3.3 mm. Performing the same analysis as before with the new weights, we can see how the errors descend much faster reaching the 10% threshold at just 𝐻𝑝=11. The times seem to increase ever so slightly, with the 10 ms mark being crossed at 𝐻𝑝=16 now. This leaves a reasonable range of possibilities for the horizon to obtain quick and accurate simulations, even if with such low values of 𝐻𝑝and fast sampling time, the total prediction time goes down from the original 0.3 s to around half of that, at 0.15 s if we choose 𝐻𝑝=15. These analyses show several important points to consider. Firstly, tuning is key and influential on errors and times. Secondly, the proposed redesign, even minimizing pure control signals, can reach mean iteration times under 𝑇𝑠=10 ms while only reducing the horizon 10 steps, barely increasing the tracking error. Using the slew rates can prove a bit more resource intensive, needing even lower horizons, but this is compensated by the tracking errors being noticeably lower too. All in all, with the applied changes, an implementation in real time now seems completely feasible, even before increasing 𝑇𝑠. Nonetheless, total prediction times have been considerably reduced from an already small starting point, so increasing the time step (after finding the new parameters) will still be useful to increase this value and take full advantage of a predictive controller. 2.5.6 Simulating Closer to Reality All the major and minor changes until now have been introduced to be able to convert the control scheme implemented in simulation into a real system. This means generalizing what initially was a very specific case or adapting to the needs of a real case. However, in all the executed simulations, the ideal case was considered, where all the data was known and used by all the connected parts of the control scheme with no issue. In a real implementation, even with a sensor like a camera that provides the full state vector, we will have noise, disturbances, and other unexpected behaviors. This section is dedicated to preparing for these situations, even if we are still in simulation. First of all, the main problems we can find when dealing with real signals between different components are disturbances and noise. As we saw in Fig. 2.2 and in (2.4) for a linear system, disturbances 𝑑(𝑘) affect the difference equation through 𝐵𝑑, and thus the evolution of the states themselves, while noise 𝑛(𝑘)is only found in the output signal. While the referenced equations are for a linear case expressed in state space, this distinction holds in general, also for nonlinear systems. To give an example related to the studied case, a disturbance could be the force of the wind or someone touching the cloth piece during the execution of the task, unforeseen interactions that change the actual state of the system in ways not predicted by the controller. Noise comes from the sensors, and could be the uncertainty or maximum precision a camera and its computer vision algorithms produce. 64 CHAPTER 2. MODEL PREDICTIVE CONTROL FOR CLOTH MANIPULATION Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 65 3. Reinforcement Learning Enhancements The present chapter is centered around the applications of Reinforcement Learning (RL) used in this Thesis to improve the newly redesigned Model Predictive Controller even further. Section 3.1 introduces the topic of RL, takes a look at the state of the art, and explains some theoretical concepts needed to understand the specific methods and algorithms used. Section 3.2 describes the first application of RL in the project, to find the optimal parameters of the linear cloth model so it behaves like the real cloth, while Section 3.3 shows how RL can be applied to tune the parameters of the controller itself. 3.1 Theoretical Background Reinforcement Learning (RL) is a wide and constantly evolving topic of research under the even larger umbrellas of Machine Learning (ML) and Artificial Intelligence (AI). As such, it revolves around the idea of creating an algorithm capable of understanding data and acting in consequence, without being explicitly programmed. Specifically, RL focuses on obtaining these situation-to-action relations, or more simply, learning how to behave, through trial and error interactions with the environment [32]. The common diagram for a RL scenario is shown in Fig. 3.1. In it, we find an agent, which is the entity that must learn, with a policy and the RL algorithm itself inside. The policy is what will be learnt, a decision (or sequence of them) that must be made depending on the current situation, or state. Observing the environment, the agent will receive this state, and act according to its policy. The action will produce a change in the environment, from where we can observe the next state and also extract a reward, measuring the performance of the action taking. All this information (action, previous and next states, reward and policy) is taken by the RL algorithm, which also saves previous data, and will keep updating the policy in order to maximize the obtained rewards. Figure 3.1: Basic diagram of Reinforcement Learning 66 CHAPTER 3. REINFORCEMENT LEARNING ENHANCEMENTS In the shown diagram, the reward is extracted from the environment directly, but another common way of representing this situation is to have an Actor-Critic structure. The actor would be the agent of the shown diagram, and the critic evaluates the taken action giving a reward depending on the evolution of the states. These names give a pretty visual image of someone acting according to what they have rehearsed, while receiving performance reviews from someone else, and improving thanks to this feedback. Even if this is the common situation for RL, this topic alone covers a wide variety of problems and different algorithms to solve them, considering deterministic models where taking a determined action on a given state will always result in the same evolution, stochastic ones where there is a chance to have a different evolution even with the same conditions and actions, model-free approaches, situations where the reward depends on just the present or considers possible future events, and the list goes on with several approaches to each specific situation too [33], [34]. Looking at the applications of RL in the Robotics field, this variety does not diminish. In the recent years, research has started focusing on robots not only on enclosed industrial environments, but also in everyday situations, where the surroundings can be more varied, the variety of tasks is bigger, and, usually, there are humans nearby, as assistive robotic applications have seen a radical increase too. In these applications, RL is usually found to learn how to perform an arbitrary task, sometimes assisted by Learning from Demonstration (LfD), where the task is first done by a human expert so the robot has a starting point and some data on successful outcomes [35], [36], and RL can be applied to find policies that reach performance levels even beyond those of the human expert. A prominent approach to RL in Robotics are the policy search algorithms. In them, a given policy is parametrized and the focus is learning the parameters that produce the optimal outcome [37]. However, in policy gradient approaches, information is lost upon updating and improving the policy, and they are either heavily influenced by starting conditions, usually optimizing to a local minimum without exploring further, or produce infeasible solutions. This motivated the creation of a new algorithm, named Relative Entropy Policy Search (REPS), which estimates new policies based on data distributions, bounding the information loss [38]. Research at the IRI has used this method successfully in robotic applications, for example parametrizing movements with Probabilistic Movement Primitives (ProMPs), which capture the variance of a set of demonstrations [39], and even improving the method with the proposed Dual form (DREPS) seen in [40]. During Chapter 2, the need for a learning approach to find the parameters of the used cloth linear model was made evident, especially in Subsection 2.5.4, where their dependency on mesh size and sampling time is shown, and using a general optimization failed, due to all the mesh nodes through the entire trajectory depending on the parameters of the model, so only reduced meshes and trajectories worked without reaching computational limits, yielding poor results. It is clear how having a trajectory depending on some parameters is the same case as the base of Policy Search, parametrizing a policy. Among these methods, REPS was chosen as the most suitable one, given the proven results in situations similar to the one studied in this Thesis, and the experience gained from that research at the IRI itself. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 67 Once delving into the world of RL, we can think beyond the previous application and try to improve the Model Predictive Controller itself with learning techniques. In several parts of Chapter 2 (e.g., when redesigning the controller, when using other models, or when minimizing the slew rate), it was clear how tuning the MPC was key to improve the performance, and the conditions also affect this tuning, which had to be performed manually until the best outcome is found. Furthermore, we were limiting the weighting matrices (with size 6×6) to have the same values for all components, when they could be different for each one. This amount of parameters and their effect on the control problem justify also applying RL to better tune the MPC. Even if their fields are independent and they have evolved separately, RL and MPC share some similarities, which have already been reported in literature [41]. Mainly, both feature an optimization problem at their base. In fact, a much greater similarity can be found between RL and Adaptive Control, which also tries to find the value of some parameters to improve the control inputs applied based on past data [34]. Adaptive Control is focused on estimating the model from input and output data, and updating the controller with progressively better estimations. This is a more mature field, using parameter estimation techniques on models with fixed structures, but even then, it has its shortcomings, as it works together in two opposing ways: good estimations need rich data, with varied values, to have more information about the input-output dynamics of the system and how it behaves, so an estimator needs a constantly moving output signal, but a controller wants exactly the opposite, usually wanting to lead the system to a steady state. This problem is reduced with an accurate initial identification of the model using rich data before closing the loop, and this type of control is widely implemented, with one particular case being of course Adaptive MPC, where new data helps tune the COM [42]. Changing this scheme to one using RL is not such a big leap, only needing to swap the method to obtain the improved model. However, this learning is also focused on the model, not tuning the controller. Other interesting combinations of RL and MPC can be found in the literature. It is common to see structures like the one found in [43], where there are several control layers, with a direct feedback control, an MPC on top, and then a learning layer. The application shown in [44] is very notable, as it actually uses an Actor-Critic algorithm to learn the objective function of the MPC itself (and the Related Work section is very illustrative too). This idea is much closer to the mentioned objective for our application, letting the RL handle the MPC tuning. With this last approach, we are once again in a situation where Policy Search can be applied. The parametrization of the policy here is a bit more complex, as the parameters are found in the objective function of the MPC, but the concept is the same, changing the parameters changes the outcome of the policy (if the optimization weights change, the control signals will be different). This affects the state of the environment, and a reward can be computed accordingly, for example, corresponding to tracking performance or also accounting for computational time. Given we are in the same type of situation as before, the chosen learning algorithm will be REPS too, for the same reasons as before. With the state of the art analyzed and the algorithm to be used chosen (both considered applications will use REPS), what remains of this section will be dedicated to explaining the basics of this method. 68 CHAPTER 3. REINFORCEMENT LEARNING ENHANCEMENTS The entirety of the algorithm will not be replicated here, as it can be found in [38], but some remarks can be added in combination with [40] to adapt it to our described applications, where we want to find the parameters that describe the optimal policy, be it the model parameters so its behavior is similar to the real cloth or the objective function weights to obtain the best tracking. With this reasoning in mind, a policy 𝜋(𝜃)is represented by a multivariate Normal distribution, with a vector of means 𝜇and covariance matrix Σ. The samples generated can be written as 𝜃∼𝑁(𝜇, Σ). In the REPS algorithm, the Kullback-Liebler (KL) divergence is used. This indicates the difference between two probability distributions (𝑝, 𝑞) over a random variable (𝑥), as seen in (3.1). KL(𝑝||𝑞)=∫𝑝(𝑥)log 𝑝(𝑥) 𝑞(𝑥)𝑑𝑥 (3.1) This indicator is used in REPS to give a divergence bound 𝜖to the difference between a given previous policy, 𝑞(𝜃), and a resulting one, 𝜋(𝜃). With known associated rewards 𝑅(𝜃), the optimal policy 𝜋∗is found as the solution of the optimization shown in (3.2), which is proven to be (3.3) in the literature. 𝜋∗(𝜃)=arg max 𝜋∫𝜋(𝜃)𝑅(𝜃)𝑑𝜃 s.t. 𝜖≥KL(𝜋(𝜃)||𝑞(𝜃)) 1=∫𝜋(𝜃)𝑑𝜃 (3.2) 𝜋∗(𝜃) ∝ 𝑞(𝜃)exp 𝑅(𝜃) 𝜂(3.3) When performing policy iteration with several samples, the exponential term of the solution acts as a weight for all the samples 𝑞(𝜃𝑘)of the previous policy, in order to obtain the new policy considering the corresponding reward for each one (𝑅(𝜃𝑘), or 𝑟𝑘). The new value 𝜂is the Lagrange multiplier for the KL bound, obtained by optimizing (3.4), the dual function of the problem in (3.2). ℎ(𝜂)=𝜂𝜖 +𝜂log  1 𝑁 𝑁 Õ 𝑘=1 exp 𝑟𝑘 𝜂 (3.4) To put all these concepts together into an algorithm we can use to obtain the optimal parameters, we show the final step by step process in Algorithm 2. In fact, the application of 𝑑𝑘mentioned in line 8 corresponds to (3.3), and this weighting can be done with a simple weighted average of the executed samples, or using Weighted Maximum Likelihood Estimation, for example. In the update, we obtain the new means and covariance matrix of the parameters that describe the policy, so after several policy updates (epochs), we have the optimal parameters, which is what we are actually looking for. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 69 Algorithm 2 Policy iteration using REPS Require: 𝑁, 𝜖, 𝑞(𝜃) 1: for each policy update or epoch 𝑖do 2: for 𝑘=1· · · 𝑁do 3: Perform an experiment with 𝜃𝑘, a sample from current policy 𝑞(𝜃𝑘) 4: Compute reward 𝑟𝑘 5: end for 6: Optimize the dual function ℎto find 𝜂 7: Find relative weights 𝑑𝑘=exp 𝑟𝑘 𝜂 8: Apply 𝑑𝑘to the corresponding 𝑞(𝜃𝑘)to find the new policy 𝜋(𝜃) 9: 𝑞(𝜃) ← 𝜋(𝜃) 10: end for In case completing the experiment for each sample is slow, and the number of samples or rollouts per epoch 𝑁is low in consequence to have learning processes over reasonable amounts of time, we can help this process by considering experiments from past epochs, and computing relative weights for 2𝑁samples and rewards. This will be especially useful in Section 3.3, where every experiment is a closed-loop simulation with the MPC controller in action. The considered experiments are very sensitive to the parameters they depend on. When learning the parameters of the linear model specifically (Section 3.2), even a slight increase in stiffness can result in the model being unstable, and the rewards can be several orders of magnitude lower, even being considered as −∞ or Not a Number (NaN) by the program executing the algorithm. Even filtering these extreme cases, having rewards that differ so much makes the relative weights 𝑑𝑘basically be either 0 or 1 (simply indicating “went wrong” or “nice”), which removes the differences between the successful, stable experiments that would make the policy update move towards optimal values. This can be fixed by adding a minimum reward limit, where results under this value saturate to it, indicating any reward under this value should be considered as equally wrong. The value itself must be low enough so the relative weights of the corresponding experiments are almost if not exactly 0, but not too low so the rest of them lose their proportional differences, and the new means can evolve towards the best values among them, not only those that worked. This modification can accelerate the learning process greatly, and reduce the cases where, even after several updates, the number of successful (stable) samples is low. Even then, the first epoch might result in all the samples being unstable given some initial distribution, in which case the policy is not updated and the epoch is repeated. Another slight modification done to the basic REPS algorithm is to add a factor 𝜆to the diagonal of the covariance matrix Σ, in order to increase exploration a bit (making it more likely to choose sample parameters a bit further away from the mean) and reduce the chance of falling in a local optimum and never leaving it. This factor can be obtained, for example, as proportional to the mean of the Singular Values of Σ, so it is larger on the first epochs, and decreases as learning progresses, and the variance goes down to have samples closer and closer to the optimal policy. 70 CHAPTER 3. REINFORCEMENT LEARNING ENHANCEMENTS With the used algorithm and the performed modifications explained, we can now focus on the work done for each one of the two applications of RL developed in this Thesis. 3.2 Learning the Parameters of the Linear Model The first implementation of RL was to find the parameters of the linear cloth model shown in Section 2.3 when changing mesh size and/or the time step (𝑇𝑠). Expressed in a local base, there are a total of seven parameters to be found, stiffness and damping for each coordinate plus the super-elastic correction parameter, Δ𝑙0𝑧. The parameters of the original model, a mesh of size 4×4and discretized with 𝑇𝑠=0.01 s, were found through optimization, minimizing the difference between its evolution and that of a nonlinear model. In fact, only the difference between the positions of the lower corners was considered, and not the rest of the mesh nor the velocities. This was possible given its small proportions, and using a relatively short input trajectory, as the optimizer could not reach a feasible solution in other situations (even increasing the maximum allowed time, the resulting parameters produced unstable evolutions). Following this idea, the first implementations of REPS to find the linear model parameters were also done comparing its evolution to the nonlinear model, or more specifically both versions used, explained in Section 2.3 and Subsection 2.5.4. The obtained rewards, however, depended on the distance in position from the complete linear mesh, instead of just the lower corners. It is important to note that these comparisons are open-loop simulations, simply registering inputoutput data of the models. The input is a trajectory for the upper corners of the cloth, computed arbitrarily beforehand, and applied to both models at the same time to save their respective evolutions and obtain the differences between them. This is mentioned now because this is also much easier to implement in a real scenario than the full closed-loop scheme, and thus data from the real cloth could be gathered quite early on in the duration of this Thesis, with just a Cartesian controller, the robot and a Vision algorithm, and no need to connect the MPC. Even the processing of the Vision data could be done offline after the experiments. Having access to real data meant being able to completely skip the nonlinear model, as it was a step between the linear one and the real cloth, and make the linear model, which is the one used in the controller, behave directly like the real cloth, given the same input trajectories. With this in mind, the rest of this section is divided in the following sequential steps to obtain the optimal model parameters for different sizes and sampling times, using only data obtained from the real cloth, and setting aside the previous experiments done with the nonlinear model, as it is now an unnecessary step between what will actually be used in a real implementation (even adding the “backup” SOM mentioned through Section 2.5 due to the feedback being slow, this model would be linear to be computationally fast). Each one of these steps is explained in its own subsection: first, the creation of the input trajectories and offline processing of the gathered real data in Subsection 3.2.1, then the learning process itself using this data in Subsection 3.2.2, and finally an analysis of the different learning experiments, and the definitive linear model parameters in Subsection 3.2.3. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 71 3.2.1 Gathering and Processing Real Data To begin the data gathering process, first we need to create a dataset of input trajectories that are executable by the real robot. With the modification introduced at the end of Subsection 2.5.6, once knowing the real robot was going to be a WAM by Barrett, the feasibility could be checked in simulation. Some trajectories, seven to be precise, were already created for the experiments performed with the nonlinear model, but they either had some unreachable points or were situated in a region of the workspace that had obstacles in the real scenario at the IRI laboratory. Some of these trajectories were variations on the one used in the optimization to obtain the original parameters, making it longer, faster and/or with more resolution. To keep this trend, trajectory 8 is simply the final one of the previous set moved to a feasible and free region of the workspace. Trajectory 9 follows the same trajectory but adds a final fast rise at the end. Trajectory 10 is radically different, with a semicircular motion contained in the 𝑌 𝑍 plane, and trajectory 11 has a fast diagonal movement and then a circular motion (quarter of a circle) in the 𝑋𝑌 plane. We have kept the original numbering instead of starting from 1 for the ones shown in this document in case a reader or future researcher wants to check the codes attached or linked in the Appendices, so they can see direct correspondence with the files. These four trajectories are shown in Fig. 3.2, where the WAM model is included for the first one. Figure 3.2: 3D plots of all the used input TCP trajectories 72 CHAPTER 3. REINFORCEMENT LEARNING ENHANCEMENTS In theory, if we want versatile model parameters that can be used with any trajectory, we could learn them executing always the same one, including movements in all possible directions. However, if the learnt parameters are significantly different depending on the situation, it might be better to apply a combination for each specific case, learning which are these cases too, instead of a one-fits-all set. These input trajectories are for the TCP, which is higher (offset Δℎ) than the upper corners of the cloth. They also contain a full pose for each point, not only a position. It is important to note that even the circular motions are still translations, and the orientation of the TCP is always constant, with 𝑍𝑇 𝐶𝑃 pointing vertically down like shown in the figure. This is because the parameters depend on the orientation of the cloth, and will be applied to its local base (for example, as seen in Subsection 2.5.2, stiffness in the direction perpendicular to the cloth plane is much lower, and this property must be kept), so it is easier to keep one orientation for the whole learning process, learn the parameters easily when the change of base is the identity (𝑅𝑐𝑙𝑜𝑡ℎ =𝐼), and apply rotations afterwards. In fact, two additional trajectories, 12 and 13, were created including rotations, but were used to validate the behavior of the system in this situation, and not to learn. The details on the method used to execute the trajectories can be seen in Section 4.3, concretely the “Read Node” explained there. The process to obtain output data through Vision is explained in detail on Section 4.4, including how to calibrate and change from camera coordinates to robot or world ones. Even if for the first sessions some of the processing code was not ready and running online, so the gathered data was raw, in coordinates relative to the camera, the processing done offline is exactly the same, so we refer to that section to avoid repeating the same processes. However, this data (Fig. 3.3 shows an example) must go through some extra steps to be used for learning, and those will of course be explained here. Figure 3.3: Raw data gathered through Computer Vision Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 73 First of all, we can clearly see how noisy the output is, which makes the obtained evolution not representative of that of the real cloth, and not fit to be our behavior to be learnt. Having the full evolution, we can discard outliers right away. These are all points far away from their two neighbors in time (previous and next points), which will be substituted by the mean between those. We can also apply filtering techniques depending on previous and future data points, which would cause a delay in a realtime implementation, but here we are just processing data to remove the noise on the full evolution. For example, we can apply a Gaussian filter with symmetric coefficients obtained with 𝑤=exp (−Δ𝑡2/2𝜎2) for a fixed 𝜎and amount of time neighbors considered. The original point (Δ𝑡=0) receives a weight of 1, and this value decreases for points before or after it. The new point is obtained with a weighted average inside the considered window, resulting in a much smoother trajectory, as shown in Fig. 3.4. Figure 3.4: Comparison between raw and filtered data for the lower left corner After this, the data must be regularized in time, as the output not only is slow in processing the vision data and giving the positions for all the nodes in the mesh, but it also returns data at irregular intervals. This can be changed easily with a simple interpolation in time, to obtain points at equally spaced intervals, as shown in Fig. 3.5, where the result has points every 20 ms. This is one of the most crucial steps, because the chosen time here will be the 𝑇𝑠at which the simulations must run during learning, and we must remember that the parameters to be learnt depend on this value. Figure 3.5: Zoomed in comparison between original data points and regularized ones 80 CHAPTER 3. REINFORCEMENT LEARNING ENHANCEMENTS A decrease in reward going to a bigger size was to be expected, as the errors are added for all nodes and not averaged, giving a worse reward in absolute value on larger meshes by default, but, first, these are the final rewards after the whole learning process, and second, some of them are near the rewards for some 𝑛=4experiments, so the most surprising result is the range these rewards take. Besides this difference, increasing the sampling time has almost no effect for the smaller size, with the rewards having just a slight downwards tendency, but this is greatly amplified for 𝑛=7, where the rewards clearly go down as 𝑇𝑠increases. We will still consider these combinations, as the best results still have reasonable associated rewards (some blue points actually have multiple experiments behind, with very similar rewards), but we must not forget this plot, and keep in mind that the parameters for the larger size and longer time steps might not be the most reliable either. A possible justification for this is that we are using a linear model, which was created to work under very specific conditions (𝑇𝑠,𝑛and trajectory), with the modifications to work with multiple sizes being made in this Thesis (Subsection 2.5.1) and not originally. Even if this is not a typical model linearization around a working point, as we move away from the original conditions, it is not strange to find that the model does not adapt as well to them. To get a unique set of parameters for each combination, we could just average the results or use a median, that is more robust to extreme values, but being able to access the rewards for each experiment, we can use this information again to compute a weighted average, giving more weight to the best results. A representative example of this situation is shown in Fig. 3.10, for 𝑛=4, 𝑇𝑠=10 ms. Here the vertical axes are the values of the parameters themselves, separated in 3 plots according to their nature, as their ranges are quite different. They are colored with the same scale according to the final cost of the experiment they come from, and thus the color is proportional to the weight they have in the computation of the values of the final parameters. Figure 3.10: Example of computing the final model parameters for 𝑛=4, 𝑇𝑠=10 ms Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 81 In this case, we can see how there are many results in blue, indicating good rewards (around 5 mm of average error, if we consider size and weights), and just one or two outliers, with parameter values quite far away and the worst associated rewards. The final weighted average, which is actually done using relative weights 𝑑𝑘computed exactly like the ones in the REPS algorithm, gives a combination of final parameters (marked with a magenta circle) near the concentrations of blue points. The final learnt parameters for all the considered cases are shown in Table 3.1. We can see how, for the original case of 𝑛=4, 𝑇𝑠=10 ms, stiffness and damping in the 𝑌-axis have increased significantly in absolute value (they were -13.4221 and -2.5735, respectively), but they are still the lowest of their respective kind. It is also interesting to see how the stiffness and damping parameters are closer to zero both with the larger mesh size and with longer time steps, which makes sense physically too, as with more nodes the springs are shorter and do not need to be so rigid, and with more time the displacements allowed will be larger. Additionally, these parameters are multiplied by 𝑇𝑠in the construction of the model, so they seem to try to compensate. The length correction parameter goes in the opposite sense, being larger for bigger sizes and time steps. This can be related to the stiffness in 𝑍being lower, as this change will increase the super-elastic effect of the cloth being stretched under its own weight, and thus the correction factor must be more prominent. However, this length increases to the point of being larger than the cloth side length for 𝑛=7, 𝑇𝑠=25, which contributes to the point made about the results for this combination being the least reliable ones, together with the fact that the stiffness in 𝑋is even lower in absolute value than the one of the 𝑌-axis for these situations. Even then, simulations can be run with these parameters, for multiple and varied trajectories moving in all directions, and they do finish with a stable evolution and reasonably good tracking. Table 3.1: Final linear model parameters, depending on size and 𝑇𝑠 Size 𝑇𝑠[ms] 𝑘𝑥𝑘𝑦𝑘𝑧𝑏𝑥𝑏𝑦𝑏𝑧Δ𝑙0𝑧 4 10 -325.8100 -166.2000 -395.4000 -4.4870 -3.4089 -4.8105 0.01493 4 15 -123.2300 -84.5230 -185.4100 -2.8007 -2.3387 -3.2631 0.03234 4 20 -41.8000 -39.1880 -91.4970 -1.7286 -1.3309 -2.2554 0.07095 4 25 -25.2390 -28.0140 -56.5340 -1.2255 -1.2131 -1.6997 0.11614 7 10 -158.6600 -101.5600 -295.8400 -3.2529 -2.9498 -3.9572 0.09723 7 15 -49.6420 -110.2600 -186.1100 -1.8422 -2.2991 -3.0405 0.14927 7 20 -33.7230 -61.2510 -112.6700 -1.3052 -1.5734 -2.3744 0.24107 7 25 -35.4460 -38.5820 -71.7900 -1.3422 -1.1800 -1.9108 0.38143 As a final comment, with all this data obtained from the learning experiments, we could think of building a model that fits it (using regression, for example), returning a combination of parameters given some 𝑛and 𝑇𝑠between the tested values or even extrapolating further from them. However, given the proven complexity of this system, the nuances of such a regression model, and that we have obtained a sufficient variety of possibilities for the experiments needed to be done for this project, this is left out of the scope of the Thesis, and as possible future research work. 82 CHAPTER 3. REINFORCEMENT LEARNING ENHANCEMENTS 3.3 Learning the Parameters of the Controller This section has been made possible thanks to the changes introduced to the controller in Chapter 2 and the new parameters of the linear model found in Section 3.2 for different conditions. The new controller has now a large number of parameters to be tuned, if we count weights, bounds and all the different categorical possibilities described in Subsection 2.5.7, which, even if some of them make more sense theoretically in one way for a better tracking, we are trying to strike a balance between minimal error and computational time, which act against each other, so we can apply RL to the controller and obtain results for all these options. One example of the new possibilities we can simulate (and later execute on the real robot) is shown in Fig. 3.11. This simulation was executed with 𝑇𝑠=20 ms and a COM side size of 𝑛=7, which is one of the most demanding situations for the model, and we can see how even then, the tracking is impressively good. The KPIs demonstrate so, with a tracking RMSE (“Norm” method) of 11.3 mm and an average computational time per step of 28.9 ms. Figure 3.11: Simulation results with a longer trajectory in more demanding conditions The increase in error with respect to the previously executed simulations can be justified by the new trajectory and conditions, and we can see how the differences between reference and lower corner trajectories are localized in the hard, fast turns when returning back to the initial position, a result that matches the expected behavior when minimizing slew rates. However, one might think the increase in computational time is unreasonable, given how times lower than 10 ms were reached before. This is actually not the case, and brings us back to the topic of tuning and the amount of parameters that can change. These almost 30 ms were obtained using 𝐻𝑝=25, and for a 𝑇𝑠of 20 ms correspond to about Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 83 1.5 times the step, which is equivalent to the 15 ms obtained with 𝑇𝑠=10 ms and 𝐻𝑝=30 seen back in Table 2.1. We can see how the proportion is similar with also similar prediction horizons, with the difference being the total prediction time, increased from 0.3 to 0.5 s with the new conditions. With just reducing the horizon to 𝐻𝑝=20, the average time goes down to about 14.4 ms, already under the time step of this simulation, with almost no increase in error (less than 1 mm) and with a total prediction time still over the original one (and much more than the 0.15 s needed to get times under 𝑇𝑠when this value is 10 ms, as analyzed in Subsection 2.5.5). And, of course, these are not the only values that affect the results. For example, both weighting matrices of the objective function, 𝑄and 𝑅, had to be tuned again. Done manually, this process takes long and is based on trial and error until finding the combination that yields the best results, so it is the perfect opportunity to apply RL. This section is divided in three parts. The first one (Subsection 3.3.1), describes all the parameters on a closed-loop simulation that could be potentially learnt. The second part, Subsection 3.3.2, contains an analysis of the categorical parameters deciding the exact structure of the controller, and finally, the best option is then tested to obtain the optimal overall results in Subsection 3.3.3. 3.3.1 Initial Considerations Before starting the learning experiments, we can take a look at all the parameters that affect the outcome of the closed-loop simulations, identify them and their nature, and decide the best course of action. First, we have the mentioned sampling time 𝑇𝑠and prediction horizon 𝐻𝑝. They can only take discrete values, as the linear model depends on the former and we only have learnt parameters for a select few values, and the latter can only be an integer (number of steps). For this reason, and also because their final values will depend on the conditions of the real implementation, we will not use them as parameters to be learnt, but instead we will try to learn the best overall combination of the rest of parameters applicable to all possible situations. In fact, a dedicated analysis has been made comparing real results obtained for multiple combinations of these parameters, shown in Section 5.3. Next, we have the objective function weighting matrices, 𝑄and 𝑅. Until now, for 𝑄, we have used the adaptive weight 𝑄𝑎, depending on the direction of the reference inside the horizon, multiplied by a constant value 𝑄𝑘, but we could try disabling 𝑄𝑎and check its effects. Furthermore, both matrices multiply vectors of 6 components (3 coordinates of the 2 corresponding corners), so they have size 6×6, but until now we were using the same weight for all components. Weighting the same direction on each corner differently makes no sense in the considered application, which wants to track both trajectories and moves the cloth with one robot, but we could still consider 3 different weights on each matrix, one per Cartesian space coordinate. Theoretically, this approach could make the optimal weights depend on the executed trajectory, so it might go against finding the best overall case. However, it could be a faster substitute for the adaptive weights, and we could apply different pre-computed weights depending on the situation checking longer sections of the future reference. Of course, this would be better only if the computations are actually faster, there is a clear relationship between trajectory and optimal weights to select from few optimal cases, and these cases offer significantly better results if applied in their specific scenario rather than using a more general combination of weights. 84 CHAPTER 3. REINFORCEMENT LEARNING ENHANCEMENTS An important remark about these weighting matrices is that what actually matters are the relative weights, the proportions between these values. Having 𝑄=2, 𝑅 =1makes all components inside each matrix matter the same, with the errors having double the cost of the control signals (or slew rates), and the same can be said about 𝑄=20, 𝑅 =10. With them being in a function to be minimized, all constant factors do not affect the final optimal solution, only the value of the objective or cost function at that point. If we multiply all weights by a constant, the optimal value of the objective function will be multiplied by that constant, but the optimal solution itself will not change. This is important now, because if we consider all the weights as independent parameters to be learnt, we would be considering an additional Degree of Freedom (DoF) that is not actually there, and we might get optimal results for a wide variety of combinations (high variances), sliding over the free parameter. Therefore, we need to apply some restrictions to these values. When they have the same weight for all components (each matrix is a constant times the identity), we can represent them in a 2D plot to have a clear visual representation of their behavior. Fig. 3.12 shows multiple ways to represent the two values where once one is fixed, the other will be too, so the free variable is their relative value. This proportion is constant along any line crossing the origin, like the shown red one, so one initial way of having a unique free parameter would be to use the angle 𝜃with the horizontal axis (in this case 𝑄), ranging from 0 to 𝜋/2rad, with 𝑄=cos 𝜃,𝑅=sin 𝜃. A very similar but simpler alternative is the one depicted in (a), where 𝑄+𝑅=1and they are always in the shown diagonal line. While this is linear and we can easily see which parameter is larger, the exact proportion is not directly clear. One way to solve this is fixing one of the two to always be 1, like in (b), where 𝑄=1 and 𝑅is free, either smaller or larger. The only disadvantage of this is that the non-fixed parameter is not bounded, and this can become a problem if the optimal proportion is very large (orders of magnitude apart), as depending on the starting Normal distribution, we might not reach it in a reasonable amount of iterations. This is solved by forcing to 1 the maximum of the two values on every case, instead of always the same one, like shown in (c). This way, the parameters are always sliding in two edges of a finite square from 0 to 1, where their proportion is still clear to us without needing to operate, and with even extreme proportions being easily reachable for the learning algorithm, setting the lowest weight around 0. The final considered option, (d), forces the minimum of the two to 1 instead, but this results in two infinite edges, with the same problems as (b). Following this reasoning, the weights of the objective function will always be set according to (c), with the largest one always being 1. Figure 3.12: Different possibilities to represent two weights with the same proportion Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 85 More specifically, this will be true for all the generic cases: 𝑄and 𝑅being two scalars times the identity, considering different weights per coordinate but with 𝑄=𝑅(3 weights in total, but one will be forced to 1 and 2 will be really free, between 0 and 1, in any given situation), with one matrix being a scalar times the identity and the other having 3 different weights (4 in total, 3 free), or with all 6 weights being different (5 free). With 3 total parameters, we can think of them as being in the three faces of a unitary cube defined by a weight being 1 and the rest equal or lower to 1. A final case we can consider is having different weights per coordinate and 𝑄∝𝑅instead, with a linear proportion to be learnt too. Moving on, after the weighting matrices we have another categorical binary variable representing whether we consider minimizing the control inputs 𝑢or the slew rates Δ𝑢. The benefits of using the slew rates have been discussed, but with the new vast array of options, there might be some cases where not using them benefits greatly, for example with lower computational times. Finally, we have the bounds considered in the optimization problem. All the optimization variables are control signals representing displacements in a time step. The bounds are only there to ensure the linear model does not extend past a reasonable value unreachable by the nonlinear model or the real cloth. The exact limit or expected maximum value depends on the maximum slope of the reference trajectory, as the cloth must move approximately at the same rate (considering the reference is on the lower corners and the control signals on the upper ones), and there is no need for much more. Not only that, but if the simulation unstabilizes, even with bounds as low as the required slope, the evolution can be chaotic nonetheless, and if the simulation is stable, these bounds only increase the computational time if not quite relaxed. The other bounds in the problem come from the constraints, concretely the constant distance between the upper corners imposed by using only one robot and introduced in Subsection 2.5.3. These are equality constraints, so the bounds are set to 0. Relaxing this to obtain faster computations comes at the price of precision, and considering small relative movements between the two controlled corners, which will affect the resulting TCP pose. In the end, both bounds will be fixed and not considered as parameters to be learnt for the aforementioned reasons. To summarize everything explained in this subsection in a short list of parameters, we can consider: •The time step 𝑇𝑠and prediction horizon 𝐻𝑝, but they will be analyzed with experimental results due to their discrete nature and how results depend on the real conditions. •Using the adaptive 𝑄𝑎or not to obtain 𝑄. •Minimizing either the control signals 𝑢or the slew rates Δ𝑢. •The structure of matrices 𝑄(or 𝑄𝑘) and 𝑅: a scalar times the identity, a different weight per coordinate and equal or proportional to one another, or another combination. •The weights of these matrices themselves. •The variable bounds, but they only increase either error or time and have been discarded. 86 CHAPTER 3. REINFORCEMENT LEARNING ENHANCEMENTS 3.3.2 Obtaining the Most Suitable Structure From the list at the end of the previous subsection, there are several categorical and binary variables, that rather than being parameters to be learnt, must be analyzed comparing learning experiments using all possible combinations. To that end, the several possibilities for structures of the weighting matrices were reduced to just 3: two scalars (multiplied by the identity), 𝑄with different weights per coordinate but with 𝑅being the same for all of them (4 weights in total), and having different weights for each coordinate for both matrices but with 𝑄∝𝑅(also 4 weights in total). A first round of experiments were conducted using 𝑛=4,𝑇𝑠=20 ms, 𝐻𝑝=25 and following the same reference trajectory. At this point in the Thesis, the real implementation had progressed enough to test some closed-loop trajectories too, and these conditions had one of the best observed behaviors. Additionally, to reduce the time needed to execute these much slower closed-loop simulations, the simulations will be done with a linear SOM. This means that the conditions are also closer to those of the real implementation, where there is a linear “backup” SOM. The triple model scheme in simulation, that would be the closest to reality, is sadly the slowest option, as the evolutions of the nonlinear model, even if discounted from the computational time because they will not be present in reality, are the ones that take the longest to compute. However, a few simulations have been executed to check that the tendencies are the same and the results of the performed analysis apply all the same. Given that we might get times over 𝑇𝑠, two reward functions were tested. One with just the negative RMSE (“Norm” method) with respect to the reference trajectory, as an equivalent for our usual first KPI, and another also penalizing the time spent over the maximum allowed, but not times under 𝑇𝑠, called TOV. The reward was not changed to this second option without testing, in case adding a secondary objective with no additional relative weight was detrimental to learning the parameters that yield the minimum tracking errors, which is the primary objective. All these options give us a total of 24 experiments to complete in this first round: two options for 𝑄𝑎 times two for Δ𝑢times three for the structures of the weighting matrices times these newly introduced two reward function variants. The parameters to be learnt iteratively using REPS are the weights inside the matrices themselves, either 2 or 4 values depending on the structure. The purpose of this first round is to analyze the differences mainly between the two first binary categories, and check if the best option between them is consistent among all types of weights and reward functions, while trying to notice any kind of trend or pattern in the resulting optimal weights for each case. To analyze the differences between weight structures, we need to execute multiple different trajectories, as mentioned before, so this requires further experiments in the following rounds. However, this first one can also help us to decide the best reward function to use between the two candidates. As mentioned previously, the learning algorithm used is REPS, explained at the end of Section 3.1, with codes adapted for the new situation (closed-loop simulation files converted to functions depending on the input parameters and returning a reward), but the situation is completely analogous to that of Section 3.2 where we learnt the parameters of the linear cloth model. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 87 We can first divide the results of this first round depending on the KPI we are analyzing, RMSE or computational time. Then, we can organize them in groups of 4 with the same weight structure and reward function, and make comparisons purely based on the use of Δ𝑢and 𝑄𝑎. This is shown in Fig. 3.13 and Fig. 3.14, with results colored according to tracking RMSE in the former and computational time (relative to 𝑇𝑠) in the latter. Figure 3.13: Results of the first round of learning experiments (RMSE) Figure 3.14: Results of the first round of learning experiments (time) First of all, looking at the tracking errors, we can see how regardless of weight structure and reward function, the worst results are always found using 𝑄𝑎and 𝑢, and the opposite selection, with only 𝑄𝑘but with Δ𝑢, yields the best outcome. Looking at the time plots, the most common effect is that computational times increase slightly with the use of Δ𝑢. However, as we can see in the marked values of the color scale on the right, absolutely all optimal results have a computational time lower than 𝑇𝑠(20 ms in this case), making the negative impact negligible, and the best choice clear: the best results are obtained without using 𝑄𝑎but minimizing Δ𝑢. 88 CHAPTER 3. REINFORCEMENT LEARNING ENHANCEMENTS In fact, having all results with computational times under 𝑇𝑠means the reward functions were the same in most cases, with the term penalizing times over 𝑇𝑠not being active. Indeed, comparing the resulting learnt weights obtained with the two functions for each of the three structures, we can see how they are really similar, as shown in Table 3.2. The obtained rewards, as well as the KPIs separately, are also almost the same (about 0.3% of difference). This is not to say that the new reward function has absolutely no effect at any point, as the starting epochs have samples with combinations of parameters that increase the computational time to levels above 𝑇𝑠, and the additional term helps avoiding these values, potentially reaching optimal results sooner. However, the time measurements are sensitive to memory and CPU usage, and for such a minimal improvement, considering how long the learning experiments take anyway, the reward function penalizing time is discarded for future experiments. Additionally, if we used it for a general case, we would have to study the relative influence of each term, and weigh the two objectives they represent, better tracking or faster optimizations, to find the best balance. For now, we can focus in obtaining the best possible tracking, as that is the primary objective of the controller, after all. Table 3.2: Optimal 𝑄, 𝑅 weights, using Δ𝑢and no 𝑄𝑎, for all other combinations Reward Weights 𝑄𝑥𝑄𝑦𝑄𝑧𝑅𝑥𝑅𝑦𝑅𝑧 −RMSE 𝑞𝐼, 𝑟𝐼 1 1 1 0.1707 0.1707 0.1707 −RMSE −TOV 𝑞𝐼, 𝑟𝐼 1 1 1 0.1695 0.1695 0.1695 −RMSE 𝑄𝑥𝑦𝑧, 𝑟𝐼 0.9159 0.7587 1 0.1885 0.1885 0.1885 −RMSE −TOV 𝑄𝑥𝑦𝑧, 𝑟𝐼 1 0.7215 0.9576 0.1676 0.1676 0.1676 −RMSE 𝑄𝑥𝑦𝑧 =𝑘𝑅𝑥𝑦𝑧 1 0.8892 0.4362 0.1791 0.1593 0.0781 −RMSE −TOV 𝑄𝑥𝑦𝑧 =𝑘𝑅𝑥𝑦𝑧 1 0.8057 0.6741 0.1823 0.1469 0.1229 Another interesting result obtained with these experiments is that the final tracking error is approximately the same regardless of the chosen weight structure (less than 0.5% difference between best and worst), and the minimum errors are actually obtained with the first and simplest option. This means that the more complex structures, with weights depending on direction, actually do not adapt better to the trajectory leading to smaller errors. In fact, looking at the results of the first round, the consistency of obtaining worse results using an adaptive 𝑄𝑎is also an indicator that adapting these weights to the current situation might not be the best decision overall. Of course, only one trajectory has been tested, but this reasoning, the slightly better results, and especially simplicity (not having to switch between optimal weights for specific situations, if optimal solutions really depended on the trajectory) push us to select the first structure, with just two weights. Now we must check that the chosen options, which have been analyzed with results obtained using a linear SOM, are also the best for the real triple model scheme, and that all the decisions made are the best course of action for the real implementation too. To that end, and keeping in mind that the 24 experiments were done with a linear SOM to reduce the overall execution time, we will not repeat all 24 again for the different scheme, but just 4, for the first weight structure and the original reward function, to check the effects of 𝑄𝑎and Δ𝑢. The results for both KPIs are shown in Fig. 3.15. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 89 Figure 3.15: Results of the learning experiments using the triple model scheme It is clear how the results match with the ones obtained previously. The best combination to minimize the tracking error is still using Δ𝑢with no 𝑄𝑎, even if these errors are larger overall due to the evolution of the nonlinear model representing the real cloth not matching the one predicted by the linear COM as well as before. Regarding computational times, the most notable impact is still the increase when using Δ𝑢, with the worst case going slightly over the maximum allotted time of 20 ms. Thankfully, not using 𝑄𝑎reduces this time a bit, but still being on the very edge of 𝑇𝑠. This could probably be improved using the reward function with the TOV term, and in any case we must remember the times are approximate and dependant on the tasks running in parallel with the learning experiments, so the option with the best tracking is still the chosen one, and the correspondence with the previous simulations is clear. At this point, we can confidently say which is the best structure for the controller: it does not use 𝑄𝑎, the objective function penalizes Δ𝑢, and each one of the weighting matrices 𝑄and 𝑅has the same values for all coordinates, so there are only two weight values. This consistently gives the best tracking results without increasing computational times over the limit, and is the simplest alternative without sacrificing performance. However, the work is not completely done, as we have not analyzed the final optimal weights themselves and the specific tuning of the controller, and how it is affected by different trajectories and conditions. This is done in the following subsection. 3.3.3 Tuning the Resulting Controller After the process described in the previous subsection, we have settled between all the possible options for the controller, but still have not analyzed how different conditions, like multiple trajectories, 𝐻𝑝, or 𝑇𝑠, affect the tuning itself. This is different than the analysis done with experimental data on Section 5.3, where the changes in the latter two variables are compared against the tracking KPI. Here, the intention is to find if the optimal tuning is the same regardless of these conditions, or if they affect the optimal 𝑄and 𝑅weights. Luckily, the final structure of the weighting matrices is a scalar times the identity for each one of them, so we only have two values to learn. In fact, what is important, as discussed in Subsection 3.3.1, is only the proportion between these two values, so there is just one value to be learnt. 96 CHAPTER 4. REAL IMPLEMENTATION Sadly, not long after starting to look into it, several compatibility issues were found, not just with the newly developed code but also with CasADi in its C++ version. This library, which, as mentioned during the development and redesign of the MPC in Chapter 2, is used to code fast and efficient optimizations, and was chosen also due to its availability in both Matlab and C++, uses some functions only available from C++11 onward, and trying to find older versions had no positive outcome, and neither did trying the opposite, as the older codes in the IRI library could not be compiled with C++11, and splitting the library to compile only the new codes with C++11 was not an option, both technically and practically, as we would not be able to connect the MPC system with any others. In the end, the implementation on the real setup is done using Robotic Operating System (ROS) in its Kinetic Kame version [45], for Ubuntu 16. ROS is flexible, robust, modular and versatile, meaning the code used for this specific scenario using a WAM can be then applied to any other kind of robot without any effort, as the MPC itself is encapsulated in what is known as a node, and different nodes can be connected easily, even ones coded in different languages (Python and C++ are compatible) and executed from different computers. For the case at hand, using ROS made it easier to connect with the Computer Vision algorithms for cloth detection, as will be detailed in Section 4.4, but it also meant that, to connect to the robot using codes developed at the IRI, a Cartesian controller had to be wrapped as a ROS node, as explained in Section 4.3. Without going into specific details of how ROS works and all of its possibilities, to understand the final implementation we need to know that there are nodes, which are modules of code that can be run simultaneously, and that they can be connected using messages. These messages, which have their structure and requirements, must to be published by one node onto a topic, and then any other node can subscribe to this topic and receive all the messages published there. Sending commands to the robot is also achieved in a similar way, specifically using what is called an action. The nodes can also be synchronized and all refer to the same common time variable, which is important in real-time applications. Knowing these basics, we now show the full closed-loop scheme of the real implementation using ROS in Fig. 4.1, where each block is a node, which will be all explained one by one in the following sections. The specific topics used to connect them will also be detailed there, to keep this diagram simple. Figure 4.1: Full diagram of the final implementation in ROS Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 97 To conclude this section, we show the final situation of the real setup utilized to execute all the experiments. Fig. 4.2 shows one of the two WAMs found in the Perception and Manipulation laboratory at the IRI, the one used in all executions. Additionally, it shows a detail of the end-effector, with a custom 3D-printed plastic adaptor. The TCP is actually inside the central hole, on the metallic plate underneath, as specified by the manufacturer. The plastic part is used to connect with the red one shown on the lower right, which is attached to the rigid link between robot and both upper corners of the cloth. Finally, Fig. 4.3 shows a picture of the complete setup during an experiment, with camera and computer in frame. Figure 4.2: Picture of the WAM used in the real setup, and piece that connects to the cloth Figure 4.3: Picture of the full real setup during an experiment 98 CHAPTER 4. REAL IMPLEMENTATION 4.2 Adapting the Model Predictive Controller This section is dedicated to explaining the process to transform the designed MPC and existing closedloop simulation written in Matlab into a code executable in C++, to then adapt this translated code into ROS. Each one of these two steps is explained separately in the following subsections. 4.2.1 Code Translation to C++ From the starting point of the Thesis, it was clear that the codes developed in Matlab would have to be translated. Even while considering using iri_libbarrett to connect to the WAM and all other involved systems, this library is written in C++. Once it was clear the final structure would be using ROS, the programming language itself did not change, as C++ is one of the two supported languages, along with Python, with the added bonus of being a compiled language, potentially faster in real executions. There exist some automatic tools available that convert Matlab programs into C++. However, for codes involving multiple files and functions, they create complex structures of data to communicate between them, which make it harder for a human to understand them and use, tweak or even debug the resulting code, and can also result in higher iteration times with several calls to smaller functions creating other auxiliary variables. The automatic method was discarded quite early on due to these disadvantages, and knowing that the translated code would have to be adapted and easy to debug. Translation was thus done manually, and progressively from a simple function to the whole closed-loop simulation. Matlab is optimized for mathematical operations, especially vectorized ones, where vectors of different sizes and classes can be used with scalars, concatenated to themselves to grow on consecutive iterations, and other user-friendly applications. To be able to perform all the operations in C++, even after declaring and initializing all variables with fixed size, a specific library was needed. In the end, Eigen was the chosen one, as it allows for a wide variety of matrix operations using simple syntax, and the executions are still fast when the programs are compiled [46]. Besides this library, and as mentioned multiple times throughout this document, CasADi was used to code the optimization problem in C++ too. Given it is the same toolbox as in Matlab, the syntax is similar, and there are tutorials, documentation and examples in all languages [47], easing the translation process. The most notable change, in fact, was having to deal with explicit variable classes and not being able to do operations with different types together, so there is an increase of intermediate steps to change from CasADi symbolic variables to Eigen matrices and vice versa. Like in Matlab, CasADi is not the optimizer itself, but a tool to connect with one, among several options, so that the users do not need to think about low-level details, and optimizations can be fast even if the problems are formulated using high-level functions. The optimizer used is IPOPT, as in the Matlab code, to keep the same properties and structures already coded, and make the translation process smoother. As mentioned in Section 2.4, this availability in both programming languages was a key feature that lead into using CasADi and IPOPT. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 99 Even before adapting the codes into ROS, it was clear that in the real setup, the codes would run in a computer using Ubuntu 16 as the operating system. This is mentioned now because even though we are using the same tools as before, they needed to be installed in the computer connected with the robot. For the most part, following the installation steps from the CasADi GitHub page was enough, but it is worth mentioning, in case this process has to be done again, or a reader wants to use the developed codes, that to install IPOPT its checkmark had to be manually ticked using cmake GUI (Graphical User Interface) during the installation, besides everything mentioned in the tutorial. Once all these initial requirements were fulfilled, the translation process began. To achieve a complete closed-loop simulation in C++, not only the main simulation code from Matlab had to be translated, but also all the functions related to initializing and operating with the linear cloth model. This is why not everything was coded into a single file, but several separate function files were made, with their corresponding header files. Even if there were multiple subsequent changes and improvements made to the C++ code, even after adapting it to ROS, making the first translation and all the codes before the final version obsolete (hence why only this final version is attached in the Appendices), the overall structure in different files was maintained, with a file for general purpose functions, like reading from and saving to a CSV (Comma-Separated Values) file, another with functions related to the linear model, and a third one for functions more specific to the MPC, besides the main closed-loop code. Once reached this point, a closed-loop simulation was available as a standalone C++ code, and the resulting SOM evolutions could be saved and exported to check their correctness. The SOM used in C++ is always a linear model, which in this case was the same as the COM, as the learning experiments from Section 3.2 were not finished at the time. Plotting these results in Matlab, we obtained Fig. 4.4, where we can see how the used reference is correctly tracked. The next step was clear: adapt the code into ROS. Figure 4.4: Results obtained with the translated closed-loop simulation in C++ 100 CHAPTER 4. REAL IMPLEMENTATION 4.2.2 Integration in the ROS Structure The second main step in the process of transforming the MPC into the real scenario is implementing the now translated simulation into ROS. We start with an independent executable and a set of C++ files, and the goal is a node that can be connected in the overall scheme. For that, we need to define all the necessary inputs and outputs of this node, knowing its contents and place in the diagram. The node itself contains two distinct parts inside: the optimizer, which uses CasADi and has the linear COM inside, and an additional linear SOM. In the simulations, both in Matlab and C++, the SOM is used to represent the real system, but in a real application, that is of course no longer needed. Instead, the linear SOM will be used as a backup, needed to simulate evolutions at a fast rate and give new initial states to the optimizer even when the feedback signals are slow. As will be explained in Section 4.4, this part of the scheme is actually the slowest one, so the linear SOM is essential for the correct performance of the system. Knowing all this, the connections are clear: the MPC node requires the reference trajectory and the most recent state vector captured from the real cloth, and outputs a control signal. In fact, this control signal is not a displacement, as the pure 𝑢returned by the optimizer, and not even an absolute position, but a full pose of the TCP at every instant, including its orientation. In fact, on the final implementation, the MPC node publishes 3 different topics, as seen in Lst. 4.1. This is just a snippet to show the syntax of these definitions, and mention the purpose of each of them. The first one is the absolute positions of the upper corners, i.e., the control signals 𝑢added with the previous upper corner positions. This is published with debugging purposes and in case a nonlinear SOM or another node needing only cloth information is connected in the future. The second publisher sends TCP pose messages on its topic. This is the information required by a Cartesian controller to operate, but in the specific setup used, it is not expressed with the right structure, and a third publisher is needed to output Cartesian commands into the topic where the Cartesian controller is subscribed. The second publisher is not removed, once again, for flexibility and debugging purposes, as we can see (or “echo”) these messages manually from a terminal. Listing 4.1: Definition of publishers and subscribers of the MPC node 1// Define Publishers 2ros::Publisher pub_usom = rosnh.advertise<mpc_pkg::TwoPoints> 3("mpc_controller/u_SOM",1000); 4ros::Publisher pub_utcp = rosnh.advertise<geometry_msgs::PoseStamped> 5("mpc_controller/u_TCP",1000); 6ros::Publisher pub_uwam = rosnh.advertise<cartesian_msgs::CartesianCommand> 7("iri_wam_controller/CartesianControllerNewGoal",1000); 8 9// Define Subscribers 10 ros::Subscriber sub_somstate = rosnh.subscribe("mpc_controller/state_SOM", 11 1000, &somstateReceived); Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 101 The attached snippet also shows the line needed to define the receiving part of this node, its subscription. In this case, it is a state vector coming from the Vision feedback, when ready, that will update the states of the linear SOM and immediately the initial state of the COM on the next optimizer call. The reference trajectories of the lower corners are loaded from separate CSV files according to the selected trajectory number. After some initial tests with this number and some other parameters being hard-coded into the source files, they were finally moved into what is known as a launch file, to make changing them significantly easier. All changes done to a .cpp source file (inside the “src” directory) must be compiled, and changing a parameter implied changing the main source code itself. Moving all the variable parameters into a launch file makes it so the compiled code always looks up the values given by the launch file, and these can be changed easily without needing to compile afterwards. The only downside to this is that the nodes can no longer be run on their own, as they always need a launch file to get the parameters from. On the other hand, node parameters can be set through arguments in launch files, which have a default coded value but can be changed on execution on the command line itself, meaning there is not even a need to explicitly change any code at all. The parameters moved to .launch files (in the “launch” directory, fittingly) are the path where the reference trajectory files can be found, the number of the chosen reference to be followed, COM and SOM side sizes nCOM,nSOM, the prediction horizon Hp, the time step Ts and a weight representing the confidence we have in the data received from the Vision feedback, Wv, which will be explained in detail in Section 4.4, as even if it is applied inside this node, it is in the callback function called whenever a feedback signal is received from the subscription, and all the operations there are related to the specifics of the Vision-related nodes. ROS has built-in ways to check time, which is common to all nodes launched under the same master or core, compute durations and iterate at fixed rates. Theoretically, the rate of the execution is defined by 𝑇𝑠, as the frequency is just the inverse of the period. However, in this first implementation in ROS, as also happened in the previous simulations, both in C++ and Matlab, the codes do not run at true real time, as the optimizer is placed in the middle of an iteration, blocking the execution of the subsequent lines of code until a solution is found. If the computational time is lower than 𝑇𝑠, then we can force a sleep time of the remaining time until that value, but if the optimizer takes longer than the theoretical limit, the whole system waits for it to finish. This is of course not an acceptable situation in the final implementation, and will be changed with the modifications explained in Section 4.5. They are left for the end of this chapter instead of explaining them now with the implementation of the MPC node to mirror the progress done during the development of this Thesis. Given how some tests were made without the MPC working at true real time, connecting with the Cartesian controller first, and then closing the loop with processed Vision data before implementing the necessary changes to work in real time, enough data was gathered from each situation, and a comparative analysis is shown in Section 5.1 of the Experimental Results chapter. Without adding this change yet, the MPC node was ready to run, and could also be tested on its own as if it was a simulation thanks to the linear backup SOM, with positive results. 102 CHAPTER 4. REAL IMPLEMENTATION 4.3 The Cartesian Controller Nodes Describing and especially controlling the movements of a robot are not simple tasks, with their own dedicated fields of study. In simple terms, robot motions can be expressed in two completely different ways. The first one is more natural to the manipulator itself, and uses only information coming from its joints, like positions, velocities, accelerations and torques, or what is known as the joint space. Controlling a robot like this does not require of conversions, the robot receives inputs in the joint space and outputs joint information too. However, for humans to plan the movements and understand the output information, it can be difficult to understand, as the correspondence with the natural, Cartesian 3D space is not direct. As a solution to this, we have the second option, expressing the movements in Cartesian space, with positions and orientations that are easily understandable by human users. However, this requires a conversion from Cartesian to joint space to connect with the robot and give it inputs, and then convert the joint information back to Cartesian space to check results or add some kind of feedback. A Cartesian controller is precisely this intermediate step needed to connect with the robot when the commands being used are expressed in Cartesian space. The process of going from TCP poses to joint values is known as Inverse Kinematics (IK), as the opposite and much easier process of obtaining a Cartesian pose knowing the joint positions is known as Forward Kinematics (FK). It is clear how we need a Cartesian controller between the output of the designed MPC and the WAM, so that the robot can follow the computed control signals. In fact, Cartesian control, besides the necessary IK, can also involve other considerations related to task and motion planning, like collision detection or path finding. To perform both FK and IK, the dimensions and types of joints are needed for the specific robot being used, as, for example, a joint displacement of a certain angle would move the End-Effector (EE) a different distance depending on the lengths of the links between them. The standard way of expressing these specifications are the Denavit-Hartenberg (DH) parameters [48], a set of four values for each joint that indicate their nature and how they are connected. For the used 7-DoF WAM, the dimensions are public on their website [49], and they were also available at the IRI, given that robot has been used before. The resulting DH parameters are shown in Table 4.1, where the joint variables 𝑞𝑖have been added too, to show that all of them are revolute (affecting the angles 𝜃𝑖), and none is prismatic. Table 4.1: DH parameters of the Barrett WAM Joint 𝜃𝑖[rad] 𝑑𝑖[m] 𝑎𝑖[m] 𝛼𝑖[rad] 10+𝑞10 0 −𝜋/2 20+𝑞20 0 𝜋/2 30+𝑞30.55 0.045 −𝜋/2 40+𝑞40−0.045 𝜋/2 50+𝑞50.30 0 −𝜋/2 60+𝑞60 0 𝜋/2 70+𝑞70.06 0 0 Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 103 With these parameters known, we could use a generic Cartesian controller from the literature that works well for our case. Several recent publications have furthered the research on compliant control, important in tasks with human-robot interaction or humans in the workspace [50], while others also apply learning techniques in the Cartesian controllers [51], so there are a lot of viable options. To keep it simple and use work previously developed at the IRI itself, the Cartesian controller developed in [52], [53] will be used. There are several important details about this controller that must be mentioned. First, it is not implemented in ROS, so it must be wrapped in a node to be compatible with the rest of the scheme. Secondly, it includes several versions, from the most simple one with constant gains to a Variable Impedance Control application, also designed for environments with humans nearby. Also, for a trajectory tracking problem, it needs the entire trajectory at once, to process it as a whole and then execute it point by point. Finally, it considers the redundancy present in the WAM to avoid singularities and increase robustness. This last point has another implication, combined with the fact that, commonly, an IK problem does not have a unique solution, as, in general, the same TCP pose can be reached with multiple combinations of joint positions. For any given pose, this Cartesian controller computes all the possible joint movements that produce no end-effector displacement (null space), and reduces the strength of the joints in those cases, enabling the possibility of an outside force, like a human, moving some parts of the robot (commonly the elbow) without displacing the TCP from the desired position. This is a feature of all the different modes of the controller, including the most basic one using constant gains and no variable impedance. Wrapping the Cartesian controller into a ROS node is out of the scope of this Thesis, and the process will thus not be detailed in this document, provided that it was carried out at the IRI. While a simple, standalone version with the bare minimum to function could have been coded from scratch without using the described base, it was important that the used controller could be utilized with the WAM at the Perception and Manipulation laboratory without compromising any of the previous work and connections already in place, both using ROS and libbarrett. After some time, the final product was ready to use, adapted to work in real time knowing only the most recent point to follow (updated by subscribing to the corresponding topic) and not the whole trajectory. However, only the version with constant gains was available, separated to be able to test the controller as soon as possible with the minimum needed to test the trajectory tracking application. Luckily, even this version includes the redundancy considerations mentioned earlier. Combining MPC with Variable Impedance Control can be an interesting topic of future research, possible when the full controller is available in ROS. Besides the Cartesian controller node itself, which is subscribed to a Cartesian goal topic (the one published by the MPC node shown in Lst. 4.1), and uses an action to connect with the WAM, a separate node was created to work in open-loop applications, like gathering data from the cloth mesh to start the learning process of Section 3.2. For these cases, the only required nodes are the Cartesian controller and the Vision feedback ones, but the Cartesian controller needs a specific message, and reference trajectories, even the ones containing just TCP poses, were saved in CSV files with a specific format, compatible with the original version of the controller that loaded the whole trajectory at once, but not in ROS. 104 CHAPTER 4. REAL IMPLEMENTATION This new node, dubbed “Read Node” due to its main function, first loads the trajectory contained in a given CSV file with the following format: one point on each row, time steps on the first column, then 7 columns corresponding to joint positions (of which only the first one was used to initialize the robot on the initial Cartesian controller, and now is not considered), 3 more columns for the positions of the TCP in 𝑋,𝑌and 𝑍, and finally 4 columns corresponding to the orientation expressed as a quaternion, with the scalar component first. Once the file is processed and the Cartesian trajectory is fully saved on a variable, the node enters a loop with a rate defined by the time steps of the first column of the input file, where it picks the next point in the trajectory, formats it as a Cartesian command message, and publishes it on the topic the Cartesian controller is subscribed to. The structure and nodes involved in an open-loop execution like gathering cloth evolution data is shown in Fig. 4.5. Figure 4.5: Diagram of an open-loop implementation in ROS using the Read Node In this case, instead of closing the loop, the data coming from the Vision nodes is saved together with the input commands into a rosbag file, directly capturing the messages being published in the specified topics, which allows for a reproduction of them afterwards to simulate data being published without using the real setup, and of course they can be converted to TXT or CSV files (one per topic) to later process the data using Matlab or any other software. The main code for this Read node can be found following the link to the full source code attached in the Appendices document. Additionally, a complete guide is included both to use the wrapped Cartesian controller, initializing the action and sending goals via terminal, and to edit it, for example, to change the gains, as the process has multiple steps in different directories that must be done every time. 4.4 The Vision Feedback Nodes The final section of the full scheme to be described is the one including the feedback nodes, with the steps required to go from a camera back into the MPC. This section is divided in three parts, corresponding to distinct nodes needed to perform these steps and to enable an actual closed-loop implementation. The first step is of course capturing the data and applying the required Computer Vision algorithms to identify the cloth and output a mesh. This is described in Subsection 4.4.1. The positions of the nodes of this mesh are given in the coordinate frame of the camera, so a calibration process is needed, as explained in Subsection 4.4.2, to know how to perform a change of base. Finally, Subsection 4.4.3 details all the changes necessary to close the loop and update the initial state of the MPC using real feedback data. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 105 4.4.1 Obtaining the Cloth Mesh Converting image data given by a camera into a cloth mesh where we know the positions of all the nodes is not an easy task, let alone converting an image to the required state vector. Processing images and their data is part of a whole dedicated field known as Computer Vision, and the algorithms required to extract the required information, identify the cloth piece from an arbitrary picture to then build a mesh according to the detected shape and output are far beyond the scope of this Thesis. Of course, a code that performs these tasks is required to close the loop with real data, but instead of developing one from scratch, we can use an already existing ROS node. The perfect candidate is the Real-time multi-cloth point cloud segmentation ROS package [54], developed by M. Arduengo et al. at the IRI itself to detect cloth pieces, separate multiple ones by color, and finally merge points that are close both in distance and color into point clouds representing the cloth pieces. The ROS node is coded in Python, but as mentioned in Section 4.1, both languages can be used in the same structure or full scheme including multiple nodes, the only requirement is to have a working connection, for example, a subscriber in a node written in C++ receiving messages published by this node on a certain topic. An important note about this node is that it uses a You Look Only Once (YOLO) object detection algorithm with a neural network which is resource intensive and needs a powerful GPU, not found in the computer connected to the WAM at the laboratory. This is however not a problem when using ROS, as several different computers can be connected under the same core or master, and launch nodes from different machines onto the same structure to work together. This Cloth Segmentation node can be launched in one computer with the camera connected to it, publish the mesh data into a topic, and the node subscribed to this information can be in another computer, the one actually connected to the WAM. While the node publishes several topics, the one needed for our application is cloth_segmentation/cloth_mesh, containing the positions of all the nodes of the obtained mesh. To compute these positions in the three axes of space, the camera needs to output a Red-Green-Blue-Depth (RGB-D) image, and the one used for the real setup is a Kinect camera like the one shown in Fig. 4.6. These positions are expressed in a local base relative to the camera, and must be converted to the global reference frame to be able to update the initial state of the MPC. To do that, first we need to know where the camera is located, with the calibration process described in the next subsection. Figure 4.6: A Kinect camera, used to capture RGB-D images and obtain current mesh positions 112 CHAPTER 4. REAL IMPLEMENTATION Finally, after updating the feedback data to current time and filtering distant results, the backup SOM state vector is updated to the received one, and the initial state of the COM that goes into the optimizer is obtained with the updated value too. However, the received data still maintains some variance added from the noise and cloth detection processes, and trusting it blindly updating the state vector directly resulted in nervous evolutions and worse errors. This is why a final parameter, a Vision weight 𝑊𝑉, was added, to consider the reliability of the feedback data versus the state vector obtained simulating the backup SOM. This way, the update that finally closes the control loop follows (4.5). 𝑥𝑆𝑂𝑀 ←𝑊𝑉𝑥𝑉 𝑖𝑠𝑖𝑜𝑛 + (1−𝑊𝑉)𝑥𝑆𝑂𝑀 (4.5) The effects of this weight, together with the sampling time 𝑇𝑠and the prediction horizon 𝐻𝑝, are analyzed using experimental data in Section 5.3. As a summary of all the steps the captured Vision data must go through to close the loop, we now show Algorithm 3. Algorithm 3 Steps to close the loop with Vision feedback data Require: Camera publishing RGB-D images 1: for each image captured at time 𝑡𝑐do 2: Use the Cloth Segmentation node to get the mesh positions 𝑝(𝑡𝑐) 3: Apply a filtering process: FMA, SMA, WMA or EMA 4: Order the node positions from left to right, bottom to top 5: Apply a base change from camera to world reference 6: Publish mesh using original acquisition time stamp, 𝑥𝑉(𝑡𝑐) 7: Receive mesh on MPC node: callback function 8: Compute delay Δ𝑡=𝑡−𝑡𝑐 9: if Δ𝑡 > Δ𝑡𝑚𝑎𝑥 then 10: Discard feedback data 11: Exit callback function 12: else 13: Update data simulating Δ𝑡/𝑇𝑠steps, obtain 𝑥𝑉(𝑡) 14: if 𝑥𝑉(𝑡) − 𝑥𝑆𝑂𝑀 (𝑡)>Δ𝑑𝑚𝑎𝑥 then 15: Discard feedback data 16: Exit callback function 17: else 18: 𝑥𝑆𝑂𝑀 (𝑡) ← 𝑊𝑉𝑥𝑉+ (1−𝑊𝑉)𝑥𝑆𝑂𝑀 ⊲New initial state for the MPC 19: end if 20: end if 21: Exit callback function 22: end for Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 113 4.5 Implementation in Real Time Until now, all the closed-loop implementations had the optimizer as a blocking step in the MPC, both in simulation and in the MPC node. This was not a problem in simulation, as after any arbitrary time needed by the optimizer to find the solution, we can force a step of 𝑇𝑠to simulate the SOM, get the feedback signal and continue with the next iteration. Even if each iteration takes a different amount of time to complete, the evolutions are virtually simulated at constant step times. This is not the case in a real implementation, where the optimization time is added to the computational time of all other commands in an iteration of the MPC node, and if the total time is greater than 𝑇𝑠, each step cannot be completed in time, causing the output control signals to be delayed, and the incoming feedback signals to be incorporated later too. This is why we need to change the structure of the MPC node to avoid having the optimizer as a blocking step, and ensure each iteration is completed within 𝑇𝑠and there is always an updated control signal being outputted at a constant rate. The simplest way to have two codes running in parallel and not blocking each other in ROS is using two separate nodes. This is the main change to the overall ROS structure needed for a real time implementation, and thus a new diagram is shown in Fig. 4.9. The old MPC node is now divided into a node dedicated only to the optimization problem (Opti Node) and another one (RT Node) with just the linear SOM that ensures a constant output rate of TCP poses, one every 𝑇𝑠, and also handles the feedback signals. Figure 4.9: Diagram of the ROS implementation in real time In this diagram, we have marked with a solid line the messages that are published at a fixed rate, one every 𝑇𝑠, and with dashed lines all the others that are not. The separation between dashes is proportional to how slow these messages are published, with the optimizer being the fastest among them (sometimes will finish before a full time step, sometimes it will take longer, depending on initial conditions and reference). The camera outputs RGB-D images at 30 Hz (33.3 ms between frames), but the Vision node needs around 100 ms to process this data and publish the mesh positions, thus its output and the output of the processing node are the slowest ones. Additionally, the contents of each message of the nodes that have not changed now are indicated using the notation described in Section 4.4, with 𝑝𝐶 𝑉being the positions of all nodes expressed relative to the reference frame of the camera, 𝑥𝑊 𝑉being the state vector obtained with Vision data in world base, and P𝑊 𝑇 𝐶𝑃 is the pose of the TCP in global coordinates (same information as a transform matrix 𝑇, but expressed as position and quaternion). 114 CHAPTER 4. REAL IMPLEMENTATION The new RT Node keeps the same callback function to process the feedback data and update the SOM state. In each iteration, at a constant rate 1/𝑇𝑠, right after checking for new feedback messages, the initial state for the optimizer is extracted from the updated SOM. This state is joined together with the updated horizon for the reference and the previous applied control signal in a new variable of initial parameters, 𝑃0. This is what the RT node publishes with a new custom message to update the initial conditions of the optimization problem in the Opti Node. The optimization node processes this data in a callback function and starts optimizing. While the theoretical rate is also set to 1/𝑇𝑠, the optimization can take longer than that to complete. Once it is done, instead of publishing the first control input to apply, 𝑢(𝑘=0), the full sequence of control inputs from 𝑘=0to 𝑘=𝐻𝑝is saved in a new variable, 𝑈𝐻 𝑝, and published for the RT node to receive. This is done precisely due to the fact that an optimization might take more than a time step to complete, and thus the RT Node might go through multiple iterations before the Opti Node publishes an updated control input. This way, instead of applying the same control input several times until the next one is obtained, which would be no improvement over publishing the new ones only once, we use consecutive predictions in the horizon and keep the backup SOM updated in real time. With just these changes, a problem still persists with the optimizer node. The published control signals allow it to take more than a 𝑇𝑠to complete an optimization if needed, but the published concatenation of control vectors is finite, with 𝐻𝑝steps as an absolute maximum time before a new sequence is required by the RT node, and there is no imposed limit to how long an optimization can take. For extreme cases in the real setup, with noise, a demanding trajectory and a noisy initial state, if not stopped, an optimization can take even seconds to complete, when the maximum prediction time (𝑇𝑠·𝐻𝑝) used is 750 ms, and in the majority of experiments it is around half a second. The simplest solution to avoid the solver getting stuck and taking more time than allowed is adding a maximum optimization time, or a timeout, so that if this time is reached, the solver stops regardless of its situation and progress of finding the optimal solution. Then the optimizer node can look for updated initial conditions, and restart the optimization with these new parameters, hopefully finding a solution before the maximum time. The easiest way of implementing a timeout in the developed code is to add a max_cpu_time termination condition to IPOPT, the solver used in conjunction with CasADi. This time threshold must be lower than the total prediction time, and even lower than half this time to ensure that at least two optimizations have been tried before the end of the horizon. In the end, the maximum time was set to 𝑇𝑠𝐻𝑝/4empirically, as it was enough to complete optimizations regularly, and only produced timeouts in scenarios where the initial parameters were demanding and the optimization would have taken seconds. With new parameters being received every 𝑇𝑠, new ones are always used on retries, and consecutive timeouts were only obtained in situations where the cloth had orientations where the mesh was hard to compute by the Vision node, in very fast movements, and for extreme combinations of low 𝑇𝑠and high 𝐻𝑝, for example 30 steps or more at 10 ms. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 115 As a final remark for this chapter, we must reiterate that the changes needed for a real-time implementation have been left for the final section to be loyal to the order in which the changes were implemented in the real setup, and because, with gathered data of all the steps (no feedback, feedback without real-time and everything included), we can show and compare the results in the following chapter, concretely in Section 5.1, to show the importance of each modification. Additionally, no experimental results using the final real setup have been shown in this chapter to clearly separate its description and implementation (the overall ROS structure, all the nodes and necessary changes to obtain its definitive version) from the experiments carried out in this setup, the results obtained and the analyses derived from them, which deserve their own dedicated chapter. 116 CHAPTER 4. REAL IMPLEMENTATION Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 117 5. Experimental Results Once the full closed-loop control scheme was implemented in ROS to work with a real setup, experiments could be done to test each modification and analyze different alternatives to find the one that yields a better reference tracking. This chapter contains the outcome of these experiments, starting with Section 5.1, where the differences between having a blocking optimizer and working in real time are shown. After the real setup is actually running in real time, a filter was needed for the Vision feedback. The selection of the most suitable one is shown in Section 5.2. The new closed-loop scheme works with multiple values for the control parameters, namely 𝑇𝑠,𝐻𝑝and 𝑊𝑉, that can affect the tracking performance. An analysis to find the best combination is shown in Section 5.3. Finally, the system was tested in demanding situations, with fast trajectories, rotations and disturbances, as seen in Section 5.4. 5.1 Effects of Working in Real Time The implementation in the real setup was progressive, and the changes to work in real time were added after testing the MPC node by itself and after closing the loop with real Vision feedback. Each change was tested, and data was gathered, meaning we can clearly show the effects of each one step by step. The first results, shown in Fig. 5.1, correspond to the MPC node working without real feedback and with the optimizer as a blocking step in the iterations. In the lower left corner, the trajectory at a constant rate is shown, in this case 𝑇𝑠=0.025 s, finishing after 22.5 s. Clearly, the results are extended in time, and in the lower right corner, the reference has been stretched 2.72 times to reach the end of the execution. Figure 5.1: Experimental results without Vision feedback nor running in real time 118 CHAPTER 5. EXPERIMENTAL RESULTS These results were to be expected, as this was implemented with an early version of the MPC without some of the improvements and tuning explained in the previous chapters (specifically the final structure and weights detailed in Section 3.3), and of course each iteration had to wait for the optimizer to finish before proceeding. Additionally, all the debugging commands, timing processes and printing data through terminal, were enabled to check the correct functioning of the system, adding more time per iteration. Besides the extended time, the evolution of the cloth and the tracking results were satisfactory even with just the MPC node working with optimizer and backup SOM, so the implementation continued adding feedback coming from the Vision and Processing nodes, as explained in Section 4.4. Figure 5.2 shows an example of the results obtained in this situation, with a full closed-loop scheme but still without running in real time. Figure 5.2: Experimental results with Vision feedback but not running in real time In this case, and all the experiments carried out in this scenario, we observe another particular behavior. The shown example was done using 𝑇𝑠=0.020 s, so the trajectory should theoretically end after 18 s, as represented in the lower left corner. However, not only does it take much longer (ending around 6.47 times later, at 116.5 s), but in some iterations, the optimizer took times in the order of seconds or tens of seconds to compute the next control signal, halting the execution completely during that time. This is indicated in the lower right corner plot, with two zones without any reference between vertical separators. Even with stretched and paused executions, the tracking itself was still correct, with the overall shape following the reference trajectories when extended in time. These results lead to the changes explained in Section 4.5 to guarantee that, first, the optimizer did not block the overall execution, and second, if the optimizer got stuck due to noisy or demanding new initial states, it would timeout before it was too late and start over with new initial parameters. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 119 In other words, it was empirically proven that explicit changes were needed to run in real time. Once those were in place, new experiments were performed to prove the new setup worked correctly. Fig. 5.3 shows an example of these experiments, where a longer and more demanding trajectory was tested (in fact, it can be seen how it is an extension of the previous one, as the first movements are the same as the full previous trajectory). Figure 5.3: Experimental results with Vision feedback and running in real time In the shown case, the time step was 𝑇𝑠=0.015 s, to prove how even with faster rates than the ones shown before, the system could keep up and finish the execution exactly at the same time as the theoretical end of the reference trajectory, at 18.75 s. We have kept the two different references in the plots to show how the “extended” version is exactly the same as the one following a constant rate. Furthermore, the tracking performance is still unaltered, even with faster rates and trajectories. In fact, checking the main KPI for tracking performance, the resulting RMSE, it is clear to see how once the Vision feedback was implemented, it slightly but consistently increased in all experiments (comparing against stretched and paused references), but with the new real time implementation, it improves again. For the shown experiments, without feedback it had 3.9 cm of RMSE, the second case with feedback but without real time went up to 6.7 cm, and the final version, even with a much more demanding situation, goes down to 5.5 cm. Even if other experiments reached the same levels of error as without adding feedback, this one is shown to demonstrate the improvements allow executing more adverse conditions. In any case, we can see how the output data for these experiments is quite noisy, and does not exactly represent the evolution of the cloth, leading to larger errors. These experiments were all carried out without any kind of filtering on the feedback data, as the filter selection had to be done once the scheme ran in real time. 120 CHAPTER 5. EXPERIMENTAL RESULTS 5.2 Vision Feedback Filter Selection Sending the raw mesh data, with the only processing being an ordering and a change of base, directly to the MPC to update the initial state for the optimizer can lead to it getting stuck and not finding an optimal solution in time, and reaching a timeout. Even when the data is combined with the evolution of the backup SOM using the 𝑊𝑉weight to reduce this effect, and distances over a certain threshold are discarded immediately as spurious signals, after some iterations, it was observed that this effect still persisted in real experiments. This is why a filter was required, even if it was at the cost of some time delay on the feedback signal. The considered alternatives are explained in Subsection 4.4.3. They are all based in the concept of Moving Average using previous data, except the “Future” version, FMA, which waits for the next sample and then applies weights to the previous and next points to get a filtered value for the central one. After some experiments testing different alternatives, it was seen how considering only the previous one or two points had resulted in no significant improvement, while more than four increased the delay between capturing the image and the data reaching the MPC too much. The definition of the EMA is recursive, as seen in (4.4d), which results in a filtered output depending on all values from before in exponentially decreasing weights. To ensure data from too much into the past did not affect the current filtered values, the expression was truncated to the past three points too, resulting in the same filter as the WMA shown in (4.4c) with exponentially decreasing weights 𝑤𝑖. This is why only three filters were compared: FMA, SMA and EMA (truncated, same as WMA), and all of them considered three data points in total to compute the filtered output. Fig. 5.4 shows the evolution of the 𝑋coordinate of the lower right corner of the cloth executing the same trajectory with each one of the considered alternatives. Figure 5.4: Evolution of the 𝑋-axis position of the lower right corner for all considered filters Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 121 First of all, these are different executions, done one after the other, changing the enabled filter in the processing node, and the EMA used 𝛼=0.66 (which, truncated at the three most recent samples, corresponds approximately to weights 0.66, 0.23 and 0.11). In the plots, focused on one coordinate to avoid cluttering the image, but representative of all of them, we can see how the unfiltered evolution is the most noisy, and how each filter reduces the variability, with SMA and EMA being the ones that produce the most noticeable changes, qualitatively, in the shown results. In fact, the evolution in the 𝑋direction of this corner has been chosen specifically because, besides the regular noise present in all coordinates and nodes, we can notice jumps between the actual corner position and another value, always at approximately the same distance from the real position. Investigating the resource of this specific noise, it was found that even without the cloth moving, the Computer Vision algorithms used to detect the cloth and produce a complete mesh sometimes did not detect the lower corners precisely. This resulted in the corner node being placed close to its neighbor on the same row, producing an error mostly in 𝑋. This effect was more severe and common in the right lower corner than in the left one, producing jumps between the two situations shown in Fig. 5.5. This is also the reason behind the persistent offset between the obtained 𝑋of this corner and the reference seen in Fig. 5.3. Figure 5.5: Detail of the found situation where the corners of the Vision mesh are not placed accurately As this problem was originated in the Vision node, solving it on the root is out of the scope of this Thesis, but a way to work around it had to be found. Given the obtained data jumps between the correct value and an offset, and that the cloth is inextensible and the reference trajectories are placed at a constant distance equal to the side length (30 cm in the performed experiments), a good filter can reduce the effects of this phenomenon. Furthermore, if the output data captured by the camera and processed by the Vision node shows an offset between reference and results in only one coordinate of one corner, we know the cloth has a constant length, and thus this difference is either because the cloth is folded, not completely extended, hiding the real corner from the camera, or an error coming from the described source. In the performed experiments, where the cloth piece was held vertically and always extended, it could only be the latter. Even if this data is not completely filtered and reaches the optimizer, it will assume the corner is slightly folded and compute the control signal correctly, just with some added tracking error not present in the actual cloth. Finally, given the Vision node was developed at the IRI, this issue is now known, and the node was updated to also feature SMA and EMA filters before publishing the mesh (in a new topic, cloth_mesh_filtered), which was proven to produce less total delay experimentally. 128 CHAPTER 5. EXPERIMENTAL RESULTS This is just an example with a rotation around the 𝑍𝐸axis of the TCP (which matches the global 𝑍-axis, in this case), but similar ones were executed around the axis perpendicular to the cloth plane with similar results. It is clear how the trajectory is tracked correctly, with the expected noise coming from the real feedback data, especially visible in the 3D plot. In this plot we can also see how some of the harder corners with immediate direction changes affecting all coordinates are also cut more smoothly in the resulting evolution. Having a translation first and a rotation after, we can clearly see how the TCP evolution keeps a constant orientation while moving and then the same position while rotating back and forth. Once again, in the TCP evolutions, we show the pose command sent to the Cartesian controller, P𝑇 𝐶 𝑃 (with both position and orientation) and the actual evolution as given by the messages published on the ROS topic iri_wam_controller/libbarrett_link_tcp. In summary, we can see how the robot tracks trajectories with rotations correctly, even those that set the cloth in positions challenging for the camera to see. Next, we can analyze the effects of changing the time step with regards to speed. The chosen method of execution was to have trajectories without time stamps on them, and simply use the next point on each iteration. This results in the same trajectory being faster or slower depending on 𝑇𝑠. To compare the effects of changing 𝑇𝑠while moving at the same speed, some new trajectories were made, exactly as previous ones, but keeping only every other point, making them twice as fast in practice. As an example, Fig. 5.11 shows a comparison between the original trajectory executed at 𝑇𝑠=10 and 20 ms, and the new fast trajectory at 20 ms. For the latter, two values of 𝐻𝑝were tested, as increasing 𝑇𝑠and keeping 𝐻𝑝the same increases the total prediction time. We can see how the original slower trajectory lasts 12.5 s at 10 ms per step, while it takes double this amount, 25 s, at double the period. Meanwhile, the fast trajectory takes 12.5 s when executed at 20 ms per step. Figure 5.11: Comparison between evolutions of the same cloth corner executing at different speeds Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 129 Executing at 𝑇𝑠=10 ms (Fig. 5.11a), we get several timeouts around 4 s, which results in the robot not receiving updated commands and not moving, followed by a fast movement around the 6 s mark. The timeouts happen because the solver cannot find the optimal solution in the short amount of allowed time. This is solved using 𝑇𝑠=20 ms (Fig. 5.11b), but with the same amount of steps, we now have a trajectory that takes twice as long to complete, meaning not only the solver had more time to finish, but movements were slower, making it easier for the Vision node to capture, process and publish the data. Between these two cases, even with different total prediction times (0.25 vs 0.5 s), the total displacement from the initial state of the prediction until the end of the horizon is kept constant, and so are the displacements per step, as we have the same displacement in the same steps, the only difference is that the steps are longer. If we keep 𝑇𝑠=20 ms and 𝐻𝑝=25 but change to the trajectory using only half the samples, the obtain the plot on the lower right corner (Fig. 5.11d), where we can see that the 𝑌coordinate is behind the reference (a fast movement makes the lower corners “lag” behind the top ones) until it changes direction. Additionally, we can see how some back and forth motions are cut short both in 𝑋and 𝑌near the end. In this case, each optimization problem has the same number of steps into the future as before, but we actually get two times the displacements from before, as half of the reference steps are gone. In other words, this execution has the same speed as the original one, moving along the same trajectory in the same time, but each optimization can see more into the future. To have a completely fair comparison, we can reduce the number of steps in the horizon to match the original total prediction time. That would be cutting 𝐻𝑝in half, but that results in 12.5. We know that very low horizons, in the order of 𝐻𝑝=10, were too short and resulted in larger errors, so we can use 𝐻𝑝=15 for a total prediction time of 0.3 s instead of the original 0.25 s. This is the lower left plot, Fig. 5.11c, where we can see a better tracking, without the original timeouts, slowing the trajectory down, or cutting corners due to seeing too many steps ahead. In fact, this yields the minimum tracking error, at 3.6 cm, compared to the original 4.6 cm. While having trajectories based on points and simply picking the next sample (sliding the window one sample) every step is a valid approach, keeps the same displacements per step, and makes higher 𝑇𝑠more favorable in the real setup, as they make the Vision node relatively faster (instead of 10 times slower, it becomes 5 times slower when going from 𝑇𝑠=10 to 20 ms), it also makes the duration of the trajectory depend on the time between samples, which is usually not desired in real applications, where each position has an associated time to reach it, set from the start. The developed codes assume the references have no time stamps, but for future applications, this can be changed to consider them. For example, a reference parser could keep track of the current time since starting the trajectory and skip samples with stamps prior to the current time, giving the MPC the next sample to consider, keeping its side unaltered. The base for a Reference node was actually created following this idea, but its development was discontinued not long after, as the majority of experiments had already been done with the previous approach, and the effects of speed could be tested with new trajectories created by simply skipping samples of the old ones regularly, as in the shown example. Finally, once the used 𝑇𝑠is known, even trajectories with time stamps can be adapted to the used approach easily, so there was no urgent need of changing it. 130 CHAPTER 5. EXPERIMENTAL RESULTS The final part of this section is dedicated to experiments executed with disturbances and other situations coming from external agents. Concretely, two different cases were studied: blocking the camera, either by walking between it and the WAM robot or by covering it with a solid object, and applying forces to the robot arm itself during execution. The results shown in Fig. 5.12 correspond to the first case, where, during execution, a human walked between the camera and the cloth, covering its view for about two seconds, and then walked back in the opposite way after a while covering the cloth again for approximately another 2 s. Figure 5.12: Cloth corners and TCP evolutions blocking the camera in two instants Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 131 The shaded areas correspond to the instants where the camera was blocked, as indicated in the figure. These times have been obtained thanks to a video recording of the experiment, but can also be double checked with the recorded ROS messages published by the Vision node, which either stop or have strange behaviors. When covering the cloth only partially, the Vision node can still detect its visible part, but the resulting mesh can have incorrect positions on all nodes and coordinates. Usually, depth (global 𝑌) is the most reliable output in these situations, as it sometimes tries to assign an entire mesh to just the visible part, moving the other two coordinates around. This can be clearly seen in the plots, where on the first blocking, around the 5 second mark, a new mesh was published after about a second with none, but it had completely incorrect 𝑋and 𝑍values for all corners. A similar behavior can be observed on the second block, past the 15 second mark, but is only pronounced in the right corners. This sudden change with a high distance is caught by the feedback processing done before updating the state of the linear models, and discarded completely, so even if we see it happen in the raw and even filtered feedback data, it does not affect the controller. A proof of this is that we can see how the evolutions of the TCP are completely smooth even during these moments. Of course, this can only be done thanks to the backup SOM inside the controller, which accurately simulates the real evolution of the cloth during the moments where the Vision feedback cannot provide updated data. Before continuing, Fig. 5.13 shows a picture of one of this situations, extracted from the aforementioned video of the experiment. Figure 5.13: Human walking in the setup, creating a barrier between camera and cloth The second experiment with this kind of disturbances coming from external agents also included an instance of a person walking and momentarily blocking the camera and Vision feedback for a couple of seconds, but it also involved more direct interactions with the WAM. As mentioned in Section 4.3, the used Cartesian controller allows movements in the null space of the WAM without offering much resistance. This space comes from the fact that there are multiple joint configurations that result in the same End Effector pose, meaning the IK problem does not have a unique solution, and thus some joint motions exist such that, combined, do not alter P𝑇 𝐶𝑃. This has multiple applications, but for the current setup and tracking problem, we can also use it as a source of disturbances. 132 CHAPTER 5. EXPERIMENTAL RESULTS For example, during execution, a human can pull and push the elbow joint like shown in Fig. 5.14 with almost no effort, and while theoretically these movements do not change the TCP pose, as the Cartesian controller keeps it in place and only allows a movement in the null space, during a real experiment, the are slight displacements and forces that are transmitted from the arm into the cloth, producing disturbances that are picked up by the camera. Figure 5.14: Human agent pulling and pushing the elbow joint during an execution In fact, the second captured experiment not only had this kind of movements in the null space causing slight disturbances, but after some seconds, the rigid piece connecting the upper corners of the cloth was pressed down on one end, causing a slight rotation and a vibration when released. This situation can be seen in Fig. 5.15. These last two pictures are frames extracted from the same video recording. Figure 5.15: Disturbance created by a person poking the cloth Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 133 With the conditions of this second experiment explained, we can now show its results, which are in Fig. 5.16. The shaded region in magenta corresponds to the interval when the camera was blocked, as in the previous experiment, while the region shaded in gray corresponds to the time when a human agent was interacting with the robot arm. Figure 5.16: Cloth corners and TCP evolutions blocking the camera and with human interaction Blocking the camera has the same effects as before, but now we see different results with the new actions. Moving the elbow of the arm (from around 11 to 15 seconds) produces slight movements in the TCP, which are also visible in the upper corners, and are propagated to have a much greater effect in the lower corners due to the non-rigid nature of the cloth. When the rigid piece between the upper corners is pressed and released around the 15 second mark, we can see how it creates a ripple effect which is very clear in the upper corners, and somewhat mitigated on the lower ones. Even in these situations, however, we can see how the reference trajectory is tracked correctly. 134 CHAPTER 5. EXPERIMENTAL RESULTS All in all, the developed final implementation on the real robot has been proven to work with demanding conditions, in multiple situations, and even under the effects of an external agent creating disturbances while keeping a good trajectory tracking performance, with these last two experiments having RMSEs of 7.1 and 8.0 cm, respectively. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 135 6. Budget and Impact This chapter focuses on the resources used during the course of this Thesis and its effects in society, both in the economical sense, with the budget discussed in Section 6.1, and the ecological sense, with the environmental impact detailed in Section 6.2. 6.1 Project Budget The total budget can be divided in material costs, including hardware, its amortization and software licenses on one side, and personal costs on the other. For the latter, we first have to consider that the present Thesis had a total duration of 7 months, from mid-February to September 2021. It is important to note that while this Thesis was awarded with a “María de Maeztu Unit of Excellence” Research Initiation grant, this only covered four of these months, from March to June, with an endowment of 525 €per month, assuming 20 hours of work per week. However, while this is the real funding obtained, it does not cover the total costs of the project, and not even only the personal costs if we compute the budget as if this Thesis was an independent, private project. The present Thesis corresponds to the final work of a Double Master’s Degree, and therefore has an assigned total of 24 ECTS (European Credit Transfer and Accumulation System) credits. At the UPC, each credit corresponds to 25 to 30 hours of work. While the usual for regular subjects is 25, in Thesis and internships it can be increased to 30 h/credit [56], thus this value will be used for the budget computations, resulting in a total of 720 hours. In addition to the hours dedicated by the student, which will be assumed as a junior engineer with a wage around 20 €/h, it has been estimated that the principal researchers devoted 100 hours to the project during the course of this Thesis, with a unit cost of 38.21 €/h. With regards to material costs, there were no purchases of new equipment needed, as all hardware was already present at the IRI, except the additional personal computer used, which was also not bought specifically for this Thesis. However, the usage of this material carries a depreciation, which must be accounted for in the total budget. The details of amortization periods and values were facilitated by the organization office at the IRI itself. We can start with the WAM robot, which was bought for 120000 €and has an estimated amortization period of 20 years, corresponding to around 500 €/month. However, this robot also needs some periodic maintenance, and historic data show that these tasks plus the regular changes of wires and other parts represent an additional cost around 18 €/month. 136 CHAPTER 6. BUDGET AND IMPACT The robot is connected to a computer at the laboratory, bought for 600 €, which for an estimated lifespan of 10 years, corresponds to 5 €/month. Additionally, for the experiments using the Vision ROS node, an additional computer was used with a better graphics card, with an estimated depreciation of 6€/month. For these experiments, the camera used was a Kinect (Xbox 360 model), bought originally for 150 €. It is expected to last for 5 years, which supposes 2.5 €/month. Finally, a personal HP laptop was used, with Windows 10 as the operating system. It had a price of 1700 €, which for a maximum lifespan of 10 years, means 14.16 €/month of depreciation. Besides the used hardware, the software used to develop this Thesis included Ubuntu 16.04, ROS Kinetic, the CasADi and Eigen libraries, and Latex (Overleaf), all open source and free. For some figures, Figma was used in its free of charge Starter plan. Other additional codes were developed at the IRI, as mentioned and cited through this document, also without any monetary cost. The only licensed software used was Matlab, and its Academic License, which can be used for research projects, costs 250 €/year. Table 6.1 shows the full budget of this final Thesis. For the amortization times, we have considered the actual periods in which each piece of hardware was used. For example, the robot and its PC were used during 5 months (March-July 2021, both included), but the experiments using the camera and requiring an additional PC were all performed in the span of only 3 months (May-July 2021). The final total cost of the project is 21106.50 €. Table 6.1: Complete project budget Concept Units Unit cost Total cost Personnel Student 720 h 20.00 €/h 14400.00 € Principal researchers 100 h 38.21 €/h 3821.00 € Amortization WAM 5 months 500.00 €/month 2500.00 € WAM Maintenance 5 months 18.00 €/month 90.00 € WAM PC 5 months 5.00 €/month 25.00 € Vision PC 3 months 6.00 €/month 18.00 € Kinect Camera 3 months 2.50 €/month 7.50 € Personal Laptop 7 months 14.16 €/month 99.17 € Software Matlab Academic License 7 months 20.83 €/month 145.83 € TOTAL 21106.50 € Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 137 6.2 Environmental Impact Developing a project always has impacts on the environment, either positive or negative, which must be studied in detail. Due to the nature of the research project developed during this Thesis, however, the total ecological impact is minimal, as no directly pollutant activity was carried out. The used WAM robot, even if not purchased specifically for this Thesis, was built with Aluminum and PVC, which are both recyclable once the robot ends its useful life. The only negative impacts on the environment are derived from the energy of these recycling processes. This robot, and all the hardware used for this project, complies with the Restriction of Hazardous Substances (RoHS) regulation of the European Union [57], which forbids usage of certain elements, and especially Lead, in electric and electronic devices. A negative impact on the environment can be derived from the consumption of electrical energy and the CO2emissions necessary to obtain it. We can assume that at least one computer was turned on during the time dedicated to the project, which is a total of 720 hours. The power rating of the personal laptop is 120 W [58]. With this value, we obtain a total consumption of 86.4 kWh during the course of this Thesis. Additionally, the WAM operates at least at 60 W [59], and was used around 400 h (5 months, but only 20 h/week), which results in 24 kWh consumed. The total electrical energy consumption rises to a total of 110.4 kWh. According to the official data published by the Spanish Government [60], a total of 0.357 kg of CO2are emitted to the atmosphere for each consumed kWh. With this factor, we can obtain that the total CO2emissions caused by the energy used during this project are of 39.41 kg CO2. The specific positive impacts of this Thesis are clearly defined with the possible applications of the developed controller, mainly in assistive robotics, as it has been developed under the CLOTHILDE project at the IRI [1]. For example, this control scheme could be used to aid a person with reduced mobility in daily tasks like getting dressed, folding clothes and putting them in the closet, prepare tablecloths, etc. While this can be a considerable impact in society and the lives of multiple people, the positive effects on the environment from an ecological point of view are more limited, and would depend on the energy and emissions saved using the developed scheme instead of another alternative on each specific application. 144 ACKNOWLEDGEMENTS Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 145 Bibliography [1] I. de Robòtica i Informàtica Industrial, “CLOTHILDE: Cloth manipulation learning from demonstration.” https://clothilde.iri.upc.edu/, 2018. [Online; accessed July 2021]. [2] I. de Robòtica i Informàtica Industrial, “MP-CloL: Learning Robotic Cloth Manipulation based on Physics Models and Model Predictive Control.” https://www.iri.upc.edu/project/show/ 245, 2020. [Online; accessed July 2021]. [3] J. Rawlings, D. Mayne, and M. Diehl, Model Predictive Control: Theory, Computation, and Design. Nob Hill Publishing, 2017. [4] A. Colomé and C. Torras, “Dimensionality reduction for dynamic movement primitives and application to bimanual manipulation of clothes,” IEEE Transactions on Robotics, vol. 34, no. 3, pp. 602–615, 2018. [5] C. E. Garcia, D. M. Prett, and M. Morari, “Model Predictive Control: Theory and practice—A survey,” Automatica, vol. 25, no. 3, pp. 335–348, 1989. [6] E. F. Camacho and C. B. Alba, Model predictive control. Springer science & business media, 2013. [7] P. Tøndel, T. A. Johansen, and A. Bemporad, “An algorithm for multi-parametric quadratic programming and explicit mpc solutions,” Automatica, vol. 39, no. 3, pp. 489–497, 2003. [8] J. M. Maciejowski, Predictive control with constraints. Pearson education, 2002. [9] R. Findeisen, L. Imsland, F. Allgower, and B. A. Foss, “State and output feedback nonlinear model predictive control: An overview,” European journal of control, vol. 9, no. 2-3, pp. 190–206, 2003. [10] G. Goodwin, M. M. Seron, and J. A. De Doná, Constrained control and estimation: an optimisation approach. Springer Science & Business Media, 2006. [11] D. A. Allan and J. B. Rawlings, “Moving horizon estimation,” in Handbook of Model Predictive Control, pp. 99–124, Springer, 2019. [12] J. Hedengren and T. Edgar, “Moving horizon estimation-the explicit solution,” in Proceedings of chemical process control (CPC) VII conference. Lake Louise, Alberta, Canada, 2006. [13] X. Man and C. Swan, “A mathematical modeling framework for analysis of functional clothing,” Journal of Engineered Fibers and Fabrics, vol. 2, 11 2007. [14] F. Coltraro, “Experimental validation of an inextensible cloth model.” Technical Report, 2020. 146 BIBLIOGRAPHY [15] F. Coltraro, J. Amorós, M. Alberich-Carramiñana, and C. Torras, “An Inextensible Model for Robotic Simulations of Textiles.” https://arxiv.org/abs/2103.09586, 2021. [16] D. Parent, A. Colomé, C. Ocampo-Martinez, and C. Torras, “Robotic Cloth Manipulation: Using Model Predictive Control for Cloth State Tracking.” Technical Report, 2021. [17] D. Baraff and A. Witkin, “Large steps in cloth simulation,” in Proceedings of the 25th Annual Conference on Computer Graphics and Interactive Techniques, SIGGRAPH ’98, p. 43–54, Association for Computing Machinery, 1998. [18] Y. Bai, W. Yu, and C. K. Liu, “Dexterous manipulation of cloth,” Computer Graphics Forum, vol. 35, no. 2, pp. 523–532, 2016. [19] Z. Erickson, H. M. Clever, G. Turk, C. K. Liu, and C. C. Kemp, “Deep haptic model predictive control for robot-assisted dressing,” in 2018 IEEE international conference on robotics and automation (ICRA), pp. 4437–4444, IEEE, 2018. [20] The MathWorks Inc., MATLAB Version 9.7.0 (R2019b). 2019. [21] J. A. E. Andersson, J. Gillis, G. Horn, J. B. Rawlings, and M. Diehl, “CasADi – A software framework for nonlinear optimization and optimal control,” Mathematical Programming Computation, In Press, 2018. [22] A. Wächter and L. T. Biegler, “On the implementation of an interior-point filter line-search algorithm for large-scale nonlinear programming,” Mathematical programming, vol. 106, no. 1, pp. 25–57, 2006. [23] S. Boyd, S. P. Boyd, and L. Vandenberghe, Convex optimization. Cambridge university press, 2004. [24] S. Sarabandi and F. Thomas, “Accurate computation of quaternions from rotation matrices,” in International Symposium on Advances in Robot Kinematics, pp. 39–46, Springer, 2018. [25] J. M. Maciejowski, Predictive control: with constraints. Pearson education, 2002. [26] L. Magni, D. Pala, and R. Scattolini, “Stochastic model predictive control of constrained linear systems with additive uncertainty,” in 2009 European Control Conference (ECC), pp. 2235–2240, IEEE, 2009. [27] D. Mayne, “Robust and stochastic model predictive control: Are we going in the right direction?,” Annual Reviews in Control, vol. 41, pp. 184–192, 2016. [28] P. Velarde, L. Valverde, J. M. Maestre, C. Ocampo-Martínez, and C. Bordons, “On the comparison of stochastic model predictive control strategies applied to a hydrogen-based microgrid,” Journal of Power Sources, vol. 343, pp. 161–173, 2017. Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 147 [29] J. M. Grosso, J. M. Maestre, C. Ocampo-Martinez, and V. Puig, “On the assessment of tree-based and chance-constrained predictive control approaches applied to drinking water networks,” IFAC Proceedings Volumes, vol. 47, no. 3, pp. 6240–6245, 2014. [30] D. C. Montgomery and G. C. Runger, Applied statistics and probability for engineers. 2010. [31] P. I. Corke, Robotics, Vision & Control: Fundamental Algorithms in MATLAB. Springer, second ed., 2017. ISBN 978-3-319-54413-7. [32] R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction. MIT press, 2018. [33] M. A. Wiering and M. Van Otterlo, “Reinforcement learning,” Adaptation, learning, and optimization, vol. 12, no. 3, 2012. [34] L. P. Kaelbling, M. L. Littman, and A. W. Moore, “Reinforcement learning: A survey,” Journal of artificial intelligence research, vol. 4, pp. 237–285, 1996. [35] S. Schaal et al., “Learning from demonstration,” Advances in neural information processing systems, pp. 1040–1046, 1997. [36] J. Kober and J. Peters, “Policy search for motor primitives in robotics,” Machine learning, vol. 84, no. 1-2, pp. 171–203, 2011. [37] M. P. Deisenroth, G. Neumann, J. Peters, et al., “A survey on policy search for robotics,” Foundations and trends in Robotics, vol. 2, no. 1-2, pp. 388–403, 2013. [38] J. Peters, K. Mulling, and Y. Altun, “Relative entropy policy search,” in Twenty-Fourth AAAI Conference on Artificial Intelligence, 2010. [39] A. Colomé, G. Neumann, J. Peters, and C. Torras, “Dimensionality reduction for probabilistic movement primitives,” in IEEE-RAS International Conference on Humanoid Robots, 2014. [40] A. Colomé and C. Torras, “Dual reps: A generalization of relative entropy policy search exploiting bad experiences,” IEEE Transactions on Robotics, vol. 33, no. 4, pp. 978–985, 2017. [41] D. Görges, “Relations between model predictive control and reinforcement learning,” IFACPapersOnLine, vol. 50, no. 1, pp. 4920–4928, 2017. [42] M. A. Bermeo-Ayerbe, C. Ocampo-Martínez, and J. Diaz-Rozo, “Adaptive predictive control for peripheral equipment management to enhance energy efficiency in smart manufacturing systems,” Journal of Cleaner Production, vol. 291, p. 125556, 2021. [43] J. M. Grosso, C. Ocampo-Martínez, and V. Puig, “Learning-based tuning of supervisory model predictive control for drinking water networks,” Engineering Applications of Artificial Intelligence, vol. 26, no. 7, pp. 1741–1750, 2013. 148 BIBLIOGRAPHY [44] N. Karnchanachari, M. I. Valls, D. Hoeller, and M. Hutter, “Practical reinforcement learning for mpc: Learning from sparse objectives in under an hour on a real robot,” in Learning for Dynamics and Control, pp. 211–224, PMLR, 2020. [45] Stanford Artificial Intelligence Laboratory, “Robot Operating System - Kinetic Kame.” https: //wiki.ros.org/kinetic. [Online; accessed July 2021]. [46] G. Guennebaud, B. Jacob, et al., “Eigen v3.” http://eigen.tuxfamily.org, 2010. [Online; accessed July 2021]. [47] J. A. E. Andersson, J. Gillis, G. Horn, J. B. Rawlings, and M. Diehl, “CasADi Documentation.” https://web.casadi.org/docs/. [Online; accessed July 2021]. [48] J. Denavit and R. S. Hartenberg, “A kinematic notation for lower-pair mechanisms based on matrices,” 1955. [49] Barrett Technology, Inc., “Barrett WAM 7-DOF Joints & Frames.” https://web.barrett.com/ files/B2576_RevAC-00.pdf, 2005. [Online; accessed July 2021]. [50] S. G. Khan, G. Herrmann, M. Al Grafi, T. Pipe, and C. Melhuish, “Compliance control and human– robot interaction: Part 1—survey,” International journal of humanoid robotics, vol. 11, no. 03, p. 1430001, 2014. [51] S. Harding and J. F. Miller, “Evolution of robot controller using cartesian genetic programming,” in European Conference on Genetic Programming, pp. 62–73, Springer, 2005. [52] D. Parent Alonso, “Control d’impedància variable amb primitives de moviment en espais latents amb comportament dòcil,” B.S. thesis, Universitat Politècnica de Catalunya, 2019. [53] D. Parent, A. Colomé, and C. Torras, “Variable impedance control in cartesian latent space while avoiding obstacles in null space,” in 2020 IEEE International Conference on Robotics and Automation (ICRA), pp. 9888–9894, IEEE, 2020. [54] M. Arduengo, C. X. Zheng, A. Colomé, and C. Torras, “Cloth Point Cloud Segmentation.” https: //github.com/MiguelARD/cloth_point_cloud_segmentation, 2021. [55] S. W. Smith, “Chapter 15 - moving average filters,” in Digital Signal Processing (S. W. Smith, ed.), pp. 277–284, Boston: Newnes, 2003. [56] U. P. de Catalunya, “El model docent: Què és un crèdit ECTS?.” https://www.upc.edu/ca/ graus/faqs/model-docent, 2021. [Online; accessed September 2021]. [57] European Parliament, “DIRECTIVA 2002/95/CE del Parlamento Europeo y del Consejo de 27 de enero de 2003 sobre restricciones a la utilización de determinadas sustancias peligrosas en aparatos Robotic Cloth Manipulation: Real Implementation with MPC and RL - Memory 149 eléctricos y electrónicos.” https://www.boe.es/doue/2003/037/L00019-00023.pdf, 2003. [Online; accessed September 2021]. [58] Hewlett-Packard, Inc., “Workstation ZBook Power G7 Datasheet.” https://www8.hp.com/ h20195/V2/GetPDF.aspx/c06908501, 2021. [Online; accessed July 2021]. [59] Barrett Technology, Inc., “WAM Arm Datasheet.” https://web.barrett.com/files/ WAMDataSheet_02.2011.pdf, 2011. [Online; accessed September 2021]. [60] Gobierno de España, “Factores de emisión de CO2 y coeficientes de paso a energía primaria de diferentes fuentes de enería final consumidas en el sector de edificios en España.” https://energia.gob.es/desarrollo/EficienciaEnergetica/RITE/Reconocidos/ Reconocidos/Otros%20documentos/Factores_emision_CO2.pdf, 2016. [Online; accessed September 2021].