scieee AI-readable full text Open interactive document viewer

Distributed Deep Reinforcement Learning in an HPC system and deployment to the Cloud

Escobar Castells, Miquel

Abstract

Combinar l'aprenentatge per reforç amb l'aprenentatge profund és, a dia d'avui, un dels reptes més grans en el sector d'investigació en intel·ligència artificial. Escalar aquest tipus d'aplicacions mitjançant supercomputadors o serveis al núvol és crucial per avançar en l'ús massiu d'aquestes tecnologies. L'objectiu d'aquest projecte és desenvolupar i testejar vàries implementacions d'algorismes de Deep Reinforcement Learning (DRL) sobre el cas d'estudi seleccionat, utilitzant el clúster CTE-POWER del Barcelona Supercomputing Center (BSC), així com fer un desplegament al núvol dels entrenaments per analitzar els seus costos i viabilitat en un entorn de producció.

Full text

Bachelor’s Degree in Data Science and Engineering Final Thesis Report Thesis defense: July 1, 2021 Distributed Deep Reinforcement Learning in an HPC system and deployment to the Cloud Author :Miquel Escobar Castells Director :Jordi Torres Vi˜ nals Co-director :Ra´ ul Garc´ ıa Fuentes June 20, 2021 Abstract Combining Reinforcement Learning and Deep Learning is the most challenging Artificial Intelligence research and development area at present. Scaling these types of applications in an HPC infrastructure available in the cloud will be crucial for advancing the massive use of these technologies. The purpose of this project is to develop and test various implementations of Deep RL algorithms given a selected case study, using Barcelona Supercomputing Center’s (BSC) CTE-POWER cluster, and make a deployment to the cloud of the training pipeline to analyze its costs and viability in a production environment. Keywords: Reinforcement Learning, RL, Deep Reinforcement Learning, DRL, HPC, cloud, Ray, AWS, SageMaker. Resumen Combinar el aprendizaje por refuerzo con el aprendizaje profundo es, a d´ıa de hoy, el desaf´ıo m´as grande en el sector de desarrollo e investigaci´on en inteligencia artificial. Escalar este tipo de aplicaciones mediante supercomputadores o servicios en la nube es crucial para avanzar en el uso masivo de estas tecnolog´ıas. El objetivo de este proyecto es desarrollar y testear varias implementaciones de algoritmos de Deep Reinforcement Learning (DRL) sobre el caso de estudio seleccionado, usando el cluster CTE-POWER del Barcelona Supercomputing Center (BSC), as´ı como hacer un despliegue a la nube de los entrenamientos para analizar sus costes y viabiliadad en un entorno de producci´on. Palabras clave: aaprendizaje por refuerzo, RL, aprendizaje por refuerzo profundo, DRL, HPC, cloud, Ray, AWS, SageMaker. Resum Combinar l’aprenentatge per refor¸c amb l’aprenentatge profund ´es, a dia d’avui, un dels reptes m´es grans en el sector d’investigaci´o en intel·lig`encia artificial. Escalar aquest tipus d’aplicacions mitjan¸cant supercomputadors o serveis al n´uvol ´es crucial per avan¸car en l’´us massiu d’aquestes tecnologies. L’objectiu d’aquest projecte ´es desenvolupar i testejar v`aries implementacions d’algorismes de Deep Reinforcement Learning (DRL) sobre el cas d’estudi seleccionat, utilitzant el cl´uster CTE-POWER del Barcelona Supercomputing Center (BSC), aix´ı com fer un desplegament al n´uvol dels entrenaments per analitzar els seus costos i viabilitat en un entorn de producci´o. Paraules clau: aprenentatge per refor¸c, RL, aprenentatge per refor¸c profund, DRL, HPC, cloud, Ray, AWS, SageMaker. 1 Acknowledgements I would first like to thank my director, Jordi Torres, for the proposal of this project and for giving me the opportunity to join such an incredible research group, as well as for your guidance all the way through. I would also like to express my most sincere gratitude to my colleagues at the Emerging Technologies for Artficial Intelligence group: to Ra´ul, for your direct contributions and support during my time in the group; to Juan Luis and Oriol, who provided feedback and counsel along the way. I would also like to thank my parents, my grandmother, my brother and sister for your unconditional support and wise counsel. Finally, I would like to end by thanking my friends Jordi A., V´ıctor, Pau, Pol, Jacobo, Lucas, Francesc, Jordi C., Arnau, and Luc´ıa, for their encouragement and support all through my studies. 2 Contents Abstract 1 Acknowledgements 2 1 Introduction 6 1.1 Context ......................................... 6 1.2 Background ....................................... 6 1.3 Justification ....................................... 6 1.4 Objectives ........................................ 7 1.5 Stakeholders ....................................... 7 2 Project planning 8 2.1 Methodology ...................................... 8 2.1.1 Management .................................. 8 2.1.2 Development .................................. 8 2.1.3 Documentation ................................. 8 2.2 Risks ........................................... 8 2.3 Work packages ..................................... 9 2.4 Milestones ........................................ 10 2.5 Gantt diagram ..................................... 12 3 Problem definition 13 3.1 Selection of the software framework .......................... 13 3.2 Selection of the cloud services provider ........................ 13 3.3 Selection of the environment .............................. 14 4 Previous work 17 4.1 Reinforcement Learning ................................ 17 4.1.1 Criterion of optimality. ............................ 18 4.1.2 Exploration vs exploitation .......................... 20 4.1.3 Off-policy vs on-policy ............................. 20 4.1.4 Model-based vs model-free ........................... 21 4.1.5 Deep Reinforcement Learning ......................... 21 4.1.6 Algorithms ................................... 22 4.1.6.1 Deep Q-Network (DQN) ...................... 22 4.1.6.2 Deep Deterministic Policy Gradient (DDPG) ........... 22 4.1.6.3 Twin Delayed DDPG (TD3) .................... 23 4.1.6.4 Soft Actor Critic (SAC) ....................... 23 4.1.6.5 Proximal Policy Optimization (PPO) ............... 24 4.2 Software frameworks .................................. 24 4.2.1 Ray ....................................... 24 4.2.1.1 RLlib ................................. 25 4.2.1.2 Tune ................................. 25 4.2.2 Tensorflow ................................... 25 4.2.3 PyTorch ..................................... 25 4.2.4 Docker ...................................... 25 4.3 CTE-POWER Cluster ................................. 26 4.4 Amazon Web Services ................................. 26 3 4.4.1 Amazon SageMaker .............................. 27 4.4.2 Amazon Simple Storage Service (S3) ..................... 27 4.4.3 Amazon Elastic Computing (EC2) ...................... 28 4.4.4 Elastic Container Registry (ECR) ...................... 28 4.4.5 Amazon CloudWatch ............................. 28 5 Implementation 30 5.1 Training in RLlib .................................... 31 5.2 Implementation in CTE-POWER Cluster ...................... 32 5.3 Implementation in Amazon Web Services ...................... 33 6 Evaluation 35 6.1 Training results ..................................... 35 6.1.1 TD3 ....................................... 35 6.1.2 SAC ....................................... 36 6.1.3 PPO ....................................... 37 6.2 Execution time ..................................... 38 6.2.1 TD3 ....................................... 39 6.2.2 SAC ....................................... 40 6.2.3 PPO ....................................... 40 6.3 Examples ........................................ 42 7 Conclusions 44 7.1 Personal conclusions .................................. 44 7.2 Objectives fulfilment .................................. 44 7.3 Future work ....................................... 44 References 46 A Environment specificatios 49 A.1 Observation space ................................... 49 A.2 Action space ...................................... 50 4 List of Figures 1 The sample of Unity’s environments rendered. Source: Unity Technologies . . . . 15 2 A sample of PyBullet’s environments rendered. Source: PyBullet ......... 15 3 Renderization of the humanoid environment at a random time step. ....... 16 4 Another renderization of the humanoid environment at a random time step. . . . 16 5 The Reinforcement Learning cycle. The agent performs an action, and the environment returns a reward together with the new state. ............... 18 6 High level abstraction of a Docker application life cycle. Source: Docker documentation ........................................ 26 7 Cloud market share, as of the end of 2020. ...................... 27 8 Architecture of a Ray training process, for TD3, SAC and PPO algorithms. . . . 31 9 Architecture of distributed computing in Ray. .................... 32 10 The architecture of the solution implementation in CTE-POWER cluster. . . . . 32 11 The architecture of the solution implementation in AWS SageMaker. ....... 34 12 TD3 algorithm average reward per time step for the best performing configuration. 36 13 SAC algorithm average reward per time step for the best performing configuration. 37 14 PPO algorithm average reward per time step for the best performing configuration. 38 15 Training iteration time by number of workers (TD3). ................ 39 16 Training iteration time by number of workers, with GPU (TD3). ......... 39 17 Training iteration time by number of workers (SAC). ................ 40 18 Training iteration time by number of workers, with GPU (SAC). ......... 40 19 Training iteration time by number of workers (PPO). ................ 41 20 Training iteration time by number of workers, with GPU (PPO). ......... 41 21 PPO training iteration time by task. ......................... 42 22 PPO training iteration time by task, with GPU. .................. 42 23 Sequence in which the agent becomes not alive in the last frame. ......... 42 24 Sequence in which the agent uses the arms to maintain equilibrium. ....... 42 25 Sequence in which the agent compensates its inertia by bending the head back. . 43 List of Tables 1 Detailed work packages and tasks. .......................... 10 2 Detailed milestones of the project. .......................... 11 3 Hyperparameters with best results for TD3. ..................... 35 4 Hyperparameters with best results for SAC. ..................... 37 5 Hyperparameters with best results for PPO. ..................... 38 6 Iteration time and speedup per resource configuration (TD3). ........... 39 7 Iteration time and speedup per resource configuration (SAC). ........... 40 8 Iteration time and speedup per resource configuration (PPO). ........... 41 9 Properties of all dimensions in the observation space. ............... 50 10 Properties of all dimensions in the action space. .................. 51 5 1 Introduction 1.1 Context The project is carried out at Barcelona Supercomputing Center (BSC-CNS), as part of the research in Reinforcement Learning conducted in the Emerging Technologies for Artificial Intelligence group. The investigation group has been focusing the last few years in the field of Deep Learning, and is now shifting into the study and application of the thriving area of Reinforcement Learning. In fact, as a direct result of the new line of research, the group manager recently published a book on the topic [38], which has been profoundly utilized for the development of this project. The purpose of this project is to use BSC’s CTE-Power cluster to develop and test various implementations of Deep Reinforcement Learning algorithms given a case study, and posteriorly perform a deployment to the cloud of the obtained models to analyze its costs and viability in a production environment. 1.2 Background In order to allow determining the originality and scope of the contributions made by the project author, the origin of the main ideas is described in this section. The project idea comes from the general line of research conducted in BSC’s Emerging Technologies for Artificial Intelligence investigation department, which is focused on state-ofthe-art Reinforcement Learning techniques, specifically in the scope of Deep Reinforcement Learning. Thus, the project initial ideas were mostly proposed by the group manager and director of this thesis. The project in itself starts from scratch, even though the knowledge that the group has already acquired in the field of Reinforcement Learning, both in terms of theoretical and mathematical background, as well as in relation to various software frameworks, will take part in the development of the project. 1.3 Justification Reinforcement Learning is a domain in Artificial Intelligence which has been gaining popularity among multiple fields, all the way from robotics to Natural Language Processing or self-driving cars. In fact, Deepmind [35] contemplates the possibility that all intelligence can always be represented as the maximization of reward, and that Reinforcement Learning is enough to reach a more human-like general AI. Even so, in Reinforcement Learning it can be difficult to fit some problems and their data to the corresponding environment and agent/s, as these are extremely data-hungry, sensitive to hyperparameter tuning and relatively unstable. If not carefully managed, this can easily lead to overcomplicating simple problems. Consequently, one of the major challenges that Deep Reinforcement Learning encounters is computation costs and training time, as very complex problems can require enormous amounts of iterations, executions and hyperparameter tuning. That is why, in order to tackle this, High 6 Performance Computing and cloud services come into play, as they can help reduce the computer costs and the training time of the more complex algorithms. 1.4 Objectives The main goals of the project are listed below. •Successfully train and test Deep Reinforcement Learning algorithms for the selected case study (also referred as environment). •Implement and train the chosen algorithms in an HPC system (the CTE-Power cluster), mainly focusing on distributed and parallel computing. •Perform a benchmarking in terms of results and training time in the CTE-Power cluster, comparing and analyzing the use of CPU, GPU and distributed computing. •Develop and deploy a training pipeline using the chosen cloud service provider, and given the Reinforcement Learning environment. 1.5 Stakeholders The main beneficiary of this work is Barcelona Supercomputing Center, more specifically the Emerging Technologies for Artificial Intelligence research group [4]. A secondary beneficiary is the Reinforcement Learning community, given that the code and results obtained from the work are public and available to the scientific community, containing clear implementations for different training methodologies in a supercomputer, as well as a complete pipeline for training and inference in the cloud. 7 three of them were considered, but the final selection was Amazon Web Services, based on the following grounds: •A robust service for machine learning in AWS Sagemaker with Reinforcement Learning specific implementations. •Its support and documentation for Ray. •The knowledge and experience acquired by the research group using AWS in previous projects. •A free-tier that provides some of its services for free up to a certain limit. Furthermore, an AWS credit of 300 USD was received from Amazon for the development of the project. 3.3 Selection of the environment Before implementing the solution, the problem must be carefully and detailedly defined. Since the goal of this project is to successfully train DRL algorithms both in a supercomputing environment (see 4.3) and a cloud service (see 4.4), and provide a solution easily scalable and reproducible given other problems, the case study or environment must be selected taking into consideration various factors. One of them is the degree of difficulty: it should be a complex enough environment such that only some algorithms and hyperparameter configurations are able to solve it, but at the same time already proven solvable in a reasonable amount of time (due to both the project’s time limitations and the intention of performing numerous trainings). Other factors to consider are the use it has in the scientific community, the quality of the implementation, the viability of its software dependencies and the compatibility with the selected frameworks. Provided that we are using Python and the Open AI Gym toolkit [3] for the RL environments, the selected environment must be compatible with it. There exist multiple open source environments implementations that are compatible with Open AI Gym, even the toolkit itself has some built-in environments, which are the first that will be explored. Apart from this, other toolkits have also been explored and tested. All of them are listed below, with a description of their characteristics. Open AI Gym. The built-in environments in Open AI Gym go all the way from Atari games, to simulated robotics or continuous control tasks. Unfortunately, the better suited environments for this project were the continuous control tasks running on Mujoco physics simulator, which only allows a free trial of 30 days before having to pay for a license. Unity Machine Learning Agents Toolkit (ML-Agents). The second explored toolkit was the Unity Machine Learning Agents [17]. It is an open source project with defined games and simulations using Unity’s software, and also has a wrapper that converts the environments to gym (the format required by Open AI Gym). The quality of the implementations and the difficulty and variety of the environments is great, and the fact that Unity is used in a huge amount of games and applications makes it a great fit. But on the downside, the execution of the environment simulation must be done with either the Unity Hub app running or by transforming the environment into a binary file (that must be OS specific) that executes it. Unfortunately, 14 the former is not feasible when using servers and the latter depends on the operating system, which rises a lot of compatibility issues. Figure 1: The sample of Unity’s environments rendered. Source: Unity Technologies PyBullet Gymperium. [11] This toolkit tackles the problem mentioned in Open AI Gym, providing an open source and free implementations of the continuous control Mujoco environments using the physics simulation Bullet. On the downside, it is not an official PyBullet [9] repository, and the provided renderization functions are very limited resulting in quite poor images. PyBullet Physics SDK. This toolkit is very similar to PyBullet Gymperium, but it actually is the official Python SDK for the physics simulator Bullet. In it, we can find open source and free implementations, among which we find some of the continuous control Mujoco environments. Given that the renderization functions are much better, this toolkit was the final selection. Figure 2: A sample of PyBullet’s environments rendered. Source: PyBullet Taking into consideration all the points mentioned above, the final selection was PyBullet’s Physics SDK HumanoidBulletEnv-v0 environment implentation, which is shown in Figures 3and 4. It is one of the most complex out of Mujoco’s continuous control environments. It is also broadly used in the scientific community together with all the other Mujoco environments, mainly for benchmarking purposes. The agent consists of a 3D humanoid figure composed by 13 links and 17 joints, and the goal of the agent is to advance in the x dimensions as fast as possible, without falling. The 15 action and observation spaces are described below. Figure 3: Renderization of the humanoid environment at a random time step. Figure 4: Another renderization of the humanoid environment at a random time step. Observation space. The observation space is formed by 44 variables: 3 of them correspond to the orientation of the agent for each spatial dimension, 3 other to the velocity of the torso, 2 are the Euler angles for the x and y dimension (also known as roll and pitch), 17 for the relative x positions for each of the joints, another 17 for the velocities on the x dimension for each of the joints, and 2 variables indicating whether or not each feet is touching the floor. See Table 9in appendix A.1 for more details on the observation space variables. Action space. The action space is composed of 17 motors, one for each joint in the humanoid figure. See Table 10 in appendix A.2 for more details on the agent actuators. Reward function. The reward function must be defined according to the environment goal, which is to advance in the x direction. It does also take accountability for states in which the robot falls (it is considered not alive), the energy cost of the performed actions, joints getting stuck between each other are and whether or not the feet are colliding with any other body part. The reward function is computed given the current and previous states as shown in Equation 1. R(st, st−1) = alive bonus(st) + progress(st, st−1) + energy cost(st) + joints at limit cost(st) + feet collision cost(st)(1) The alive bonus(st) function returns −1 when the state of the environment indicates that there is no physical possibility that the robot does not fall, and 1 otherwise. The progress(st, st−1) function measures the speed of the robot, as target dist(st)−target dist(st−1) time . The energy cost(st) is a penalty that measures the amount of energy the robot uses, which is measured by adding the powers used at all the joints: this aims to reduce noise and unnecessary movement. Finally, the feet collision cost(st) returns −1 when at least one of the two feet is colliding with another joint, link or object, and 1 otherwise. 16 4 Previous work 4.1 Reinforcement Learning Reinforcement Learning is a domain in artificial intelligence, and more specifically, in machine learning, that seeks to train a learner (the agent) to achieve a defined goal (the reward) by performing a series of actions within a defined environment. At any given time step tthe agent selects an action atgiven an observation of the environment’s current state st, in order to maximize the cumulative reward rt, which is returned by the environment. Intuitively, the main idea behind Reinforcement Learning is to let the agent acquire knowledge and learn from experience, that is, by allowing it to interact with the environment through its available actions, and obtaining a reward which is directly related that the goal that wants to be achieved. The core concepts of Reinforcement Learning are the following: Agent. The agent Ais the solution to the problem. It is a function that learns to take the optimal actions given the environment’s state. Environment. The environment Eis the representation of the problem the agent tries to solve. It is formed by a set of states S, known as state space, which can be infinite. It must have a transition function defined, which given an action atperformed by the agent and the current state streturns the next state st+1, and a reward function, which maps the action to the reward rt;. State. At any given time step tthe environment is represented by a set of variables. A state is each possible combination of values for this set of variables. Observation. An observation is the information of a state stthat the agent can observe and has knowledge of. From this point onwards, we will assume that the observation is equal to the state (that is, the agent can observe all the available information) for simplicity purposes. Action. The agent can select an action aout of the environment’s action space A(s) for each specific state. It is represented as a set of variables, which can be discrete or continuous. The action taken by the agent affects the environment, which updates its state st+1 and returns a reward rt. Reward. The reward ris the feedback given by the environment to the agent when the latter performs an action. It is computed by the reward function R(s, a), which takes as an input the state and the action. The agent seeks to maximize this value. The definition of the reward function can be complex, and must return high values for positive behaviour and lower values for negative behaviour. We must also take into consideration the concept of delayed reward, that is, when we only can know if the taken action is correct in a future time step. Policy. The policy πis the criterion or method that the agent utilizes in order to select the optimal action (the action that maximizes the reward) given the environment’s state. It can be either deterministic π(s) = aor stochastic π(a|s) = P(at=a|st=s). 17 Figure 5: The Reinforcement Learning cycle. The agent performs an action, and the environment returns a reward together with the new state. Mathematically speaking, a lot of environments can be represented as a Markov Decision Process (MDP): it is formed by a finite set of states Sand the transition probabilities P(st+1|st, a) = P(s0|s, a) that depend upon the actions A(s) the agent can take. Thereafter, the agent’s task lies on estimating the expected cumulative reward Gof its actions in order to select the one that maximizes it. 4.1.1 Criterion of optimality. There are three main models of optimal behaviour, that is, the function we are looking to optimize in order to maximize the notion of cumulative reward, also known as return, which basically establishes how far into the future the policy must take into account when selecting the present action. A common approach is making use of a discount factor γ, to discount the weight of future rewards with time, as shown in Equation 2. Gt0=E"t=∞ X t=t0 γtrt#(2) Another usual approach is the finite-horizon, which only considers the future hsteps, all with the same weight, as shown in Equation 3. Gt0=E"t=t0+h X t=t0 rt#(3) An alternative is the average-reward model [https://arxiv.org/abs/2010.08920], which considers all future rewards (infinite horizon) by taking the average, as shown in Equation 4. Gt0= lim h→∞ E"1 h t=t0+h X t=t0 rt#(4) State-value function. Given a policy we can define the state-value function Vπ(s) as the expected reward given a state, as shown in Equation 5. Vπ(s) = Eπ[Gt|St=s] (5) 18 Action-value function. And the action-value function Qπ(s, a), also known as Q-function, as the expected reward of performing an action agiven a state sand a policy π, as shown in Equation 6. As explained in the algorithms section, the Q-function can be customized to represent a modification of the obtained reward depending on the algorithm, for instance with the intention of encouraging exploration of the agent. Qπ(s, a) = Eπ[Gt|St=s, At=a] (6) Bellman equations. Assuming an infinite-horizon with discount factor γmodel, from Equation 2we obtain: Gt= T X i=0 γirt+i=rt+γGt+1 (7) And thus, applying (7) to (5) and (6), we obtain: Vπ(s) = Eπ[Gt|St=s] = X a π(a|s)·X s0∈S P(s0|s, a)·[R(s, a) + γVπ(s0)] (8) Qπ(s, a) = Eπ[Gt|St=s, At=a] = X s0∈S P(s0|s, a)·[R(s, a) + γVπ(s0)] (9) We can rewrite then (8) as: Vπ(s) = Eπ[Gt|St=s] = X a π(a|s)·Qπ(s, a) (10) Policy optimization. Conceptually, the policy that achieves the optimal value and action-value functions V∗and Q∗is the following: V∗(s) = max πVπ(s) (11) Q∗(s, a) = max πQπ(s, a) (12) Advantage function. A common approach for most algorithms in DRL consists on using the advantage concept in the objective function. The advantage function, defined in Equation 13, compares the Qvalue of an action to the average Q-value of all the actions given the same state, extracting the difference. Temporal difference methods [36], which are used by most of DRL algorithms, seek to estimate the advantage function and use this for the objective function to train the agent. A(s, a) = Q(s, a)−V(s) (13) 19 4.1.2 Exploration vs exploitation The training algorithms differ from supervised and unsupervised learning methodologies in that they do not use any pre-labeled data as the former, neither unlabeled data as the latter, as there is no previous data. They actually learn by interacting with the environment and analyzing the obtained rewards given their corresponding state and performed action, which is used as the training data. That is in fact why one of the biggest dilemmas in Reinforcement Learning is the trade-off between exploitation (using the already acquired knowledge) and exploration (of non-inspected options) [18]. On the one hand side, the agent must perform actions that based upon prior experience it knows will return the highest rewards, in order to find the optimal path. But on the other end, it must explore alternative and unknown options with the aim of collecting this prior experience. Since both strategies are exclusive, that is, at a given time step the agent must chose one or the other, it makes it a very difficult task to find the ideal balance. In order to tackle this, there exist various methods to take the decision for the agent. The most broadly used exploration strategy is the -Greedy Exploration, which consists on selecting, at each time step, an exploratory (random) action with probability , where  < 1, and selecting the learned optimal action with probability 1 −, as shown in Equation 14. (random action from A(s) if ξ <  maxaQ∗(s, a) otherwise,(14) There exist some other alternatives to this method that play around with the ideas of reducing the value of with time [5] [37]. The simplest approach to this is the decaying -Greedy, which uses the decay rate γto reduce over time, as shown in Equation 15. t=γt0,where 0 < γ < 1 (15) Another common method consist on using softmax [29], which sets the policy’s probability of selection for each action by ranking the value-action function values using a Boltzmann distribution, as shown in Equation 16. π(s, a) = eQ(s,a) τ Pa0 eQ(s,a0) τ (16) 4.1.3 Off-policy vs on-policy One of the most common classifications of Reinforcement Learning algorithms consists of dividing them into off-policy and on-policy methods. Off-policy algorithms estimate the Q-value function independently from their policy. That is, they perform actions (they are exploring, as the actions are not selected by the policy) and store the retrieved data (state-action pairs and the resulting rewards). Then, separately, a policy is trained based upon the observed Q-values. 20 On the other hand, on-policy algorithms update the Q-values directly using the current policy’s action. The same policy that is being trained to select the optimal actions is also responsible for the retrieval of the state-action pairs and corresponding rewards that are used for its own training. This actually translates into off-policy algorithms being more sensible to hyperparameters, biased, and prone to divergence, but at the same time they have lower variance and thus are more sample-efficient. 4.1.4 Model-based vs model-free Another usual classification of Reinforcement Learning methods consists on separating them contingent on if they are model-based or model-free. Model-based algorithms are those that, in order to train their policy, they model the environment, that is, they try to approximate the transition P(s0|s, a) and reward R(s, a) functions, and then use these approximations to estimate the optimal policy. Contrarily, a model-free algorithm does not seek to approximate a model of the environment (the transition and reward functions), but instead estimates the policy directly from experience. It trains a value function that given a state returns the optimal action to perform. 4.1.5 Deep Reinforcement Learning Deep Reinforcement Learning (DRL) is a subarea of Reinforcement Learning that combines deep neural networks with the reinforcement learning architecture. Given an environment in which the amount of states were too large (or even infinite), the estimation of the V-function, Q-value table or the policy is not feasible. That is when deep neural networks come into play: they can behave as function approximators, capable of receiving a flexible input (which represents the observation and/or action spaces) and output an estimation of the value. In DRL algorithms, we find that the neural networks can behave as estimators of: •State-value function (Vπ(s)): also known as V-function (see Equation 5), which receives as input an state sand returns the expected reward given a policy π. •State-action value function (Qπ(s, a)): also known as Q-function, which receives as input an action aand state sand returns the expected reward given a policy. •Policy (π(s)): the value function, which receives as input an state and returns the optimal action in order to maximize the cumulative reward. This approach allows RL algorithms to be applied to much more complex environments with successful results. Most of the research and RL applications as of today require the use of deep neural networks, as it is the case with this project. For instance, environments with continuous actions are more suitable to a DRL approach and would require a discretization of the actions’ domains if a dynamic programming approach were to be taken. 21 4.1.6 Algorithms Several algorithms are used in this project, each of them being different and might be more suitable for certain environments or more sensitive to some of the hyperparameters. In this section, a detailed description is made for each one of them. 4.1.6.1 Deep Q-Network (DQN) The Deep Q-Network algorithm was introduced by [28], and consists of approximating the Q-function, that returns a value for each of the environment state-action pairs (which is the expected reward by selecting that action in that specific state) by means of a deep neural network. Thus, the optimal deterministic action to perform can be determined following the policy shown in Equation 17, which can be directly obtained from the Q-function estimation. π(s) = arg max aQθ(s, a) (17) In order to do so, the algorithm uses two networks: the main one (called Q-network), and a target neural network, that is a copy of the former and acts as the agent to perform episode steps and obtain rewards. This data is then employed to update the weights of the main neural network after each episode, with samples being drawn uniformly from the replay buffer, and the gradient is computed using mini-batches in order to update the Q-network. Every ksteps (which is an hyperparameter), the target network is updated to the Q-network. This is done to minimize the volatility of the neural network used for the exploration and compilation of training data by only updating it every ksteps, maintaining stability for the periods in which it retrieves the data. The network is trained to minimize the loss function shown in Equation 18, which is the squared error between the network’s estimated Q-value and the observed reward. L= (rt+γarg max aQθ(st+1, at+1)−Qθ(s, a))2(18) On the downside, the DQN algorithm is only applicable to discrete action spaces, as it requires an efficient evaluation of the Q-function approximation which is not achievable in continuous action spaces unless a discretization is performed. 4.1.6.2 Deep Deterministic Policy Gradient (DDPG) The Deep Deterministic Policy Gradient algorithm [23] uses an actor-critic architecture. The critic Qφis responsible of learning the Q-function, while the actor πθdefines the policy that is used to select the optimal actions. The loss function of the critic is the same one used in DQN, as shown in Equation 19. LC= (rt+γQφ(st+1, πθ(st+1)) −Qφ(s, a))2(19) To train the actor, the loss function uses the critic network Qφwhich estimates the Q-value of the state and the action selected by the policy, as shown in Equation 20. LA=−Qφ(s, πθ(s)) (20) 22 4.1.6.3 Twin Delayed DDPG (TD3) Twin Delayed DDPG (TD3) [13] is an off-policy algorithm. It can be described as the improved version of DDPG, which can be unstable and prone to brittling with respect to hyperparameters, due to the critic network overestimating Q-values. TD3 tackles this issue by reducing DDPG’s overestimation bias in Q-values, by doing the following: •The addition of a second critic network (contrary to DDPG’s only having one). TD3 takes the minimum out of the two estimates of the Q-value, as shown in Equation 21, where d= 1 whenever st+1 is a terminal state. Q(st, at) = rt+γ(1 −d) min i=1,2Qφi (st+1, πθ(st+1)) (21) •The delay of the policy updates. TD3 updates the actor network less frequently than the critics (by default, one actor network update per every two critic updates). The policy is learned by maximizing only Qφ1, as shown in Equation 22. max θ E[Qφi(s, πθ(s))],∀s∈S(22) •Adding noise to the target action. This prevents the actor from taking advantage from Qvalues estimation errors. As shown in Equation 23, clipped noise is added to the policy’s selected action, with the hyperparameters cand σ, and takes into account the domain [amin, amax] of the action. A(s) = clip (πθ(s) + clip(, −c, c), amin, amax),where ∼N(0, σ) (23) These modifications translate into a significantly better performance than the DDPG baseline. 4.1.6.4 Soft Actor Critic (SAC) Soft Actor Critic (SAC) [16] [15] is an off-policy algorithm that stands out for not only looking for a maximization of the reward, but also seeking to maximize the entropy, which is basically a measure of the randomness within the policy. In other words, it seeks to obtain the highest possible reward while acting as randomly as possible. At each time step, the reward is modified by adding a bonus proportional to the entropy of the policy. Thus, the new cumulative reward is defined as shown in Equation 24. The hyperparameter αcontrols the explore-exploit tradeoff, a higher value translating into more exploration while a lower value indicates more exploitation. Gt=rt+γ(rt+1 +αH(π(·|st+1))) (24) And by applying the definition of entropy to 24: Gt=rt+γ(rt+1 −αlog(π(at+1|st+1))) (25) The algorithm is composed by two critic networks Qφ1and Qφ2, and a policy πθ. Learning Q-values. The Q-value function is defined then as Qφi(st, at) = Gt. Then the 23 5 Implementation There are two independent and clearly separated implementations in this work, one corresponding to CTE-POWER’s training scripts, and the other corresponding to the AWS cloud scripts and configuration. All the code is in a git repository, which is hosted in GitHub at https://github.com/miquelescobar/ddrl. In it, one can find all the code, results, figures, plots, videos and documentation generated during the development of the project. The README file provides a comprehensive guide for reproducibility, explaining each of the directories in detail as well as the dependencies installation process. One of the aims of this work is to deeply understand the differences between the implementation in a private computing cluster (in this case, the CTE-POWER cluster) and the implementation in an external cloud services provider (in this case, Amazon Web Services). Listed below are the main properties, advantages and disadvantages for each of the two options. CTE-POWER cluster + Very little learning overhead for using it, as it only requires extra knowledge of Slurm Workload Manager [40]. + It is ”free” to use for the members of research groups in BSC. −Cluster maintenance breaks prevent the users from accessing and using the machines, and can last from a few hours to a few days. −The operating system, Red Hat Enterprise Linux Server 7.5 is the same for all the machines, and might cause problems with some dependencies that work differently in other OS. Furthermore, the average user does not have access to the configuration of the machine and all the dependencies must be installed and maintained by a specialized professional. −There is no open access to the Internet from the machines. This precludes the setup of any web service or communication that could be needed in an application, and thus requires another machine to host the setup. Amazon Web Services −It requires a considerable amount of previous work to familiarize and learn how AWS works, and more specifically, how to use AWS SageMaker. −All the offered services are billed, thus an efficient use is required to not go over budget. + There is an extensive variety of server instances to use, with all possible resource combinations. + There is also a great variety of operating systems and pre-built Docker images that include some of the most common dependencies installed. Anyhow, these can always be customized by the user by means of the Docker image making the service compatible with basically any system. + The SageMaker service is fully integrated with other AWS services, making it easy to monitor, store data or deploy the resulting model, among others. 30 In summary, the two options have very different properties but at the same time each of them fulfils a clear and differentiated purpose. From the point of view of research, the CTE-POWER cluster is very adequate to carry out trials executing as many jobs as necessary, experiment with as many algorithms, environments and configurations as desired, since the costs are so low. On the other hand, AWS can be used for those cases in which not much experimentation is required (the task is clearly defined and the algorithm is known previously or is already trained in CTE-POWER, for instance) but the model or results must be integrated into an application, and thus requires some communication and coordination with other modules. 5.1 Training in RLlib Both approaches have in common the framework used with the built-in algorithms, which is Ray. The algorithms can be trained following different architectures. The algorithms selected for the chosen environment (TD3, SAC, PPO) all follow the same architecture, shown in Figure 8. In terms of resource configuration, Ray easily allows to specify the amount and properties of the rollout workers, that is: •The amount of workers for the Trainer module. •The amount of sampling rollout workers (responsible for sampling new data). •The number of CPUs and GPUs for each of the worker types. , Figure 8: Architecture of a Ray training process, for TD3, SAC and PPO algorithms. Whenever the required resources exceed those of one node, we can recur to distributed computing. The training algorithm maintains the architecture shown in Figure 8, but there is one more step required. A master process must initialize Ray in one head node (must know the addresses of the worker nodes to communicate with all of them) and one or more worker nodes (all must know the address of the head node to communicate with it) that initialize the required Ray workers as usually. Then, the head node is responsible for executing the training script, storing the checkpoints and metrics. The Ray workers in the worker nodes are responsible for executing their given tasks (learning or sampling). Figure 9shows a high-level representation of this architecture. 31 Figure 9: Architecture of distributed computing in Ray. Thus, the mention of the training scripts in the following sections refer to a script that executes this architecture, and the configuration of the training includes the resource allocation. 5.2 Implementation in CTE-POWER Cluster For training the algorithms using the CTE-POWER, Python scripts are implemented and posteriorly executed using the Slurm Workload Manager [40], with which the allocated resources and properties of the job can be configured. It is very important that the allocated resources are enough for those demanded in the configuration of the training script, otherwise the script will detect that there are not enough resources and the execution will result in an error. Furthermore, all the required dependencies are satisfied by loading the corresponding environment modules in the cluster, which include the required versions of the necessary software frameworks and/or packages, for both Reinforcement Learning algorithms and training (Ray, Tensorflow, Pytorch) and Reinforcement Learning environments implementations (OpenAI Gym, PyBullet). At the same time, the training scripts allow parametrization for choosing the environment, algorithm and hyperparameters (there can be more than one parallel combination, see 4.2.1.2). The results are stored at the defined directory, as well as the configured checkpoints during training. The schema in Figure 10 shows the implemented workflow for the training process in the CTE-POWER cluster. Figure 10: The architecture of the solution implementation in CTE-POWER cluster. 32 A shown in the schema, a training job is defined and executed from a bash script, which requires four independent modules. The sbatch options consist of a series of parameters that configure the job (from the resources allocation to the output of the logs, among others). The environment modules refer to the layers that the job requires to operate, which are software packages (the Python version, the required libraries, etc.) defined by the system administrators. The training script is the command that will be executed, and must handle all the training process. The Ray Tune configuration file are the parameters that will be passed to the training, such as the environment, the algorithm, stopping conditions and hyperparameters. When the four modules above are defined or referred to in the bash file, the training job can be executed using the sbatch command. During the execution, the training job stores the corresponding checkpoints and training metrics. Once the job finishes, we are left with a final model as well as with a complete metrics files with an iteration granularity level. In order to generate the videos of the interactions with the trained model, the corresponding renderization scripts can be executed in a computer with the required dependencies installed, since CTE-POWER computation nodes do not have the capabilities to generate videos from the renderization of environment states. 5.3 Implementation in Amazon Web Services Some of the trainings made in CTE-POWER, mainly the most successful ones, are reproduced in the cloud for comparison purposes, using the AWS SageMaker service. In order to train an algorithm for a given environment, the instance type, number of instances, timeout, dependencies and training script can be specified to create an AWS SageMaker training job, which can then be executed. During the execution, configurable checkpoints and renderized videos are stored in the provided S3 bucket, and the configured metrics can be tracked live using either AWS CloudWatch or by setting up a TensorBoard. Once the execution is finsihed, due to any of the predefined stopping conditions, the final model artifact is uploaded to that same S3 bucket. 33 Figure 11: The architecture of the solution implementation in AWS SageMaker. As can be seen in Figure 11, the training job requires the helper code, which configures the job, uploads the necessary files, builds and pushes the Docker container and starts the execution. From there onwards, the control of the pipeline is left completely to the SageMaker service. Each of the modules that appear in the architecture schema are explained below. The helper code consists of a script (in a Jupyter Notebook) with a series of commands that configure the parameters for the training job (see RLEstimator class documentation in AWS Sagemaker Python SDK). It is also responsible for builidng and pushing the corresponding Docker image to ECR. Finally, the script submits the training job in order for it to start. The helper code also must upload the training script, which is the entry point of the training job, as well as the required dependencies to S3 (only Python code, any other dependency must be installed in the Docker image). The training job instance retrieves the Docker image from ECR and runs the Docker container process. The process downloads the training script and dependencies from S3 and uses them as the source code to execute. During the execution, the output generated by the job is the checkpoints folder, as well as the renderized videos of the agent at each checkpoint and the CloudWatch metrics, which were previously configured in the helper code. At the end of the job, a compressed file is generated with the model artifact, that is, all the properties and the weights of the final model. Finally, the model can be deployed into inference by referencing the resulting model artifact. SageMaker provides an easy interface to automatically create the endpoint, formatting the responses and setting up the serverless service responsible for listening and answering the received petitions at the endpoint. This inference environment can be configured to suit the availability and desired requirements of the inference service. 34 6 Evaluation In this section the results of all the executed trainings for the selected environment are analyzed. The focus is made in analyzing the execution time (we observe how parallelism and distributed computing reduces it significantly) and the obtained rewards per each algorithm and hyperparameter combination. Refer to the GitHub repository at https://github.com/miquelescobar/ddrl/results/, where the plotting scripts, figures, renderized videos and summarized training metrics can be found. 6.1 Training results The trained algorithms are TD3, SAC and PP0. This section shows the trainings with best results for each of the algorithms. 6.1.1 TD3 A detailed explanation of the algorithm can be found in 4.1.6.3. Even though tries to solve the drastic overestimation problem of the DDPG algorithm, it is still considered an unstable algorithm that is prone to overestimate Q-values, since it does nothing explicit to control the policy updates. The best performing training is presented in Figure 12, and corresponds to the hyperparameter configuration shown in Table 3. Hyperparameter Symbol Value Tau τ0.002 Horizon h1000 Actor learning rate 0.003 Critic learning rate 0.003 Discount factor γ0.99 Batch size 256 Actor hidden layers [512, 512] Critic hidden layers [512, 512] Activation function ReLU Target noise 0.2 Target noise clip c0.5 Table 3: Hyperparameters with best results for TD3. We can observe in Figure 12 some sporadic spikes in volatility as well as three clear performance collapses (at around timesteps 1.5e6,3e6 and 4e6. The latter are probably caused by the algorithm visiting what are known as catastrophic states [24], which are states of the environment that the agent has never visited before or it did but a long time ago (known as catastrophic forgetting), and thus the Q-value estimation is erroneous which at the same time can lead with a high probability to new catastrophic states. Even so, we observe that the algorithm is quick to recover by training the newly found states. 35 This is a clear example of one of the biggest problems in RL, the trade-off between exploration and exploitation (see 4.1.2). Since TD3 does not tackle this issue explicitly, the consequences are greater than in other algorithms (such as SAC or PPO). Figure 12: TD3 algorithm average reward per time step for the best performing configuration. 6.1.2 SAC A detailed explanation of the algorithm can be found in 4.1.6.4. Similarly to TD3, SAC uses two critic networks and one actor, but the main difference is in the definition of the Q-value. We have seen that TD3 can suffer from overestimation problems, and in order to sort this problem out, SAC modifies the Q-value by adding to the cumulative reward the policy’s entropy for the given state as a bonus, thus encouraging policies with higher entropy, that is, that select actions in a more ”random” manner. As we observe in the plot in Figure 13, that shows the best performing SAC training (with the hyperparameters configuration presented in Table 4), the training seems a lot more smoother and less volatile. 36 Figure 13: SAC algorithm average reward per time step for the best performing configuration. Hyperparameter Symbol Value Tau τ0.005 Horizon h1000 Actor learning rate 0.001 Critic learning rate 0.001 Discount factor γ0.99 Batch size 256 Actor hidden layers [512, 512] Critics hidden layers [256, 256] Activation function ReLU Initial alpha α1.0 Table 4: Hyperparameters with best results for SAC. 6.1.3 PPO A detailed explanation of the algorithm can be found in 4.1.6.5. Differently from TD3 and SAC, PPO tackles the stability issue with a much more invasive methodology: it reduces the objective by clipping the ratio πθ(a|s) πθnew(a|s), and thus the update of the policy is constrained at each iteration. This translates into a much more stable training, as can be seen in Figure 14, and increases the probability of convergence. Observing the reward evolution plot in Figure 14, it is noticeable that the amount of time steps is much higher than in the TD3 and SAC trainings, more specifically one order of mag37 nitude higher. That is due to the lower sample efficiency of on-policy methods, that directly translates into a higher variance, as explained in Section 4.1.3. Even so, the average time it takes per time step is considerably lower than TD3 and SAC, thus compensating the low sample efficiency. Figure 14: PPO algorithm average reward per time step for the best performing configuration. The hyperparameters that obtained the highest reward are shown in Figure 5. Hyperparameter Symbol Value Horizon h1000 Actor GAE lambda λ1.0 Learning rate 0.0005 Discount factor γ0.99 Batch size 4000 SGD minibatch size 256 Clip parameter 0.3 Table 5: Hyperparameters with best results for PPO. 6.2 Execution time The execution time is a critical metric to monitorize during training of algorithms. Time is always related to cost: in the CTE-POWER cluster resources are limited and a training job does not free them until it is finished, thus reducing execution time is key; in AWS SageMaker you are actually paying for all the seconds that the training job servers are up and running. 38 Consequently the best resource configuration is the one that achieves the best time and cost efficiencies, that is, the best ratio between the total execution time and the total cost of the execution. In Figure 8we observe the architecture and resource allocation for the training jobs execution. In this section, we analyze the effects that different resource allocation and workers configurations have in the training execution time. 6.2.1 TD3 For the TD3 algorithm, we can observe in Figures 15 and 16 the execution time as a function of the number of workers used. As the number of workers increases, there is a clear reduction of the time spent per iteration. The speedup obtained between the execution with just one worker and the one with 64 and a GPU is around ×13. Figure 15: Training iteration time by number of workers (TD3). Figure 16: Training iteration time by number of workers, with GPU (TD3). Num workers GPU CPUs / worker Iteration time (s) Speedup 1 0 1 25.56 1.00 1 1 1 21.69 1.18 2 0 1 14.38 1.78 2 1 1 11.58 2.21 4 0 1 13.10 1.95 4 1 1 10.09 2.53 8 0 1 8.12 3.15 8 1 1 6.76 3.78 15 0 1 3.63 7.04 16 0 1 5.20 4.92 16 1 1 4.24 6.04 32 0 1 2.68 9.52 32 1 1 2.19 11.69 64 0 1 2.35 10.88 64 1 1 1.99 12.86 Table 6: Iteration time and speedup per resource configuration (TD3). 39 References [1] Takuya Akiba et al. “Optuna: A Next-generation Hyperparameter Optimization Framework”. In: Proceedings of the 25rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2019. [2] J. Bergstra, D. Yamins y D. D. Cox. “Making a Science of Model Search: Hyperparameter Optimization in Hundreds of Dimensions for Vision Architectures”. In: Proceedings of the 30th International Conference on International Conference on Machine Learning - Volume 28. ICML’13. Atlanta, GA, USA: JMLR.org, 2013, I–115–I–123. [3] Greg Brockman et al. OpenAI Gym. 2016. eprint: arXiv:1606.01540. [4] BSC. Emerging Technologies for Artificial Intelligence research group.https://www.bsc .es/discover-bsc/organisation/scientific-structure/emerging -technologies -artificial-intelligence. [5] Olivier Caelen y Gianluca Bontempi. Improving the Exploration Strategy in Bandit Algorithms. Dec. 2007. doi:10.1007/978-3-540-92695-5 5. [6] Tianqi Chen y Carlos Guestrin. “XGBoost: A Scalable Tree Boosting System”. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. KDD ’16. San Francisco, California, USA: ACM, 2016, pp. 785–794. isbn: 978-1-4503-4232-2. doi:10.1145/2939672.2939785.url:http://doi.acm.org/ 10.1145/2939672.2939785. [7] Tianqi Chen et al. MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems. 2015. arXiv: 1512.01274 [cs.DC]. [8] R. Collobert, K. Kavukcuoglu y C. Farabet. “Torch7: A Matlab-like Environment for Machine Learning”. In: BigLearn, NIPS Workshop. 2011. [9] Erwin Coumans y Yunfei Bai. PyBullet, a Python module for physics simulation for games, robotics and machine learning.http://pybullet.org. 2016–2021. [10] Alexander I Cowen-Rivers et al. “HEBO: Heteroscedastic Evolutionary Bayesian Optimisation”. In: arXiv preprint arXiv:2012.03826 (2020). winning submission to the NeurIPS 2020 Black Box Optimisation Challenge. [11] Benjamin Ellenberger. PyBullet Gymperium.https://github.com/benelot/pybullet -gym. 2018–2019. [12] Stefan Falkner, Aaron Klein y Frank Hutter. “BOHB: Robust and Efficient Hyperparameter Optimization at Scale”. In: Proceedings of the 35th International Conference on Machine Learning. Ed. by Jennifer Dy y Andreas Krause. Vol. 80. Proceedings of Machine Learning Research. PMLR, Oct. 2018, pp. 1437–1446. url:http://proceedings.mlr .press/v80/falkner18a.html. [13] Scott Fujimoto, Herke van Hoof y David Meger. Addressing Function Approximation Error in Actor-Critic Methods. 2018. arXiv: 1802.09477 [cs.AI]. [14] Synergy Research Group. Cloud Market Ends 2020 on a High while Microsoft Continues to Gain Ground on Amazon. Tech. rep. Synergy Research Group, February 2, 2021. url: https://www .srgresearch .com/articles/cloud -market -ends -2020 -high -while -microsoft-continues-gain-ground-amazon. [15] Tuomas Haarnoja et al. Soft Actor-Critic Algorithms and Applications. 2019. arXiv: 1812 .05905 [cs.LG]. [16] Tuomas Haarnoja et al. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. 2018. arXiv: 1801.01290 [cs.LG]. 46 [17] Arthur Juliani et al. Unity: A General Platform for Intelligent Agents. 2020. arXiv: 1809 .02627 [cs.LG]. [18] Leslie Pack Kaelbling, Michael L. Littman y Andrew W. Moore. “Reinforcement Learning: A Survey”. In: CoRR cs.AI/9605103 (1996). url:https://arxiv.org/abs/cs/9605103. [19] Kirthevasan Kandasamy et al. Tuning Hyperparameters without Grad Students: Scalable and Robust Bayesian Optimisation with Dragonfly. 2020. arXiv: 1903.06694 [stat.ML]. [20] Lisha Li et al. Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization. 2018. arXiv: 1603.06560 [cs.LG]. [21] Eric Liang et al. RLlib: Abstractions for Distributed Reinforcement Learning. 2018. arXiv: 1712.09381 [cs.AI]. [22] Richard Liaw et al. Tune: A Research Platform for Distributed Model Selection and Training. 2018. arXiv: 1807.05118 [cs.LG]. [23] Timothy P. Lillicrap et al. Continuous control with deep reinforcement learning. 2019. arXiv: 1509.02971 [cs.LG]. [24] Zachary C. Lipton et al. Combating Reinforcement Learning’s Sisyphean Curse with Intrinsic Fear. 2018. arXiv: 1611.01211 [cs.LG]. [25] Yu-Ren Liu et al. ZOOpt: Toolbox for Derivative-Free Optimization. 2018. arXiv: 1801 .00329 [cs.LG]. [26] Martın Abadi et al. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. Software available from tensorflow.org. 2015. url:https://www.tensorflow.org/. [27] Michael Mesnier, Gregory Ganger y Erik Riedel. “Object-based storage”. In: Communications Magazine, IEEE 41 (Sept. 2003), pp. 84–90. doi:10.1109/MCOM.2003.1222722. [28] Volodymyr Mnih et al. Playing Atari with Deep Reinforcement Learning. 2013. arXiv: 1312.5602 [cs.LG]. [29] Ling Pan et al. Reinforcement Learning with Dynamic Boltzmann Softmax Updates. 2019. arXiv: 1903.05926 [cs.LG]. [30] Adam Paszke et al. “PyTorch: An Imperative Style, High-Performance Deep Learning Library”. In: Advances in Neural Information Processing Systems 32. Curran Associates, Inc., 2019, pp. 8024–8035. url:http://papers.neurips.cc/paper/9015-pytorch-an -imperative-style-high-performance-deep-learning-library.pdf. [31] Fabian Pedregosa et al. “Scikit-learn: Machine learning in Python”. In: Journal of machine learning research 12.Oct (2011), pp. 2825–2830. [32] J. Rapin y O. Teytaud. Nevergrad - A gradient-free optimization platform.https:// GitHub.com/FacebookResearch/Nevergrad. 2018. [33] Matthew Rocklin. “Dask: Parallel computation with blocked algorithms and task scheduling”. In: Proceedings of the 14th python in science conference. 130-136. Citeseer. 2015. [34] John Schulman et al. Proximal Policy Optimization Algorithms. 2017. arXiv: 1707.06347 [cs.LG]. [35] David Silver et al. “Reward is enough”. In: Artificial Intelligence 299 (2021), p. 103535. issn: 0004-3702. doi:https : / / doi .org / 10 .1016 / j .artint .2021 .103535.url: https://www.sciencedirect.com/science/article/pii/S0004370221000862. [36] Richard Sutton. “Learning to Predict by the Method of Temporal Differences”. In: Machine Learning 3 (Aug. 1988), pp. 9–44. doi:10.1007/BF00115009. [37] Michel Tokic. Adaptive -Greedy Exploration in Reinforcement Learning Based on Value Differences. Sept. 2010. doi:10.1007/978-3-642-16111-7 23. 47 [38] Jordi Torres. Introducci´on al aprendizaje por refuerzo profundo. Teor´ıa y pr´actica en Python. Barcelona: WATCH THIS SPACE Book Series. Kindle Direct Publishing, 2021. isbn: 9798599775416. [39] Wikipedia. Euler angles — Wikipedia, The Free Encyclopedia.http://en .wikipedia .org/ w /index .php ?title = Euler% 20angles & oldid= 1028147008. [Online; accessed 20-June-2021]. 2021. [40] Andy B. Yoo, Morris A. Jette y Mark Grondona. “SLURM: Simple Linux Utility for Resource Management”. In: Job Scheduling Strategies for Parallel Processing. Ed. by Dror Feitelson, Larry Rudolph y Uwe Schwiegelshohn. Berlin, Heidelberg: Springer Berlin Heidelberg, 2003, pp. 44–60. isbn: 978-3-540-39727-4. [41] Matei Zaharia et al. “Apache Spark: A Unified Engine for Big Data Processing”. In: Commun. ACM 59.11 (Oct. 2016), pp. 56–65. issn: 0001-0782. doi:10.1145/2934664. url:https://doi.org/10.1145/2934664. [42] Chiyuan Zhang et al. A Study on Overfitting in Deep Reinforcement Learning. 2018. arXiv: 1804.06893 [cs.LG]. 48 A Environment specificatios A.1 Observation space The observation space is composed by the Euler angles αand β[39], the differences between the current and target values (x, y, z) which are constant (in fact, target zis defined as the initial value of z), the collision of the feet, and finally the relative positions and the speeds of all the joints. Name Type Interval Description αfloat [−π, π] The Euler angle α βfloat [−π, π] The Euler angle β x target difference float [−π, π] The difference between current xthe constant target x y target difference float [−π, π] The difference between current ythe constant target y z target difference float [−π, π] The difference between current zthe initial value of z velocity x float [−1,1] The velocity of the robot at the xaxis (comparing stto st−1positions) velocity y float [−1,1] The velocity of the robot at the yaxis (comparing stto st−1positions) velocity z float [−1,1] The velocity of the robot at the zaxis (comparing stto st−1positions) feet1 contact bool {0,1}Whether or not the right feet is in contact with any other link, joint or object. feet2 contact bool {0,1}Whether or not the left feet is in contact with any other link, joint or object. abdomen x position float [−1,1] The relative position to the body of abdomen x joint. abdomen y position float [−1,1] The relative position to the body of abdomen y joint. abdomen z position float [−1,1] The relative position to the body of abdomen z joint. right hip x position float [−1,1] The relative position to the body of right hip x joint. right hip y position float [−1,1] The relative position to the body of right hip y joint. right hip z position float [−1,1] The relative position to the body of right hip z joint. left hip x position float [−1,1] The relative position to the body of left hip x joint. left hip y position float [−1,1] The relative position to the body of left hip y joint. left hip z position float [−1,1] The relative position to the body of left hip z joint. right knee position float [−1,1] The relative position to the body of right knee joint. left knee position float [−1,1] The relative position to the body of left knee joint. 49 right shoulder1 position float [−1,1] The relative position to the body of right shoulder1 joint. right shoulder2 position float [−1,1] The relative position to the body of right shoulder2 joint. left shoulder1 position float [−1,1] The relative position to the body of left shoulder1 joint. left shoulder2 position float [−1,1] The relative position to the body of left shoulder2 joint. right elbow position float [−1,1] The relative position to the body of right elbow joint. left elbow position float [−1,1] The relative position to the body of left elbow joint. abdomen x speed float [−1,1] The speed of the joint, mapped between −1 and 1. abdomen y speed float [−1,1] The speed of the joint, mapped between −1 and 1. abdomen z speed float [−1,1] The speed of the joint, mapped between −1 and 1. right hip x speed float [−1,1] The speed of the joint, mapped between −1 and 1. right hip y speed float [−1,1] The speed of the joint, mapped between −1 and 1. right hip z speed float [−1,1] The speed of the joint, mapped between −1 and 1. left hip z speed float [−1,1] The speed of the joint, mapped between −1 and 1. right knee speed float [−1,1] The speed of the joint, mapped between −1 and 1. left knee speed float [−1,1] The speed of the joint, mapped between −1 and 1. right shoulder1 speed float [−1,1] The speed of the joint, mapped between −1 and 1. right shoulder2 speed float [−1,1] The speed of the joint, mapped between −1 and 1. left shoulder1 speed float [−1,1] The speed of the joint, mapped between −1 and 1. left shoulder2 speed float [−1,1] The speed of the joint, mapped between −1 and 1. right elbow speed float [−1,1] The speed of the joint, mapped between −1 and 1. left elbow speed float [−1,1] The speed of the joint, mapped between −1 and 1. Table 9: Properties of all dimensions in the observation space. A.2 Action space The action space is formed by 17 actuators, each corresponding to one joint of the robot. All of them can be set to any value in the [−1,1] interval. The power is the maximum force that the joint can apply (for instance, a knee is stronger than the shoulder). The resulting applied 50 force to each joint is thus computed as joint input ×power. Name Type Domain interval Power abdomen x float [−1,1] 100 abdomen y float [−1,1] 100 abdomen z float [−1,1] 100 right hip x float [−1,1] 100 right hip y float [−1,1] 300 right hip z float [−1,1] 100 left hip x float [−1,1] 100 left hip y float [−1,1] 300 left hip z float [−1,1] 100 right knee float [−1,1] 200 left knee float [−1,1] 200 right shoulder1 float [−1,1] 75 right shoulder2 float [−1,1] 75 left shoulder1 float [−1,1] 75 left shoulder2 float [−1,1] 75 right elbow float [−1,1] 75 left elbow float [−1,1] 75 Table 10: Properties of all dimensions in the action space. 51