Full text
id188972 DEVELOPING COORDINATION METHODS FOR AI AGENTS IN OPEN-ENDED ENVIRONMENTS ORIOL MIRÓ LÓPEZ-FELIU Thesis supervisor VÍCTORGIMÉNEZÁBALOS(DepartmentofComputerScience) Thesis co-supervisor ADRIANTORMOSLLORENTE Degree Bachelor'sDegreeinInformaticsEngineering(Computing) Bachelor's thesis Facultat d'Informàtica de Barcelona (FIB) Universitat Politècnica de Catalunya (UPC) - BarcelonaTech
Acknowledgmeents I would like to express my sincere gratitude to my supervisor, Víctor Giménez Ábalos, for his invaluable guidance, support, and encouragement throughout the course of this research. His expertise and insights have been instrumental in shaping the direction and quality of this work. I am also deeply grateful to my co-supervisor, Adrián Tormos Llorente, for his support, dedication, and constructive feedback, as his contributions have greatly enhanced the depth and clarity of this work. Finally, I would like to extend my heartfelt appreciation to my partner, Aitana Conejero Moreno, for their unwavering support, love, patience, and understanding during the entire research process. Their constant encouragement and belief in me have been a source of strength and motivation.
Abstract Effective coordination among intelligent agents remains a significant challenge, especially in open-ended environments. Existing methods often lack scalability, generalizability, and efficiency, hindering progress in various domains. Moreover, the growing reliance on Large Language Models (LLMs) introduces bottlenecks, resulting in slow and costly systems. We propose a novel Reinforcement Learning (RL)-based coordination method that enables message exchange across agents while treating their original architectures as nearly black-box systems. Our method was successfully integrated into STEVE-1, a State-of-the-art (SOTA) short-term instruction following model, and evaluated in the Multiagent Minedojo environment. Experiments show promising results in scenarios requiring coordination, although resource constraints limited conclusive findings in situations where coordination was beneficial but not always necessary, likely due to the complexity of the tested tasks. The proposed method demonstrated remarkable flexibility, requiring no task-specific adaptations, and the results suggested that agents utilize our approach in a manner tailored to the specific requirements of each task. These findings have significant implications for the applicability of single agents to multi-agent settings, for enabling coordination when LLMs are infeasible, and when used in conjunction with LLMs to produce more cost-effective and responsive systems. All project code can be found at https://github.com/oriolmirolf/TFG , and supplementary materials, such as example videos, are available at https://drive.google.com/drive/folders/1_Ez795BllPY_ BptTfXNeE0BXsGHt2IPX?usp=sharing.
Resumen La coordinación efectiva entre agentes inteligentes sigue siendo un reto considerable, especialmente en entornos abiertos. Los métodos actuales a menudo carecen de escalabilidad, generalizabilidad y eficiencia, limitando el progreso en varios dominios. Además, la dependencia creciente en LLMs introduce cuellos de botella, lo que resulta en sistemas lentos y costosos. Proponemos un nuevo método de coordinación basado en RL que permite el intercambio de mensajes entre agentes, tratando sus arquitecturas casi como cajas negras. Nuestro método ha sido integrado con éxito a STEVE-1, un modelo estado del arte capaz de seguir instrucciones a corto plazo, y ha sido evaluado en el entorno Multiagent Minedojo. Los experimentos muestran resultados prometedores en los escenarios que requieren coordinación, aunque los pocos recursos disponibles han limitado los resultados en aquellas situaciones donde la coordinación era beneficiosa pero no siempre necesaria, probablemente debido a la complejidad de las tareas evaluadas. El método propuesto ha demostrado una flexibilidad destacable, sin requerir adaptaciones específicas según la tarea, y los resultados sugieren que los agentes adaptan el uso del método a los requerimientos específicos de cada tarea. Estos hallazgos tienen implicaciones significativas respecto a la aplicabilidad de agentes individuales a sistemas multiagente, a la hora de posibilitar la coordinación cuando el uso de LLMs no es factible, y para producir sistemas más rápidos y rentables cuando se utilizan en conjunción con LLMs. Todo el código del proyecto se puede encontrar en https://github.com/ oriolmirolf/TFG , y los materiales complementarios, como vídeos de ejemplo, están disponibles en https://drive.google.com/drive/folders/1_ Ez795BllPY_BptTfXNeE0BXsGHt2IPX?usp=sharing.
Resum La coordinació efectiva entre agents intel·ligents continua sent un repte considerable, especialment en entorns oberts. Els mètodes actuals sovint manquen d’escalabilitat, generalitzabilitat i eficiència, limitant el progrés en diversos sectors. A més, la dependència creixent en LLMs introdueix colls d’ampolla, la qual cosa resulta en sistemes lents i costosos. Proposem un nou mètode de coordinació basat en RL que permet l’intercanvi de missatges entre agents, tractant les seves arquitectures gairebé com a caixes negres. El nostre mètode ha estat integrat amb èxit a STEVE-1, un model estat de l’art capaç de seguir instruccions a curt termini, i ha estat avaluat en l’entorn Multiagent Minedojo. Els experiments mostren resultats prometedors en els escenaris que requereixen coordinació, tot i que els pocs recursos disponibles han limitat els resultats en aquelles situacions on la coordinació era beneficiosa però no sempre necessària, probablement a causa de la complexitat de les tasques avaluades. El mètode proposat ha demostrat una flexibilitat destacable, sense requerir adaptacions específiques segons la tasca, i els resultats suggereixen que els agents adapten l’ús del mètode als requisits específics de cada tasca. Aquestes troballes tenen implicacions significatives respecte a l’aplicabilitat d’agents individuals a sistemes multiagent, a l’hora de possibilitar la coordinació quan l’ús de LLMs no és factible, i per produir sistemes més ràpids i rendibles quan s’utilitzen en conjunció amb LLMs. Tot el codi del projecte es pot trobar a https://github.com/oriolmirolf/TFG , i els materials complementaris, com vídeos d’exemple, estan disponibles a https://drive.google.com/drive/folders/1_Ez795BllPY_ BptTfXNeE0BXsGHt2IPX?usp=sharing.
Contents I GEP Module 1 1 Context and Scope 1 1.1 Context ......................................... 1 1.1.1 Preliminaries ................................. 2 1.1.2 Problemtobesolved............................. 5 1.1.3 Stakeholders ................................. 6 1.2 Justification...................................... 6 1.2.1 Previousstudies ............................... 6 1.2.2 Justification.................................. 6 1.3 Scope ......................................... 7 1.3.1 Objectives................................... 7 1.3.2 Requirements................................. 8 1.3.3 Potential risks and obstacles . . . . . . . . . . . . . . . . . . . . . . . . . 8 1.4 MethodologyandRigour............................... 8 1.4.1 Methodology................................. 9 1.4.2 Monitoring tools and validation . . . . . . . . . . . . . . . . . . . . . . . 9 1.5 LawsandRegulations................................. 9 II MVP 10 2 Environment and Model Preparations 10 2.1 Multiagent Minedojo Adaptations . . . . . . . . . . . . . . . . . . . . . . . . . . 10 2.2 STEVE-1AdaptationS ................................. 11 2.2.1 Issues and limitations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 3 Treasurehunt Task 11 3.1 Escapemode ..................................... 12 3.2 Dungeonlayout.................................... 13 3.3 RewardShaping.................................... 13 4 Method 14 4.1 TrainingAlgorithm .................................. 14 4.1.1 Foundations of single-network PPG . . . . . . . . . . . . . . . . . . . . . 14 4.1.2 Multi-agent Implementation . . . . . . . . . . . . . . . . . . . . . . . . . 16 4.2 CoordinationMethod................................. 17 4.2.1 Scalability .................................. 19 4.2.2 Generalizability................................ 20 5 Optimizations 20 5.1 BaselineandChallenges ............................... 20 5.2 TrainingOptimizations ................................ 20 5.3 Smart memory management . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 5.4 Environment optimization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 5.5 Results......................................... 22 6 Experiments 22 6.1 STEVE-1Calibration................................. 22 6.1.1 Fighter and Collector Calibration . . . . . . . . . . . . . . . . . . . . . . 23 6.1.2 Results .................................... 23 6.2 BaselineExperiments................................. 24 6.3 Coordination Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 6.4 ResultsandAnalysis ................................. 26 6.4.1 Studying Coordination Patterns . . . . . . . . . . . . . . . . . . . . . . . 27 6.4.2 Potential Reasons for Suboptimal Coordination Performance . . . . . . . . 28
III Extension Cycles 29 7 Improving the Treasurehunt Task 29 7.1 Baseline and Coordination Training . . . . . . . . . . . . . . . . . . . . . . . . . 29 7.2 ResultsAnalysis.................................... 30 7.3 Conclusion and Coordination Patterns . . . . . . . . . . . . . . . . . . . . . . . . . 31 8 Blind Guiding Task 32 8.1 TheBlindGuidingTask................................ 32 8.1.1 MazeLayout ................................. 32 8.1.2 RewardShaping ............................... 32 8.2 Experiments...................................... 33 8.2.1 STEVE-1 Calibration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33 8.2.2 Baseline and Coordination Training . . . . . . . . . . . . . . . . . . . . . 33 8.2.3 Results and Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34 8.2.4 Coordination Patterns and Conclusion . . . . . . . . . . . . . . . . . . . . 35 IV Discussion and Conclusions 36 9 Discussion 36 10 Conclusion 36 11 Future Work 37 12 Revisiting Project Planning 37 12.1 Reviewing Initial Objectives . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37 12.2 Reviewing Temporal Planning, Budget and Sustainability . . . . . . . . . . . . . . 38 V Appendices 44 A Temporal Planning 44 A.1 Descriptionoftasks.................................. 44 A.1.1 Rolesinvolved ................................ 44 A.1.2 Taskdefinition ................................ 44 A.1.3 Summaryofthetasks............................. 47 A.1.4 Resources................................... 48 A.2 GanttDiagram .................................... 49 A.3 RiskManagement................................... 49 A.3.1 Researcher’s inexperience . . . . . . . . . . . . . . . . . . . . . . . . . . 49 A.3.2 Insufficient Computational Resources . . . . . . . . . . . . . . . . . . . . 50 A.3.3 Complexity of Multi-agent Coordination . . . . . . . . . . . . . . . . . . 50 A.3.4 Complexity of Algorithm Implementation . . . . . . . . . . . . . . . . . . 50 A.3.5 Technical Limitations of Models . . . . . . . . . . . . . . . . . . . . . . . 50 B Budget 51 B.1 Identificationofcosts.................................. 51 B.2 Costestimates..................................... 52 B.3 Managementcontrol ................................. 53 C Sustainability 54 C.1 Selfreflection..................................... 54 C.2 EconomicDimension................................. 55 C.3 Environmental Dimension . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55 C.4 SocialDimension................................... 56 D Evaluation of current tasks 56
Developing Coordination Methods for AI Agents in Open-Ended Environments E Adaptation Details 57 E.1 Multiagent Minedojo Adaptation Details . . . . . . . . . . . . . . . . . . . . . . . 57 E.2 STEVE-1 Adaptation Details . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 57 E.2.1 Adaptation of MineRL to Minedojo Spaces . . . . . . . . . . . . . . . . . 58 E.2.2 Conditional Scale Adaptations . . . . . . . . . . . . . . . . . . . . . . . . 59 F Treasurehunt Task Design Details 59 F.1 EscapeMethodDetails ................................ 59 F.2 DungeonLayoutDetails ............................... 60 G Optimizations Details 61 G.1 Environment Optimization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61 H Additional Training Details 62 H.1 Original Treasurehunt Task . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62 H.1.1 Baseline Training for the Original Treasurehunt Task Details . . . . . . . . 62 H.1.2 Coordination Training for the Original Treasurehunt Task Details . . . . . 63 H.2 Improved Treasurehunt Task, Original Treasurehunt, Coordination . . . . . . . . . 64 H.3 BlindGuidingTask.................................. 64 I Studying Coordination Patterns Details 65 I.1 Original Treasurehunt Task . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 65 I.2 Improved Treasurehunt Task . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 65 I.3 BlindGuidingTask.................................. 66 J Kolmogorov-Smirnov Test Details 67 K Normality Analysis for Blind Guiding Task’s Distributions 68 Context and Scope Oriol Miró López-Feliu 8
List of Figures 1 Examples of Minecraft Gameplay, Survival Game Mode[27] . . . . . . . . . . . . 3 2 Basic Diagram of Reinforcement Learning[32] . . . . . . . . . . . . . . . . . . . 4 3 Basic Diagram of Multi-agent Reinforcement Learning[18] . . . . . . . . . . . . . 4 4 VPTMethodOverview[13].............................. 5 5 MineCLIPOverview[23]............................... 5 6 STEVE-1Architecture[38].............................. 5 7 Percentage of successful exits per exit mode. Each experiment consisted of 100 episodes......................................... 12 8 Diagram of Layout 3 for the Treasurehunt dungeon: In blue are the agents’ starting positions, and on green are the possible locations for the exit. . . . . . . . . . . . . 13 9 Percentage of Successful Exits and Items Collected per Layout. Each experiment consistedof100episodes. .............................. 13 10 Single-network Phasic Policy Gradient (PPG) architecture. The policy and value function share a torso network for feature extraction, and separate heads map the latent features to policy and value outputs[60]. . . . . . . . . . . . . . . . . . . . . 16 11 Modified STEVE-1 architecture [Designed with draw.io] . . . . . . . . . . . . . . 19 12 Visualisation of how the matrix of dungeon looks like. Limited by the low render distance of the environment. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 13 Average time per epoch for different optimization methods. Error bars represent the standarddeviation. .................................. 22 14 Collector agent parameter calibration. Prioritizing escape performance, we chose the second prompt with a conditional scale of four. . . . . . . . . . . . . . . . . . . . 24 15 Fighter agent parameter calibration. Focusing on the fighter’s interaction with the environment through inflicted damage, we would select a conditional scale of zero; nevertheless, this would result in the fighter ignoring the coordination messages, therefore we select the first prompt, with a conditional scale of one. . . . . . . . . 24 16 Training losses comparison across agents, phases, and parameter variations (smoothed). 25 17 Losses per learning rate comparison for the first coordination method (smoothed). We selected 2×10−7given it was the highest LR that resulted in stable training . . 26 18 Cumulative reward comparison between the baseline and coordination methods for each agent over N= 100 episodes........................... 26 19 Histograms of cumulative rewards for the baseline and coordination methods, displaying a double Gaussian distribution. . . . . . . . . . . . . . . . . . . . . . . . . 26 20 UMAP-reduced messages for Agent 1. Colored by the subsequent action taken (binary). Successful vs Unsuccesful episodes. . . . . . . . . . . . . . . . . . . . . 27 21 UMAP-reduced messages for Agent 2. Colored by the subsequent action taken (binary). Successful vs Unsuccesful episodes. . . . . . . . . . . . . . . . . . . . . 27 22 F1 score difference on high reward vs low reward episodes, when attempting to classify the action taken based on the (UMAP-reduced) coordination message. . . . 28 23 Training losses for the baseline and coordination methods on the improved Treasurehunttask. ....................................... 30 24 Cumulative reward comparison between the baseline and coordination methods for each agent over N= 300 episodes in the improved Treasurehunt task. . . . . . . . 30 25 Histograms of cumulative rewards for the baseline and coordination methods in the improved Treasurehunt task, displaying a trimodal distribution. . . . . . . . . . . . 30 26 Partial diagram for the Blind Guiding maze. In black: maze borders, in cyan: guide starting position, and in blue: blind starting position. . . . . . . . . . . . . . . . . 32 27 Guide Agent Calibrations. We selected the second prompt, and a conditional scale of 10............................................ 33 28 Blind Agent Calibrations. We selected the first prompt, and a conditional scale of 10. 33 29 Training losses for the baseline and coordination methods on the Blind Guiding Task. 34 30 Box plot of cumulative rewards for the baseline, unidirectional coordination, and bidirectional coordination methods for each agent in the Blind Guiding task. . . . . 34 31 Histograms of cumulative rewards for the baseline, unidirectional coordination, and bidirectional coordination methods in the Blind Guiding task. The prominent peaks in Agent 1’s Unidirectional and Bidirectional Coordination histograms coincide with thefirst“turn”inthemaze............................... 34
Developing Coordination Methods for AI Agents in Open-Ended Environments Figure 2: Basic Diagram of Reinforcement Learning[32] Figure 3: Basic Diagram of Multi-agent Reinforcement Learning[18] 1.1.1.4 STEVE-1 STEVE-1[ 38 ] is the base model that this thesis will use for research, and it is therefore fundamental to understand. First, STEVE-1 is a generative model for Minecraft that can follow open-ended text and visual instructions. It is a SOTA model capable of performing a wide variety of short-term tasks, perfectly aligning with this research. Given how much data was used for its training, the latent capabilities of the model are not yet fully known, including its application to MAS. The technological foundations of STEVE-1 are Video PreTraining (VPT)[13] and MineCLIP[23]: • VPT is an innovative approach developed by OpenAI in the domain of action learning from videos. It focuses on learning how to perform tasks by watching and learning from originally unlabelled online videos in the context of Minecraft. This is an interesting use case, as there are thousands of hours of Minecraft gameplay online, which is rarely annotated. It was first trained using a small dataset of labelled contractor data to obtain an Inverse Dynamics Model (IDM), a model that can be used to predict what actions are taken at each time step in a video. At the same time, there was an important process of data filtering, where 70000 hours of clean (no watermarks, same game mode, etc.) but unlabeled Minecraft gameplay was obtained. The IDM was then used to label the clean set of videos, and the resulting dataset was used for training a model through Behavioural Cloning (BC)[ 48 ]: a technique in which a model learns to perform tasks by directly mimicking the actions demonstrated to it. Figure 4 illustrates this process. • MineCLIP is an innovative agent learning algorithm that uses large pre-trained videolanguage models as a learned reward function for Minecraft. It is, specifically, a contrastive video-language model that correlates video snippets with natural language descriptions, resulting in a “correlation score”. The original paper showed that the correlation score can be used effectively as an open-vocabulary, massively multi-task reward function for RL training, which is incredibly interesting for the topic at hand. Figure 5 shows how MineCLIP works. Moving on to the details of STEVE-1, the VPT model was fine-tuned to achieve visual goals embedded in the MineCLIP latent space, and a prior model was trained to translate text instructions into MineCLIP visual embeddings. The architecture can be seen in Figure 6. The process of fine-tuning VPT involved BC with hindsight relabeling. This is a useful technique in multi-goal RL, where any trajectory taken by an agent can be treated as a suboptimal demonstration to reach a final state, regardless of the original goal[ 66 ]. In the case of Minecraft, given how many possible tasks there are and how challenging it is to achieve a specific one, this approach is particularly useful. The Prior is composed of a Conditional Variational Autoencoder (CVAE)[ 53 ] that interfaces with a Text Encoder and a Gaussian Prior. A CVAE is a type of generative model that learns to encode input data into a latent (hidden) representation and then generates output data conditioned on some additional information. The Text Encoder processes textual instructions (e.g., “chop a tree”) using frozen MineCLIP embeddings to produce a latent text representation zy . These embeddings are GEP Module Oriol Miró López-Feliu 4
Developing Coordination Methods for AI Agents in Open-Ended Environments Figure 4: VPT Method Overview[13] Figure 5: MineCLIP Overview[23] aligned with the latent representations of the visual target goal zTgoal through a CLIP[ 45 ] Objective, which ensures compatibility within the same latent space. The CVAE receives the latent text representation zy and, informed by the Gaussian Prior, generates a predicted latent visual goal zTgoal . This Gaussian Prior is a probabilistic model that encourages the production of visual embeddings that are statistically coherent with the given text embeddings. Figure 6: STEVE-1 Architecture [38] 1.1.2 Problem to be solved Although many methods for agent-to-agent coordination exist in the literature, these methods often rely on assumptions that seldom scale to open-ended problems, require manual modification for new problems, or produce inefficient results. The lack of succinct and off-the-shelf coordination methods is hindering many tasks, such as autonomous driving or the management of smart city infrastructure. We translate this problem to solving tasks in a game environment, on architectures and methods that could, in the future, be extrapolated to new domains. GEP Module Oriol Miró López-Feliu 5
Developing Coordination Methods for AI Agents in Open-Ended Environments 1.1.3 Stakeholders This project involves multiple parties, each with varying levels of involvement and potential benefits. We can categorize these stakeholders into two distinct groups based on their interaction with the project. The first group consists of those directly involved in the research, including the supervisors and the researcher. Víctor Giménez Ábalos, serving as the primary supervisor, and Adrián Tormos, the co-supervisor, bring their expertise in MAS to the project. Their research interests align with exploring the Minecraft environment, and they will provide guidance and mentorship to ensure the project’s successful execution. The researcher, Oriol Miró López-Feliu, is responsible for planning, executing, and documenting the project, as well as conducting experiments, analyzing results, and drawing conclusions. The second group of stakeholders, while not directly participating in the research, stand to benefit from its outcomes. This group can be further divided into two subgroups: companies specializing in the development of AI agents for MAS, who could potentially incorporate the project’s findings into their models, and the scientific community at large. The latter will gain access to the study and its results, which can serve as a foundation for future research endeavors in this field. 1.2 Justification Having established a comprehensive understanding of the theoretical foundations pertinent to this research, we now transition towards articulating its justification within the broader academic context. To do so, we will first introduce relevant previous studies (Section 1.2.1), discussing the current literature and recent advances. This will be followed by an argument for the need for the current research and a justification of the chosen methods and directions (Section 1.2.2). 1.2.1 Previous studies In the field of multi-agent systems[ 64 ], a significant portion of research focuses on the use of LLMs to improve agent coordination and communication[ 26 ][ 17 ][ 7 ]. Previous studies have demonstrated the effectiveness of LLMs in tasks such as watch-and-help (WAH) or in scenarios that require multiple intermediate steps. In contrast, our thesis adopts a different approach, emphasizing direct agent-toagent coordination. This approach has been explored through various methodologies[ 58 ][ 21 ], and despite the existence of research in open environments[ 6 ], it remains very limited in open-ended environments like Minecraft[44]. In the interest of testing MAS in Minecraft, some studies, such as MarLÖ[ 44 ] and VillagerBench[ 10 ], define a series of tasks and benchmarks. Although they do not directly apply to our research, they serve as inspiration for our work. 1.2.2 Justification The primary motivation for this thesis is the necessity of direct coordination mechanisms in MAS, especially in open environments like Minecraft. While LLMs can be used for complex real-time decision-making, their resource-intensive nature limits their practicality in our research. Conversely, direct agent-to-agent coordination provides a more feasible approach, enabling simplified, faster, and more responsive interactions between agents. This is crucial in dynamic and unpredictable environments, such as Minecraft, where delayed responses can result in suboptimal outcomes. Moreover, the findings from this research could potentially be combined with LLMs, leveraging their exceptional reasoning and long-term planning capabilities with the more reflexive and synchronized nature of direct coordination, leading to models that excel at both short and long-term tasks. Minecraft serves as an ideal environment for this research due to its open-world nature, offering thousands of possible tasks and real-world analogous constraints, such as resource management, unpredictable terrain, and the need for strategic planning and adaptation. Furthermore, the growing research interest in this environment enhances the relevance of the potential results. Given the resource constraints, selecting a base agent was crucial, as training a model from scratch in complex environments is exceptionally challenging due to the prolonged learning curve resulting from intricate action and observation spaces. STEVE-1, a SOTA model, was chosen for several key GEP Module Oriol Miró López-Feliu 6
Developing Coordination Methods for AI Agents in Open-Ended Environments reasons. Firstly, using a SOTA model ensures the relevance of the research conclusions. Secondly, STEVE-1’s pre-training on a wide range of tasks significantly reduces computational costs compared to starting from scratch. This is particularly important when solving tasks with high effective horizons, where the agent needs to make a long sequence of correct actions before receiving feedback. In such cases, the probability of discovering the optimal policy through random exploration is extremely low. However, a pre-trained model is more likely to have learned useful features and behaviors that can guide its exploration, making it more feasible to solve tasks with high effective horizons. This is exemplified by the VPT research, where training from a randomly initialized policy failed to achieve almost any reward, while RL fine-tuning from the VPT foundation model performed substantially better. Moreover, STEVE-1’s architecture includes a text embedding component, providing a natural way to incorporate inter-agent communication, which is a central objective of this research. Choosing a modern, SOTA model like STEVE-1 aligns well with the research goals, as it provides a strong foundation for investigating communication methods in a realistic and up-to-date setting. 1.3 Scope Building upon the justification outlined earlier (Section 1.2.2), this Scope section will outline the main goals and sub-objectives, clarifying the extent of this research (Section 1.3.1). It will also detail both the functional and non-functional requirements needed to achieve such goals (Section 1.3.2), and recognize potential risks and obstacles (Section 1.3.3), such as the complexity of algorithm implementation or the possible limitations of models. 1.3.1 Objectives The primary goal of this thesis is to develop direct coordination mechanisms for MAS in complex dynamic environments, with a specific focus on the Minecraft environment. To achieve this overarching goal, we have divided the project into several detailed objectives, each with its associated sub-objectives: 1. Theoretical Understanding and Framework Evaluation: (a) Comprehensive understanding of the “Adaptation of the MineDojo Framework for Use in Multi-agent Reinforcement Learning” and its implications for MAS in Minecraft. (b) Detailed examination of the STEVE-1 model’s architecture and functionality, with an emphasis on adaptability and efficiency in multi-agent scenarios. (c) In-depth analysis of the current SOTA in MAS coordination, focusing on trends, challenges, and opportunities in the Minecraft environment. 2. System Development and Implementation: (a) Effective integration of Multi-agent MineDojo in the Minecraft setting, emphasizing compatibility and efficiency. (b) Adaptation of the STEVE-1 model to the Multi-agent MineDojo environment. (c) Implementation of diverse tasks within the MineDojo environment to facilitate a comprehensive evaluation of MAS capabilities. 3. Development of Coordination Mechanisms: (a) Design of various reward functions to enhance agent coordination. (b) Exploration of multiple architectural approaches for multi-agent coordination, focusing on different structures and interaction models. (c) Refinement of techniques based on performance metrics and outcomes, through model training on defined tasks. (d) In-depth performance evaluation of the system across different tasks, with a focus on coordination mechanism effectiveness. 4. Comparative and Performance Analysis: (a) Comparative assessment of MAS performance against existing SOTA models, with a focus on improvement areas. GEP Module Oriol Miró López-Feliu 7
Developing Coordination Methods for AI Agents in Open-Ended Environments (b) Analysis of the efficiency, scalability, and robustness of the coordination mechanisms developed. (c) Synthesis of findings to formulate conclusions on the effectiveness of developed methods and to propose future research directions. 1.3.2 Requirements To ensure the quality of the thesis, certain requirements must be met. These can be divided into functional and non-functional requirements. Non-Functional Requirements • Use of good programming practices, with a readable style and a maintainable code structure. • Ensure an efficient implementation at every step of the project. • Need for high computing power for training. • Ensure training is correct and agents are learning to coordinate. Functional Requirements • Support a variety of tasks, testing different aspects of agent coordination. • Provide mechanisms for agent-to-agent communication. • Facilitate continuous learning and performance improvement for agents. • If external data is used at any point, ensure its application is correct and serves its purpose. 1.3.3 Potential risks and obstacles Throughout the development of the thesis, several potential risks and obstacles may need to be addressed. • Researcher’s inexperience: As this thesis is the researcher’s first work involving RL and many other technologies, the learning curve may slow down progress more than initially contemplated. • Insufficient Computational Resources: Training multiple agents simultaneously in a complex environment like Minecraft is a very computationally expensive task, which could result in delays at various stages of the project. • Complexity of Multi-agent Coordination: Coordination in MAS is complex, presenting a risk in terms of designing and implementing effective mechanisms and ensuring correct interaction and learning. • Complexity of Algorithm Implementation: Designing and fine-tuning algorithms for a complex task, such as multi-agent coordination, is inherently challenging. There is a risk of encountering unforeseen issues, resulting in delays. • Technical Limitations of Models: There is a risk that the chosen model, STEVE-1, may have technical limitations or may not perform as expected when adapted, which could require significant rework or adjustments. 1.4 Methodology and Rigour To ensure the correct development of the thesis, this final section of Context and Scope will cover the methodology that will be employed. This will concretely be the creation of a MVP, followed by Extension Cycles (Section 1.4.1). Then, to control whether the work is on the correct path and to validate the results, tools and validation methods will be discussed, and an argument will be made on the difficulty of validation in such an exploratory thesis (Section 1.4.2). GEP Module Oriol Miró López-Feliu 8
Developing Coordination Methods for AI Agents in Open-Ended Environments 1.4.1 Methodology Taking into account the investigative character of this thesis, we will adopt an MVP with Extension Cycles methodology, which can be divided into the following: • Initial Development Stage: The project begins with the creation of a MVP, a fundamental version of agent coordination. • Iterative Extension Cycles: After the MVP, the project will enter a series of extension cycles (duration to be decided). These are dedicated to incremental enhancement of agent coordination. Each extension cycle concludes with a thorough evaluation of the newly integrated aspects. This methodology aligns well with the project given how the direction to be explored regarding methods for agent coordination is yet unknown, as it assures having something tangible midway through and allows for creative and innovative approaches afterward. 1.4.2 Monitoring tools and validation We will use a GitHub repository for version control, facilitating the tracking of each developmental stage and protecting the project against possible data loss. The repository will be structured to include a main branch for stable, tested code, and a development branch for ongoing enhancements. The code will be available to supervisors, allowing them to monitor progress at any time and ensure that the work progresses as expected. To assess and validate the results, a series of benchmarks are typically employed. The innovative essence of this work makes this difficult, as there are no direct previous examples to test against. Therefore, in the early stages of the project, we will define reward functions for each task, and these will later be used to compare the different results obtained and validate that the methods are working and that there is improvement. It is fundamental that these reward functions correctly reflect if the agents are coordinating, otherwise all future decisions could be misguided. As per the initial benchmark, it will be calculated before implementing any coordination mechanism; this will ensure that there is a clear starting point to know how well the coordination mechanisms really works in contrast with the agent’s base performance. On top of this, we will arrange a meeting every two weeks with the thesis supervisor and co-supervisor. These meetings will serve to discuss the project’s progress, address any challenges, and ensure that the project remains aligned with its objectives. 1.5 Laws and Regulations In the context of this research project, it is important to consider the relevant laws and regulations that may apply. However, due to the project’s research nature, its non-reliance on external datasets, and the absence of potentially dangerous AI applications, many of these legal frameworks have limited direct applicability. The European Union’s Artificial Intelligence Act (EU AI Act) is a proposed regulation that aims to create a comprehensive legal framework for the development, deployment, and use of AI systems within the EU [ 1 ]. The EU AI Act has just been approved, on the 21st of May 2024, therefore we consider its potential implications. The EU AI Act classifies AI systems into different risk categories, with varying levels of regulatory requirements. Given that this project focuses on fundamental research and does not involve high-risk AI applications, such as those used in critical sectors like education, employment, or law enforcement, it would likely fall under the “minimal” or “limited” risk category. This category only enforces transparency obligations, such as informing a person of their interaction with an AI system, and we will make sure to follow these requirements. Another potential point where regulations may be relevant is the use of pre-existing models and frameworks. Our base works, Multiagent Minedojo and STEVE-1, are based on Minedojo and VPT respectiely, both of which are licensed unde the MIT License [ 3 , 4 ], which allows for the use, modification, and distribution of the software, provided that the original copyright notice and permission notice are included[ 5 ]. The project will ensure compliance with these licensing terms and provide proper attribution to the original creators. GEP Module Oriol Miró López-Feliu 9
Part II MVP This section builds upon the project management preparations to explore the development and evaluation of a coordination technology applied to a single task in the Multiagent Minedojo environment, constituting the MVP. To enable this research, we first address the necessary adaptations and modifications made to the Multiagent Minedojo environment and the STEVE-1 model, ensuring compatibility and resolving initial issues (Section 2). Furthermore, we identify a limitation in the current Multiagent Minedojo tasks, which lack scenarios that truly require active collaboration between agents. To overcome this challenge, we introduce a new task called Treasurehunt, specifically designed to assess multi-agent coordination (Section 3). Through a series of iterative experiments, we construct the task environment and develop an appropriate reward system to guide the agents towards the desired collaborative behaviors. Given the computational complexity of working in the Multiagent Minedojo environment, we focus on selecting an efficient training algorithm based on PPG (Section 4). We propose a novel coordination mechanism that enables effective collaboration between two STEVE-1 agents by allowing them to communicate and condition their behavior on information from each other. However, the combination of these steps results in a prohibitive computational cost, prompting us to devote significant efforts to optimizing our approach (Section 5). Finally, we enter the experiments section, where we calibrate the agents, adjust parameters by training a baseline model, and test our coordination method to validate the effectiveness of the proposed solution (Section 6). The experiments serve as a crucial step in assessing the viability of our coordination technology and its potential for facilitating multi-agent collaboration in complex environments. 2 Environment and Model Preparations We initially believed that both Multiagent Minedojo and STEVE-1 could be utilized with minimal adaptations, requiring only the incorporation of a few additional features. However, we encountered challenges in extrapolating these projects beyond their specific use cases, necessitating extensive modifications and significant time investment. We begin by examining the adaptations made to Multiagent Minedojo and the issues encountered (Section 2.1). Subsequently, we explore the adaptations made to our base model, the STEVE-1, along with the problems discovered and the unforeseen limitations that emerged throughout the course of our research (Section 2.2). 2.1 Multiagent Minedojo Adaptations Although the environment accommodated multiple agents, it did so in a limited manner, not allowing, for example, the issuing of different rewards to individual agents. This restriction constrained the number of possible tasks and specifically precluded the task we propose in Section 3; therefore, we decided to implement a new feature: support for asymmetrical roles within the agents. At first, we believed this was the only adaptation needed, however as we progressed with the project, several obscure issues emerged, producing bizarre behaviors and puzzling us for weeks. The main challenge was that these issues did not display any error messages, making them difficult to track down. For instance, there were several bugs with the internal mechanisms of the environment that caused changes to reflect on the first agent but not on the second one. The most perplexing issue was a memory leak that led to an ever-increasing use of memory, ultimately resulting in a crash. We first suspected that the problem lay within our training code and spent nearly a month experimenting with various approaches and second-guessing our overall strategy, which unfortunately resulted in a significant amount of lost work. However, we eventually identified the issue and promptly resolved it. Appendix E.1 provides a more detailed account of these issues and the steps taken to address them. 10
Developing Coordination Methods for AI Agents in Open-Ended Environments 2.2 STEVE-1 AdaptationS STEVE-1 is a model developed for the MineRL environment[ 28 ], which although similar to Minedojo, exhibits notable differences in the Observation and Action spaces. Fortunately, MineDojo was built on top of MineRL, extending and modifying various aspects of the environment, but maintaining all original functionalities. This results in an injective mapping between the two spaces, where MineRL’s spaces are of a lower dimension, allowing a conversion with no lost information. With this, STEVE-1 was able to function on Minedojo seamlessly, with little internal tweaking. For a more detailed explanation of the space mapping process, and the small tweaks made, please refer to Appendix E.2.1. The model also employs a technique called conditional scaling; as a basic explanation, this involves computing both conditional and unconditional predictions for an agent’s action and combining them using a scaling parameter. Its application proved extremely powerful, improving the performance of the model up to a factor of 15 on certain tasks[ 38 ]. However, a limitation was identified in the original code, as this technique was only implemented for inference, and could not be used for training (it only supported batch size of 1, and in a non gradient-enabled way). Considering the potential benefits, we decided to invest significant efforts on adapting this technique, and succeeded, which proved very beneficial to our research as will be shown later, on Section 6.1. The modifications made were extremely challenging, given the complex internal structure of the model, and the details can be found in Appendix E.2.2 2.2.1 Issues and limitations As happened with the environment, several issues emerged once the project was in a more advanced state. Most of these problems stemmed from the fact that STEVE-1 was trained on top of VPT using a different methodology than ours. While they employed BC fine-tuning, we utilized RL fine-tuning, which meant that certain parts of VPT that we needed were not adapted by STEVE-1. Although many of these issues were straightforward to identify and resolve, some proved to be more challenging, particularly those that did not display any errors. One notable example involved two distinct functions within the STEVE-1 code responsible for obtaining the model’s actions. The first function had gradients disabled for episode collection, while the second had gradients enabled for training. At a certain point, as will be detailed in the training algorithm (Section 4.1), we needed to work with the outputs of both functions. However, deep within the code, one of the functions underwent an additional processing step. This discrepancy led to occasional model losses that were several orders of magnitude greater than expected, causing instability in the training process. Initially, we attributed this issue to our training code and spent two weeks adjusting hyperparameters, only to later discover that the problem originated from the STEVE-1 code itself. After an exhaustive investigation, we identified the error, which was then easily resolved. Despite our efforts to adapt STEVE-1 to our requirements, we found that the model was less “plugand-play” than initially anticipated. Its behavior was heavily dependent on prior training, which led to unexpected outcomes in certain scenarios. For instance, as discussed in Section 3.1, when prompted to collect seeds the model would often disregard its observations and randomly punch the ground, even when plants were in close proximity. Furthermore, STEVE-1 demonstrated poor generalization capabilities when presented with novel prompts. Consequently, we were compelled to modify our tasks to align with the model’s proven abilities, as detailed in Section 3. 3 Treasurehunt Task The current Multiagent Minedojo lacks tasks that necessitate active participation from multiple agents, as demonstrated in Appendix D. To address this limitation we introduce Treasurehunt, a collaborative task inspired by the MarLÖ competition. Treasurehunt involves two agents with asymmetric roles working together to explore a randomized dungeon layout, gather treasures (1 to 3 per run, represented as objects on the floor), and reach the exit while the fighter protects the vulnerable collector from enemies (1 to 2 per run, either a Zombie or a Skeleton). The agents are given a maximum of 250 steps to complete the task, to reduce memory constraints. Treasurehunt evaluates coordination by assessing the fighter’s capacity to maintain spatial proximity to the collector, react swiftly to threats, and prioritize immediate dangers. The randomized environment MVP Oriol Miró López-Feliu 11
Developing Coordination Methods for AI Agents in Open-Ended Environments ensures the development of adaptable coordination strategies rather than reliance on memorization. Success is determined by the collector’s ability to gather treasures and reach the exit safely, and the fighter’s effectiveness in providing protection and eliminating threats. We focus on the task environment design, selecting an appropriate escape mode and dungeon layout, and evaluating the collector’s performance to validate the task feasibility (Sections 3.1 and 3.2). Furthermore, we address reward shaping to guide the agents towards desired behaviors and promote collaboration (Section 3.3). 3.1 Escape mode Choosing an appropriate exit is key for our layout, for several reasons: 1. Our task is centered around the collector’s search for the exit, and we reward or punish our agents according to their performance; were the collector largely incapable of recognising the exit, agents would fail to learn further. 2. The collector must have a clear understanding of it for any coordination mechanisms to be meaningful; as will be detailed in Section 4.2, we base our approach on “messages” between agents. A poor understanding of its objective will lead to poor and misleading communication, possible resulting in the unlearning of known behaviours (garbage in, garbage out). Figure 7: Percentage of successful exits per exit mode. Each experiment consisted of 100 episodes. The exit should be something STEVE-1 recognizes and can achieve with ease. Given the importance of this decision, we decided to test different possible exits modes based on tasks proved achievable on the STEVE-1 paper:[ 38 ] “break dirt”, “find water”, “gather wood”, “break seeds”, “break leaves”. Ideally we would experiment with other parameters too, but given the combinatorial complexity this would entail, we instead chose the best parameters and prompts for the equivalent tasks in the STEVE-1 paper. For our tests, We designed a test dungeon and allowed a low number of steps per episode, to avoid the agent from accidentally achieving success. The results can be seen in Figure 7: The results clearly indicate that the “gather wood” mode of escape was the most successful. Interestingly, the performance on tasks previously shown to be attainable in the STEVE-1 paper was remarkably subpar; visually inspecting the results this was due to the model’s over reliance on prior knowledge. For example, for “break dirt”, the agent completely ignored its observations and instead directly started breaking the floor (which was not dirt). Fortunately, this overreliance is not ubiquitous across all tasks, and in the “gather wood” scenario the model demonstrates the ability to actively search for and locate its objective. Extra details for this subsection can be found in Appendix F.1 MVP Oriol Miró López-Feliu 12
Developing Coordination Methods for AI Agents in Open-Ended Environments 3.2 Dungeon layout Another key aspect of our task is the design of the dungeon layout, which must strike a balance between being sufficiently challenging to necessitate agent coordination and being achievable enough to allow the agents to receive rewards and learn effectively. We experimented extensively with this, beginning with a complex layout, and iteratively constructing a new, simpler one if the agents struggled too much. While it would have been possible to train the agents to solve more challenging dungeon layouts, we decided against it to maintain our focus on agent coordination rather than task-specific solutions. Additionally, our limited computational resources restricted the amount of training we could perform, leading us to prioritize training on coordination methods over task-solving itself. For our tests, we measured a layout’s difficulty by the collector’s capacity to complete it. Our task requires coordination, therefore we do not spawn the fighter nor enemies, to simulate an “ideal scenario”, where the fighter perfectly protects the collector. For completeness, we also record the percentage of treasure the collector manages to collect. Our objective was to achieve a ≥50% success rate, for which we had to construct three separate dungeons. The selected layout and the results can be see in Figures 8 and 9 Figure 8: Diagram of Layout 3 for the Treasurehunt dungeon: In blue are the agents’ starting positions, and on green are the possible locations for the exit. Figure 9: Percentage of Successful Exits and Items Collected per Layout. Each experiment consisted of 100 episodes. The data clearly demonstrates the progressive improvement in the collector’s performance as the dungeon layout was simplified, and our desired success threshold was met on the third layout. To demonstrate that this layout still required coordination and was not overly simplistic, we also incorporated the results for the third layout in a complete setting, which included the fighter agent and lurking enemies. 3.3 Reward Shaping We conclude the construction of the Treasurehunt task by addressing the different rewards agents receive for their interactions with the environment. This is a critical component of RL, and can significantly impacts training speed and results; important factors are the density of reward[ 41 ], or how well these reflect desired behaviour (in our case, cooperation between agents). With these principles in mind, we designed the reward functions shown in Tables 1 and 2 (different for each agent, due to their asymmetrical roles). A major obstacle was aligning the fighter’s actions with its overarching goal of supporting the collector; early experiments showed the fighter often wandered, fighting mobs not attacking the collector, while the latter struggled. To address this, we introduced rewards to encourage the fighter to stay close to the collector, and a scaling parameter decay factor, used to weight the fighter’s rewards based on their impact to the collector. As a proxy, we define the importance as the distance from the object of the reward to the collector, and compute the decay factor as: decay factor = exp−λ×d(1) MVP Oriol Miró López-Feliu 13
Developing Coordination Methods for AI Agents in Open-Ended Environments The graph-based Coordination Module takes the latent outputs from all agents Ht+T={hi t+T|i∈ {1, . . . , N}} and the adjacency matrix A as input, and processes them using graph convolution operations[33]: H′ t+T←GraphConvϕ(Ht+T, A) The updated latent output for each agent is then processed by a separate feed-forward neural network to generate the message vector: mk t+1 ←FFNϕk(h′k t+T) This method offers improved scalability, as the computational complexity is determined by the number of edges in the graph rather than the number of agents, and provides greater flexibility in modeling various communication topologies. Furthermore, the edges can be dynamic, allowing for adaptability to changing conditions. For instance, in the context of autonomous vehicles, an edge could be contingent upon the distance between the vehicles. It is important to note that the authors’ understanding of graph neural networks is limited, and as such, the proposed approach may be subject to inaccuracies or oversights. 4.2.2 Generalizability The proposed coordination method takes advantage of the existing STEVE-1 architecture, significantly reducing the computational resources required for training and testing by utilizing the pre-trained weights of STEVE-1’s linear layer. This approach allows for efficient incorporation of the message from the other agent without extensive fine-tuning, enabling focus on the development and optimization of the coordination module. Nevertheless, the fundamental concepts and design principles of the coordination method are transferable and can serve as a foundation for various architectures and domains; just as STEVE-1 successfully fine-tuned VPT to be conditioned by prompts from scratch, it is conceivable that, given adequate computational resources, another model could be fine-tuned to be exclusively conditioned using our proposed method. 5 Optimizations Optimization was crucial in this research due to the limited computational resources available and the heavy computational requirements of the Multiagent Minedojo environment. During our initial training attempts, we encountered a significant computational barrier that hindered our research progress, with each training epoch requiring over six minutes to complete. This made it impractical to conduct the thousands of epochs necessary for our study. With multiple agents, the computational load multiplies, making it essential to find ways to reduce resource usage and improve efficiency. Moreover, considering the environmental impact of high computational demands, optimizing the system also contributes to sustainability by reducing electricity consumption. Given these constraints and potential benefits, we decided to invest time in optimizing various aspects of the system to make the research feasible. After an in-depth study of our approach, we identified three potential points of optimization: our training algorithm, memory management, and environment loading. These, and their respective optimizations, will be discussed in detail below, followed by an analysis of the results. 5.1 Baseline and Challenges Initial training experiments revealed that approximately 15GB of GPU memory would be required to accommodate all necessary models and data. However, with only 8GB of GPU memory available, this resulted prohibitive. As a naive solution, some components were moved to the CPU, allowing training to proceed. Unfortunately, this approach resulted in very slow performance, as models and data were frequently accessed from outside the GPU; This approach marks our baseline. 5.2 Training Optimizations To simplify the computational load and reduce memory usage, we made some changes to the training procedure. First, we employed a single-network architecture for the PPG algorithm, instead of using separate networks actor and critic within an agent. This approach, previously described in Section MVP Oriol Miró López-Feliu 20
Developing Coordination Methods for AI Agents in Open-Ended Environments 4.1.1, allowed us to utilise one network per agent instead of two. Another optimization was regarding the previously introduced KL divergence penalty, which helps ensure policies do not deviate too much from the original. It is specially important in the case of fine-tuning, given we assume our original model to be, up to a great extent, “correct’. As in the original VPT paper, we computed the KL divergence by maintaining an “original model” per agent. Given this “original model” does not “learn”, and we start off the same model for both agents in our fine-tuning, another training optimization was using the same “original model” for both agents, therefore reducing the needed networks from 2 to 1. Overall, with these training optimizations, we reduced the number of networks used from 6 (actor, critic, original, per agent) to 3 (single-network actor-critic per agent, shared original). Each STEVE-1 model has 248M parameters[ 38 ], and so this results in 248M×3 = 744M less parameters in memory, and were we updating the whole model when fine-tuning, 248M×2 = 496M less parameters to update. This makes training more efficient, allowing us to train for more episodes, and simpler, hopefully allowing our models to converge faster. 5.3 Smart memory management Despite reducing the amount of models we needed, our whole training procedure still did not fit into our GPU memory. This made training very slow, because data and models outside of the GPU memory are much slower to operate with. Therefore, we thought that by smartly moving components (data and models) around, so they are on the GPU when needed, we could greatly improve performance. To effectively implement this memory management strategy, we utilized memory profiling tools to monitor the memory usage at different stages of our code, identifying when precious GPU memory was being wasted. We then moved data and models that were not actively being used to the much slower CPU, and only transferred back to the GPU when needed for computation. This approach required careful planning and tracking of the data flow to minimize the number of transfers between GPU and CPU, as these transfers can be time-consuming and incur a high number of overheads. For the last part, we also restructured the code when possible, so the parts requiring the same components were closer, resulting in less data transfers. We also manually delete variables and executed the garbage collector, to ensure no unnecessary data is stored. 5.4 Environment optimization Figure 12: Visualisation of how the matrix of dungeon looks like. Limited by the low render distance of the environment. Our previous optimization still proved slow, and after much thought we identified another bottleneck: resetting the environment. This process is incredibly time-consuming, and can sometimes take up to several minutes. The time requirement is especially significant in Multiagent Minedojo, having to create a separate Minecraft instance for each agent. For this reason, Minedojo provided a mechanism for a faster reset which is reportedly 100 times faster than normal reset[ 23 ], albeit at the cost of several disadvantages, such as structures being destroyed. Given the inescapable frequency resetting the environment (when collecting episodes for training, calculating performance, ...), we decided to invest some time making “fast reset” feasible. The key to our optimized “fast reset” lay in the amount of dungeons we create during every normal reset: instead of creating creating a world with a single dungeon, we generate a matrix of dungeons within the world, and teleport agents around when an episode finishes. The technical details of this are complex, and are included in Appendix G for reference. The resulting world structure resembles a Rubik’s Cube, as illustrated in Figure 12. In our implementation, we generate a total of 125 unique dungeons within each world. Although it is possible to create a larger number of dungeons, we encountered internal library errors that necessitated resetting the environment approximately every 100 episodes. Consequently, generating additional dungeons MVP Oriol Miró López-Feliu 21
Developing Coordination Methods for AI Agents in Open-Ended Environments beyond this limit would not be beneficial. In the event that the agent reaches the final generated structure, a manual reset is triggered. 5.5 Results The optimizations implemented in this research have significantly improved the efficiency and feasibility of training. To evaluate the impact of each optimization, we tested the average speed per training epoch across 100 epochs for each method, and present the results on Figure13. Figure 13: Average time per epoch for different optimization methods. Error bars represent the standard deviation. The cumulative effect of these optimizations is evident in the nearly 10-fold reduction in the average time per epoch, from 366.41 seconds in the original implementation to just 39.51 seconds with all optimizations applied. To put these results in perspective, with the original method it would take 101.78 hours, over four days, to execute 1000 training epochs, while with the fully optimized version, it takes merely 10.98 hours. Moreover, the consistently lower standard deviations, which we attribute to the more better control of memory resources, indicate that these not only reduce the average training time but also contribute to more stable and predictable performance. 6 Experiments Having completed the necessary preparations, we now proceed with the experiments. The first step is to calibrate some STEVE-1 parameters for each agent (Section 6.1). This is important to avoid unnecessarily training our models on tasks they could already perform well. Next, to validate our results, we establish a baseline by fine-tuning both agents with the same parameters on the same task, but without any coordination technology (Section 6.2). Despite our focus being on coordination methods, we must train a baseline model to guarantee any increase of performance is exclusively due to our coordination methods, and not because of the extra training. In this section, we also experiment with learning parameters, as our single-network PPG methodology introduces new hyperparameters and modifies others. However, most of these parameters should be translatable from the VPT paper. We then move on to training our agents using the best configuration from the baseline, incorporating the coordination technology (Section 6.3), where we also study the coordination language between the agents, and finally we provide an overall analysis of the results (Section 6.4). 6.1 STEVE-1 Calibration Calibrating the prompt and conditional scale is crucial for optimal performance in the STEVE-1 model. The prompt guides the model’s behavior, and the STEVE-1 paper demonstrated that effective prompt engineering can lead to significant performance improvements, such as a three-fold increase MVP Oriol Miró López-Feliu 22
Developing Coordination Methods for AI Agents in Open-Ended Environments for some tasks like obtaining dirt. We adopt a similar approach to maximize our model’s capabilities. The conditional scale ( λ ) balances the model’s goal-conditioned and prior behaviors during inference. The authors show that tuning λ can improve performance by up to 15 times for some tasks, with the best results obtained for values between 5.0 and 7.0. We will also tune λ to ensure optimal results, considering the differences in our tasks’ nature. Ideally, we would calibrate both agents together, but this is computationally prohibitive. For each agent, we test λ in the range [0,10] with four different prompts, resulting in 40 experiments per agent, each 100 episodes long. This large experiment, consisting of 8000 episodes in total, took nearly two days to run and is at the limit of our computational capacity. Calibrating both agents together would require (10 ×4)2= 160,000 episodes, which is infeasible for us. 6.1.1 Fighter and Collector Calibration The collector agent’s task is described as collecting items from the floor then breaking a tree. In previous work [ 38 ], it was shown that redundancy and diversity of expression in the description of the task used in calibration impacts positively in performance, which is believed to be caused by natural language’s diversity. We propose prompts 1 through 4 for collector calibration, each more extensive than the last: 1. “pick up items on floor, collect things around you, grab stuff off ground then look for tree, find tree for chopping, break tree for wood after gathering items” 2. “collect treasure, collect dropped items, then chop down the tree, gather wood, pick up wood, chop it down, break tree” 3. “gather the objects, collect the items you see, pick up the things scattered around, then look for a tree and chop it down, find a tree and break it to get wood, punch a tree to gather wood” 4. “collect what’s on the floor, pick up the objects scattered around, gather the items you find on the ground, then search for a tree and chop it, locate a tree and break it for wood, find a tree and punch it to get wood” The fighter agent’s calibration is similar to the collector’s, with the addition of investigating whether the model can recognize the other agent, which has not been tested in previous work [ 23 , 38 ]. If successful, this could be advantageous for communication between agents. To test this hypothesis, we crafted four prompts, organized into two pairs (1, 3) and (2, 4), where one prompt in each pair references the other agent, and the other does not: 1. “punch mobs, punch enemies, look around for enemies, fight zombies and skeletons” 2. “punch zombies, punch skeletons, attack hostile mobs, kill zombies, kill skeletons” 3. “punch mobs, punch enemies, look around for enemies, fight zombies and skeletons, defend your teammate, defend the other player, protect the other player” 4. “punch zombies, punch skeletons, attack hostile mobs, kill zombies, kill skeletons, defend your teammate, defend the other player, protect the other player” 6.1.2 Results To evaluate the effectiveness of each prompt and conditional scale combination, we measured the percentage of successful exits achieved by the collector agent and the percentage of items collected along the way. As our goal is to optimize the collector’s isolated actions, we do not spawn any enemies and instruct the fighter agent to “wander around”, to still account for its interference, up to an extent. For the fighter agent, we computed three different metrics: damage received by the collector, damage dealt by the fighter, and damage dealt to the collector by the fighter. In this case, we enable enemy spawning and instruct the collector to “wander around” to simulate its search for the exit. The results are presented in Figures 14 and 15, which depict the optimal prompt and conditional scale for the collector and fighter agents, respectively. Interestingly, our findings diverge from those reported in the STEVE-1 paper, where researchers consistently observed improved performance with higher conditional scales. The collector agent’s escape performance improved significantly with the chosen prompt and conditional scale combination, achieving a 61% average escape success rate, a 10-point improvement over previous experiments, while the fighter agent’s results were more stochastic. MVP Oriol Miró López-Feliu 23
Developing Coordination Methods for AI Agents in Open-Ended Environments Figure 14: Collector agent parameter calibration. Prioritizing escape performance, we chose the second prompt with a conditional scale of four. Figure 15: Fighter agent parameter calibration. Focusing on the fighter’s interaction with the environment through inflicted damage, we would select a conditional scale of zero; nevertheless, this would result in the fighter ignoring the coordination messages, therefore we select the first prompt, with a conditional scale of one. 6.2 Baseline Experiments To validate our results and establish a point of comparison, we set up a baseline by fine-tuning both agents with the same parameters on the same task, without any coordination technology. We also experiment with learning parameters, as although most of the parameters from the VPT paper should be translatable, our single-network PPG methodology introduces new hyperparameters and modifies others. Our preliminary experiments, which were based on techniques from the VPT paper (as detailed in Appendix H.1.1), encountered challenges such as exploding gradients and a lack of learning progress. We hypothesize that these difficulties may stem from the extrapolation of the model to a multi-agent setting and a fundamentally different task, necessitating a more cautious approach to model updates. To address these issues, we introduced two key techniques: partial network freezing and learning rate scheduling. The concept of freezing was inspired by a recent paper on parameter-efficient fine-tuning [ 57 ], which demonstrated the effectiveness of fine-tuning only specific components (such as LayerNorm) of a pre-trained transformer-based model. However, this approach still led to some instability, prompting the introduction of the learning rate scheduler. To validate our findings, we trained our models on 1,000 episodes of data, and the results are presented in Figure 16. As can be observed, the initial two training experiments resulted in complete failure due to exploding gradients. However, in the third experiment, we successfully stabilized the losses, with the loss exhibiting apparent convergence. MVP Oriol Miró López-Feliu 24
Developing Coordination Methods for AI Agents in Open-Ended Environments Figure 16: Training losses comparison across agents, phases, and parameter variations (smoothed). 6.3 Coordination Experiments We now introduce the coordination technology and extend the training to 2000 episodes; although the baseline converges in fewer episodes (1000), preliminary experiments with the coordination technology revealed that the models required significantly more episodes. As we modified STEVE-1’s conditioning mechanism, a primary concern was the potential for our changes to interfere with its functionality. To address this issue, we implemented the following measures: • The weights of the coordination module were initialized to zero to mitigate any early destructive effects on the model’s performance. • A penalty was introduced for large messages to encourage the model to filter out noise and prioritize the transmission of relevant information. The penalty, however, resulted in the agents not communicating at all, as it resulted in a complete absence of messages at the end of training, which made us discard this technique. During the the experimentation, a series of issues not related to the communication per-say were discovered. This included things such as some parameters causing messages to be ignored, and the learning rate for the baseline and coordination model causing severe instability. Eventually, we only explored the LR as a hyperparameter (Figure 17), as further testing proved too costly (30h per training run). We consider the selected combination good enough for our purposes, as our objective is not to solve the task but test the effects of coordination mechanisms in working (not necessarily optimal) models. The training details, including losses and hyperparameters, are provided in Appendix H.1.2. Although some of these details were mentioned in the previous section to justify our design choices, we have omitted them from the main text to maintain focus on the key aspects of our approach. MVP Oriol Miró López-Feliu 25
Developing Coordination Methods for AI Agents in Open-Ended Environments Figure 17: Losses per learning rate comparison for the first coordination method (smoothed). We selected 2×10−7given it was the highest LR that resulted in stable training 6.4 Results and Analysis To compare the baseline and coordination methods, we ran both configurations for 300 episodes and plotted the cumulative reward results (Figure 18), which show the reward distribution for each agent under both settings. In addition, we plotted the histograms of the cumulative rewards to better understand their distribution characteristics (Figure 19). The histograms reveal a bimodal distribution, with the two modes separated by the reward difference between successful and unsuccessful episode completions. Figure 18: Cumulative reward comparison between the baseline and coordination methods for each agent over N= 100 episodes. Figure 19: Histograms of cumulative rewards for the baseline and coordination methods, displaying a double Gaussian distribution. To assess whether the coordination method offers a statistically significant improvement, we conduct a one-sided two-sample Kolmogorov-Smirnov (KS) test, as it is suitable for comparing bimodal distributions. The KS test compares the Empirical Cumulative Distribution Functions (ECDFs) of the two samples and is sensitive to differences in both location and shape of the distributions. The null hypothesis ( H0 ) states that the coordination method does not lead to higher rewards, while the alternative hypothesis ( H1 ) states that it does. Please see Appendix J for the mathematical details. The KS test results and effect sizes are summarized in Table 3. Based on the results, with p-values greater than the Bonferroni-adjusted significance level α/2 = 0.025 for both agents, we fail to reject the null hypothesis ( H0 ) at a 95% confidence level. This indicates insufficient evidence to conclude that the coordination method leads to significantly higher cumulative rewards compared to the baseline. The medium effect sizes for Agent 1 (0.461) and Agent 2 (0.321), according to Cohen’s guidelines, suggest that the differences between the baseline and coordination methods are moderate. MVP Oriol Miró López-Feliu 26
Developing Coordination Methods for AI Agents in Open-Ended Environments Table 3: Kolmogorov-Smirnov test results and effect sizes for Agent 1 and Agent 2 (Bonferroniadjusted) Agent Test Statistic p-value Effect Size Significance Agent 1 0.0167 1.8403 0.461 NS Agent 2 0.0067 1.9736 0.321 NS NS: Not significant at Bonferroni-adjusted significance level α/2 = 0.025. Although the coordination method did not demonstrate a significant advantage over the baseline, it did not perform worse either. In the next section we will examine the nature of this communication protocol and its impact on the agents’ behavior in-depth. 6.4.1 Studying Coordination Patterns We hypothesize that the bimodal distribution in Figure 18 is due to two distinct coordination patterns, potentially influenced by factors such as mutual gaze between agents. To investigate this hypothesis and determine if message content differs significantly between successful and unsuccessful coordination, we conducted a comparative analysis between high-reward (successful) and low-reward (unsuccessful) episodes. Evaluating message content is challenging due to the multifaceted nature of the information they may contain. As an indirect method, we investigated whether agents take actions based on received messages. Using UMAP for dimensionality reduction (found to be the most interpretable) and binary visualization for each action type (agents can perform multiple actions each step), we found distinct patterns in high-reward episodes as evidenced by clusters of points , while low-reward episodes exhibited only noise (Figures 20 and 21) Figure 20: UMAP-reduced messages for Agent 1. Colored by the subsequent action taken (binary). Successful vs Unsuccesful episodes. Figure 21: UMAP-reduced messages for Agent 2. Colored by the subsequent action taken (binary). Successful vs Unsuccesful episodes. To rigourously test our hypothesis, we classified actions based on coordination messages and compared F1-scores between low-reward and high-reward episodes. We used UMAP to reduce coordination messages to a 5-dimensional space and employed a Random Forest classifier (chosen for its non-linear capacity). The results (Figure 22) showed significantly higher F1 scores in high-reward episodes across most actions for both agents, indicating that coordination messages during successful collaborations contain more meaningful information. The consistency in F1 scores for the collector’s “attack” and the fighter’s “movement” actions suggests stable core behaviors. However, differences in F1 scores for the collector’s “camera” and the fighter’s “attack” actions in high-reward episodes suggest potential strategic information sharing. The minimal impact on infrequent actions implies prioritization of mission-critical task communication. These findings support the hypothesis of two distinct coordination patterns, with high-reward episodes featuring more purposeful and effective communication, although further analysis is needed to identify the causes of each pattern. To complement our analysis, we also explored other patterns, such as the evolution of coordination messages throughout the training process. However, these investigations did not yield results as compelling as those presented in the main text. For the sake of completeness and transparency, we have included these additional findings in Appendix X. MVP Oriol Miró López-Feliu 27
Developing Coordination Methods for AI Agents in Open-Ended Environments Figure 22: F1 score difference on high reward vs low reward episodes, when attempting to classify the action taken based on the (UMAP-reduced) coordination message. 6.4.2 Potential Reasons for Suboptimal Coordination Performance Given the suboptimal results, we explore four potential reasons for the lack of improvement in the agents’ collaborative capabilities. Firstly, the computational power required to train the agents with the coordination technology may be insufficient due to the increased complexity introduced by the coordination module. Our current setup, which already pushed our PC to its limits, may not be able to handle the noise generated by the agents learning to generate and interpret messages effectively. Secondly, upon closer examination, it appears that the Treasurehunt task itself may have inherent issues that hinder the agents’ performance. The task contains numerous sources of randomness, such as a variable number of treasures, which can make it challenging for the agents to develop consistent strategies. Additionally, the tight time constraint imposed by the task may prevent the agents from exploring more sophisticated collaborative behaviors. This hypothesis is supported by a visual inspection of the results, which revealed that the agents’ primary strategy was to head directly for the exit, likely due to the fear of running out of time. Thirdly, the inherent complexity of our task may be too high, with rewards being too sparse or the number of reward signals too large. The agents need to navigate the dungeon, collect treasure, combat enemies, and coordinate their actions to achieve the ultimate goal of escaping. The presence of multiple reward signals, such as treasure collection, damage received, damage inflicted, successful escape, etc., may overwhelm the agents. In such cases, a curriculum learning approach, or a simpler task, might be more suitable for facilitating the learning process. Fourth, our coordination method may interfere with STEVE-1’s prompt conditioning, potentially leading to catastrophic forgetting or causing our messages to be filtered out. Interfering with the pre-trained STEVE-1 model’s heavily tuned prompt conditioning risks disrupting its performance. Due to our current setup, computational constraints, and personal circumstances, addressing the lack of computational power (problem 1) is not feasible. Sourcing outside compute would require extra financial investment and significant time and effort to prepare the complex Multiagent Minedojo environment with its numerous package versions and dependencies, which could be further complicated by potential constraints imposed by the outside compute provider. Therefore, in the upcoming extension cycles, we will focus on resolving the issues in increasing order of complexity, tackling as many as time allows. Due to the limited timeframe, we were only able to address the inherent limitations of the Treasurehunt task (problem 2) and task complexity (problem 3), as discussed in sections X and Y, respectively. MVP Oriol Miró López-Feliu 28
Part III Extension Cycles The MVP exposed deficiencies in agent coordination within the Treasurehunt task. We propose two extension cycles to address these limitations. First, we will modify the Treasurehunt task’s time limit, reward structure, and treasure distribution to foster improved coordination (Section 7. Second, we introduce the Blind Guiding Task, a novel task designed to test and enhance agent coordination in a controlled setting (Section 8. 7 Improving the Treasurehunt Task In our analysis of the Treasurehunt task’s potential deficiencies (Section 6.4), we identified that the imposed time limit per episode was overly restrictive, hindering the agents’ ability to effectively coordinate. This led to the collector agent prioritizing reaching the escape point over collecting treasure towards the end of the training process. Moreover, the agents’ lack of awareness regarding the remaining time contributed to unstable training, further exacerbated by the variable number of treasures (1 to 3 per run), resulting in higher rewards for some runs despite similar agent behavior. We implemented three modifications to tackle the challenges: 1. Increased the time per run from 250 to 450 steps, the maximum allowed by our memory constraints, enabling agents to explore more strategies and potentially develop active coordination. 2. Modified the reward structure to encourage consistent treasure collection throughout the episode while considering time limitations by: • Increasing the reward per piece of treasure collected. • Scaling the reward based on the percentage of time remaining in the episode, discouraging agents from disregarding treasure due to time concerns. 3. Stabilized the rewards by fixing the amount of treasure per episode to 2, ensuring a consistent and predictable reward environment. 7.1 Baseline and Coordination Training The baseline and coordination training experiments were conducted without modifications to the established methodology, using the same hyperparameters and architectures described in Sections 6.2 and 6.3. Analysis of the training losses (Figure 23) reveals that the baseline model converged, while the coordination model did not, due to differences in learning rates. It is important to note that, despite more attempts on our part, the coordination method’s learning rate could not be increased further without causing the collapse of the STEVE-1 conditioning mechanism. The apparent lack of convergence in the coordination method’s overall training losses may not necessarily indicate a lack of learning or improvement. A closer examination of the local behavior reveals a logarithmic decrease, suggesting that the coordination model is still learning, albeit at a slower pace compared to the baseline. Given sufficient training time and computational resources, the coordination model could potentially achieve convergence and exhibit enhanced performance. Interestingly, the policy function exhibited better convergence for the baseline model, while the value function seemed to perform better with the coordination technology. This could indicate that the coordination method is very useful estimating the value of states and actions in a collaborative setting, but there are still challenges to overcome in order to learn a stable policy. 29
Part IV Discussion and Conclusions 9 Discussion These results show our method’s capacity to enable successful inter-agent communication in scenarios where coordination was necessary, severely outperforming the baseline, while employing pure RL techniques. We successfully integrated our method into STEVE-1, a SOTA model with no original capabilities to coordinate, with only slight modifications to its architecture. However, results were inconclusive when coordination was only beneficial to the tasks’ outcome, a discrepancy we attribute to the tested tasks’ inherent complexity, coupled with our limited resources, as we likely necessitate further training and task convergence. We believe slight modifications to our method, such as a preliminary layer which ascertains whether coordination is necessary at each moment, would mitigate these issues. Perhaps the most notable feature of the proposed method is its flexibility and adaptability, requiring no domain nor task-specific modifications, and as some evidence suggested, shaping communication patterns in task-specific ways. Yet, as is often the case, such flexibility comes at the expense of interpretability and control, as the agents’ language remains obscure. We hypothesize that messages contain information such as the current state, goals, intentions, or planned actions. However, without a means to interpret the messages, it remains challenging to verify this hypothesis. Some results even suggested agents could be utilizing different communication patterns within the same task, albeit to varied success, and further work is needed to determine what situations, or task characteristics, lead to successful communication. Intriguingly, our findings indicate that bidirectional coordination, where both agents send and receive messages, may enhance coordination compared to unidirectional communication, where there are designated senders and receivers, even in scenarios where the necessity for bidirectional coordination remains ambiguous. This suggests that agents engage in more complex interactions than merely reacting to received messages. It is plausible that agents provide feedback, request assistance, or discern when their counterparts require guidance. However, these interpretations are speculative and warrant further investigation to elucidate the precise nature of the agents’ interactions. This research also presents intriguing results regarding the applicability of SAS to multi-agent scenarios, as the step in between is of exponential difficulty. We speculate our method could allow the transition of SAS to MAS (as we did with STEVE-1), even seeing the former’s training as a stepping stone to the latter, in the same way BC is often the start of SAS’ training. Other results also suggest our method could be beneficial when training MAS using actor-critic paradigms, as we found it to be especially beneficial for estimating the value of each action. This should be compared to other MAS training methodologies, such as centralised training and decentralised execution, where observations are shared, but we hypothesize our method offers a more flexible and refined acknowledgement between agents. It is important to note the training challenges inherent to multi-agent systems. The baseline and communication-enabled agents may have different training requirements, leading to skepticism about the fairness of their comparison. We recommend exploring different training strategies, such as training the policy first and then the communication, or training both policy and communication simultaneously, to determine which approach yields better results. Further research is needed to establish the best practices for training multi-agent systems with inter-agent communication capabilities. 10 Conclusion We have argued that simple communication between agents can be performed through allowing internetwork embedding passing, which we applied by adding the information from a “communicationhead” of an agent to a received task-description textual embedding of the other agent. This can be seen as a modifier of the task description received, a feature used in many recent models [ 40 , 59 , 61 ], and is a rather intuitive way of procuring communication. Furthermore, it allows agents to share 36
Developing Coordination Methods for AI Agents in Open-Ended Environments information and coordinate their actions without requiring extensive modifications to their underlying architectures. The results indicate that our method is highly effective in scenarios where coordination is necessary, although they were less conclusive when coordination is only beneficial but not crucial. We believe that such a gap in performance is due to the intrinsic randomness and complexity of the task tested, which would probably ameliorate with further task convergence and training, or slight modifications to our method, but that remains to be seen. Despite these limitations, the flexibility and adaptability of our method are promising, as it required no task-specific modifications and the exchanged messages were used in varied ways by the agents. We believe our approach could be applied to a wide range of multi-agent scenarios, such as enabling swarms of drones to coordinate their actions for search and rescue missions, facilitating communication between autonomous vehicles to optimize traffic flow and prevent collisions, or allowing robots in a factory setting to collaborate on complex assembly tasks. Moreover, it could enable diverse SAS to function in a multi-agent setting. However, albeit the results are indicative of good performance through this simple methodology, further experimentation with other architectures and tasks, as well as analysis on the passed messages is required to fully understand the capabilities and limitations of our method, as we will discuss ahead. 11 Future Work The research presented in this thesis has opened up several potential directions for future exploration. We discussed many of these in the previous Discussion (Section 9), such as our method’s potential usefulness for actor-critic paradigms, proposals of slight modifications to ensure our method’s always used in a “needed” way, etc. Even then, we believe there could be other interesting applications, albeit outside the scope of our thesis, as we will present in the following paragraphs. Despite our method’s focus on not using LLMs, it would be interesting to explore its use in conjunction: using our method for low-level coordination while LLMs issue higher-level orders (such as modifying STEVE-1’s prompts). Such an approach would encompass the best of both worlds, providing efficient and flexible coordination,together with high-level guidance and adaptation. Another promising research direction is the scalability of our approach to >2 agents. We unfortunately could not explore this further, but several proposals have been discussed in Section 4.2.1. We find our graph-based method to have the most potential for large systems, as it allows defining explicit relationships between agents, although this would have to be investigated. Finally, the applicability of our method to a broader range of tasks, environments, and architectures remains to be explored. While our approach has demonstrated promising results in the Multiagent Minedojo environment, its applicability to other domains, such as robotics, game-playing, or realworld coordination scenarios is yet to be determined. Future research should test the generalizability of our method across different task types, agent architectures, and environmental conditions, as the limited scope of our project only allowed for two tasks to be implemented. 12 Revisiting Project Planning In this section, we revisit the project planning established at the outset of this research, evaluating our adherence to the initial objectives, temporal planning, budget, and sustainability considerations. 12.1 Reviewing Initial Objectives We will now revisit the objectives established at the outset of this research (Section 1.3.1); some of these resulted misaligned, stemming from our initial lack of knowledge, nevertheless they were all completed successfully. We defend each one and cite evidence: 1. Theoretical Understanding and Framework Evaluation: (a) Completed , as evidenced by our analysis of the current tasks (Appendix D) and our overall correct usage of the environment, with many tasks implemented within. Discussion and Conclusions Oriol Miró López-Feliu 37
Developing Coordination Methods for AI Agents in Open-Ended Environments (b) Completed, as proven by how well integrated our method is into the model’s architecture (Section 4.2), and our understanding of different factors that made the model work (Sectors 6.1 and 8.2.1) (c) Completed , seen throughout the work but specially during the development of our efficient training algorithm (Section 4.1), where we relied heavily in existing literature, as well as our analysis of the project’s place within research (Section 1.1.2 2. System Development and Implementation: (a) Completed, as stated in Section 2.1. We also fixed many errors and bugs. (b) Completed , as shown in Section 2.2, together with many error fixes and new feature integrations. (c) Completed , as we introduced two tasks (Sections 3 and 8), testing situations where coordination was beneficial or necessary, respectively 3. Development of Coordination Mechanisms: (a) Completed , although this this objective was not well aligned with the research’s needs; we proposed reward structures for each task, justifying our decisions, and improving them when needed (Sections 3.3, 7, 8.1.2). (b) Completed , as we discussed within Section 4.2; overall, we experimented with regularisation for our penalty (Section 6.2), and unidirectional vs bidirectional coordination (Section 8.2.2). Moreover, we discussed potential architecture modifications to scale our method to >2agents (Section 4.2.1). (c) Completed , as we experimented with several training methods and hyperparameters, to stabilise training and achieve better performance (Sections 6.2, 6.3, and 8.2.2). (d) Completed , evidenced by our several analysis (Sections 6.4, 7.2, and 8.2.3), and our explorations of coordination patterns (6.4.1). 4. Comparative and Performance Analysis: (a) Completed, with our analysis against the baseline (Sections 6.4, 7.2, and 8.2.3) (b) Completed , with severe optimizations (Section 5), analysis of the scalability (Section 4.2.1), and analysis of the performance and the patterns involved (Sections 6.4, 7.2, 8.2.3, and 6.4.1). (c) Completed, as shown in Sections 9, 10 and 11). 12.2 Reviewing Temporal Planning, Budget and Sustainability Following the initial challenges, we proposed an updated schedule on the “Fita de Seguiment”; after this, the project progressed as anticipated, with no further significant deviations. For the sake of completeness, the updated schedule can be found in Figure 32 Figure 32: Final temporal planning of the project, same as on the “Fita de Seguiment” Discussion and Conclusions Oriol Miró López-Feliu 38
Developing Coordination Methods for AI Agents in Open-Ended Environments Regarding the project budget, it remained consistent with the initial estimates, with no significant deviations observed. Therefore, the original budget set on Table 9 (on Appendix B resulted final. Sustainability was a key consideration throughout the project, and as was discussed in Section C, its biggest impact in the current project regards electricity consumption during training. To numerically estimate the impact of our work, we perform a gross estimate of the energy consumed; this is not fully accurate as we still used our PC for other tasks, but these were much less resource-intensive. Accounting for all training experiments (failed and successful), we conducted approximately t= 300 hours of training on an NVIDIA GeForce RTX 3070 GPU, which has a power consumption of P= 220 W [55]. The total energy consumption Ecan be calculated as follows: E=P×t= 220 W×300 h= 66 kWh (10) To put this into perspective, the average annual electricity consumption per capita in Spain is around ≈4900 kWh [ 2 ], meaning that our training experiments consumed ≈1.35% of the average annual per-capita electricity. It is worth noting that these figures are all thanks to the optimizations discussed in Section 5, where we achieved a 9.27x reduction in training time. Without these, the total energy consumption would have been approximately Eunoptimized = 66 kWh×9.27 ≈612 kWh , or ≈12.5% of Spain’s per-capita consumption, a significant reduction. Our personal conclusions on the project’s sustainability are that the energy consumption was ultimately negligible compared to the average per-capita consumption, and that the implemented optimizations further minimized its environmental impact. Discussion and Conclusions Oriol Miró López-Feliu 39
References [1] Artificial Intelligence Act: deal on comprehensive rules for trustworthy AI | News | European Parliament — europarl.europa.eu. https: //www.europarl.europa.eu/news/en/press-room/20231206IPR15699/ artificial-intelligence-act-deal-on-comprehensive-rules-for-trustworthy-ai . [Accessed 23-05-2024]. [2] Energy consumption in Spain — worlddata.info. https://www.worlddata.info/europe/ spain/energy-consumption.php. [Accessed 18-06-2024]. [3] GitHub - MineDojo/MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge — github.com. https://github.com/MineDojo/MineDojo . [Accessed 23-052024]. [4] GitHub - openai/Video-Pre-Training: Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos — github.com. https://github.com/openai/ Video-Pre-Training. [Accessed 23-05-2024]. [5] MIT License — choosealicense.com. https://choosealicense.com/licenses/mit/ . [Accessed 23-05-2024]. [6] Myriam Abramson and Ranjeev Mittu. Multi-Agent Coordination in Open Environments, pages 217–229. Springer US, Boston, MA, 2006. [7] Saaket Agashe, Yue Fan, and Xin Eric Wang. Evaluating multi-agent coordination abilities in large language models, 2023. [8] Alhussein Alhussein Fawzi, Matej Balog, Bernardino Romera-Paredes, Demis Hassabis, and Pushmeet Kohli. Discovering novel algorithms with alphatensor, Oct 2022. [9] Amazon Web Services. What is reinforcement learning? https://aws.amazon.com/ what-is/reinforcement-learning/, 2024. Accessed: 2024-02-25. [10] Anonymous. Villagerbench: Benchmarking multi-agent collaboration in minecraft, Feb 2024. [11] Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization, 2016. [12] Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron Courville, and Yoshua Bengio. An actor-critic algorithm for sequence prediction, 2017. [13] Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune. Video pretraining (vpt): Learning to act by watching unlabeled online videos, Jun 2022. [14] Dennis Barrios-Aranibar and Luiz Marcos Garcia Goncalves. Learning coordination in multiagent systems using influence value reinforcement learning. In Seventh International Conference on Intelligent Systems Design and Applications (ISDA 2007), pages 471–478, 2007. [15] Sebastiaan Bollaart. The environmental cost of llms: A call for efficiency, May 2023. [16] M. A. Bucci, O. Semeraro, A. Allauzen, G. Wisniewski, L. Cordier, and L. Mathelin. Control of chaotic systems by deep reinforcement learning. Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sciences, 475(2231):20190351, 2019. [17] Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chi-Min Chan, Heyang Yu, Yaxi Lu, Yi-Hsin Hung, Chen Qian, Yujia Qin, Xin Cong, Ruobing Xie, Zhiyuan Liu, Maosong Sun, and Jie Zhou. Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors, 2023. [18] Michele Chincoli and Antonio Liotta. Self-learning power control in wireless sensor networks. Sensors, 18:375, 01 2018. [19] Karl Cobbe, Jacob Hilton, Oleg Klimov, and John Schulman. Phasic policy gradient, 2020. 40
[20] Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. Transformer-XL: Attentive language models beyond a fixed-length context. In Anna Korhonen, David Traum, and Lluís Màrquez, editors, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2978–2988, Florence, Italy, July 2019. Association for Computational Linguistics. [21] Nayana Dasgupta and Mirco Musolesi. Investigating the impact of direct punishment on the emergence of cooperation in multi-agent reinforcement learning systems, 2023. [22] Elhadji Amadou Oury Diallo, Ayumi Sugiyama, and Toshiharu Sugawara. Coordinated behavior of cooperative agents using deep reinforcement learning. Neurocomputing, 396:230–240, 2020. [23] Linxi Fan, Guanzhi Wang, Yunfan Jiang, Ajay Mandlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De-An Huang, Yuke Zhu, and Anima Anandkumar. Minedojo: Building open-ended embodied agents with internet-scale knowledge. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022. [24] Glassdoor. Average salaries per position and country. Accessed: 04-03-2024. [25] Kailash Gogineni, Peng Wei, Tian Lan, and Guru Venkataramani. Scalability bottlenecks in multi-agent reinforcement learning systems, 2023. [26] Ran Gong, Qiuyuan Huang, Xiaojian Ma, Hoi Vo, Zane Durante, Yusuke Noda, Zilong Zheng, Song-Chun Zhu, Demetri Terzopoulos, Li Fei-Fei, and Jianfeng Gao. Mindagent: Emergent gaming interaction, 2023. [27] greg065. Minecraft 4-player coop on one pc, 60fps. https://www.reddit.com/r/ localmultiplayergames/comments/bppf3s/minecraft_4player_coop_on_one_pc_ 60fps/, 2019. Accessed: 25/02/2024. [28] William H. Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden Codel, Manuela Veloso, and Ruslan Salakhutdinov. Minerl: A large-scale dataset of minecraft demonstrations, 2019. [29] Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, 2018. [30] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016. [31] Jiechuan Jiang, Kefan Su, and Zongqing Lu. Fully decentralized cooperative multi-agent reinforcement learning: A survey, 2024. [32] Prathima Kadari. What is reinforcement learning and how does it work (updated 2024), Feb 2024. [33] Thomas N. Kipf and Max Welling. Semi-supervised classification with graph convolutional networks, 2017. [34] James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences, 114(13):3521–3526, March 2017. [35] Vijay Konda and John Tsitsiklis. Actor-critic algorithms. In S. Solla, T. Leen, and K. Müller, editors, Advances in Neural Information Processing Systems, volume 12. MIT Press, 1999. [36] Dhireesha Kudithipudi, Mario Aguilar-Simon, Jonathan Babb, Maxim Bazhenov, Douglas Blackiston, Josh Bongard, Andrew P. Brna, Suraj Chakravarthi Raja, Nick Cheney, Jeff Clune, Anurag Daram, Stefano Fusi, Peter Helfer, Leslie Kay, Nicholas Ketz, Zsolt Kira, Soheil Kolouri, Jeffrey L. Krichmar, Sam Kriegman, Michael Levin, Sandeep Madireddy, Santosh Manicka, Ali Marjaninejad, Bruce McNaughton, Risto Miikkulainen, Zaneta Navratilova, Tej Pandit, Alice 41
Parker, Praveen K. Pilly, Sebastian Risi, Terrence J. Sejnowski, Andrea Soltoggio, Nicholas Soures, Andreas S. Tolias, Darío Urbina-Meléndez, Francisco J. Valero-Cuevas, Gido M. van de Ven, Joshua T. Vogelstein, Felix Wang, Ron Weiss, Angel Yanguas-Gil, Xinyun Zou, and Hava Siegelmann. Biological underpinnings for lifelong learning machines. Nature Machine Intelligence, 4(3):196–210, 3 2022. [37] Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel. End-to-end training of deep visuomotor policies. Journal of Machine Learning Research, 17(39):1–40, 2016. [38] Shalev Lifshitz, Keiran Paster, Harris Chan, Jimmy Ba, and Sheila McIlraith. Steve-1: A generative model for text-to-behavior in minecraft. 2023. [39] Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning, 2019. [40] Sabrina McCallum, Max Taylor-Davies, Stefano Albrecht, and Alessandro Suglia. Is feedback all you need? leveraging natural language feedback in goal-conditioned RL. In NeurIPS 2023 Workshop on Goal-Conditioned Reinforcement Learning, 2023. [41] A. Mohtasib, G. Neumann, and H. Cuayáhuitl. A study on dense and sparse (visual) rewards in robot policy learning. pages 3–13, 2021. [42] Maria Viorela Muntean. Multi-agent system for intelligent urban traffic management using wireless sensor networks data. Sensors, 22(1), 2022. [43] Nasim Nezamoddini and Amirhosein Gholami. A survey of adaptive multi-agent networks and their applications in smart cities. Smart Cities, 5(1):318–347, 2022. [44] Diego Perez-Liebana, Katja Hofmann, Sharada Prasanna Mohanty, Noburu Kuno, Andre Kramer, Sam Devlin, Raluca D. Gaina, and Daniel Ionita. The multi-agent reinforcement learning in malmö (marlö) competition, 2019. [45] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision, 2021. [46] Abu Rayhan. The future of work: How ai and automation will transform industries, 07 2023. [47] Iván González Reguera. daptation of the minedojo framework for use in multiagent reinforcement learning. Bachelor’s thesis, Universitat Politècnica de Catalunya (UPC) - BarcelonaTech, Barcelona, 1 2024. Graduate in Computer Engineering (Computing), Facultat d’Informàtica de Barcelona (FIB). [48] Caude Sammut. Behavioral Cloning, pages 93–97. Springer US, Boston, MA, 2010. [49] John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel. Highdimensional continuous control using generalized advantage estimation, 2018. [50] John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms, 2017. [51] John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms, 2017. [52] David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. Mastering the game of go without human knowledge. Nature, 550(7676):354–359, 10 2017. [53] Konstantin Sofeikov. Implementing conditional variational auto-encoders(cvae) from scratch, Apr 2023. 42
[54] Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour. Policy gradient methods for reinforcement learning with function approximation. In S. Solla, T. Leen, and K. Müller, editors, Advances in Neural Information Processing Systems, volume 12. MIT Press, 1999. [55] LAR Systems. GeForce RTX 3070 Folding@Home PPD Averages, Power Consumption & Research Projects. https://folding.lar.systems/gpu_ppd/brands/nvidia/folding_ profile/ga104_geforce_rtx_3070, 2024. Accessed: 18-06-2024. [56] Massimo Tipaldi, Raffaele Iervolino, and Paolo Roberto Massenio. Reinforcement learning in spacecraft control applications: Advances, prospects, and challenges. Annual Reviews in Control, 54:1–23, 2022. [57] Taha ValizadehAslani and Hualou Liang. Layernorm: A key component in parameter-efficient fine-tuning, 2024. [58] T. van der Heiden, C. Salge, E. Gavves, and H. van Hoof. Robust multi-agent reinforcement learning with social empowerment for coordination and communication, 2020. [59] Mariana Vargas Vieyra and Pierre Ménard. Learning generative models with goal-conditioned reinforcement learning, 2023. [60] Kaixin Wang, Daquan Zhou, Jiashi Feng, and Shie Mannor. Ppg reloaded: an empirical study on what matters in phasic policy gradient. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org, 2023. [61] Mianchu Wang, Rui Yang, Xi Chen, and Meng Fang. GOPlan: Goal-conditioned offline reinforcement learning by planning with learned models. In NeurIPS 2023 Workshop on Goal-Conditioned Reinforcement Learning, 2023. [62] Wikipedia contributors. List of best-selling video games. https://en.wikipedia.org/ wiki/List_of_best-selling_video_games, 2024. Accessed: 2024-02-25. [63] Jenny Yang, Andrew A S Soltan, David W Eyre, and David A Clifton. Algorithmic fairness and bias mitigation for clinical machine learning with deep reinforcement learning. Nat Mach Intell, 5(8):884–894, July 2023. [64] Lei Yuan, Ziqian Zhang, Lihe Li, Cong Guan, and Yang Yu. A survey of progress on cooperative multi-agent reinforcement learning in open environment, 2023. [65] Kaiqing Zhang, Zhuoran Yang, Han Liu, Tong Zhang, and Tamer Ba¸sar. Fully decentralized multi-agent reinforcement learning with networked agents, 2018. [66] Lunjun Zhang and Bradly C. Stadie. Understanding hindsight goal relabeling from a divergence minimization perspective, 2023. [67] Y. Zhang, Q. Yang, D. An, and C. Zhang. Coordination between individual agents in multiagent reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 11387–11394. AAAI, 2021. 43
Part V Appendices A Temporal Planning A.1 Description of tasks Early and precise temporal planning is fundamental to ensure success on such a large project, justifying the need for this section. The thesis’ temporal dimension is bound by the semester’s start date, 12th of February, and the beginning of the reading turns, 23rd of June (this day corresponds to the Sunday before they start). A hopefully very productive extra day is available, given that 2024 is a leap year, thus allowing 135 days to finish this project, exactly 19 weeks. The researcher can dedicate 28 hours per week to the project, which means that the work must be completed in around 532 hours1. A.1.1 Roles involved In order to complete the different tasks that will be proposed, the following roles were identified: • Project Manager: As the name indicates, this role is responsible for planning the project and ensuring that it is carried out as expected until completion. The tasks involving this role are the GEP module and its less technical documentation, communications with the supervisors, and the preparation of the final oral presentation. • Technical Writer: A substantial part of the work involves writing quality technical documentation of the research produced, and this role is in charge of that. • AI Researcher: Since there is an important part of previous study in this work, it will be completed by the researcher specializing in AI. • Data Scientist: Everything else falls under this umbrella role, who will be in charge of investigation, model implementation, training, and analysis of results. A.1.2 Task definition This project involves four large groups of tasks. The first group is project planning, where the work will be defined, planned, and many other aspects considered. This will be followed by a previous study, allowing the researcher time to gather background on the different technologies to be applied. Next, all components of the practical side of the thesis will be introduced, and finally time will be allocated to the documentation of research results and other miscellaneous tasks. A.1.2.1 Project Planning The first part of the project covers GEP, the mandatory project management module intended to guide the work. The tasks required to do this can be divided into the following: •Contextualization and Project Scope (P1): – Definition of the scope of the project, its contextualization, and its relevance to the area of study. –Expected duration: 20 hours, given how much previous research is also involved to properly contextualize the work. –Role: Project manager. •Temporal Planning (P2): – Planning of the total execution of work, providing a description of the tasks and their requirements. 128h / week * 19 weeks = 532h 44
–Expected duration: 7.5 hours, since it is more straightforward once the previous section has been completed. –Role: Project manager. •Economic Management and Sustainability (P3): –Analysis of the sustainable and economic aspects of the project. –Expected duration: 7.5 hours, because it requires just calculating the costs of what was previously detailed. –Role: Project manager. •Development of the Final Document (P4): – Compiling all previous parts, correcting their mistakes, and producing a final project management document. –Expected duration: 10 hours, as it is estimated this is how long it will take to incorporate the GEP tutor’s and supervisors’ feedback. –Role: Project manager. A.1.2.2 Previous Study Since this work originates from and builds upon two other projects, Multi-agent Minedojo and STEVE-1, a solid understanding of them is vital to avoid errors due to ignorance. Furthermore, the inexperience of the researcher and the exploratory nature of this thesis requires a study of the concepts involved, including the SOTA and the applicable literature, to ensure that the results are relevant and well-founded. This previous study can be subdivided into the following. •Study Base Work 1 (S1): – Study of “Adaptation of the MineDojo Framework for Use in Multi-agent Reinforcement Learning”, its code and the techniques involved –Expected duration: 5 hours, as the original document is 165 pages long and its technical. The hours allocated also account for the future consulting that will be done. –Role: AI Researcher. •Study Base Work 2 (S2): –Study of STEVE-1, its code, and the involved techniques. –Expected duration: 10 hours, as the researcher is very inexperienced with many of the complex technologies employed. –Role: AI Researcher. •Study of RL and MARL techniques (S3): – Study RL and MARL techniques, to gain an understatement and intuition of methods that could be developed for the particular case of this work. –Expected duration: 40 hours, some of them will be conducted at the beginning to gather background, and the rest throughout the project, when implementing methods. –Role: AI Researcher. •Study SOTA and Literature (S4): –Study the current SOTA and the literature relevant to this thesis. –Expected duration: 40 hours, approximately 20 hours to understand the current context, situation and trends, and the rest to be completed throughout the project, in accordance with the study of RL and MARL techniques, to implement methods and evaluate results. –Role: AI Researcher. A.1.2.3 Practical Work With that out of the way, the researcher will be ready to begin the most substantial portion of the thesis, which encompasses: 45
Activity Import (C) Comments P1 - Contextualization and Project Scope 564.00 Project Manager, 20 hours P2 - Temporal Planning 211.50 Project Manager, 7.5 hours P3 - Economic Management and Sustainability 211.50 Project Manager, 7.5 hours P4 - Development of the Final Document 282.00 Project Manager, 10 hours S1 - Study Base Work 1 (Multi-agent Minedojo) 135.40 AI Researcher, 5 hours S2 - Study Base Work 2 (STEVE-1) 270.80 AI Researcher, 10 hours S3 - Study RL and MARL techniques 1,083.20 AI Researcher, 40 hours S4 - Study SOTA and Literature 1,083.20 AI Researcher, 40 hours W1 - Testing of Multi-agent Minedojo 264.75 Data Scientist, 5 hours W2 - Adaptation of STEVE-1 1,323.75 Data Scientist, 25 hours W3 - Task definition and implementation 2,118.00 Data Scientist, 40 hours W4 - MVP Definition Phase 529.50 Data Scientist, 10 hours W5 - MVP Implementation Phase 1,853.25 Data Scientist, 35 hours W6 - MVP Training Phase 1,853.25 Data Scientist, 35 hours W7 - MVP Evaluation Phase 529.50 Data Scientist, 10 hours W8 - Iteration of Extension Cycles 6,618.75 Data Scientist, 125 hours M1 - Research Documentation 1,712.25 Technical Writer, 75 hours M2 - Meetings and Communications 564.00 Project Manager, 20 hours M3 - Oral Presentation 338.40 Project Manager, 12 hours Total CPA (Costs per activity) 21,547.00 Software Overleaf Software 0.00 Free to use TeamGannt 0.00 Free to use LibreOffice Calc 0.00 Free to use Python & Libraries 0.00 Free to use Github 0.00 Free to use VS Code 0.00 Free to use STEVE-1 0.00 Hardware PC 73.48 532h of use, base cost = 1,436.76C Display 15.35 532h of use, base cost = 300C Keyboard 2.56 532h of use, base cost = 50C Mouse 1.53 532h of use, base cost = 30C Space Rent 810 16.2C/m²/month * 10m² * 5 months Furniture 7.67 Base cost = 300C, ammortised to 10 years Electricity 225 45C / month, 5 months Internet Access 200 40C / month, 5 months Total GC (General Costs) 1335.59 Total Costs (Total CPA + Total GC) 22,882.59 Contingency 3432.3885 Contingency margin = 15% Total CPA+CG+Contingency 26,314.98 Cluster Usage 30 Cost = 100C. Risk = 30% Delays in practical work (half a week) 222.39 Cost = Data Scientist, 28h. Risk = 30% Delays in practical work (a week) 148.26 Cost= Data Scientist, 48h. Risk = 10% Total Incidentals 400.65 PROJECT TOTAL 26,715.63 Table 9: Budget structure B.2 Cost estimates CPA are calculated by assigning previously detailed roles to their matching tasks in the Gannt diagram (Figure 33), and multiplying their dedicated hours by their cost per hour. These estimates of the hourly rate were taken from GlassDoor[ 24 ] and because the platform only offered the annual average salary, this was later converted to the hourly rate, estimating 40 hours worked per week. A summary of the costs of the specific role can be found in Table 10, and the estimated total cost of CPA is e21,547.00.
Role Cost/hour ( e ) Cost/hour with SS (e)Total hours Role total cost (e) Project Manager 21.69 28.20 77 2,171.40 Technical Writer 17.56 22.83 75 1,712.25 AI Researcher 20.83 27.08 95 2,572.60 Data Scientist 40.73 52.95 285 15,090.75 Total 532 21,547.00 Table 10: Role Costs Summary GC require taking into account the used software, hardware, and space-related costs. All software used is free, resulting in a cost of e 0. To calculate the cost of hardware resources, their respective cost is amortized as the fraction of their lifetime that they will be used for, better described by the following equation: Amortized cost =Total cost of the product ×Total project hours Total hours in expected life (11) The total hours in a resource’s expected life are assuming a normal workload of 40 hours per week during 5 years. That is, 40 hours / week ×52 weeks / year ×5years = 10,400 hours . The results of these calculations can be seen in Table 11. Item Cost (e) Hours Used Expected Life Hours Amortized Cost ( e ) PC 1,436.37 532 10,400 73.48 Display 300 532 10,400 15.35 Keyboard 50 532 10,400 2.56 Mouse 30 532 10,400 1.53 Total 92.92 Table 11: Hardware Amortized Cost Finishing GC costs, space resources include the rent of the room from which the researcher will work from (calculated as the average rent per square meter in their area multiplied by the size of the room), the amortized cost of the furniture (assuming a life expectancy of 10 years), and the average electricity and internet bill (assuming 19 weeks ≈ 5 months of work). The details of the amortized furniture costs are the following: Base cost =e300 , Expected life = 10 years = 20800 hours , Amortized cost =e7.67 . The details of the cost of space have not been included in a separate table to avoid redundancy and can be found in Table 9. The total of GC results in e1335.59. Adding up CPA and GC results in e 22,882.59. A contingency margin of 15% was used, and therefore the total reached e26,314.98. Finally, the incidence costs took into account the possible risks that could impact the project’s budget. Their cost is calculated as what they would cost multiplied by their risk. Incidence costs included the possibility of having to use outside compute power and the results of delays in different parts of practical work (which is the part with the greatest risk of delay). The delays of both a half-week and a full-week have been accounted for. The total cost of incidents is e 400.65, and the total cost of the project adds up to e26,715.63. B.3 Management control To control the project budget, a numerical indicator will be calculated for each task performed. This indicator will be the cost deviation Cd , which is the difference between the estimated resources consumed Re and the real resources consumed Rr , times the cost per resource consumed Cr . The equation describing this is: Cd= (Re−Rr)×Cr(12) The deviation cost will indicate exactly how far the estimations shown in Table 9 are from the real values (in euros, as it is the currency used in this project). There are three possible scenarios: 53
•Cd>0:The task cost less than expected. •Cd<0:The task cost more than expected. •Cd= 0:The task cost exactly as expected. The total cost deviation will be the sum of the cost deviations of all the resources in the project. If this metric is negative, part of the contingency cost will be used to cover such expenses. However, if it is positive, the money can be reallocated to, for example, incidences. C Sustainability For the last part of project planning, sustainability considerations must be made, where self-reflection will be included, followed by the analysis of the economic, environmental, and social dimensions of this project. C.1 Self reflection Due to the nature of this specific subsection, first-person writing will be used. This will be the only part of the project where this style is used. Completing the sustainability self-evaluation made me think and assess many things. First, before starting this project, I personally did not know that sustainability included economic and social dimensions. I believed the first to be irrelevant to sustainability and the latter to be a part of the field of ethics instead. It was also unknown to me the amount of metrics and indicators available and considerations one must make, and there were many points of view that I honestly had not considered before, making me feel a bit ashamed. During the duration of my degree, sustainability was a transversal competence, assessed in three separate courses. Even then, they were an afterthought; In most cases, this was something “extra” and was not given much importance. In all of them, you could still get an excellent score ( ≥9 ) on the course, while completely ignoring sustainability aspects. Reviewing my scores for this competence, I achieved an A in three separate evaluations. Given the reflection made on the previous paragraph, this highlights how there is still a long way to go in order to form engineers with a sustainable mindset. Looking forward, I will try to do better. I wholeheartedly believe that sustainability should not be an obstacle that only obstructs projects and limits what we can do, but an end in itself and should be considered as one of the ultimate goals of any project. I will try to make the relevant considerations before starting the project and maintaining them throughout its duration. When possible, I will also apply the relevant indicators, although if the self-evaluation made another thing clear, it is how uninformed I am about them. I myself have a long way to go and I will try to learn more. 54
C.2 Economic Dimension Regarding PPP: Reflection on the cost you have estimated for the completion of the project: The cost analysis, which can be found in Section B, accurately takes into account human and material resources, and other possible costs (such as contingencies or incidents) have also been taken into account. Most of the project cost is the salary of the Data Scientist, as this role will work for 285h and each hour is billed at e 52.95. Something that could be improved from the budget is to possibly break down this role’s tasks further; some of its work, such as testing implementation or adapting things, could be performed by a software engineer with an hourly salary of e 23.77 (as per Glassdoor[ 24 ]), which is much more feasible. Parts of the project where budget was saved include the use of the researcher’s local PC instead of sourcing outside compute, despite the latter being much better in an ideal case. Overall, the budget feels reasonable and could be completed in a real-life scenario, albeit some cuts to reduce the cost. Regarding Useful Life: How are currently solved economic issues (costs...) related to the problem that you want to address (state of the art)? Currently, there is not much research on the specific problem this thesis will attempt to tackle, as the Justification section (Section 1.2) explained. The problem tends to be solved with the use of LLMs, and this work will not employ them. How will your solution improve economic issues (costs ...) with respect other existing solutions? A possible economic improvement is the reduced cost of not using LLMs. These are extremely expensive to train, and therefore the only viable option is to use a trained one, but most options are pay-to-use; therefore, developing an approach that does not require LLMs is inherently cheaper. Moreover, if the results of this thesis were employed together with LLMs, optimizing short-term tasks will lead to fewer errors and therefore a reduced number of calls to LLMs, consequently reducing the cost. It is challenging to measure the exact improvement in cost, but it will be attempted after the project is finalized. It will be attempted by comparing the cost efficiency of this solution compared to other SOTA models, and trying to quantify the number of LLMs calls avoided, together with their average cost. C.3 Environmental Dimension Regarding PPP: Have you estimated the environmental impact of the project? The environmental impact of this project will be low, and mostly result from electricity use, as a GPU will be running for many hours. To empirically measure this, CO2eq, a standardized measure to quantify carbon emissions associated with ML model training, will be calculated when training the model, or were it not possible for any reason, other metrics such as electricity consumption will be considered. Regarding PPP: Did you plan to minimise its impact, for example, by reusing resources? Yes. This project will build on a previous model, STEVE-1. This approach conserves many resources, since a significant amount of computational power is typically expended during the initial phases of model development. The researcher’s PC will be used instead of outside compute as far as possible, also. Moreover, software to track the energy consumption of experiments will be considered and the code will be thoroughly tested to avoid running experiments that could contain errors. Regarding Useful Life: How is currently solved the problem that you want to address (state of the art)? As was said in a previous answer, the exact research direction is not yet addressed in complex environments such as Minecraft, and the most similar research directions use LLMs. How will your solution improve the environment with respect other existing solutions? 55
Training and use of LLMs is also extremely costly for the environment, with an immense carbon footprint. Training in GPT3, for example, consumed as much power as the average Dutch household in 9 years[ 15 ]. Therefore, avoiding the use of LLMs or providing alternatives, as in the current work, results in a great improvement with respect to environmental costs. C.4 Social Dimension Regarding PPP: What do you think you will achieve -in terms of personal growthfrom doing this project? This project will mean a lot of growth for the researcher. It will introduce them to the research world and enrich their perspective on their professional options. It will teach them how to manage a large project throughout its lifetime (project planning to defense of results). Finally, it will allow them to learn about the topic at hand, which will enrich them intellectually. Regarding Useful Life: How is currently solved the problem that you want to address (state of the art)? The same responses as before, mostly through LLMs. How will your solution improve the quality of life (social dimension) with respect other existing solutions? It is difficult to assess how the solution will improve quality of life in the short term. However, in the long term, it could have great potential in helping coordinate autonomous systems, which could improve efficiency or safety in many fields, for example, in traffic control systems. Regarding Useful Life: Is there a real need for the project? Yes. Currently, there might not be many direct applications that strictly need this project, but this project is working in a direction that will surely be needed in the future. With AI and automation taking over, it is only a matter of time that these two come together, and when it happens, efficient direct coordination of AI agents will be fundamental. D Evaluation of current tasks Upon evaluating the current tasks offered by Minedojo, it becomes evident that they are not wellsuited for testing multi-agent coordination. The tasks can be categorized into three main groups: Creative, Playthrough, and Programmatic. However, each of these categories presents significant limitations when it comes to assessing the coordination abilities of multiple agents. The Creative category, which includes tasks such as building replicas of famous structures or creating themed houses, poses significant challenges in terms of evaluating success and training agents for such complex objectives. These tasks are inherently subjective and lack clear metrics for measuring coordination between agents. Moreover, the complexity of these tasks makes them unsuitable for the scope of this research, which aims to focus on more short-term, well-defined objectives. The Playthrough category, consisting of a single task - defeating the Ender Dragon and obtaining the trophy dragon egg - is an incredibly complex and long-term objective. Even single-agent systems have not yet achieved this task, making it an unrealistic goal for testing multi-agent coordination. Lastly, The Programmatic category, while offering a wider range of tasks, still falls short in providing a suitable challenge for testing multi-agent coordination. This category can be further divided into four subcategories: • Survival: These tasks involve staying alive for as many days as possible and could potentially test adaptability and fast response in a multi-agent setting. However, the long-term nature of these tasks and the time constraints of this project make them less feasible for adaptation. • Harvest: Revolving around collecting resources, these tasks are highly parallelizable and do not necessarily require coordination between agents. Agents could work independently and still achieve high evaluation scores, rendering these tasks ineffective for testing coordination. 56
• Tech Tree: These tasks involve the creation and use of a hierarchy of tools and armor. Despite their potential for testing agent coordination in planning and assigning subtasks, they are too complex and extend beyond the scope of this project. • Combat: While offering fast-paced challenges, these tasks are uninteresting: both agents comprise similar roles, and their task is too simple. Coordination would likely be rendered to mere “help signals”, therefore it is not suitable. Overall, current tasks are not well suited to test coordinatiion due to their limitations in complexity, parallelizability, and lack of clear metrics for evaluating coordination. To effectively assess the coordination abilities of multiple agents, it is necessary to develop new tasks specifically designed for this purpose, focusing on short-term objectives that require clear communication, adaptability, and fast response among agents. E Adaptation Details This section provides a comprehensive overview of the adaptations and error corrections implemented in the Multiagent Minedojo environment and the STEVE-1 model. E.1 Multiagent Minedojo Adaptation Details During the initial stages of utilizing the Multiagent Minedojo environment, it appeared to function as intended. However, as the training process progressed, several critical issues emerged, some of which were attributed to the asymmetrical roles of the agents in our task, as described in Section 3. These errors could not have been accounted for by the previous researcher, as their objective was to enable the environment, and the errors arose from our specific tasks. The first necessary adaptation addressed the incorrect formatting of commands for the second agent. In the original code, when attempting to teleport the first agent, the correct “/tp” command was used; however, for the second agent, this was erroneously extrapolated to a nonexistent “/tp2” command. This issue was discovered while attempting to optimize the environment’s reset function, as described in Section 5. Although this did not present as an error, it produced confusing outcomes that puzzled us for a considerable time, such as teleporting only one of the agents or causing other unexpected behaviors with other commands. To resolve this, we meticulously reviewed and corrected all commands related to the second agent, ensuring adherence to the correct format. Another issue identified was the unnecessary limitation on command execution. In the original code, commands such as the previously mentioned “tp” could only be executed within the first three steps after resetting the environment. This restriction, deeply embedded within the code, was discovered while attempting to optimize the environment and did not display an error. However, it rendered changes in the environment’s parameters ineffective. We thoroughly examined the code and made necessary modifications to allow command execution throughout the training process. Furthermore, the lack of consideration for asymmetrical roles in the creation of Multiagent Minedojo resulted in rewards being issued to both agents, regardless of their individual actions. To address this, we modified the internal structure of the library to support asymmetrical rewards. The most challenging issue encountered was a memory leak that caused an ever-increasing use of memory, eventually leading to a crash during training. Initially, we suspected that inefficiencies in our own training code were the cause. However, upon further investigation, we discovered that the memory leak caused old instances to persist after resetting the environment, filling up the memory. This issue halted training every 10 iterations or so, which was significantly lower than our requirements. Resolving this issue necessitated the implementation of a mechanism to terminate all Minecraft instances when performing a reset. E.2 STEVE-1 Adaptation Details This subsection presents the adaptations made to the STEVE-1 model to ensure compatibility with the Minedojo environment and to enable conditional scaling. 57
E.2.1 Adaptation of MineRL to Minedojo Spaces To adapt STEVE-1, originally designed for the MineRL environment, to the Minedojo environment, we needed to address the differences in their observation and action spaces. E.2.1.1 Observation Space Adaptation The MineRL observation space consists solely of an RGB image, which aims to impose the same conditions on agents as on humans when “playing” the game. In contrast, MineDojo expands the observation space to include a richer set of environmental data, which includes: •RGB frames: Equivalent to MineRL’s observation. •Equipment and Inventory: Detailed information about items the agent carries and wears. •Inventory Changes: Alerts the agent to changes in inventory. •Voxels (Surrounding Blocks): 3x3x3 surrounding blocks around the agent. •Life and Location Statistics: Health and positioning data. • Nearby Tools and Damage Source: Information about tools in proximity and sources of damage. As STEVE-1 can directly utilize Minedojo’s RGB frames, which are identical to MineRL’s, minimal adaptation is required for this space. It is worth noting that the channel order for the image observation differs between the two environments: Minedojo’s is (3, height, width), while MineRL’s is (height, width, 3). E.2.1.2 Action Space Adaptation The action spaces in MineRL and MineDojo differ significantly, as shown in Tables 12 and 13. The MineDojo action space is more granular and includes additional actions such as crafting, equipping, placing, and destroying items, along with their respective arguments. In contrast, MineRL’s action space only allows actions that a human player could perform with a keyboard or a controller. Action Human Action Description forward W key Move forward back S key Move backward left A key Strafe left right D key Strafe right jump Space key Jump inventory E key Open/close inventory and crafting grid sneak Shift key Modifier for careful movement and inventory interaction sprint Ctrl key Fast movement attack Left mouse button Attack or interact in inventory use Right mouse button Use item or interact in inventory drop Q key Drop items from inventory hotbar.1-9 Keys 1-9 Switch active item camera Mouse (-180, 180) for x and y, continous Table 12: MineRL Action Space. Human actions are indicated to show how similar this action space is to that of a real Minecraft player. All actions are discrete (binary) except camera 58
Action Details Description Discrete Values 0 0: noop, 1: forward, 2: back Forward and backward 3 1 0: noop, 1: left, 2: right Move left and right 3 2 0: noop, 1: jump, 2: sneak, 3: sprint Jump, sneak, sprint 4 3 0: -180, ..., 24: +180 degrees Camera delta pitch 25 4 0: -180, ..., 24: +180 degrees Camera delta yaw 25 5 0: noop, 1: use, 2: drop, 3: attack, 4: craft, 5: equip, 6: place, 7: destroy Functional actions 8 6 Specific craftable items Argument for “craft” 244 7 Inventory slot indices Argument for “equip/place/destroy” 36 Table 13: MineDojo Action Space. All actions are discrete, as shown in the column. To bridge this gap, we established a mapping between the MineRL actions and their MineDojo counterparts. While every MineRL action can be mapped to a corresponding MineDojo action, not all MineDojo actions have a direct equivalent in MineRL. However, this discrepancy is not a significant issue, as the agents in MineDojo can still perform these actions using more natural, real-life-like interactions, such as using the mouse to click and navigate the inventory. This approach is more intuitive and realistic compared to the domain-specific representations in MineDojo for actions like “craft” or “equip”, which are executed through specific action codes and hinder generalization across other games. We mapped all actions with a direct translation directly and left the special Minedojo actions “craft” and “equip / place / destroy” as inaccessible. For the camera, we modified the internal code for Minedojo, allowing continuous values, thus avoiding the loss of information incurred when translating a continuous space into a discrete one. E.2.2 Conditional Scale Adaptations To enable conditional scaling in STEVE-1, we modified two functions: the function used to collect episodes and the gradient-enabled function used for training. The main challenge was dealing with both conditioned and unconditioned model states and with large batch numbers, as the structure of the model states is very complex, making it difficult to extend the current method to batch sizes ≥1 . The original code for episode collection keeps track of the model’s state by using a twice-as-long state, half of which is for conditional logits, and the other half for unconditional logits. We modified this by instead keeping track of two states internally, conditional and unconditional, and slightly modified the rest of the code to allow this. Separating the model states made it much easier to prepare them for batching, and this allowed us to implement conditional scaling on the gradient-enabled function used to obtain predictions during training. These modifications resulted in improved instruction following while allowing much faster and more diverse training using batches. It is important to note that all changes required careful consideration to ensure that the original functionality and performance of STEVE-1 were maintained. We thoroughly tested and validated the changes to ensure that the agent’s behavior remained consistent with the original implementation. F Treasurehunt Task Design Details This section provides a detailed description of the of the Treasurehunt task’s design, including the escape method and dungeon layout. F.1 Escape Method Details In this test, we utilize only the collector agent, while the fighter is teleported far away to avoid interference, as Multiagent Minedojo always requires two agents. To ensure the agent does not 59
achieve exit through luck, the steps per episode are limited to 150. We employ the layout seen in Figure 34, where the possible exit locations are indicated. Each experiment consists of 100 episodes. The exits selected for testing were based on tasks proven achievable in the STEVE-1 paper, and the results can be seen in Figure 7: 1. Break dirt: A block of dirt is placed, and the collector escapes upon breaking it. 2. Find water: A block of water is set as a possible exit, and the collector escapes when touching it. 3. Gather wood: A column of logs resembling a tree is made, and the agent escapes through breaking any log. 4. Break seeds: A plant is placed that the agent should break to escape. 5. Break leaves: The agent must break leaves placed close to the ceiling to resemble leaves from a tree. Agents position Possible Exits Figure 34: Diagram of Layout for testing exit in the Treasurehunt task F.2 Dungeon Layout Details The first layout created, as seen in Figure 35, proved overly ambitious, with the collector failing to achieve almost anything. The agent’s starting positions are in the bottom left of the dungeon across all layouts, while other elements, such as the possible locations of the exit and obstacles, were randomized. The exit was a single block out of all the possible ones, far away from the agents to force exploration. The obstacles were columns and lava pools, with each column being a 2×2 set of blocks from the floor to the ceiling and each lava pool being a 5×5 one block deep hole. Each column appeared with a probability of 50% to block the agent’s vision and incite exploration. Only one of the four possible lava pools was spawned to increase danger and make agents take different paths across the dungeon every episode. The dungeon ranged from (-10, -10) to (10, 10) on the (x, z) axis, with each position accounting for a block, including 0. With a height of 5 blocks, this layout had a total area of 21 ×21 ×5 = 2205 blocks. We spawned 5-6 treasure per episode and allowed 400 steps. 60
Agents position Columns Lava Exit Figure 35: Diagram of Layout 1 for Treasurehunt dungeon Visual inspection of the agent’s performance revealed that the dungeon was too large with too many obstacles. Moreover, STEVE-1 often confused the columns for trees and attempted to break them, indicating that it searched for tall, narrow structures rather than specifically trees. To facilitate the task, the dungeon size was reduced to 11 ×11 ×5 = 605 blocks, columns were eliminated as obstacles, and the number of lava pools was reduced. The single block exit and single lava pool appearance were maintained. This new layout can be seen in Figure 36. For this layout, 3-4 treasure were spawned per episode, and 300 steps were allowed. Agents position Lava Exit Figure 36: Diagram of Layout 2 for the Treasurehunt dungeon Despite these simplifications, the agent still struggled, falling into lava while not looking or not having enough time to find and break the tree. To achieve a layout the collector could solve at least half the time, the dungeon was further simplified, making it closer to the one used to determine the exit mode. Layout 3 can be observed in Figure 8. The layout size was reduced to 7×7×5 = 245 blocks, all obstacles were eliminated, 1-3 pieces of treasure were spawned, and 250 steps were allowed to reach the exit. This design proved simple enough, as shown by the results in Figure 9. G Optimizations Details G.1 Environment Optimization Internally, ‘fast reset” differs from a “hard reset” by instead of destroying all instances and regenerating the world, simply killing the agents (in game, not their instances) and teleporting them elsewhere. The nuances of this approach are: 61
Cohen’s d is used to calculate the effect size, which provides a standardized measure of the magnitude of the difference between the two distributions. It is independent of the sample sizes and allows for comparison across different studies or datasets. Cohen’s d is calculated using the following formula: d=¯x1−¯x2 sp where ¯x1 and ¯x2 are the sample means of the baseline and coordination methods, respectively, and sp is the pooled standard deviation, given by: sp=s(n1−1)s2 1+ (n2−1)s2 2 n1+n2−2 Here, n1 and n2 are the sample sizes, and s1 and s2 are the sample standard deviations of the baseline and coordination methods, respectively. In general, effect sizes can be interpreted as follows: • Small effect size: 0.2≤d < 0.5 • Medium effect size: 0.5≤d < 0.8 • Large effect size: d≥0.8 In our analysis, the effect sizes for both agents (0.461 and 0.321) are medium, indicating that the differences between the baseline and coordination methods are moderate. These medium effect sizes, along with the non-significant p-values, provide insufficient evidence to conclude that the coordination method leads to significantly higher cumulative rewards compared to the baseline. K Normality Analysis for Blind Guiding Task’s Distributions Figure 47 shows the QQ-plots of the distributions, as we argue they can not all be considered normal: Figure 47: QQ-plots for the distributions in the Blind Guiding task results. 68