scieee AI-readable full text Open interactive document viewer

Energy hardware and workload aware job scheduling towards interconnected HPC environments

D'Amico, Marco,Corbalán González, Julita

Abstract

New HPC machines are getting close to the exascale. Power consumption for those machines has been increasing, and researchers are studying ways to reduce it. A second trend is HPC machines' growing complexity, with increasing heterogeneous hardware components and different clusters architectures cooperating in the same machine. We refer to these environments with the term heterogeneous multi-cluster environments. With the aim of optimizing performance and energy consumption in these environments, this paper proposes an Energy-Aware-Multi-Cluster (EAMC) job scheduling policy. EAMC-policy is able to optimize the scheduling and placement of jobs by predicting performance and energy consumption of arriving jobs for different hardware architectures and processor frequencies, reducing workload's energy consumption, makespan, and response time. The policy assigns a different priority to each job-resource combination so that the most efficient ones are favored, while less efficient ones are still considered on a variable degree, reducing response time and increasing cluster utilization. We implemented EAMC-policy in Slurm, and we evaluated a scenario in which two CPU clusters collaborate in the same machine. Simulations of workloads running applications modeled from real-world show a reduction of response time and makespan by up to 25% and 6% while saving up to 20% of total energy consumed when compared to policies minimizing runtime, and by 49%, 26%, and 6% compared to policies minimizing energy.

Full text

1 Energy hardware and workload aware job scheduling towards interconnected HPC environments Marco D’Amico , and Julita Corbal´ an Abstract—New HPC machines are getting close to the exascale. Power consumption for those machines has been increasing, and researchers are studying ways to reduce it. A second trend is HPC machines’ growing complexity, with increasing heterogeneous hardware components and different clusters architectures cooperating in the same machine. We refer to these environments with the term heterogeneous multi-cluster environments. With the aim of optimizing performance and energy consumption in these environments, this paper proposes an Energy-Aware-Multi-Cluster (EAMC) job scheduling policy. EAMC-policy is able to optimize the scheduling and placement of jobs by predicting performance and energy consumption of arriving jobs for different hardware architectures and processor frequencies, reducing workload’s energy consumption, makespan, and response time. The policy assigns a different priority to each job-resource combination so that the most efficient ones are favored, while less efficient ones are still considered on a variable degree, reducing response time and increasing cluster utilization. We implemented EAMC-policy in Slurm, and we evaluated a scenario in which two CPU clusters collaborate in the same machine. Simulations of workloads running applications modeled from real-world show a reduction of response time and makespan by up to 25% and 6% while saving up to 20% of total energy consumed when compared to policies minimizing runtime, and by 49%, 26%, and 6% compared to policies minimizing energy. Index Terms—job scheduling, energy modeling, energy-aware, HPC, DVFS, multi-cluster environment F 1 INTRODUCTION OVER the last years, we have witnessed an explosion of hardware heterogeneity. Starting from GPUs used for general purpose, then a high number of processors thanks to instruction sets like ARM [1], Power ISA [2], RISCV [3], and finally special-purpose architectures, e.g., Tensor Processing Units (TPUs) are contributing to increasing the heterogeneity and complexity of HPC machines. In HPC, cited architectures collaborate in computing nodes, which in turn are grouped in clusters. In some cases, different clusters cooperate in the same machine. We call those environments heterogeneous multi-cluster environments. A practical example is an old machine still active, like in the case of Marenostrum III [4] and Minotauro [5], or heterogeneous clusters like Power9, running aside Marenostrum IV [6]. While these machines have their own resource manager, the author hypothesizes that interconnecting the clusters under one resource manager can greatly improve overall utilization. Another example is the prototype machine developed in DEEP-EST project [7], proposing a number of heterogeneous clusters collaborating under the same machine and managed by a unique resource manager. Each cluster has general-purpose CPUs and accelerators, presenting different •M. D’Amico is with Barcelona Supercomputing Center (BSC), Barcelona, Spain. E-mail: [email protected] •J. Corbal´an is with Universitat Politecnica de Catalunya, Barcelona, Spain. E-mail: [email protected] performances and energy behaviors. From a broader perspective, with the advance of memories and interconnections, multiple machines will likely be connected and exchanging workloads, resulting in machines with huge potential but also high hardware diversity and more complex management of resources. On the other side, HPC machines are extremely power demanding, such that the US. Department of Energy set a goal of 20 MW power consumption for an exascale machine. In recent years research and industry put a high effort investigating how to stay under this constraint. At the same time, there is a significant shift from classical performanceoriented to an efficiency-oriented research mentality, with increasing awareness of green and eco-friendly concepts. Top500 [8] list now is accompanied by the Green500 [9], classifying most energy-efficient machines. New metrics are proposed to draw up the list, i.e., TGI metric [10] as an alternative to GFLOPS/W, which evaluates non-computing devices, e.g., memories and disks. Research in energy consumption moved in both hardware and software directions. Regarding hardware techniques, miniaturization plays a vital role in reducing power consumption, but this process is getting more and more to the silicon physical limit, and alternative chemical elements are being studied. Meanwhile, Dynamic Frequency and Voltage Scaling (DVFS) became a popular solution to reduce consumption by dynamically modifying the processor’s voltage and frequency. This technique is beneficial when trying to control the maximum power consumption or when dealing with memory- © 2021 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes,creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. https://dx.doi.org/10.1109/TPDS.2021.3090334 2 bounded applications. In the second case, the so-called memory wall, related to main memory limited bandwidth and high latency, prevents higher frequencies from benefiting the runtime. Regarding software techniques, we can classify them in the following categories, from a low-level to a high-level perspective of the system: •Low level and OS interfaces for DVFS, collection of power metrics. •Energy modeling and estimation for hardware and applications. •Thermal, power, and energy-aware task scheduling in multi-core processors and GPUs, node-level powercapping. •Thermal, power, and energy-aware job scheduling and resource management at cluster level: two main techniques are used to limit power consumption: (1) Overprovisioning: it is a strategy in which more hardware than the one that can be powered is bought, and part of it is selectively shut down or powered on based on the necessity. (2)Powercapping at the node, job, and system level. We identified that most of the research [11] is mainly focusing on power scheduling and controlling, with few considerations on the fact that minimizing power does not mean minimizing the energy. While power-awareness is essential, the authors’ opinion is that power-saving solutions are related to a constraint in HPC machines’ design. On the other side, an energy-aware scheduler needs to optimize the energy-performance trade-off continually. For this reason, power-awareness and energy-awareness techniques are not exclusive, and they can be efficiently combined. Furthermore, as suggested by recent research [12], this whole variety of solutions miss automatic ways of configuring the systems and user’s activity. In particular, the job scheduling and resource management layer manages all the resources and gives an interface to users to use them. At the current state of the art, we identified that it is the user’s responsibility to specify the type of resource needed, the number of resources, and the best energy settings for the submitted job. This complexity is not acceptable in heterogeneous multi-cluster environments, where several computing node types coexist, each one with different characteristics, resulting in confusion and unnecessary knowledge needed for users. Filling those gaps, we propose the Energy-Aware MultiCluster scheduling policy (EAMC-policy). Enabling perjob energy-performance characterization and comparison on heterogeneous environments at job scheduling level, EAMC automatizes jobs’ placement and selection of optimal clock frequencies, respecting a trade-off between performance, energy consumption, and response time. For this purpose, we extended an energy model from the related work [13] to characterize applications’ runtime and energy on different hardware resources. Based on the extended model, we implemented a multi-objective energy and performance classification. We used it to assign a different priority to each job-resource combination, favoring the optimal ones but keeping the non-optimal to balance the load and reduce response time. EAMC-policy reduces user’s overhead and required expertise, error-prone information, and eventual malicious behavior by automatizing the process of selecting frequency and optimal hardware. We integrated it into Slurm [14], a Distributed Resource Management System (DRMS), and the Slurm Simulator [15]. Integrating the prediction model in a job scheduling simulator, we made the Slurm Simulator workload and energyaware, capable of calculating energy consumption based on the type of application and not only the hardware like most of the job scheduling simulators. We modeled various applications and benchmarks behavior in terms of energy and performance for multiple architectures, and we distributed them in a workload generated with Cirne’s model [16]. We studied the case of a heterogeneous two-cluster environment, each one equipped with processors with different characteristics, i.e., number of cores, frequencies, and cache size. We evaluated performance by changing the distribution of jobs that favor the first and the second cluster. By running job scheduling simulations, we observed improvements in energy consumption by up to 20% while reducing response time by 25% and makespan by 6%. This document is structured as follows: Section 2 resumes the state of the art of the energy-aware job scheduling topic, Section 3 describes the preliminary work on job modeling, energy tools, and interfaces. Section 4 explains the EAMC-policy in detail, while Section 5 evaluates it and compares variants of the policy and its parameters with the standard version of the DRMS. Finally, Section 6 resume and conclude the document, with insights on the future work. 2 RELATED WORK Several surveys exist and resume the work done in the energy field for HPC. Czarnul et al. [12] give an overview of leading energy-aware HPC technologies, diversifying computing environments, device types, metrics, benchmarks, energy-saving methodology, and energy simulators. Regarding energy-aware job scheduling, Maiterth [11] resumes some of the leading HPC centers’ energy strategies and their future directions. Most centers work on powercap solutions, a few are working on overprovisioning, and two investigate energy-aware solutions. We found this is a general tendency in the research, with powercap adopted as the primary strategy. In our vision, optimizing energy consumption is crucial, while power limits are a constraint due to machines’ powering. We identified the scarcity of research on energy-aware job scheduling solutions for heterogeneous multi-cluster environments. The authors consider this is a vital field to investigate, given the potential and complexity of those environments. Netti et al. [17] explore the effect of different ways of prioritizing critical resources, the scarcest and most demanded. While this work brings improvement for heterogeneous clusters, it is not energy and workload aware. Extending this work to include energy-aware prioritization could lead to interesting results. Some policies are based on a model that is aware of manufacturing and assembling variability. Chasapis et al. [18] propose different scheduling policies considering processor manufacturing variability under a power constraint. Moore 3 et al. [19] propose a temperature-aware workload placement by detecting cooling inefficiencies that may come from places relatively distant from the temperature sensor, and it tries to minimize heat re-circulations. EAMC can use pernode energy models to model the manufacturing variability. Sarood et al. [20] combine malleability and DVFS to create a scheduling policy that adapts the workload to a strict power budget in over-provisioned systems. Similarly, in a precedent research [21], [22], we used malleability and node sharing techniques to reduce response time, makespan, and energy consumption. The integration of the policies is a promising research direction. Barry et al. [23] explore overprovisioning by powering down nodes in a controlled way using online simulation and controlling system slowdown, keeping it to acceptable values. This method works well on non highly loaded systems, but it is difficult to exploit in the opposite case. Borghesi et al. [24] implement hybrid scheduling techniques for systems under a power cap. They perform a power estimation by using machine learning techniques on historical data instead of a per-job characterization. While this interesting approach does not require the user to specify energy tags, as required by our model, it is not easy to obtain the same per-job precision as our approach. In the research of Rajagopal et al. [25], based on this work [26], researchers integrate into Slurm [14] a poweraware scheduler that allows setting powercaps and collects power metrics at runtime to adjust them. It incorporates an interface to choose a frequency based on user hints and powercap constraints. The solution works with a basic power model using hardware information and user guidance, with estimations representing an upper limit and not the actual power consumption. Our research differs since it optimizes energy, not power. Moreover, no energy input from users is required to set the best frequency, but it is predicted and set automatically using online metrics collection and historical data. In their work, DVFS is used when reaching the powercap, while in our work, we use DVFS to select the optimal performance-efficiency trade-off for each job. In future work, the two policies could be combined, i.e., by choosing optimal frequency for performance and efficiency while assuring a powercap. Auweter et al. implemented two policies: energy-tosolution and best-runtime, to select the most suitable frequency for jobs [27]. Users are required to specify tags for similar jobs to be able to use the proper job characterization. This work is the base of our research, as we implemented a similar description of runtime, energy, and power for jobs. On our side, our instrumentation allows us to verify the correct use of energy tags and recalculate optimal frequencies at runtime. We implemented it in Slurm. We extended the energy model to characterize jobs on heterogeneous clusters. We propose the EAMC-policy to prioritize the most efficient architectures following an energy-performance tradeoff strategy instead of energy-to-solution or best-runtime. 3 PRELIMINARY WORK EAMC-policy selects optimal energy configuration and places jobs based on energy and performance estimations. EAMC requires estimating the job’s runtime and power consumption at scheduling time before the job start. In this section, we describe the designed framework that allows EAMC-policy operation. 3.1 Job power and runtime modeling To estimate runtime, power, and energy consumption for different processors, frequencies, and applications, we revised available energy models from the state of the art. Various energy models and tools for energy prediction are proposed in the literature, with distinct complexity. Some of them are general, based on hardware information only, whereas the chosen [13] is able to model different applications’ behavior for different architectures. The model needs two inputs: 1) application data: a collection of metrics that characterize applications. The model uses runtime, average consumed power, average number of memory transactions per instructions (TPI), and average cycles per instruction (CPI) at the reference frequency fref. 2) hardware coefficients: a set of parameters obtained by running a learning phase based on a set of benchmarks. The chosen model uses six coefficients learned using linear regression: A, B, C, D, E, F. Equations 1, 2, and 3 describe the mathematical model that estimates power consumption Pand runtime Tat the frequency f, starting from a default frequency fref at which application data is given. P(f) = A∗P(fref ) + B∗TPI(fref ) + C(1) CP I(f) = D∗CP I(fref ) + E∗TPI(fref ) + F(2) T(f) = T(fref )∗CP I(f) CP I(fref )∗fref f(3) To implement and use the energy model in a real-world environment, we used Energy-Aware Runtime (EAR) [28]. EAR is a framework that provides an energy-efficient solution for HPC clusters. It includes monitoring, accounting of the applications’ performance, and it integrates energy optimization via DVFS at node and cluster level. It is deployed on SuperMUC-NG [29] and implements the previously described energy model. EAR was used to learn hardware coefficients A, B, C, D, E, and F to model two computing nodes equipped with general-purpose processors with a different number of cores and frequencies, as described in Section 5. In a non-simulated environment, EAR collects application data dynamically at runtime and stores them in the application database to be used by the energy model. In our test simulated environment, we use static data collected by running several applications with diversified computing and memory behaviors in a previous phase. Used applications are described in Table1. To evaluate the proposed policy in our test system, we used and extended an interface developed in the context of the DEEP-EST project [30]. 4 The interface is configurable with different machine configurations and energy models. DRMS uses it to retrieve energy, time, and power estimations. It can be used in a real environment as a standard interface that abstracts underlying energy models or to simulate tools like EAR in simulated environments. To run our simulated environment, we implemented the chosen prediction model by extending the interface. Each hardware configuration is identified by a model id, a series of hardware coefficients, a node configuration, and a range of frequencies. Different hardware configurations can exist for the same node, e.g., a node using or not using the accelerator or a processor using or not using the vector extension. 3.2 Application Database The Application Database AppDB stores application metrics collected at runtime for each hardware configuration. For each tuple, appDB contains fields as the appID, the modelID, the necessary application metrics. In a real system, application database entries for jobs are created on the fly. In the first run, there is no energy information associated with the job. Still, metrics are collected by EAR, which creates an entry in the DB associated with the user-specified appID. For the first run, optimal frequency is not estimated previously, but it is set at runtime as soon as data is available. In the case of multiple clusters but with application data related to only one of the clusters, performance and energy consumption can be estimated from data in the appDB in combination with the hardware model, or they can be evaluated when the application runs for the first time on the hardware. In our simulated environment, for the sake of simplicity, the appDB is static and available at the beginning of simulations. To obtain the application’s data, we run applications in different architectures, collect necessary metrics using the EAR accounting and store them in the database. 3.3 Model extension for multi-cluster environments When collecting metrics, while hardware counters and average power are automatically collected by EAR, the application’s runtime is potentially variable in each run. From the DRMS point of view, it is not straightforward to estimate the runtime parameter with precision, so commonly, users are asked to give a requested time for the job. These values are usually far from the real job’s runtime, and it can also be a default value in the case the user does not specify it. Having requested time as the only time-related information for the job, the energy model assumes this value relates to the primary partition. To calculate a requested time for the other partitions, we extended the energy model to learn a seventh parameter, time coefficient. time coefficient represents the ratio between the main and another partition runtime. This parameter is estimated by running the same application on both hardware architectures. Future work will still consider more complex models that do not need this phase, e.g., based on already collected hardware metrics. 3.4 User interface for job submission When submitting a job, users can avoid specifying which hardware architecture they want to use if they accept that the job will run on nodes automatically selected by the scheduler. We evaluated EAMC on processors from the same family, so we automatically consider the app can run on every available hardware architecture, thanks to binary compatibility. Considering similar architectures is not far from the case of more heterogeneous ones, given the high effort of research and industry in supporting interoperability, hardware-independent programming languages, and libraries. In the case there is no binary compatibility, users can specify different binaries for the same appID. This information can be stored in the appDB or contained in the job script in a format readable by a job script parser. Users are still in charge of specifying the number of requested nodes and processes. The number of threads or processes per node can be configured by using Slurm plugins or precedent work [21], [22] more intelligently. If a memory requirement is specified and some nodes cannot satisfy this requirement, a new number of nodes will be calculated. Using more heterogeneous architectures, on the other side, increases the complexity of parameters for the job. For example, the number of requested nodes needs to increase to maintain a reasonable runtime when switching from a fast and power-consuming resource to a slow and less demanding one. Future work will try to take into account those details by increasing the complexity of the hardware models. It is also the user’s responsibility to specify an appID in the job script to recognize jobs with the same behavior. On the other side, EAR can verify appID related history corresponds once the job starts, giving the DRMS instruments to detect and penalize fraudulent users’ behavior. Suppose runtime collected metrics do not conform to stored information in the appDB for the specified appID. In that case, the application keeps running, and a new frequency is calculated based on runtime data. If the user repeatedly uses wrong appIDs, the DRMS can detect it by analyzing his history of jobs, return a warning message, or take additional actions. 4 ENERGY-AWARE JOB SCHEDULING IN MULTICLUSTER ENVIRONMENTS EAMC-policy is based on priority backfill, with energy predictions influencing the priority of arriving jobs. 1. This section describes the implementation of the policy. In the policy presentation, for simplicity, we use the term partitions as a logical representation of resources inside the DRMS. While ideated with the introduced concept of multiclusters, the policy is generic, independent of the physical and logical representation of the hardware, either physically or non physically separated, grouped, or non grouped in partitions. 1. EAMC code is publicly available [31] 5 4.1 Problem definition We model the proposed policy as an online job scheduling problem. While this problem usually has as the objective function the minimization of the makespan, in our case, we describe it as a multi-objective optimization problem with objectives of the reduction of makespan and energy consumption. The most efficient solver for job scheduling problems in the literature is the backfill algorithm. We used a backfill version with priorities, where each job gets assigned a priority based on several factors, e.g., arrival time, requested number of nodes. To make backfill energy aware, we included a runtime and energy classification of jobs as a further factor for priorities, thus influencing the scheduling order. More precisely, we calculate a priority for each job-partition combination jp, defined as a job jrunning on a partition pin two phases using Equation 4 and 5. The first equation is used to estimate the optimal processor frequency fopt, while the second is used to classify and assign a different priority to each p for j. Firstly, in Equation 4, given a jp, with the objective of calculating the optimal frequency fopt, we define the metric distjp as the Euclidean distance from the origin in the Euclidean space representing time and energy normalized to the minimum value of the respective metrics (Eminjp, tminjp) in the prediction space of jp.t weight parameter increases or decreases runtime estimation weight over energy. distjp =minsEjp(f) Emin jp 2 +t weight ∗tjp(f) tmin jp 2 , ∀f∈[fmin, fmax](4) As a second point, in Equation 5, given multiple jobpartition entries for a job, to classify partitions, we used the Euclidean distance again to calculate part eff, in this case normalizing energy and time among all partitions of the job j(Eminj, tminj) at frequency fopt calculated for each partition with Equation 4. As a result, part eff can be used as an indicator of performance and efficiency, and it can be used to influence the jp priority. Again, t weight affects the weight of time over energy. part eff =sEp(fopt) Emin j 2 +t weight ∗tp(fopt) tmin j 2 , ∀p∈partitions (5) 4.2 Proposed policy Figure 1 and 2 represent two variants of the EAMC-policy. First, a job, represented on the top left, is submitted. At submission time, the policy assigns the priority to the arriving jobs and, in the case of multiple partitions, to each jobpartition combination. We extended this part of the DRMS with the Energy-prediction Priority Module (EPM). Once a different priority for each job-resource is assigned, an EAMC-Scheduler (EAMCS) schedules jobs, from a unique priority queue, in order of priorities. Following on, we describe EPM and EAMCS components. Listing 1: Energy Prediction Module (EPM) implementation 1for each j in submitted jobs : 2i f ( ! j −>partitions) 3j −>p a r t i t i o n s = g et p a r ti t i o n s ( 4j −>app id ) ; 5for each p in j −>partitions : 6EA−API get predictions (p , j −>app id ) ; 7p−>best f va l = get optimal freq (p , 8p−>energy , p−>time ) ; 9i f (p−>energy <emin) 10emin = p−>energy ; 11i f (p−>time <tmin ) 12tmin = p−>time ; 13for each p in j −>partitions : 14p−>p art eff = g et p a rt e fficie n c y ( 15p−>energy , p−>time , emin , tmin ) ; 16sort ( j −>p art itio ns , cmp part eff ( ) ) ; 17 18for each p in j −>partitions : 19p−>p r i o r i t y = s e t p r i o r i t y (p , policy type ) ; 20enqueue ( job queue , j ) ; 4.2.1 EAMC-Scheduler (EAMCS) and Energy-prediction Priority Module (EPM) EAMCS, based on EASY-backfill [32], tries to schedule one job-partition per time in order of priority. Using backfill allows scheduling jobs considering a trade-off between system energy, performance, and average wait time. For example, if the higher priority job cannot run due to lack of resources or a big enough time window for backfilling it, the lower priority job-partition can start if enough resources are available. While not being the best solution in terms of energy and runtime, it is optimal when considering the wait time. Finally, following the Slurm implementation, during backfill, holes in the scheduling reserved by high priority jobs cannot be used by lower priority jobs, and job-partitions from the same job do not influence each other as they request different resources. The proper equilibrium of jobs running on the favored and non-favored partition can lead to optimal performance in the time-energy dimension. Extreme behaviors, where jobs poorly performing are discarded, are managed in one of the policy alternatives described in this section, with details on EAMCS. Listing 1 describe EPM implementation in detail. In the first inner loop, EPM gets runtime and energy predictions using the energy interface and energy models (line 5). The parameter time coefficient is used as a multiplier to predict runtime and energy consumption for different partitions. Then, using Equation 4, it individuates the best frequency for each job-partition, setting it as the default value (line 6). The same loop calculates Eminjand tminj. Once frequencies are picked, the second loop evaluates and sorts partitions according to Equation 5 (lines 11-13). In the last loop, set priority() function (line 16) is in charge of assigning a priority to each partition according to the selected strategy. In the end, the job-partition couple enters the job priority queue. We developed three strategies that implement different ways of assigning priorities to job-partitions, depending on 6 job2(j2) +appID assignpriority j2p2 j2p1 j1p1 j1p2 backfillscheduler jobqueue lowprio highprio Energyprediction module apptracesHWconf energy,power, runtimemodels partition1 partition2 Fig. 1: Job submission and scheduling phases for EAMCReorder policy. The big-box represents the DRMS. Red boxes are modified parts, and green boxes are energy prediction components. With EAMC-Reorder, jobs-partition entries are reordered while keeping the same priority in the queue. how aggressive is the policy in favoring optimal partitions over arrival order of jobs: •EAMC-Reorder: it changes the order of partitions for the same job, prioritizing job-partition with better performance. As a consequence, EAMCS maintains the arrival order of jobs while reordering partitions within the job. Jobs-partition entries are queued similarly to job2 in Figure 1. •EAMC-PriorityInc: it assigns a different priority to each job-partition based on the position in the sorted partition list. As shown in Figure 2, the two jobpartition are inserted in different points of the job queue, creating two separate queues, one with preferred partitions, the others for non-optimal partitions. In this case, EAMCS favors optimal jobpartitions over non-optimal ones, giving less weight to the arrival time order. •EAMC-PriorityInc-V: similarly to PriorityInc, based on the same figure, it assigns a different priority to each job-partition, based on the sorted partition list, with the difference that the priority is lowered if the evaluated metric is below a certain threshold. We set the threshold to 35% from the optimal part eff metric. As an example, if j1p1 exceeds the threshold and j2p2 does not, j1p1 would go at the bottom of the queue. EAMCS first checks all optimal job-partition entries, then the second choices, and finally all the job-partitions below the threshold. 5 EVALUATION The evaluation is based on simulations using the BSC Slurm jobs scheduler Simulator [15]. We modeled a workload of 5000 jobs, with a makespan between 10 and 15 days, job2(j2) +appID j2p2 j1p1 j2p1 j1p2 jobqueue backfillscheduler highpriolowprio assignpriority Energyprediction module apptracesHWconf energy,power, runtimemodels partition1 partition2 Fig. 2: Job submission and scheduling phases for EAMCPriorityInc policies. The big-box represents the DRMS. Red boxes are modified parts, and green boxes are energy prediction components. With EAMC-PriorityInc and EAMCPriorityInc-V, the job queue is separated into higher priority and lower priority queues depending on energy and performance evaluations. depending on applications distribution and used frequency, using Cirne model [16] configured with ANL arrival pattern. Jobs’ average requested nodes are 18.29, and jobs’ average duration is 3.65 hours at the highest frequency. We run the workload in the following simulated clusters: 1) p1: 512 nodes equipped with: 2x Platinum 8168 CPU [2.70GHz-1.20GHz] 24C, TDP 205W and 12 x 16GB DDR4 SDRAM 2) p2: 512 nodes equipped with: 2x Gold 6254 CPU [3.10GHz-1.20GHz] 18C, TDP 200W and 12 x 32GB DDR4 SDRAM 3) p3: 512 nodes equipped with: 2x Gold 6148 [2.4GHz1GHz] 20C, TDP 150W and 12 x 16GB DDR4 SDRAM The number of modeled hardware architectures was limited by the available architectures and permissions needed to collect the necessary data. We sized the number of simulated jobs and system size according to the time limits imposed by the testing environment. We run the EAR learning phase to get coefficients for the simulated nodes on Lenox [33] cluster. We used data from eight applications, running them on one node using all the available cores 2. Table 1 describes the set of applications, presenting different memory and compute profiles, reported through average cycles per instruction (CPI), memory bandwidth (GB/s), the ratio of the runtime and power between partitions, and the ratio between the delta between the maximum and minimum runtime and the delta of frequency ∆r/∆f. The higher the value of ∆r/∆f, the more the runtime scales with the frequency, e.g., ep.D runtime 2. Energy models and applications data is publicly available [34] 7 is highly affected by changes in the processor’s frequency, while STREAM is not. We simulated two different systems: 1) A system made up of p1 and p2. 2) A system made up of p1 and p3. When comparing p1 and p2, seven out of eight applications prefer the first partition in terms of performance. In terms of power, the first partition showed slightly higher power consumption for the CPU component, more than the nominal 5 Watts, while the second partition showed up to double DRAM consumption compared to the first. From an energy perspective, applying Equations 4 and 5, for the tested t weight values, the first seven apps favor the first partition, while app 8 favors the second. For this comparison, app8, STREAM, is a weak scaling benchmark, where an amount of memory is allocated per-process, i.e., per-core. The second partition, having a reduced number of cores, has fewer data to manage, explaining the difference in runtime while showing similar hardware metrics. This is the typical behavior of some memory-bounded applications. EAMC will be aware of STREAM behavior and prioritize the partition with a lower number of cores. For the same comparison, we distributed the applications among the workload by randomly drawing samples from three distributions: 1) 13% benefits from partition 2: while the distribution is uniform among all applications, 87% of the workload favors the first partition. 2) 33% benefits from partition 2: In this 33% of jobs run app8, which favor the second partition, the remaining 67%, uniformly distributed among remaining apps, prefer the first. 3) 50% benefits from partition 2: finally, this case evaluates an even load among the two partitions in terms of the number of jobs per favored resource. 50% of jobs run app8. Comparing p1 and p3, at base frequency, five out of eight applications run optimally on partition 1, while at the optimal frequency, all the applications run optimally on the same partition. In this case, we used a strong scaling version of STREAM that uses a fixed amount of memory per computing node. For this comparison, we run the Workload 13%. The evaluation follows in this section, where we analyze the performance of the developed policies compared to Base Slurm, and policies inspired by Auweter [27], called Min energy and Min runtime. Base Slurm is configured to run jobs at default frequency, i.e., the maximum frequency, and each job is submitted only to the optimal partition. Min runtime runs jobs at the frequency that minimizes the runtime, while Min energy runs jobs at the frequency minimizing the job’s consumed energy. For those policies, jobs are submitted to both partitions, and each job-partition gets the same priority. The scheduler tries to run the job in the default order, first in the first specified partition and immediately after in the second, not establishing the favored one. Evaluated metrics: Fig. 3: Savings in terms of makespan, average response time, sum of runtimes and sum of jobs’ energy over Base Slurm for EAMC policies, Min runtime and Min energy. •Makespan: The difference between last job end time and first job arrival time. •Average response time: The average of jobs’ response time, defined as the difference between end time and arrival time •Sum of the jobs’ consumed energy: sum of individual jobs’ consumed energy, not including idle nodes. •Sum of apps runtime: the sum of individual jobs’ runtime, used to understand the impact of EAMCpolicies on the jobs’ performance. •Avg frequency: the average of picked frequencies for the whole workload, or per partition. •Percentage of apps in the optimal partition: The percentage of jobs scheduled in the favorite partition. We first analyze the system p1/p2. We compare the EAMC policies variants to Base Slurm, Min runtime, and Min energy. We then compare Min runtime and Min energy to our policy when changing the applications’ distribution in the workload. Finally, for system p1/p3, we analyze the 13% case. For both systems, p1/p2 and p1/p3, we analyze the performance when changing the t weight parameter, i.e., giving equal or 1.5x more importance to performance over energy efficiency. 5.1 Comparing to Base Slurm This section compares EAMC-policy variants’ performance with Base Slurm, where jobs are submitted only to the optimal partition, and with Min runtime and Min energy, where jobs are submitted to multiple partitions with a frequency that minimizes runtime or energy. In Base Slurm, jobs run at the default frequency, i.e., the maximum frequency, system p1/p2 is considered. Figure 3 shows the improvement in percentage over the Base Slurm for the analyzed metrics. Applications are equally distributed among the Workload 13%, and EAMC is configured with t weight=1. Analyzing makespan and response time, asking all the available partitions has a considerable impact on all the evaluated scenarios, independently from the job-partition priority, and up to 39%, and 64% for Min runtime. The poor performance obtained by Base Slurm suggests that the more the systems are interconnected, the better the load can be distributed among clusters. Besides, workload 8 AppID Name CPI GB/s Runtime (s) Power (W) ∆r/∆f (s2) 1 lu.C 0.67 / 0.66 / 0.57 70.95 / 63.73 / 59.7 52.49 / 57.45 / 62.78 386 / 367 / 331.25 2.38 / 2.87 / 4.18 2 ep.D 0.60 / 0.60 / 0.60 0.03 / 0.04 / 0.03 64.20 / 74.20 / 86.59 349 / 337 / 290.32 5.67 / 6.37 / 8.97 3 bt-mz.C 0.38 / 0.39 / 0.38 23.33 / 19.96 / 17.7 43.95 / 51.32 / 57.70 417 / 411 / 342.73 3.52 / 4.05 / 5.65 4 sp-mz.C 0.44 / 0.50 / 0.42 66.09 / 52.24 / 51.17 46.05 / 59.76 / 58.58 460 / 427 / 375.47 3.07 / 3.34 / 4.64 5 lu-mz.C 0.62 / 0.62 / 0.63 19.87 / 21.14 / 18.89 59.21 / 61.41 / 65.98 355 / 360 / 308.81 4.75 / 4.70 / 6.33 6 ua.C 0.99 / 0.95 / 0.87 51.29 / 49.05 / 45.5 46.96 / 53.11 / 54.51 368 / 360 / 324.6 1.74 / 1.84 / 2.64 7 DGEMM 0.40 / 0.40 / 0.38 66.43 / 62.76 / 50.87 47.40 / 52.60 / 61.97 491 / 497 / 366.25 3.29 / 3.08 / 4.09 8 STREAM 4.46 / 3.67 / 3.55 138.87 / 137.47 / 136.86 107.55 / 81.43 / 110.34 381 / 365 / 330.52 0.16 / 0.06 / 0.13 TABLE 1: Set of applications and their characteristics. Metrics specified in the order of partitions, in the format p1 / p2 / p3. and hardware-aware scheduling can save energy without sacrificing system performance. Base Slurm can schedule all the jobs in the optimal partition, obtaining the best runtime for each job, but at the cost of the time needed to wait for the favored architecture. Min runtime achieves excellent results for time metrics, at the cost of consumed energy, up 12% compared to Min energy. Min energy obtains 5% energy saving while not scheduling jobs in the optimal performance, but it gives up half of the response time compared to other policies. Regarding EAMC policies, we can observe that Reorder underperforms slightly compared to other variants because favoring the job’s arrival time order leads to fewer possibilities for the scheduler to favor optimal job-partitions. PriorityInc and PriorityInc-V perform similarly to Min runtime in terms of makespan and response time while increasing energy savings by 9%. Increased jobs’ runtime, given the lower processors’ frequencies, is compensated by the workloadaware scheduling of jobs. Min energy and Min runtime only schedule jobs to the first available partition, achieving a low number of running jobs in the optimal partition, as we can observe in Figure 5. EAMC policies, particularly PriorityInc versions, can schedule more jobs in the favored partition, especially in p2, running almost 100% of STREAM, achieving shorter runtime and higher energy savings. 5.2 Changing t weight parameter for three apps distributions As previously commented, seven out of eight applications in our set favor the first partition. In this evaluation, we test the different application distributions: 13% (uniform distribution), 33%, and 50% of jobs benefiting p2. The objective is to evaluate performance for different load levels per optimal partition by changing the number of jobs running the app with id 8, STREAM. We run all the EAMC policies, and we tested two values for t weight parameter for Equation 4 and 5: 1 and 1.5. The latter increases the weight of performance over energy in the EPM by 50%. This evaluation focuses on the trend between the three workload distributions, presenting more insights on the 33 and 50 scenarios, since we already analyzed Workload 13. Figure 4 reports a summary of time and energy metrics for Workload 13, 33, and 50, normalized to Min runtime, and expressed as increase in percentage over it. First, moving from distribution 13 to 33 and 50, we notice increasing performance for EAMC when comparing to the other two policies. In those cases, the scheduler can schedule jobs in the optimal partition without sacrificing the response time, as Base Slurm does. We observe a reduction of response time, by up to 2%, 10%, and 25% respectively for Workloads 13, 33, and 50, when using PriorityInc. Looking at single partitions, we notice that p2 shows the most significant improvement in response time, up to 53% in PriorityInc-V-1.5. This behavior is compensated by the increased response time in p1, given the lower processor frequency, up to 34%, but still increasing overall response time thanks to better usage of p2. While the Reorder policy does not reach other EAMC variants in total average response time, it shows more balance in response time between partitions. Gains in response time are accompanied by good results for makespan, ranging from - 5% to 6% compared to Min runtime. Min energy does not perform well in time-related metrics. It increases the response time and the makespan by up to 51% and 23% compared to Min runtime, and by up to 65% and 25% compared to PriorityInc, for Workload 33 and 50, and by 93% in Workload 13 Min runtime energy consumption increases by 17% and 21% for Workloads 33 and 50 with respect to Min energy. EAMC policies show up to 4% of improvement over Min energy and up to 20% over Min runtime. Energy savings are higher in the second partition, where app8, with almost null frequency scaling, can benefit from a reduced runtime and energy when running over it. PriorityIncV only slightly improves energy savings over PriorityInc, showing a low number of jobs that perform particularly poorly on a single partition. We identified that app2, ep.D, was affected by the policy variant’s threshold out of the eight applications. Figure 5 shows the percentage of applications scheduled in the optimal partition. Due to EAMC scheduling, the number of jobs running in the optimal partition reaches 82%, compared to 55% in the compared policies, motivating the improvements in the described results. The average frequency for EAMC policies is close to Min runtime values concerning p1 and close to Min energy in p2. With the increasing number of app8 in Workloads 33 and 50, the average frequency in p2 decreases, as this app’s optimal frequency is the lowest. As a final remark, in Workload 50, Base Slurm average response time and energy consumption improve over Min runtime by 19% and 10%, at the cost of 7% higher makespan. This is a specific situation where the load is overall balanced among partitions, but we observed that Base Slurm could not adapt to changes in the load. Besides, while the load is balanced overall, it is not balanced continuously, so EAMC can take advantage of temporary 9 (a) Time metrics for Workload 13% (b) Energy metrics for Workload 13% (c) Time metrics for Workload 33% (d) Energy metrics for Workload 33% (e) Time metrics for Workload 50% (f) Energy metrics for Workload 50% Fig. 4: Main time (left) and energy (right) metrics normalized to Min runtime policy and reported as improvement in percentage over it for the first system configuration (p1 and p2). Reported Reorder, PriorityInc-V, and Min energy. Fig. 5: The average percentage of applications scheduled in the optimal partition for Workloads 13%, 33%, and 50% for the system p1/p2. The average of the two partitions and the partition p2 are presented. unbalance, further reducing time and energy metrics. To conclude, PriorityInc-V-1.5 achieves better timerelated metrics, with little to no impact on energy consumption. Thanks to the ability to reduce the priority of jobs that have a high impact on performance when running on the secondary partition, it further improves the scheduling. This policy would be our pick for systems on which the performance has primary importance. PriorityInc-1 and PriorityInc-V-1 achieve higher energy saving at a small cost of response time. This pick will be the favorite if energy savings or power costs have the highest importance. 5.3 Changing t weight parameter for system p1/p3 As a final evaluation, Min energy and Min runtime to EAMC for the second system configuration. For this evaluation, at maximum frequency, 5 out of 8 applications perform optimally on p1, 3 on p3, while at optimal frequency, all