Improving PUE of Legacy Computer Room
Abstract
Power Usage Effectiveness (PUE) is crucial to the operation of modern data centers. We investigated possible AI technologies to predict PUE trend. Through the investigation of various AI technologies, the XGBoost provided the best model for predicting the PUE values and is serve as the base model for future optimization.
Full text
XXX-X-XXXX-XXXX-X/XX/$XX.00 ©20XX IEEE Improving PUE of Legacy Computer Room – (I) Chester Chen Data Science and Technology Division National Center for Highperformance Computing Taiwan [email protected] Yen-han Chiang High Performance Computing Division National Center for Highperformance Computing Taiwan [email protected] Yo-chin Lin Data Science and Technology Division National Center for Highperformance Computing Taiwan [email protected] Yu-bin Fang High Performance Computing Division National Center for Highperformance Computing Taiwan [email protected] Weicheng Huang High Performance Computing Division National Center for Highperformance Computing Taiwan [email protected] Abstract— The PUE performance is crucial to the operation of modern data centers. The current work, which is in its first stage of a series activity to enhance power consumption of a data center, investigated possible AI technologies to predict PUE trend. Data preprocessing and AI technologies attempted are provided in this work. Through the investigation of various AI technologies, the XGBoost provided the best model for predicting the PUE values and is serve as the base model for future optimization. Keywords— AI, PUE, Time Series Data (key words) I. INTRODUCTION Over the years, the warning of global warming has become the reality. As a result, the power saving become one of the doctrines for saving the Earth. On the other hand, Supercomputer centers, as well as Data Centers, keep growing in its capacity and thus to serve the ever growing scientific and technological community better. As a consequence, the power consumption of a data center/supercomputing center, keeps reaching new highest point and can easily go up to 45% of the operation cost of a data center [1]. For a new data center, various power saving mechanisms, such as Direct Liquid Cooling (DLC) can be introduced. However, for an computer room that exists for long time, the budget to renovate the infrastructure is hard to come by. Therefore, it is the goal of this work to investigate the possibility to fine-tune the existing infrastructure to save the power consumption. Various vendors have attempted to optimize the power consumption. For example, Google took the data-driven approach to optimize its data centers’ power performance. By applying DeepMind AI model to the 19 selected features, and data collected over at least 2 years, Google’s prediction accuracy reached 99.6% on its Power Usage Effectiveness (PUE) prediction. Followed the implementation of Artificial Intelligence (AI) enhanced controlling system, it ended up with 40% energy savings in cooling the Google’s Data Centers [3][4][5][6][7][8][9][9][10][11]. The thermal optimization solution of Siemens provides a package of facility improvement measures based on data collection and data analysis, followed by the integration of conventional physic-based modelling and machine learning techniques to optimize the space cooling [12]. The current work attempted to adopt AI technologies to improve the PUE of current infrastructure of an existing machine room of NCHC. The attempt is divided into 3 stages; 1). PUE Prediction, 2). Optimization, 3). Semi-real time adjustment. This paper will focus on the first stage to show the work of predicting the PUE based on historical data. II. METHODOLOGY A. Data collected In the application of AI techniques to the optimization of PUE, the data is essential. A year-worth of data collected over computing facility in Taichung of NCHC is used for the task. The infrastructure of the computer room that corresponds to the data collected is shown in Figure 1. It consists of 2 redundant cooling water supplying systems, CWP4/CHP4/CH5 vs. CWP5/CHP5/CH5. Only 1 set of colling water supplying system will be activated at any moment of operation time. The water temperatures at various points are monitored constantly with their data collected. Other related data, such as power consumption and frequency of fans and chillers, electricity consumption of computers, ambient temperature and humidity etc. are all collected and investigated. There are 36 features collected. The detailed list of the data collected is not presented here due to the limitation of space. The data collected was stored in MongoDB under the AIOT platform of NCHC for further data preparation.
Figure 1 : Infrastructure of the computer room corresponds to the data collected for analysis The data stored in MongoDB was further formatted in CSV format. The data was, then, ready for pre-processing for later use. The process included remove of outlier, remove of duplicated data, and so on. Different features were sampled with different frequency, therefore, resampling was performed over all the data. For current study, 5 min. was the sampling interval of all features. For those data with higher sampling frequency, the mean of 5 min. interval was used as the sampled data. The effect of re-sampling on the features is shown in Figure 2 for illustration purpose. The pattern of ambient temperature and Figure 2 : Example of data sampling of power consumption chill water temperature are shown in Figure 3 and Figure4. It can be seen that the trend of chiller water temperature and ambient temperature were consistent and thus to enhance the confidence of the data feature selected. Figure 3 : Data example of ambient temperature Figure 4 : Data example of chiller water temperature The SOP for data preprocessing is illustrated in following Figure 5. B. Feature Selection Figure 5 : SOP of data pre-processing Before the modeling training, two features selection mechanisms were implemented to further reduce the number of features to be analyzed. Through the Pearson Correlation analysis, shown in Figure 6, 25 out of 36 features remained for further analysis. Followed by a Random Forest Analysis, shown in Figure 7, was performed to further reduce the number of features from 25 to 11 and thus to make the analysis feasible. III. EXPERIMENTS A one year worth of data was used to train the AI model for predicting future PUE. Out of the full dataset, 90% of the data were used as training data while the rest 10% were used for validation. By comparing the results from LSTM, RF, ARIMA and XGboost, the result of XGBoost was the best choice with R2 as 0.9427 with data of 5 min. interval. Experimental data is shown in the following Table 1. TABLE I. R2_SCORE FOR VARIOUS AI ANALYSIS FOR PUE AI Model R2_score dataset LSTM 0.5680 11 days/5 min RF 0.9405 1 year/5 min. ARIMA 0.5600 1 year/ 1 hour XGBoost 0.9427 1 year/ 5 min. In addition, to investigate which time interval will serve the best purpose of predicting PUE, 3 experiments were conducted using XGBoost for data time intervals as 5 min., 1 hour, and 12 hours. In this comparison, the training data size vs. verification data size was 1:1, using half year data. The result was shown in Table 2, the case of 1-hour interval had lowest RMSE and the highest R2_score. However, the results are about the same. Identify applicable funding agency here. If none, delete this text box.
TABLE II. COMPARISON OF TIME INTERVAL EFFECT OF PUE PREDICTION Sampling Frequency 5 min. 1 hour 12 hours RMSE 0.00678 0.00517 0.00663 R2 0.95110 0.97032 0.94944 Figure 6 : Pearson Correlation Analysis Figure 7 : Random Forest Analysis IV. FUTURE WORK As pointed out in previous session, current work is the first stage of the investigation of PUE performance enhancement for the existing computer room. In this stage, the features as well as the PUE prediction model has been defined through the investigation performed in this work. The next step is to optimize the features to improve the PUE performance. ACKNOWLEDGMENT Current research is supported by NSTC project “Secure and Distributed Data Cloud for AI Platform between Taiwan and Japan”, NSTC-107-2923-E-492-002-MY4, and NARLabs Innovation Project “Federated Learning Framework for International Collaboration – using scheduling optimization of computing facility as example”. REFERENCES [1] “Datacenter Efficiency and PUE Measurement,” Technical Report of Delta Electronics, Inc., https://www.deltapowersolutions.com/enus/mcis/white-paper-datacenter-efficiency-and-pue-measurement.php [2] Jim Gao, Google, “Machine Learning Applications for Data Center Optimization,” https://research.google/pubs/pub42542/, https://static.googleusercontent.com/media/research.google.com/en//pub s/archive/42542.pdf, 2014. [3] https://www.google.com/about/datacenters/efficiency/ [4] https://sustainability.google/progress/projects/machine-learning/ [5] https://blog.google/inside-google/infrastructure/better-data-centersthrough-machine/ [6] https://blog.google/outreach-initiatives/environment/deepmind-aireduces-energy-used-for/ [7] https://www.linkedin.com/in/jimgao/#experience [8] https://www.deepmind.com/blog/deepmind-ai-reduces-google-datacentre-cooling-bill-by-40 [9] https://www.deepmind.com/blog/safety-first-ai-for-autonomous-datacentre-cooling-and-industrial-control [10] https://techwireasia.com/2020/05/has-google-cracked-the-data-centrecooling-problem-with-ai/ [11] Nevena Lazic, Tyler Lu, Craig Boutilier, Moonkyung Ryu, Eehern Wong, Binz Roy, Greg Imwalle, “Data center cooling using model-predictive control,” Proceedings of Advances in Neural Information Processing System 31 (NeurIPS 2018), 2018. [12] “How Artificial Intelligence is Cooling Data Center Operations,” https://assets.new.siemens.com/siemens/assets/api/uuid:fdf5b798-479e4f74-80d2-e472daf1f09a/final-wsco-white-paper-ai.pdf, 2020.