scieee AI-readable full text Open interactive document viewer

Analyzing Energy Consumption in Battery Electric Vehicles: A Statistical-based Approach

Polymeni, Sofia; Spanos, Georgios; Pitsiavas, Vasileios; Lalas, Antonios; Votis, Konstantinos; Tzovaras, Dimitrios

Abstract

The shift to battery electric vehicles (BEVs) is a key step towards reducing greenhouse gas (GHG) emissions and decreasing dependence on fossil fuels, ultimately achieving Sustainability Development Goals and more specifically, a more sustainable form of transportation. However, while BEVs help eliminate tailpipe emissions, understanding and consequently optimizing their energy efficiency is still an ongoing task, as battery life is often affected by various driving conditions, including driving behavior, terrain and sometimes weather. Traditional linear modeling techniques often lack the ability to capture the non-linear and complex relationships found in real-world scenarios. For this reason, this study introduces a comprehensive statistical framework for energy consumption modeling for BEVs, specifically developed with the purpose to explain and estimate energy consumption through an in-depth exploratory analysis. Building on the acquired knowledge from the data, two distinct energy consumption modeling approaches are developed, namely a multiple principal component regression and a custom generalized additive model, incorporating principal components and smoothing splines polynomials. The experimental results showcase that the integration of smoothing splines further refines the proposed model, achieving an accuracy of around 90% and improving performance by approximately 20% on average across all clusters, compared to the simpler linear variant.

Full text

JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 1 Analyzing Energy Consumption in Battery Electric Vehicles: A Statistical-based Approach Sofia Polymeni, Georgios Spanos, Vasileios Pitsiavas, Antonios Lalas, Konstantinos Votis, and Dimitrios Tzovaras Abstract—The shift to battery electric vehicles (BEVs) is a key step towards reducing greenhouse gas (GHG) emissions and decreasing dependence on fossil fuels, ultimately achieving Sustainability Development Goals and more specifically, a more sustainable form of transportation. However, while BEVs help eliminate tailpipe emissions, understanding and consequently optimizing their energy efficiency is still an ongoing task, as battery life is often affected by various driving conditions, including driving behavior, terrain and sometimes weather. Traditional linear modeling techniques often lack the ability to capture the non-linear and complex relationships found in real-world scenarios. For this reason, this study introduces a comprehensive statistical framework for energy consumption modeling for BEVs, specifically developed with the purpose to explain and estimate energy consumption through an in-depth exploratory analysis. Building on the acquired knowledge from the data, two distinct energy consumption modeling approaches are developed, namely a multiple principal component regression and a custom generalized additive model, incorporating principal components and smoothing splines polynomials. The experimental results showcase that the integration of smoothing splines further refines the proposed model, achieving an accuracy of around 90% and improving performance by approximately 20% on average across all clusters, compared to the simpler linear variant. Index Terms—Battery electric vehicle (BEV), energy consumption modeling, generalized additive model (GAM), k-means clustering, principal component analysis (PCA), smoothing splines. I. INTRODUCTION THE promise of reducing greenhouse gas (GHG) emissions and decreasing dependence on fossil fuels has made battery electric vehicles (BEVs), and all electric vehicles (EVs) in general, a critical component of the transition toward Sustainability Development Goals (SDGs) achievement and specifically toward more sustainable transportation [1]. Thanks to their reliance solely on electric energy [2], BEVs use rechargeable battery packs to power electric motors, thereby eliminating tailpipe emissions and significantly reducing the carbon footprint associated with transportation [3]. However, ensuring that BEVs meet their environmental promise, not only requires advancements in battery and motor technology, but also an in-depth understanding of their energy consumption patterns under diverse operating conditions, including auxiliary energy demands, terrain, outside temperature, as well as the effects of heating, ventilation and air conditioning (HVAC) in-vehicle systems [4]. Thus, The authors are with the Information Technologies Institute, Centre for Research and Technology Hellas, 57001 Thessaloniki, Greece (e-mail: [email protected]; [email protected]; vasilispitsiav[email protected]; [email protected]; kvo- [email protected]; Dimitrios.Tzov[email protected]). Manuscript received ; revised . (Corresponding author: Georgios Spanos) accurate modeling and analysis of such energy consumption influencing factors can help in energy efficiency optimization and, consequently, to optimize driving behavior on the road. So far, traditional statistical approaches, such as multiple linear regression (MLR), have often been employed for energy consumption estimation [5], [6], but are often considered biased models due to their inherent linear assumptions that fail to capture the complex relationships in real-world data [7]. On the other hand, modern statistical approaches, such as smoothing splines and additive modeling, provide a more flexible framework for analyzing energy consumption by capturing non-linear trends that are often overlooked by linear models, allowing for a deeper understanding of the relationships between each influencing variable [8]. This study aims to fill this gap by proposing an explanationfocused statistical framework for energy consumption modeling in BEVs. Instead of focusing on predictive accuracy, our primary goal is to discover and analyze the correlations between independent variables (e.g., driving behavior, environmental factors and vehicle systems) and energy consumption. For this reason, the proposed framework systematically employs principal component analysis (PCA), to identify the most influential variables and reduce data dimensionality, and then applies k-Means clustering on the established optimal principal components, to group instances, namely driving situations, where both the estimator variables and energy consumption itself exhibit similar features. Finally, in each cluster-specific and PCA-reduced dataset, smoothing splines are applied to further understand the non-linear impact of these principal components on energy consumption. The contribution of this paper is twofold and mainly related to understanding and modeling energy consumption in BEVs. Firstly, this work develops a comprehensive statistical framework that combines advanced exploratory analysis, including PCA, clustering and smoothing splines analyses, to reveal any underlying, non-linear correlations between key energy influencing factors in BEVs. In contrast to prior work, which primarily employs black-box machine learning (ML) models for modeling EV energy usage, this study emphasizes interpretability and transparency. The integration of PCA and clustering methodologies further simplifies the exploratory analysis by identifying significant data relations but in a reduced-dimensional space, offering reduced variance. On the other hand, the development of two energy consumption modeling frameworks using a multiple principal component regression approach and a custom generalized additive model, incorporating the polynomial transformations derived from the smoothing splines as estimators, offers key insights for the © 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. DOI: 10.1109/TTE.2025.3618308 JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 2 development of next-generation BEV technologies, creating an end-to-end statistical framework capable of revealing the non-linear interactions among driving behavior (e.g., speed, acceleration), environmental factors (e.g., elevation, temperature), and vehicle systems (e.g., HVAC). The remainder of this paper is organized as follows. In Section II, related energy consumption modeling approaches from the most recent literature are presented, with all EVrelated approaches using the original version of the vehicle energy dataset. Section III describes the complete methodology for developing the proposed energy consumption modeling framework, step-by-step, while in Section IV, the comparative analysis of our custom framework compared to a simpler linear approach is offered. Finally, Section V summarizes the conclusions of this work. II. RELATED WORK Proper understanding and regulation of EV energy consumption is a key enabler to their widespread adoption and efficient operation. Recent research has been focused on predicting and optimizing EV energy usage using statistical methods and ML algorithms alongside battery-level advances, including neuromorphic state of charge (SoC) estimation [9], electrothermal SoC and temperature modeling [10], transformer-based internal short-circuit detection [11], and deep learning for early battery life prediction [12]. However, the potential of more interpretable frameworks to explain the complex, non-linear data relationships has been relatively underexplored. A. Energy Consumption Management for Electric Vehicles Research on energy consumption modeling for EVs and hybrid EVs (HEVs) has gained significant attention, with many methodologies being used for accuracy, efficiency, and practical enhancement. In this context, various studies from the current literature have utilized the original energy vehicle dataset (VED) [13] to analyze energy consumption patterns and develop corresponding prediction models, with a key focus on identifying energy consumption influencing factors and developing efficient energy management strategies. Statistical methods, such as PCA and clustering, were frequently used for dimensionality reduction and data segmentation. For driving scenarios classification, more common approaches included the implementation of PCA and kMeans++ clustering, or any of its variations, to provide insights into how factors like speed, acceleration and road topography affect energy consumption in EVs [14], [15]. On the other hand, ML approaches, such as the improved soft-actor-critic reinforcement learning (RL) algorithm [14], or deep deterministic policy gradient methods [15], were also applied for energy management optimization, by combining objectives like fuel economy, battery aging costs, and state-of-charge sustainability. Probabilistic models further enhanced energy predictions by capturing uncertainties in road segment-level and route-level energy consumption, employing deep neural networks (DNNs) or long short-term memory (LSTM) models for prediction tasks [16]. Hybrid methodologies were found to be quite effective in connected BEV systems, with feature selection using PCA and random forests (RFs), followed by model comparisons involving algorithms like XGBoost, artificial NNs (ANNs) and LSTMs, helping to refine energy consumption predictions [17]. Other approaches emphasized spatiotemporal modeling, using MLR with stepwise selection and spatial autoregressive models, and incorporating geographic and temporal factors to predict trip-wise electricity demand, ultimately revealing significant variations in energy consumption across different regions and times [18]. Eco-routing optimization models integrated energy consumption estimation with traffic flow and congestion-level analysis, using backpropagation neural networks (BPNNs) combined with traffic-related metrics to evaluate energy-efficient routes for diverse vehicle types [19]. Finally, performance comparison studies explored a range of ML and ANN models to identify optimal estimators for BEV energy consumption, providing insights into variable relationships and model performance [20]. B. Generalized Additive Modeling for Energy Analysis Recently, generalized additive models (GAMs) have gained considerable attention for energy consumption modeling, particularly due to their ability to capture non-linear relationships while preserving model interpretability [21]. However, although their application in the field of EVs remains limited, various studies across the broader energy modeling domain have demonstrated their potential in providing transparent and reliable energy estimation frameworks. More specifically, in electricity demand forecasting, GAMs have been widely used to model short-term load variations by incorporating environmental and temporal predictors, including weather conditions, time-of-day and seasonal cycles, as smooth terms, to effectively predict future electricity usage [21], [22]. In such contexts, incorporating smoothing functions, such as cyclic splines for daily and weekly periodicity, allowed for more flexible representations of load behavior, contributing to enhanced forecasting accuracy and lower generalization error compared to traditional linear regression models [21]. Within building energy modeling, GAM-based techniques were employed to estimate consumption across commercial and residential facilities, taking into account non-linear interactions between variables, such as outdoor temperature, internal occupancy rates, and HVAC system load, and utilizing tensor product smooths or interaction terms to capture their complex relationships [23]. Such framework were found to be particularly effective at identifying peak load periods under varying usage scenarios. Similarly, renewable energy systems, such as wind power plants, adopted spatiotemporal GAMs to assess production yield by modeling power curves and environmental dependencies over time and across different regions [24]. In such solutions, smooth terms for wind speed, temperature, and turbine operation hours were used to address the uncertainty of energy production and optimize post-construction energy yield assessments. Hybrid approaches, combining GAMs with NN methodologies, have also been explored for high-resolution JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 3 Fig. 1: Flowchart of the proposed BEV energy consumption modeling framework. peak demand estimation, further enhancing prediction robustness by combining interpretable smooth terms with the representational capacity of deep learning [25]. C. Problem Definition and Contribution While prior studies made significant advancements in energy consumption modeling, they mainly focus on predictionoriented approaches, such as route-level forecasting [16], spatiotemporal energy estimation [18], or HEV-specific energy management [14], [15]. As a result, they typically fall into two main estimation categories, namely black-box predictive modeling using deep learning or reinforcement learning [14], [15], [17], and route-level estimation models with spatiotemporal characteristics [16], [18], which, though effective, often sacrifice model transparency and fail to provide deeper insights into the underlying factors that drive energy usage. On the other hand, despite their growing presence in energy-related domains [22]–[24], GAMs have yet to be extensively adopted for EV energy consumption modeling. For this reason, the proposed GAM-based statistical energy consumption modeling framework prioritizes data-driven exploration and model interpretability, capturing more complex, non-linear data correlations and allowing for a more indepth analysis of variable interactions under different driving scenarios. In addition, the statistical analysis performed in the extended VED dataset version [26], aids in this direction, enabling the inclusion of more diverse and granular driving data. III. ENERGY CONSUMPTION MODELING In this section, the complete methodology followed for the development of the proposed energy consumption modeling framework is analyzed step-by-step. As seen in the process flowchart in Fig. 1, the proposed modeling methodology begins with the data collection and preprocessing step, where the finalized CERTH-BEV dataset is produced from the raw data values. PCA analysis is then implemented to reduce the dimensionality of the original CERTH-BEV dataset, followed by an additional clustering analysis, applied to the optimal number of principal components to identify any operational scenarios. To capture nonlinear data relationships, smoothing splines are also utilized in the same cluster-specific and PCA-reduced datasets. Finally, building upon all these analyses, two distinct statistical energy consumption modeling approaches are developed: (i) a multiple principal component regression model, using the optimal principal components as estimators, and (b) a generalized additive model, incorporating the principal components transformed by the smoothing splines polynomials. A. Data Collection and Preprocessing While numerous datasets exist about BEVs, the VED dataset [13], collected from 383 personal cars in Ann Arbor, Michigan, stands out among them due to its extensive data scale and the inclusion of various vehicle types, including internal combustion engine (ICE) vehicles, HEVs, plug-in EVs (PHEVs), as well as BEVs. However, for the purposes of this work, the extended VED (eVED) dataset [26] was selected, which includes new additional attributes associated JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 4 TABLE I: Statistics from the original eVED [26] dataset vs. our processed versions Dataset Vehicle IDs Records Time Frame Sampling Time eVED Non-specific 3,202,800 ms Random eVED-BEV 10; 455; 541 476,308 ms Random CERTH-BEV 10; 455; 541 4,733 min 1-minute with trip GPS records, such as road speed limits, elevation, intersections, traffic signals, bus stops and railway crossings from APIs, as well as energy consumption calculations for each timestamp, contrary to the previous version. It should be noted that, given the dataset’s lack of drivetrain-specific telemetry (e.g., motor torque, regenerative braking), direct modeling of propulsion power and other auxiliaries was not feasible. For this reason, this study focuses on modeling and understanding the factors influencing total energy consumption. Like its previous version, the original eVED dataset contains trip data from various vehicle types, including three BEVs (vehicle IDs: 10, 455, and 541), which are the only data that were ultimately used for the proposed energy consumption profiling methodology. In addition, for each vehicle in this BEV dataset (eVED-BEV), separate CSV files were created, in which the acceleration of the vehicle was calculated based on both the time and speed difference between each consecutive timestamp. For each of the derived CSV files, the timestamp was then reset and recalculated to ensure temporal continuity and offer a better representation of the vehicle’s movement dynamics. Following the establishment of driving cycles, the next step involved creating fixed 1-minute intervals for the entire dataset based on the millisecond-level timestamp of each CSV file. This approach enabled a more comprehensive analysis of energy consumption patterns for each vehicle. Within each of these intervals, the vehicle’s total energy consumption was calculated by summing the existing energy consumption values from the original dataset according to the time grouping. In addition to the energy consumption calculation for each time interval, the finalized dataset (CERTH-BEV) also includes various statistical measures—mean, maximum, minimum, average, median, and standard deviation—of key factors influencing energy consumption, such as air conditioning/heating power, elevation, outside air temperature (OAT), vehicle speed, and acceleration. The calculation of these statistical measures is inspired by the works of Spanos et al. [27], [28] to better capture the distribution of these factors within a specific time interval (1 minute in this case). The statistics of the final CERTH-BEV dataset, which incorporates all the aforementioned data for all three BEVs, as opposed to the original eVED dataset, are presented in Table I. B. Principal Component Analysis Following the data preprocessing, PCA was implemented on the final CERTH-BEV dataset for dimensionality reduction [28]–[30], helping to identify the most influential variables impacting energy consumption in BEVs by also reducing (a) Explained variance (b) Cumulative explained variance Fig. 2: Scree plot representing (a) the proportion of the dataset’s total variance captured by each individual PC, and (b) the total proportion of variance explained by a certain number of PCs, added together sequentially. variance [7]. This way, all the information can be condensed into fewer principal components that still capture the majority of the data variance, while also limiting any data redundancies. To determine the appropriate number of principal components (PCs), a scree plot [31] was produced (Fig. 2a), representing the eigenvalues of each principal component in descending order, to guide the selection of the optimal number of components (p-value). Based on this plot, several p-values were evaluated, including 3, 6, 10, 16 and 22 PCs, assessing their ability on effective data representation while also preserving variance. As seen from Fig. 2, the first 16 principal components were selected as the optimal solution based on the elbow method, offering a cumulative explained variance of around 99% (Fig. 2b). Fig. 3 depicts the loadings of each one of the top-16 PCs, with each chart corresponding to a specific principal component, highlighting the variables that contribute most to the vehicles’ energy consumption explanation, with a cumulative explained variance just around 99%. As seen in the charts, PC1 variables like heater power and OAT showcase the strongest positive and negative contributions, respectively, suggesting that PC1, which explains the biggest variance, primarily captures the effects of HVAC systems on energy consumption. On the other hand, in PC2, a different set of dominant variables related to terrain and driving behavior is identified, with elevation and vehicle speed variables showcasing the highest positive contributors. Finally, the PC value of each observation in the CERTH- JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 5 Fig. 3: Loadings of the top-16 PCs from the CERTH-BEV dataset, showcasing the key variables that affect energy consumption. JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 6 Fig. 4: Contribution of the original CERTH-BEV data to the variation captured by each of the top-16 PCs. BEV dataset was calculated as the weighted sum of the variable values, with weights corresponding to their respective loadings [29], as explained in Eq. (1) as follows: PC value = n X i=1 (loadingi×variablei)(1) Fig. 4 depicts the significance of each original feature’s influence on the 16 most significant principal components, with redand blue-colored values representing strong positive and negative influence of the feature to the corresponding principal component, respectively. On the other hand, neutral colors showcase smaller loadings, indicating a smaller influence to each principal component. For instance, PC1 is mainly influenced by a combination of features like OAT and heater power (max, mean, median and mean), while the main features that influence PC2 are related to terrain and driving behavior, such as elevation and vehicle speed. On the same note, while PC1 is mainly influenced by low OAT and the vehicle’s heater power, PC6 showcases a completely opposite variation, influenced by high OAT and high A/C consumption. PC3 and PC4 are mainly influenced by the A/C power variable in combination with high and low road elevation, respectively, whereas PC5 is mostly related to driving behavior. Finally, PC9 is almost entirely defined by the vehicle’s median acceleration variable, showcasing the highest loading value observed, possibly indicating its strong connection to consistent acceleration patterns. Other later components, such as PC7, PC10, PC11 and PC13, capture more specific variations, often involving standard deviations variables (e.g., OAT[DegC] std) or HVAC usage (i.e., Air Conditioning Power[Watts] std or Heater Power[Watts] std), or particular extremes of driving behavior (e.g., Vehicle Speed[km/h] min, Acceleration[m/sˆ2] min). C. Clustering Analysis Following the PCA analysis from the previous section, clustering analysis was also performed, to classify the observations based on the converted PC variables and identify different patterns in energy consumption behaviors within the dataset. To determine the optimal number of clusters, a range of cluster numbers was tested for which the data were reduced to the reduced dimensional space of the optimal principal components, and k-Means clustering [32]–[34] was applied to assign data points to each cluster. As showcased in Fig. 5, the optimal number of clusters was found to be k= 3, as determined using the silhouette score [35], [36], creating three clusters on the PC space emphasizing combinations of variables such as HVAC usage, vehicle speed and elevation changes, which are also identified as key contributors to energy consumption from the PCA loadings in Fig. 3. Overall, the clustering results indicate that energy consumption in BEVs is highly related to certain environmental and driving conditions, like low outside temperature, high vehicle speeds and elevation. To determine which original features, or any of their combinations, are the most influential to the vehicle’s energy consumption, based on the loadings extracted from the PCA analysis, a reverse PCA analysis was performed [37], reconstructing the original dataset using only the most significant principal components. For this reason, the reconstructed data are an approximation of the original CERTH-BEV dataset and not an exact reconstruction. In addition, since the original data were scaled (standardized) prior to the PCA analysis, the inverse of that scaling was also applied to the reconstructed data, as well. To further interpret the three aforementioned clusters, Fig. 5 projects the cluster centroids back to the original feature space to highlight the contribution of each feature to each cluster, with each bar representing the standardized mean value of each feature for a direct comparison of feature intensities between the three clusters. As seen from the bar plot, the three clusters are mainly differentiated by HVAC usage patterns, ambient temperature, and, consequently, their overall energy consumption, with cluster 0 (“Predominant Heater Usage”) representing cold weather driving, with moderate heater power usage (∼0.6 above the mean) and quite low OAT values (∼0.4 below the mean), cluster 1 (“Moderate HVAC Usage & Typical Driving”) representing driving scenarios with moderate environmental conditions, with OAT values being generally around or slightly above the mean value (i.e., ∼0.1 above the mean) and both air conditioning and heater power usage being minimal, and cluster 2 (“Intensive A/C Usage”) representing hot weather driving scenarios, with significantly high air conditioning power usage (∼1.5 well above the mean) and correspondingly high OAT values (∼0.1 above the mean). JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 7 Fig. 5: k-Means cluster centroids projected back to the original CERTH-BEV feature space. Based on the above analysis, it is evident that cluster 2 showcases the highest standardized mean energy consumption (∼0.18 above the mean), confirming that heavy air conditioning power usage in high-temperature conditions leads to substantial energy draw. On the other hand, cluster 0 exhibits the lowest standardized mean energy consumption (∼0.42 below the mean), highlighting that heater power usage and low OAT values are less impactful compared to air conditioning usage in higher temperatures. Finally, cluster 1 represents operational scenarios with baseline or average standardized mean energy consumption (∼0.08 above the mean), serving as a benchmark for average energy usage under standard operating conditions. D. Smoothing Splines Analysis To provide a more complex approximation of the relations in the energy consumption data, smoothing splines [38]–[40] were implemented individually within each cluster on the PCA-reduced dataset, as a way of capturing non-linear data relationships thanks to their non-parametric nature. As a result, their implementation allows us to remove any noise from the original dataset (i.e., vehicle speed fluctuations due to brief accelerations or minor shifts in road elevation) and showcase any functional relationships between energy consumption (response) and each of the aforementioned influencing factors (estimators). The further alignment of the spline fitting with each cluster-specific energy usage profile, identified previously through the centroid reconstructions, can in turn offer more granular patterns. As depicted in Fig. 6 and 7, while the original scatter plots represent the collected values as raw data points, the overlaid smoothing splines follow the overall shape of the actual data points, capturing the general trend and revealing data relationships that otherwise were not immediately apparent. For instance, in cluster 0 (Predominant Heater Usage), PC1 and PC2 showcase clear upward trends in relation to energy consumption, confirming their positive correlation. On the other hand, PC3 and PC5 exhibit distinct non-linear effects where energy consumption increases at both low and high PC values and causes a dip at mid-range values, indicating an inflection point in the influence of mixed HVAC and terrain variables. On the other hand, in cluster 1 (Moderate HVAC Usage & Typical Driving), the smoothing spline plots showcase mostly linear or weakly non-linear curves, possibly indicating that energy consumption in this specific operational scenario is more stable and less affected by extreme variable values. The prominent dip in the mid-range values of PC1 and PC4 can possibly represent optimal energy scenarios, where, in the case of PC1, HVAC usage is minimal and OAT values are near average, while for PC4, terrain is level and HVAC systems are off or used efficiently. Finally, in cluster 2 (Intensive A/C Usage), the spline curves reveal stronger nonlinear effects, with energy consumption showcasing a sharp rise for higher PC1 variables, reflecting the significant energy draw of the A/C load. The sharp increase along the PC7 curve, particularly influenced by OAT variability and A/C power fluctuation based on its loadings, suggests that energy consumption increases equally sharply when fluctuations in temperature or air conditioning become pronounced. E. Multiple Principal Component Regression To closely fit the vehicle’s energy consumption data, using also the collected information from the exploratory analysis, a multiple principal component regression model is proposed, JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 8 (a) Cluster 0 (Predominant Heater Usage) (b) Cluster 1 (Moderate HVAC Usage & Typical Driving) Fig. 6: Scatter plots of the top-16 PCs showcasing the total energy consumption (response) against each PC variable (estimators) for (a) Cluster 0, and (b) Cluster 1. The spline curves (red) represent the continuous functions fitted to the actual data points capturing the non-linear relationship between the PC and energy consumption. JOURNAL OF L A T EX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 9 Fig. 7: Scatter plots of the top-16 PCs showcasing the total energy consumption (response) against each PC variable (estimators) for Cluster 2 (Intensive A/C Usage). The spline curves (red) represent the continuous functions fitted to the actual data points capturing the non-linear relationship between the PC and energy consumption. that extends the core MLR principles [41] by using the most important (top-16) principal components as estimators. This way, the proposed approach reduces dimensionality while retaining the most significant information collected from the original variables, as explained in Eq. (2) as follows: g(E(Y)) = β0+ 16 X j=1 βj(PCj)(2) where g() is the link function that relates the expected value of the response variable Yto the estimators, β0is the intercept, and P Cjdenotes the j-th principal component from the PCA. The slope coefficients βjare applied to each of the top-16 principal components, and the model captures linear relationships between the response and the principal components, rather than the original variables. F. Generalized Additive Modeling with Principal Components Following the development of the multiple principal component regression model that used the 16 first principal components, and taking into consideration the smoothing splines implementation on the PCA-reduced dataset within each cluster, a custom cluster-specific GAM model is also proposed, which extends the core GAM principles [8] by using the most important principal components as estimators to capture more complex relationships between the estimators and the response variable. However, instead of directly using the principal components as estimators like we did in the multiple principal component regression model, our custom GAM model uses the predefined smoothing spline basis function from Section III-D to transform each estimator, and then uses the coefficients and knots from each spline curve to approximate the smooth functions as polynomial-like transformations. Then, the transformed features are passed to the additive model based on the following equation: g(E(Y)) = β0+ 16 X j=1 dj X k=0 βj,kP Ck j(3) Here, function Pdj k=0 βj,kP Ck jrepresents the polynomial approximation of the original smooth function λj(P Cj), where djis the degree of the polynomial used for the j-th estimator, βj,k are the coefficients of the polynomial terms, and P Ck jis the k-th power of the P Cjestimator. G. Evaluation Metrics The performance of both of the proposed energy consumption modeling approaches was evaluated through widely used regression metrics, specifically the cross-validated root mean squared error (RMSE) and its normalized version (nRMSE), which is scaled by the range of actual values (max(y)− min(y)) [1]. While the RMSE metric provides an accurate estimate of the model’s average estimation error in the same unit