scieee AI-readable full text Open interactive document viewer

MAS-DR: An ML-Based Aggregation and Segmentation Framework for Residential Consumption Users to Assist DR Programs

Tzallas, Petros; Papaioannou, Alexios; Dimara, Asimina; Bezas, Napoleon; Moschos, Ioannis; Anagnostopoulos, Christos-Nikolaos; Krinidis, Stelios; Ioannidis, Dimosthenis; Tzovaras, Dimitrios

Abstract

The increasing complexity of energy grids, driven by rising demand and unpredictable residential consumption, highlights the need for efficient demand response (DR) strategies and data-driven services. This paper proposes a machine learning-based framework for DR that clusters users based on their consumption patterns and categorizes individual usage into distinct profiles using K-means, Hierarchical Agglomerative Clustering, Spectral Clustering, and DBSCAN. Key features such as statistical, temporal, and behavioral characteristics are extracted, and the novel Household Daily Load (HDL) approach is used to identify residential consumption groups. The framework also includes context analysis to detect daily variations and peak usage periods for individual users. High-impact users, identified by anomalies such as frequent consumption spikes or grid instability risks using IsolationForest and kNN, are flagged. Additionally, a classification service integrates new users into the segmented portfolio. Experiments on real-world datasets demonstrate the framework’s effectiveness in helping energy managers design tailored DR programs.

Full text

Academic Editors: Mikaeel Ahmadi and Atsushi Yona Received: 23 December 2024 Revised: 13 January 2025 Accepted: 8 February 2025 Published: 13 February 2025 Citation: Tzallas, P.; Papaioannou, A.; Dimara, A.; Bezas, N.; Moschos, I.; Anagnostopoulos, C.-N.; Krinidis, S.; Ioannidis, D.; Tzovaras, D. MAS-DR: An ML-Based Aggregation and Segmentation Framework for Residential Consumption Users to Assist DR Programs. Sustainability 2025,17, 1551. https://doi.org/ 10.3390/su17041551 Copyright: © 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/ licenses/by/4.0/). Article MAS-DR: An ML-Based Aggregation and Segmentation Framework for Residential Consumption Users to Assist DR Programs Petros Tzallas 1,2,* , Alexios Papaioannou 1,2,* , Asimina Dimara 1,3 , Napoleon Bezas 1, Ioannis Moschos 1, Christos-Nikolaos Anagnostopoulos 3 , Stelios Krinidis 1,2 , Dimosthenis Ioannidis 1 and Dimitrios Tzovaras 1 1Centre for Research and Technology Hellas, Information Technologies Institute, 57001 Thessaloniki, Greece; [email protected] (A.D.); [email protected] (N.B.); [email protected] (I.M.); [email protected] (S.K.); [email protected] (D.I.); [email protected] (D.T.) 2Management Science and Technology Department, Democritus University of Thrace (DUTh), 65404 Kavala, Greece 3 Intelligent Systems Lab, Department of Cultural Technology and Communication, University of the Aegean, 81100 Mytilene, Greece; [email protected] *Correspondence: [email protected] (P.T.); [email protected] (A.P.) Abstract: The increasing complexity of energy grids, driven by rising demand and unpredictable residential consumption, highlights the need for efficient demand response (DR) strategies and data-driven services. This paper proposes a machine learning-based framework for DR that clusters users based on their consumption patterns and categorizes individual usage into distinct profiles using K-means, Hierarchical Agglomerative Clustering, Spectral Clustering, and DBSCAN. Key features such as statistical, temporal, and behavioral characteristics are extracted, and the novel Household Daily Load (HDL) approach is used to identify residential consumption groups. The framework also includes context analysis to detect daily variations and peak usage periods for individual users. High-impact users, identified by anomalies such as frequent consumption spikes or grid instability risks using IsolationForest and kNN, are flagged. Additionally, a classification service integrates new users into the segmented portfolio. Experiments on real-world datasets demonstrate the framework’s effectiveness in helping energy managers design tailored DR programs. Keywords: energy consumption clustering; customer segmentation; consumption pattern analysis; feature selection for clustering; demand response (DR) programs 1. Introduction In recent decades, a shift towards Renewable Energy Sources (RES), like solar and wind, has been observed. The intermittent and unpredictable nature of RES, due to its dependence on weather conditions, has introduced challenges in energy grid’s management [ 1 ]. Furthermore, forecasting residential energy consumption poses additional challenges due to the high variability and fluctuations of the household usage patterns [ 2 – 4 ]. Therefore, a need has arisen for effective strategies to counterbalance these aforementioned uncertainties and, thus, enhance grid stability and flexibility. For this purpose, the concept of Demand Response (DR) strategies was birthed. DR strategies enable consumers to adjust their usage during peak periods in response to signals [ 5 , 6 ]. These strategies offer opportunities to not only achieve energy supply and demand balance effectively by aligning demand management strategies to the distinct needs and behaviors of different groups, but also Sustainability 2025,17, 1551 https://doi.org/10.3390/su17041551 Sustainability 2025,17, 1551 2 of 33 offer significant cost-saving benefits [ 7 ]. Based on the applied DR programs, consumers are able to receive personalized incentives that are tailored to their usage profiles [8]. In more recent years, the increase in the usage of smart meters in residential buildings has given a powerful tool in the hands of energy policy makers, regarding the availability of real-time energy consumption and production data [ 9 , 10 ]. In addition, the exponential increase in the usage of Artificial Intelligent (AI), and more specifically Machine Learning, in residential energy management systems enable utilities to process the available data in order to identify trends, predict consumption behaviors, forecast renewable energy units’ production accurately, and manage the storage system effectively [ 11 – 15 ]. Unsupervised machine learning, and more specifically clustering algorithms, can be very effective when used for the optimization of the DR programs. More specifically, it can provide useful insights for the energy consumers, both residential and industrial, while, additionally, grouping consumers with similar load profiles, assisting utility companies in creating more tailored and responsive DR programs that boost both efficiency and participation. Leveraging clustering algorithms such as K-means or Hierarchical Agglomerative Clustering, DR programs can optimize how policy makers engage with different consumer groups [ 16 – 19 ]. Despite the indisputable benefits of unsupervised learning in energy-related applications, there is a sizable gap in research that needs to be filled to further investigate clustering algorithms’ usefulness [ 20 ]. First and foremost, most approaches lack a systematic methodology to determine the most suitable clustering algorithms that align with specific objectives. Moreover, the inherent subjectivity and intuitive interpretation of segmentation results often undermine their reproducibility, emphasizing the critical absence of a universally accepted metric for evaluating cluster quality [ 21 ]. Adding to this, the dynamic and highdimensional nature of energy consumption data demands a robust framework capable of handling real-world complexity while delivering actionable insights. This underscores the necessity for clearly defined clustering services to ensure consistency and reliability in applications. Furthermore, integrating clustering with advanced evaluation metrics could bridge the gap between theoretical models and practical deployment, ultimately optimizing residential energy consumption strategies and supporting tailored DR programs. Within this context, this paper addresses a holistic approach that is based on a variety of commonly used clustering algorithms and metrics for aggregating and segmenting the consumption of residential users, in order to assist Demand Response (DR) programs with useful insights. The framework can be separated into two major categories based on the data used. The first category analyzes the data from several different residential users, in order to segment them into groups, based on a wide variety of features, which are carefully selected, derived from their consumption and based on the underused concept of the Household Daily Load (HDL), which will be further presented in following sections. Furthermore, based on similar features, classification techniques are explored to classify new users into already existing groups, an approach that can benefit the decision makers by providing them with implicit information on the consumption patterns of said users without the need of difficult to acquire historical data. In addition, different services in analyzing the consumption patterns of a single residential user are presented. Specifically, a novel approach is presented to identify users suitable for DR strategies by analyzing a user’s daily variation, or by finding their peak consumption in different daily periods. Additionally, by utilizing anomaly detection techniques, among others, a methodology to identify users that could have high impact on DR strategies is proposed. Overall, this paper addresses gaps in systematic clustering methodologies and practical evaluation metrics, advancing the application of machine learning in energy systems introducing the following main novelties: Sustainability 2025,17, 1551 3 of 33 • Integration of Clustering and Classification for DR Programs: This paper introduces a machine learning-based framework (MAS-DR) that combines clustering and classification techniques to segment residential users and classify new users into pre-existing clusters, optimizing DR strategies and improving their applicability to a dynamic user base. • Household Daily Load (HDL) Concept: This paper incorporates the novel Household Daily Load (HDL) approach, which analyzes seasonal and monthly consumption patterns. This allows for finer segmentation of user behavior based on temporal variations, supporting tailored interventions such as season-specific DR strategies. • Feature Selection for Enhanced Clustering: This study applies advanced feature extraction and selection techniques, including Principal Component Analysis (PCA) and iterative clustering evaluation, to ensure that only the most relevant features influence segmentation, enhancing the accuracy and efficiency of clustering. • Comprehensive Evaluation of Multiple Clustering Algorithms: The framework systematically evaluates several clustering algorithms (e.g., K-means, Hierarchical Clustering, Spectral Clustering, DBSCAN) and metrics (e.g., Silhouette Score, Davies–Bouldin Index, Calinski–Harabasz Index) to identify the optimal methods for segmenting residential energy consumption patterns. The remainder of the paper is organized as follows. Section 1offers an introduction on the concept of Demand Response and the value that Artificial Intelligence (AI) and more specifically clustering algorithms offer in building DR strategies. Section 2offers a comprehensive literature review on the related works that explore clustering models in energy related applications. Section 3presents in detail the methodology that was followed in the design of the proposed framework, together with the clustering algorithms and the evaluation metrics that were used. In Section 4, the case study that was followed in the demonstration of the framework is presented. Section 5presents the results of the evaluation of the MAS-DR framework on real-life data. Finally, Section 7offers a conclusive analysis of the results of the proposed framework. 2. Literature Review As already mentioned, in this section, a literature review of recent advancements in the clustering techniques used for gathering insights on users’ consumption is presented. In most of the studied works, the clustering algorithms are used for segmenting the users into distinguished groups. The primary focus, in many approaches, is to determine the optimal number of clusters, in order to enhance the segmentation accuracy, making use of the Silhouette Score (SIL) evaluation metric [ 22 , 23 ]. In a similar manner, Abdulnassar and Nair in their approach in [ 24 ] tried to minimize the computation cost, without sacrificing accuracy, by optimizing the initial centroids. Furthermore, feature selection techniques can be used to maximize the efficiency of the clustering algorithms [25,26]. Another common approach is the utilization of different clustering models. More specifically, Gaussian mixture models [ 25 ], Self-Attention LSTM [ 27 ], and Ant Colony Clustering [ 28 ], together with commonly used clustering algorithms, such as K-means, Hierarchical Clustering [ 29 , 30 ], and DBSCAN [ 31 ], are only some of the algorithms used to segment smart energy data effectively. These studies highlight the growing trend of using hybrid approaches in gathering insights for the consumption of the residential users. Another categorization of the available works is the clustering objectives. These vary from the most common segmentation of the users into clusters based on their consumption [ 19 , 25 – 41 ] to identifying different consumption patterns on different users’ or a single user’s consumption time series [ 25 , 26 , 28 – 35 , 37 ], to identifying the daily variation of the consumption, in order to provide helpful insights [ 26 , 28 , 33 , 40 ], and, finally, to detecting Sustainability 2025,17, 1551 4 of 33 anomalies in the consumption time series [ 28 , 33 , 41 ]. Each clustering services can offer valuable insights to decision makers regarding the energy consumption of users, whether residential or not, and their availability for applying DR solutions. Finally, in [ 29 ], Satre-Meloy et al. made use of several different classification methods to predict the cluster membership of households based on occupant activity data during evening peak electricity usage hours, highlighting the importance of classification methods in different segmentation services. While prior works have explored demand-side management frameworks, highlighting their use of game-theoretic approaches and mixed-integer programming, for scheduling of Responsive Loads in Smart Distribution System, these primarily focus on optimization-based strategies for scheduling and load management [ 42 , 43 ]. In contrast, this work emphasizes a machine learning-based framework that combines clustering, classification, and advanced feature extraction techniques to provide a datadriven segmentation of residential users. The MAS-DR framework specifically addresses the gaps in clustering methodologies by introducing the novel Household Daily Load (HDL) concept and systematically evaluating clustering algorithms and metrics for tailored Demand Response (DR) strategies. This distinction lies in our emphasis on unsupervised learning for user behavior insights, which complements but differs fundamentally from optimization-based approaches. In Table 1, an overview of the studied literature on the application of clustering and classification techniques for aggregating and segmenting energy users’ consumption is presented. This table highlights the gaps addressed by the proposed work, which aims to provide a holistic approach to the identified issue, offering valuable insights for DR specialists. Table 1. Summary of related works for clustering on energy-related data. Citation Resident. Cons. Various Models Various Metrics Optimal Cluster No Segment Users Classify Users Cons. Patterns Daily Variation Anomaly Detection Feature Selection [22]✓ ✓ ✓ [19]✓ ✓ ✓ ✓ ✓ [23]✓ ✓ ✓ [24]✓ ✓ [25]✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ [27]✓ ✓ ✓ ✓ [28]✓ ✓ ✓ ✓ ✓ ✓ [32]✓ ✓ ✓ ✓ ✓ [33]✓ ✓ ✓ ✓ ✓ ✓ ✓ [34]✓ ✓ ✓ ✓ ✓ ✓ [29]✓ ✓ ✓ ✓ ✓ ✓ ✓ [35]✓ ✓ ✓ ✓ ✓ [30]✓ ✓ ✓ ✓ ✓ ✓ [36]✓ ✓ ✓ [26]✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ [31]✓ ✓ ✓ ✓ ✓ ✓ [37]✓ ✓ ✓ ✓ ✓ [38]✓ ✓ ✓ ✓ [39]✓ ✓ ✓ [40]✓ ✓ ✓ ✓ ✓ ✓ [41]✓ ✓ ✓ ✓ ✓ This Paper ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ 3. Methodology This section presents the general methodology of the proposed framework, which is visualized in Figure 1. MAS-DR offers six different segmentation and aggregation services, which will be further discussed in the following subsections. The whole framework is designed to acquire and process raw consumption measurements from different residential houses and provide meaningful insights tailored to DR programs. Sustainability 2025,17, 1551 5 of 33 The raw data are processed through various levels of feature extraction and clustering techniques, tailored to a unique use case. The first use case explores the consumption of different users in order to divide them into groups (Service 1 and 2). The extracted features provide an overview of the users’ behavior, which is grouped and evaluated for targeted interventions using optimal clustering algorithms and feature importance analysis (Service 1.1). Based on classification methods, the cluster that a newly added user belongs to is also identified (Service 6). Furthermore, advanced thresholds approaches are used to analyze a user’s daily variation (Service 3) or identify a user’s peak consumption (Service 4), providing insights on whether the particular user is suitable for DR strategies. Anomaly detection techniques are employed to identify users that could cause issues in DR processes (Service 5). In this section, the different services that the MAS-DR framework offers to fulfill these objectives are analyzed in detailed. In addition, the different clustering algorithms and evaluation metrics are presented. Figure 1. MAS-DR framework conceptual methodology. 3.1. Aggregation and Segmentation Services The MAS-DR framework introduces a suite of aggregation and segmentation services aimed at analyzing and grouping residential users based on their energy consumption patterns. These services are designed to provide actionable insights for DR strategies, allowing energy managers to optimize interventions for individual or clustered households. The aggregation and segmentation processes leverage machine learning techniques to extract meaningful features from raw consumption data, identify usage patterns, and classify users into distinct categories. The framework addresses various segmentation goals, including identifying daily consumption trends, peak load periods, and high-impact users who could influence grid stability. It integrates advanced clustering methods, combined with Sustainability 2025,17, 1551 6 of 33 feature selection techniques while ensuring robust and efficient segmentation results. These services are critical in tailoring DR programs to user-specific behaviors while enhancing the grid’s operational resilience and energy efficiency. In the following subsections, each aggregation and segmentation service is elaborated in detail, highlighting its methodology, algorithms, and key objectives. 3.1.1. Service 1: Segment Users into Clusters After Feature Extraction The first objective of MAS-DR framework is the segmentation of different residential users into users based on features derived from their energy consumption. As already mentioned, this is a highly common objective of recent studies. This work tries to dig deeper into the consumption time series of a household by extracting a wide variety of features, which are presented in Table 2. The features are separated into two main categories, basic and extended. The basic features, namely average consumption, variance of the consumption, and key quantiles (10th and 90th), are the most commonly used in different studies and provide an initial understanding of a household’s energy use and its variability. Having the basic features as a foundation, the extended features offer a deeper analysis and can be divided into Statistical, Temporal, Behavioral, and Trend subcategories. Statistical metrics like median, standard deviation, and extreme values (max and min) represent a greater range of variability, whereas temporal features focus on specific patterns, such as workday versus weekend consumption. Behavioral features provide detailed insights regarding the type of the users that live in a household, as they can represent activities in different time of the day. Furthermore, trend analysis, using variables such as trend slope and month-to-month change, identifies long-term usage trends, whereas rate-of-change measurements reveal daily consumption volatility. These variables, combined with their hierarchical classification, provide a thorough picture of energy usage. The next step of this service is to select the most important of the aforementioned features, which ensures that the clustering procedure is based on the most important and relevant variables. Two complementary methodologies are used for this analysis. First, Principal Component Analysis (PCA) [ 44 ] is performed to determine which features account for the majority of the variance in the dataset, giving a quantitative foundation for feature prioritizing. Second, the dataset is iteratively clustered, removing one feature at a time, and the influence on the cluster Silhouette Score [45] is monitored, which evaluates clustering quality. Table 2. Features used for energy consumption segmentation. Category Subcategory Features Details Basic Statistical Average Consumption The mean value of energy consumption. Variance Consumption The variance of energy consumption. 10th quantile The consumption level above which 10% of the values lie. 90th quantile The consumption level below which 90% of the values lie. Extended Statistical Median Consumption The median value of energy consumption. Std Consumption The standard deviation of energy consumption. Max Consumption The highest recorder value of energy consumption. Min Consumption The lowest recorder value of energy consumption. 25th quantile The consumption level above which 25% of the values lie. 75th quantile The consumption level below which 75% of the values lie. Temporal Daily Average The average consumption during a day. Weekly Average The average consumption during a week. Weekend Average The average consumption during weekends. Weekday Average The average consumption during weekdays. Behavioral Midnight average The average consumption during late-night hours (00:00–05:00), typically representing idle usage. Morning average The average consumption during morning hours (05:00–12:00), typically representing start-of-day activities. Afternoon average The average consumption during midday hours (12:00–18:00), typically representing work-related activities. Evening average The average consumption during evening hours (18:00–00:00), typically representing peak household activities. Trend Trend Slope The gradient of a linear trend line fitted to the data, showing whether energy usage is increasing, decreasing, or stable over time. Month-to-month Change The percentage change in energy consumption from one month to the next, indicating seasonal or behavioral shifts. Average Daily Change The average of day-to-day changes in consumption, revealing volatility in daily energy usage patterns. Sustainability 2025,17, 1551 7 of 33 3.1.2. Service 2: Segment Users into Clusters Based on Household Daily Load (HDL) The term Household Daily Loads (HDL) was presented by Yilmaz et al. in [ 46 ] and it represents the average daily consumption of a household derived from smart meter data. The feature represents the average of hourly consumption data over a given period within a day and smooth out day-to-day variations while highlighting general consumption trends. The final result of this approach is one vector for each user that contains their average consumption for each hour of a day for a specific period, typically a month, a season, or a year. In the proposed framework, Season-based HDL and Month-based HDL is integrated into the framework. Season-based HDL investigates differences in consumption habits between seasons (e.g., winter and summer). These profiles demonstrate how seasonal elements such as weather and daylight affect residential energy consumption. Clustering based on season-specific HDL allows energy operators to adjust interventions, such as heating-focused programs in the winter or cooling-focused initiatives in the summer. On the other hand, Month-based HDL catches monthly variations, such as increased energy consumption during the holidays or changes in daylight savings. Enhancing the MAS-DR framework with HDL analysis offers a finer understanding on a households’ behavior and provides stakeholders, such as energy providers, with a tool to aggregate users that could benefit from similar DR strategies. 3.1.3. Service 3: Analyze User Daily Variation The next objective that the proposed tool focuses on is the analysis of a user’s consumption variation in a single day. This service aims to find a way to quantify the variation of the user’s daily patterns by utilizing clustering methods on their consumption. The MAS-DR follows the methodology that is highlighted in Algorithm 1. Algorithm 1 : Analysis of users’ daily energy consumption variation. function ANALIZEVARIATION(userList) ▷Define thresholds clustering_number_threshold ←predefinedThreshold sum_of_top_threshold ←topThreshold ▷Step 1: Cluster user’s daily data for user ∈userList do dailyData ←GroupByDay(user.consumptionData) optimalClusters ←FindOptimalNumberClusters(dailyData) dailyClusters ←ApplyClusteringAlgorithm(dailyData, optimalClusters) user.clusterData ←dailyClusters user.optimalClusterCount ←optimalClusters end for ▷Step 2: Analyze cluster distribution for user ∈userList do if user.optimalClusterCount <clustering_number_threshold then ▷Condition 1: Consistent behavior user.behavior ←“Consistent Behavior” else ▷Condition 2: Predictable pattern topTwoCount ←CountTopTwoClusters(user.clusterData) if topTwoCount >sum_of_top_threshold then user.behavior ←“Predictable Pattern” else user.behavior ←“Diverse Behavior” end if end if end for ▷Step 3: Assess suitability for DR programs for user ∈userList do if user.behavior ∈{“Consistent Behavior”, “Predictable Pattern”} then user.DRSuitability ←True else user.DRSuitability ←False end if end for end function After the definition of the thresholds, the first step of the methodology contains the grouping of the consumption data for each user by day to facilitate further analysis. Next, for each user, the optimal number of clusters is determined, representing the potential Sustainability 2025,17, 1551 8 of 33 patterns the user’s consumption could fall into. A clustering algorithm is then applied to group these daily consumption profiles into meaningful clusters. In the second step, the cluster distribution is analyzed to identify users with significant representation in a limited number of clusters. This involves investigating two conditions. The first condition checks whether the optimal number of clusters is below a predefined threshold, which would indicate “Consistent Behavior”. If not, the second condition evaluates whether the data points in the most populated clusters exceed a specified threshold, suggesting that a significant portion of the user’s days follow a “Predictable Pattern”. In the third and final step of the algorithm, users are assessed for suitability in DR programs, with users that meet either of the conditions being characterized as either consistent or predictable, and in turn, marked as suitable for DR services. Based on the aforementioned methodology, the MAS-DR framework can identify residential users that exhibit either consistent daily variation or are likely to have predictable consumption schedules, enabling them to respond effectively to price signals or load adjustment requests, and thus can be considered for DR programs. 3.1.4. Service 4: Finding the Consumption Peak in Daily Periods The aim of this objective is to identify the peak consumption periods within a user’s daily electricity usage by segmenting the day into distinct time periods. This can assist in assessing whether the user can participate in DR programs. This notion is based on the hypothesis that when the peak consumption is pinpointed on evening periods, this is a typical residential household; additionally, evening peaks are of particular interest due to higher electricity prices during this period, offering an opportunity for cost savings and grid optimization. The methodology that this services follows is presented in Algorithm 2. Algorithm 2 : Finding the consumption peak in daily periods. function FINDPEAKDAILY(userList) ▷Step 1: Segmentation of the Day midnightPeriod ←[00:00, 05:00] morningPeriod ←[05:00, 12:00] afternoonPeriod ←[12:00, 18:00] eveningPeriod ←[18:00, 24:00] for user ∈userList do dailyData ← GroupByTimePeriod(user.consumptionData, {midnightPeriod, morningPeriod, afternoonPeriod, eveningPeriod}) user.segmentedData ←dailyData end for ▷Step 2: Peak Consumption Analysis for user ∈userList do peakPeriod ←FindPeakConsumptionPeriod(user.segmentedData) user.peakPeriod ←peakPeriod ▷Condition: Peak occurs in the evening if peakPeriod == eveningPeriod then user.peakBehavior ←“Evening Peak” else user.peakBehavior ←“Other Peak” end if end for ▷Step 3: Assess Suitability for DR Programs for user ∈userList do if user.peakBehavior == “Evening Peak” then user.DRSuitability ←“Potentially Suitable for DR Programs” else user.DRSuitability ←“Not Suitable for DR Programs” end if end for end function In the first step of the algorithm, the day is divided into four distinct periods: midnight (00:00–05:00), morning (05:00–12:00), afternoon (12:00–18:00), and evening (18:00–00:00). Each user’s electricity consumption data are grouped into these periods to create a segmented dataset. The second step involves the identification of the time period with the highest electricity consumption (peak period) for each user and determines whether this peak consistently occurs during the evening period. If the peak is in the evening, the users’ behavior is classified as “Evening Peak”. In the third step of the methodology, each user Sustainability 2025,17, 1551 9 of 33 is assessed for suitability for DR programs. More specifically, users with an evening peak can be prioritized for DR strategies, such us load adjustments during high-demand hours, while users with with peaks in other periods may still be monitored. 3.1.5. Service 5: Identifying High-Impact Users This service offered by the MAS-DR framework aims to identify users that have a high impact on the operation of the energy grid, and could potentially cause issues when DR strategies are applied to them. The service examines the consumption of each user in order to identify their peak consumption, the variance of their consumption, and the frequency of their peaks. In addition, the number and frequency of the potential anomalies that the users’ consumption presents is detected. Key stakeholders, using this framework, can establish a threshold for each of these finding, based on their knowledge on the domain and their needs. If the threshold is surpassed, valuable insights are derived from the segmentation process of the residential consumption. Algorithm 3presents the methodology for Service 5. Algorithm 3 : Identifying High-Impact Users. function IDENTIFYHIGHIMPACTUSERS(userList) ▷Define thresholds peak_consumption_threshold ←predefinedPeakThreshold variability_threshold ←predefinedVariabilityThreshold spike_frequency_threshold ←predefinedSpikeFrequencyThreshold anomaly_threshold ←predefinedAnomalyThreshold ▷Step 1: Feature Analysis for High-Impact Users for user ∈userList do peakConsumption ←FindPeakConsumption(user.consumptionData) if peakConsumption >peak_consumption_threshold then user.isHighPeakUser ←True else user.isHighPeakUser ←False end if consumptionVariability ←CalculateVariability(user.consumptionData) if consumptionVariability >variability_threshold then user.isHighVariabilityUser ←True else user.isHighVariabilityUser ←False end if spikeFrequency ←CalculateSpikeFrequency(user.consumptionData) if spikeFrequency >spike_frequency_threshold then user.isFrequentPeakUser ←True else user.isFrequentPeakUser ←False end if if user.isHighPeakUser and user.isHighVariabilityUser and user.isFrequentPeakUser then user.isHighImpactUser ←True else user.isHighImpactUser ←False end if end for ▷Step 2: Anomaly Detection for user ∈userList do anomalies ←DetectAnomalies(user.consumptionData, method={“IsolationForest”, “kNN”}) anomalyCount ←CountAnomalies(anomalies) if anomalyCount >anomaly_threshold then user.isHighAnomalyUser ←True else user.isHighAnomalyUser ←False end if end for ▷Step 3: User Segmentation for user ∈userList do if user.isHighImpactUser and user.isHighAnomalyUser then user.segment ←“Potential Issue User” else if user.isHighImpactUser then user.segment ←“High-Impact User” else user.segment ←“Regular User” end if end for end function The algorithm for identifying high-impact users involves a three-step process. In the first step, the feature analysis phase, thresholds are predefined for peak consumption, consumption variability, and spike frequency. For each user, the algorithm calculates their Sustainability 2025,17, 1551 16 of 33 (c) Spectral (d) BIRCH Figure 3. Silhouette Score all clusters and algorithms for Service 2. From the results, it can be observed that each metric has a slight difference regarding the number of clusters. In this paper, the most voted result was followed, which corresponded to the optimal number of clusters being two. 4.2.3. Check Feature Importance The main idea when designing the framework, regarding the enhancement of Service 1, is to utilize all the available information derived form the consumption of residential users. Keeping that in mind, the efficiency of the service is of equal significance. Thus, it is valuable to further investigate the important of each feature. MAS-DR offers two methods to determine the feature importance (Service 1.1), Principal Component Analysis (PCA), to observe which features contribute most to the variance in the dataset, and iteratively clustering the dataset, excluding one feature and observing the impact on a cluster’s Silhouette Score. More specifically, regarding the first method, the first step is to cluster the data in the optimal amount of clusters using one of the clustering algorithms. Then, feature importance for each cluster is computed using the PCA analysis. A dictionary of feature importance is created, with sorted and normalized results, where the higher values are better. The result of this analysis is presented in Figure 4. The second method follows an iterate-based logic. More specifically, iteratively exclude each feature form the dataset and apply a clustering algorithm (in our case, K-means), to the reduced dataset without the feature. The next step is to compute the Silhouette Score, in each iteration, which measures how well-separated the clusters are. If the Silhouette Score is high, then the excluded feature is less important. The next step is the sorting of features based on the calculated Silhouette Scores. Finally, in order to have intuitive results, the scores are reverse and normalized, where a feature with a higher score is considered more important. The results for our dataset and for the features discussed in Section 3are presented in Figure 5. The final step of the feature importance is the combination of the two methods, based on which the feature vector that MAS-DR uses is decided. Sustainability 2025,17, 1551 17 of 33 Figure 4. Impact of each feature for all classes, based on the PCA analysis. Figure 5. Impact of each feature for all classes, based on the SIL analysis. Sustainability 2025,17, 1551 18 of 33 5. Aggregation and Segmentation Insights This section of the paper concerns the presentation of the insights derived from the MAS-DR framework. The results can be separated into three main categories regarding the objectives, namely cluster segmentation of the dataset, identification of each user’s DR insights, and the assignment of a new user to pre-defined clusters. 5.1. Segmenting Users into Clusters This section presents the results of Service 1 and Service 2, which focus on clustering residential users based on their energy consumption patterns. Service 1 utilizes extended statistical, temporal, and behavioral features, while Service 2 applies the Household Daily Load (HDL) approach for seasonal and monthly analysis. The segmentation aims to identify distinct user groups, enabling tailored Demand Response (DR) strategies. Clustering performance is evaluated using multiple algorithms and metrics to ensure robust and meaningful results. 5.1.1. Service 1 This service focuses on clustering residential users based on a variety of statistical, temporal, and behavioral features derived from their energy consumption. Multiple clustering algorithms were evaluated using three performance metrics: Silhouette Score, Davies–Bouldin Index (DBI), and Calinski–Harabasz Index (CHI). The results for Service 1, as shown in Table 6, indicate that the Spectral Clustering algorithm outperforms other methods, achieving the highest Silhouette Score (0.83) with extended features, the lowest Davies–Bouldin Index (0.47), suggesting well-separated clusters and a reasonable Calinski– Harabasz Index (1096.85). Furthermore, K-means performed well with extended features, achieving a CHI of 1348.05, but its Silhouette Score was slightly lower (0.68). Conversely, DBSCAN and Hierarchical Clustering showed competitive results but underperformed compared to Spectral Clustering. Additionally, it can be observed that in most of the approaches the extended features provide an increase in clustering efficiency. Table 6. Clustering performance metrics for Service 1: User segmentation after feature extraction. Algorithm Extended Features Silhouette Score Davies–Bouldin Index Calinski–Harabasz Index K-means False 0.67 1.02 1327.38 True 0.68 1.00 1348.05 Hierarchical False 0.56 1.59 1185.57 True 0.58 1.83 1225.88 Spectral False 0.50 1.65 501.193 True 0.83 0.47 1096.85 DBSCAN False 0.73 0.84 1067.81 True 0.75 1.49 1116.26 Mean Shift False 0.44 1.58 175.68 True −0.16 1.59 60.47 GMM False 0.33 2.03 98.34 True 0.42 2.19 86.09 Note: Bold values indicate the best-performing algorithms based on each metric. The clustering analysis revealed two primary user clusters (Cluster 0 and Cluster 1). For Cluster 0 (Figure 6a), users exhibit lower overall power consumption, as indicated by their mean and median profiles. Sustainability 2025,17, 1551 19 of 33 (a) Cluster 0 (b) Cluster 1 Figure 6. Cluster visualization for Service 1 using Spectral Clustering. Moreover, consumption remains stable with fewer extreme spikes throughout the year. For Cluster 1 (Figure 6b), users have higher average consumption with greater variability and their load profiles show occasional spikes, particularly in certain periods (e.g., winter). The 2D visualization (Figure 7) highlights the distinct separation between the two clusters. While Cluster 0 users are densely grouped (lower consumption range), Cluster 1 users are more spread out, indicating higher variability and consumption. Overall, spectral clustering delivers the best performance for user segmentation after feature extraction. The two clusters distinguish users based on consumption magnitude and variability. The extended feature set enhances clustering performance, as evident from improved metrics (Silhouette Score, DBI). This segmentation enables targeted DR strategies, allowing energy managers to optimize interventions for each user group effectively. Sustainability 2025,17, 1551 20 of 33 Figure 7. Two-dimensional visual representation of clusters for Service 1 using Spectral Clustering. 5.1.2. Service 2 This service analyzes the energy consumption of residential users by leveraging the Household Daily Load (HDL) concept, focusing on seasonal variations. The segmentation aims to uncover consumption trends that align with specific seasonal periods (winter, spring, summer, and autumn). Clustering performance was evaluated using multiple algorithms and three metrics: Silhouette Score, DBI, and CHI. As presented in Table 7, spectral Clustering consistently outperforms other algorithms across all seasons—specifically, it has the highest Silhouette Score (e.g., 0.73 in autumn), the lowest DBI (e.g., 0.17 in autumn), indicating compact and well-separated clusters, and the highest CHI values (e.g., 28.96 in autumn), reflecting strong cluster structures. Furthermore, K-means also performs reliably with a Silhouette Score of 0.71 and competitive CHI values for multiple seasons, particularly in winter and autumn, while DBSCAN and Hierarchical Clustering show subpar performance in comparison, particularly due to less compact and less clearly separated clusters. The summer season results, as visualized in Figure 8, highlight two distinct clusters. In Cluster 0 (Figure 8a), users with low and stable energy consumption are represented. The mean and median profiles remain consistently below 20 kW, with minimal fluctuations over time. In Cluster 1 (Figure 8b), users with higher and more variable energy consumption are represented. The mean and median power levels are significantly higher, with peaks occasionally exceeding 100 kW, indicating substantial energy usage variability. Spectral Clustering provides the most effective segmentation for seasonal HDL, outperforming other methods across all metrics. Moreover, clear patterns emerge in user behavior, Cluster 0 users are characterized by low and stable energy consumption. Cluster 1 users exhibit higher and more dynamic consumption patterns. Seasonal segmentation (e.g., summer analysis) highlights the impact of seasonal variations, offering actionable insights for DR programs tailored to seasonal peaks. This segmentation allows energy managers to target high-impact users and tailor interventions, such as peak-load reduction strategies during summer months when consumption is elevated. Sustainability 2025,17, 1551 21 of 33 Table 7. Clustering performance metrics for Service 2: User segmentation based on Season-HDL. Algorithm Season Silhouette Score Davies–Bouldin Index Calinski–Harabasz Index K-means winter 0.71 0.18 22.96 spring 0.71 0.18 22.96 summer 0.66 0.81 23.08 autumn 0.71 0.18 22.96 Hierarchical winter 0.33 1.50 15.42 spring 0.71 0.18 22.96 summer 0.66 0.81 23.08 autumn 0.71 0.18 22.96 Spectral winter 0.70 0.17 22.91 spring 0.72 0.17 25.45 summer 0.70 0.18 22.91 autumn 0.73 0.17 28.96 DBSCAN winter 0.42 1.57 13.43 spring 0.46 1.38 16.65 summer 0.52 1.32 16.52 autumn 0.55 1.15 20.07 Mean Shift winter 0.28 0.92 17.61 spring 0.17 0.94 13.41 summer 0.20 0.92 15.25 autumn 0.28 1.07 15.83 GMM winter 0.61 0.68 16.56 spring 0.71 0.68 17.96 summer 0.66 0.81 13.08 autumn 0.71 0.68 17.96 Note: Bold values indicate the best-performing algorithms based on each metric for each season. (a) Cluster 0 (b) Cluster 1 Figure 8. Cluster visualization for Service 2: Season-based HDL (summer) using K-means Clustering. Sustainability 2025,17, 1551 22 of 33 The 2D representation of clusters in Figure 9highlights the separation between the two identified clusters for the summer season using K-means clustering. Cluster 0 (purple) are users with low and stable energy consumption are tightly grouped, indicating minimal variance in their energy usage. Cluster 1 (green) are users with higher and more varied consumption are more dispersed, reflecting significant differences in their energy behavior. The clear spatial separation between the clusters reinforces the effectiveness of the K-means algorithm in identifying distinct user groups based on the Season-HDL feature. Figure 9. Two-dimensional visual representation of clusters for Service 2: Season-based HDL (summer) using K-means Clustering. The clustering performance across months (January, April, July, and November) was analyzed using multiple algorithms and evaluation metrics, as presented in Table 8. The results indicate the following trends: K-means consistently delivers strong performance, particularly in November, with the highest Silhouette Score of 0.72, the lowest Davies– Bouldin Index of 0.16, and the highest Calinski–Harabasz Index of 27.96, reflecting compact and well-separated clusters. Spectral Clustering also performs well, with competitive Silhouette Scores and DBI values across all months. Hierarchical Clustering and DBSCAN show weaker performance, particularly in January and April, with lower cluster compactness and higher DBI scores. The monthly analysis for January, visualized in Figure 10, reveals two primary clusters. Cluster 0 (Figure 10a) users in this cluster exhibit low and stable energy consumption, with both mean and median power values consistently below 20 kW. Consumption remains predictable, with minor fluctuations throughout the month. Cluster 1 (Figure 10b) represents higher energy consumers with occasional spikes, particularly during specific periods of the month. The mean consumption consistently exceeds 50 kW, indicating larger or more energy-intensive households. A 2D representation of clusters in Figure 11 highlights the separation between the two identified clusters. Sustainability 2025,17, 1551 23 of 33 Table 8. Clustering performance metrics for Service 2: User segmentation based on Month-HDL. Algorithm Season Silhouette Score Davies–Bouldin Index Calinski–Harabasz Index KMeans January 0.61 0.97 23.44 April 0.61 0.97 23.44 July 0.66 0.81 23.08 November 0.72 0.16 27.96 Hierarchical January 0.38 1.46 16.46 April 0.61 0.97 23.44 July 0.66 0.81 23.08 November 0.71 0.18 22.96 Spectral January 0.61 0.97 23.44 April 0.62 0.82 19.72 July 0.66 0.81 23.08 November 0.71 0.18 22.96 DBSCAN January 0.29 1.57 12.78 April 0.32 1.56 12.78 July 0.66 0.81 23.08 November 0.55 0.81 23.08 Mean Shift January 0.21 0.94 14.44 April 0.31 1.14 14.39 July 0.34 0.94 19.50 November 0.31 1.01 17.50 GMM January 0.61 0.97 23.44 April 0.61 0.97 23.44 July 0.66 0.81 23.08 November 0.71 0.18 22.96 Note: Bold values indicate the best-performing algorithms based on each metric for each month. (a) Cluster 0 (b) Cluster 1 Figure 10. Cluster visualization for Service 2: Month-based HDL (January) using K-means Clustering. Sustainability 2025,17, 1551 24 of 33 Figure 11. Two-dimensional visual representation of clusters for Service 2: Month-based HDL (January) using K-means Clustering. Cluster 1 users can be prioritized for peak load reduction strategies to address occasional spikes. Cluster 0 users represent steady, manageable consumption patterns, contributing to grid stability. The insights derived from monthly segmentation enable more granular DR interventions, ensuring that seasonal or monthly energy patterns are addressed effectively. 5.2. Identification of Users’ DR Capability This section evaluates the suitability of residential users for DR programs based on their energy consumption behaviors. Results from Service 3 (daily variation analysis), Service 4 (peak consumption identification), and Service 5 (high-impact user detection) are combined to provide a holistic view of user readiness. Users are categorized based on their behavior patterns, peak consumption periods, and potential grid impact, enabling energy managers to prioritize DR interventions effectively and ensure grid stability. 5.2.1. Service 3 Service 3 evaluates the daily energy consumption variation of individual users to assess their suitability for DR strategies. The analysis identifies users exhibiting consistent or predictable behavior based on two key thresholds are users with fewer than 3 clusters exhibit consistent behavior. Furthermore, if the top clusters cover more than 80% of daily consumption, the behavior is classified as predictable. The results in Table 9indicate that users such as user_3, user_12, and user_209 exhibit consistent behavior (fewer than 3 clusters with 100% of data in the top clusters), making them ideal candidates for DR programs. Users like user_2655 and user_4266 demonstrate predictable behavior, with the top clusters covering more than 80% of the data, also qualifying them for DR strategies. Users such as user_330 and user_450 display diverse behavior, where consumption is spread across many clusters (low top-cluster coverage), rendering them less suitable for DR. Sustainability 2025,17, 1551 25 of 33 Table 9. User daily variation analysis—indicative results. User Number of Clusters Sum of Top Clusters Behavior Suitable for DR user_3 3 100 Consistent TRUE user_12 2 100 Consistent TRUE user_68 3 100 Consistent TRUE user_209 2 100 Consistent TRUE user_330 8 63.29 Diverse FALSE user_450 11 47.12 Diverse FALSE user_1250 2 100 Consistent TRUE user_2655 4 82.47 Predictable TRUE user_3003 2 100 Consistent TRUE user_3651 2 100 Consistent TRUE user_4266 5 92.47 Predictable TRUE Out of the analyzed users, most exhibit either consistent or predictable behavior, making them suitable for targeted DR interventions. Users with diverse behavior may require further analysis or alternative energy management strategies to optimize their participation in DR programs. These results help prioritize residential users for DR implementation, focusing on those with stable and predictable energy usage patterns. 5.2.2. Service 4 Service 4 identifies the peak energy consumption period for each user and assesses their suitability for DR strategies. The day is segmented into four periods: • Midnight (00:00–05:00); • Morning (05:00–12:00); • Afternoon (12:00–18:00); • Evening (18:00–24:00). The objective is to pinpoint users whose peak consumption occurs during evening hours, as this aligns with typical residential demand spikes and offers the most potential for DR interventions. The results presented in Table 10 highlight that users such as user_516, user_933, user_4268, and user_5100 exhibit peak consumption during the evening, making them DR-ready. These users can benefit from load-shifting strategies to reduce stress on the grid. Users like user_159 (morning) and user_1162 (midnight) have peaks outside the evening window, rendering them less suitable for typical residential DR programs. Users with afternoon peaks, such as user_2383 and user_3698, may still require tailored interventions but are not prioritized for evening-focused DR strategies. Table 10. User consumption peak identification—indicative results. User Afternoon Evening Midnight Morning Peak DR Ready? user_159 0.20 0.21 0.15 0.22 morning FALSE user_516 0.01 0.02 0.01 0.01 evening TRUE user_933 0.04 0.07 0.02 0.04 evening TRUE user_1162 0.12 0.10 0.27 0.07 midnight FALSE user_2383 0.07 0.04 0.01 0.05 afternoon FALSE user_3698 0.18 0.12 0.06 0.06 afternoon FALSE user_4268 0.17 0.21 0.03 0.11 evening TRUE user_5100 0.16 0.22 0.04 0.15 evening TRUE Overall, evening peak users are approximately half of the analyzed users, which demonstrate evening peaks, making them prime candidates for load reduction or pricebased incentives during peak hours. Non-Evening Peaks are users with peaks in other time periods (i.e., morning, afternoon, and midnight) may require further analysis for customized DR solutions. The identification of peak periods allows energy managers to Sustainability 2025,17, 1551 32 of 33 30. Bimenyimana, S.; Asemota, G.; Li, L. Clustering Residential Electricity Consumption: A Case Study. In Proceedings of the EEET’18: 2018 International Conference on Electronics and Electrical Engineering Technology, Tianjin, China, 19–21 September 2018; pp. 44–51. [CrossRef] 31. Toussaint, W.; Moodley, D. Clustering Residential Electricity Consumption Data to Create Archetypes that Capture Household Behaviour in South Africa. S. Afr. Comput. J. 2020,32. [CrossRef] 32. Choi, H.; Qureshi, N.M.F.; Shin, D. Comparative Analysis of Electricity Consumption at Home through a Silhouette-score prospective. In Proceedings of the 2019 21st International Conference on Advanced Communication Technology (ICACT), PyeongChang, Republic of Korea, 17–20 February 2019; pp. 589–591. [CrossRef] 33. Bao, H.; Guo, S.; Mo, J.; Zhao, Z.; Wang, Z.; Chen, Z.; Liang, J. An analysis method for residential electricity consumption behavior based on UMAP-CRITIC feature optimization and SSA-assisted clustering. In Energy Reports, Proceedings of the 2022 3rd International Conference on Power, Energy and Electrical Engineering (PEEE 2022), Barcelona, Spain, 18–20 November 2022; Volume 9, pp. 245–254. [CrossRef] 34. Afzalan, M.; Jazizadeh, F.; Eldardiry, H. Two-Stage Clustering of Household Electricity Load Shapes for Improved Temporal Pattern Representation. IEEE Access 2021,9, 151667–151680. [CrossRef] 35. Morales, F.; Garcia Torres, M.; Velázquez, G.; Daumas-Ladouce, F.; Gardel, P.; Gómez-Vela, F.; Divina, F.; Noguera, J.; Ayala, C.; Pinto-Roa, D.; et al. Analysis of Electric Energy Consumption Profiles Using a Machine Learning Approach: A Paraguayan Case Study. Electronics 2022,11, 267. [CrossRef] 36. Hmwe, T.T.; Thein, N.Y.T.; Cho, K.M. Improving Clustering Quality Using Silhouette Score. J. Comput. Appl. Res. 2020,1, 58–62. 37. Skaif, A.; Ayache, M.; Kanaan, H. Energy consumption clustering using machine learning: K-means approach. In Proceedings of the 2021 22nd International Arab Conference on Information Technology (ACIT), Muscat, Oman, 21–23 December 2021; pp. 1–7. [CrossRef] 38. Hosseini, S.; Carli, R.; Dotoli, M. Robust Optimal Energy Management of a Residential Microgrid Under Uncertainties on Demand and Renewable Power Generation. IEEE Trans. Autom. Sci. Eng. 2020,18, 618–637. [CrossRef] 39. Albert, A.; Rajagopal, R. Smart Meter Driven Segmentation: What Your Consumption Says About You. IEEE Trans. Power Syst. 2013,28, 4019–4030. [CrossRef] 40. Toussaint, W.; Moodley, D. Automating Cluster Analysis to Generate Customer Archetypes for Residential Energy Consumers in South Africa. arXiv 2020, arXiv:2006.07197. 41. Walstad, K.; Vadlamudi, V.V. Electric Utility Customer Segmentation from Advanced Metering System Data Using K-Shape Clustering—A Norwegian Case Study. In Proceedings of the 2022 IEEE PES Innovative Smart Grid Technologies Conference Europe (ISGT-Europe), Novi Sad, Serbia, 10–12 October 2022; pp. 1–6. [CrossRef] 42. Rawat, T.; Niazi, K.R.; Gupta, N.; Sharma, S. A two stage interactive framework for demand side management in smart grid. In Proceedings of the 2019 IEEE 16th India Council International Conference (INDICON), Rajkot, India, 13–15 December 2019; IEEE: New York, NY, USA, 2019; pp. 1–4. 43. Rawat, T.; Niazi, K.; Gupta, N.; Sharma, S. A two-stage optimization framework for scheduling of responsive loads in smart distribution system. Int. J. Electr. Power Energy Syst. 2021,129, 106859. [CrossRef] 44. F.R.S., K.P. LIII. On lines and planes of closest fit to systems of points in space. Lond. Edinb. Dublin Philos. Mag. J. Sci. 1901, 2, 559–572. [CrossRef] 45. Rousseeuw, P.J. Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. J. Comput. Appl. Math. 1987, 20, 53–65. [CrossRef] 46. Yilmaz, S.; Chambers, J.; Patel, M. Comparison of clustering approaches for domestic electricity load profile characterisation– Implications for demand side management. Energy 2019,180, 665–677. [CrossRef] 47. Liu, F.T.; Ting, K.M.; Zhou, Z.H. Isolation Forest. In Proceedings of the 2008 Eighth IEEE International Conference on Data Mining, Pisa, Italy, 15–19 December 2008; pp. 413–422. [CrossRef] 48. Nizan, O.; Tal, A. k-NNN: Nearest Neighbors of Neighbors for Anomaly Detection. arXiv 2023, arXiv:2305.17695. [CrossRef] 49. Wu, J. Advances in K-Means Clustering: A Data Mining Thinking; Springer Science & Business Media: Berlin/Heidelberg, Germany, 2012. 50. Nielsen, F. Hierarchical Clustering. In Introduction to HPC with MPI for Data Science; Springer: Cham, Switzerland, 2016; pp. 195–211. [CrossRef] 51. Ng, A.Y.; Jordan, M.I.; Weiss, Y. On spectral clustering: Analysis and an algorithm. In Proceedings of the 14th International Conference on Neural Information Processing Systems: Natural and Synthetic, Vancouver, BC, Canada, 3–8 December 2001; MIT Press: Cambridge, MA, USA, 2001; NIPS’01; p. 849–856. 52. Ester, M.; Kriegel, H.P.; Sander, J.; Xu, X. A density-based algorithm for discovering clusters in large spatial databases with noise. In Proceedings of the KDD’96: Proceedings of the Second International Conference on Knowledge Discovery and Data Mining, Portland, OR, USA, 2–4 August 1996; Volume 96, pp. 226–231. 53. Wu, K.L.; Yang, M.S. Mean shift-based clustering. Pattern Recognit. 2007,40, 3035–3052. [CrossRef] Sustainability 2025,17, 1551 33 of 33 54. Reynolds, D. Gaussian Mixture Models. In Encyclopedia of Biometrics; Li, S.Z., Jain, A., Eds.; Springer: Boston, MA, USA, 2009; pp. 659–663. [CrossRef] 55. Ho, T.K. Random decision forests. In Proceedings of the 3rd International Conference on Document Analysis and Recognition, Montreal, QC, Canada, 14–16 August 1995; IEEE: New York, NY, USA, 1995; Volume 1, pp. 278–282. 56. Davies, D.L.; Bouldin, D.W. A Cluster Separation Measure. IEEE Trans. Pattern Anal. Mach. Intell. 1979,PAMI-1, 224–227. [CrossRef] 57. Cali´nski, T.; Harabasz, J. A dendrite method for cluster analysis. Commun. Stat. 1974,3, 1–27. [CrossRef] 58. Greater London Authority. Smartmeter Energy Use Data in London Households; Greater London Authority: London, UK, 2024. 59. Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V.; et al. Scikit-learn: Machine learning in Python. J. Mach. Learn. Res. 2011,12, 2825–2830. 60. McKinney, W. Data structures for statistical computing in python. In Proceedings of the 9th Python in Science Conference, Austin, TX, USA, 28 June–3 July 2010; Volume 445, pp. 51–56. 61. Plotly Technologies Inc. Collaborative Data Science; Plotly Technologies Inc.: Montreal, QC, Canada, 2015. 62. Zhang, H.; Li, Z.; Xue, Y.; Chang, X.; Su, J.; Wang, P.; Guo, Q.; Sun, H. A Stochastic Bi-level Optimal Allocation Approach of Intelligent Buildings Considering Energy Storage Sharing Services. IEEE Trans. Consum. Electron. 2024,70, 5142–5153. [CrossRef] Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.