Full text
Lappeenranta University of Technology School of Engineering Science Computational Engineering and Technical Physics Intelligent Computing Master’s Thesis Kalifa Manjang IDENTIFICATION OF CUSTOMER PROFILES FROM ELECTRICITY CONSUMPTION DATA Examiners: Prof. Lasse Lensu Assoc. Prof. Samuli Honkapuro Supervisors: Adjunct Prof., Dr. Xiao-Zhi Gao Associate Prof. Arto Kaarna Prof. Lasse Lensu
2 ABSTRACT Lappeenranta University of Technology School of Engineering Science Computational Engineering and Technical Physics Intelligent Computing Kalifa Manjang Identification of customer profiles from electricity consumption data Master’s Thesis 2018 63 pages, 23 figures, 16 tables. Examiners: Prof. Lasse Lensu Assoc. Prof. Samuli Honkapuro Keywords: K-means clustering, genetic algorithm, power user profiling, Davies-Bouldin index, Silhouette index, Calinski-Habarasz index. The electric power suppliers are interested in identifying and categorising their consumers’ profiles into different categories according to their energy consumption habits. The profiling of users can help with understanding how the users consume the energy and how the energy usage may affect the electricity distribution grid. However, the privacy of the electricity users is well protected by the current law. This study focuses on data mining methods to extract the relevant knowledge based on anonymous data obtained from smart meters. The K-means clustering algorithm was used in grouping the energy consumption data. To improve the quality of the clusters formed via the K-means clustering and to tackle the common problem of local optimum, the genetic algorithm (GA) was adopted in refining the clusters. The use of two validity indices to compare the methods showed that combining K-means and GA did indeed improve the clustering quality.
3 PREFACE Bismillahi, Rahmani, Rahmeen all praise be to Allah. I would like to thank my supervisors for their undivided attention, dedication and guidance provided to me during the course of this master’s thesis. This work would not have otherwise been achieved without their support. I thank my beautiful wife for her patience. To my parents, thank you for supporting my dreams of getting a higher education and instilling in me discipline, respect and hard work. Finally, I would like to thank the LUT administration for the scholarship I was offered to pursue a double degree program in this prestigious institution. I am grateful. Lappeenranta, August 31, 2018 Kalifa Manjang
4 CONTENTS 1 INTRODUCTION 7 1.1 Background................................. 7 1.2 Objectives and delimitations . . . . . . . . . . . . . . . . . . . . . . . . 8 1.3 Structureofthethesis............................ 9 2 ELECTRICITY CONSUMER PROFILING 10 2.1 Poweruserprofiling............................. 10 2.2 Techniques used in power user profiling . . . . . . . . . . . . . . . . . . 11 2.2.1 Neural approaches . . . . . . . . . . . . . . . . . . . . . . . . . 11 2.2.2 Clustering algorithms . . . . . . . . . . . . . . . . . . . . . . . . 11 2.2.3 Statistical approaches . . . . . . . . . . . . . . . . . . . . . . . . 12 2.2.4 Fuzzyapproaches.......................... 12 2.2.5 Hybridmethods........................... 13 2.3 Reviewoftechniques............................ 13 3 PROPOSED APPROACH FOR ELECTRICITY POWER PROFILING 17 3.1 K-meansclustering ............................. 17 3.2 Geneticalgorithm.............................. 19 4 EXPERIMENTS AND RESULTS 21 4.1 Descriptionofdata ............................. 21 4.2 Pre-processing................................ 21 4.3 Dimensionality reduction . . . . . . . . . . . . . . . . . . . . . . . . . . 22 4.4 Evaluationcriteria.............................. 23 4.4.1 Silhouetteindex .......................... 23 4.4.2 Davies–Bouldin index . . . . . . . . . . . . . . . . . . . . . . . 24 4.4.3 Calinski-Harabasz index . . . . . . . . . . . . . . . . . . . . . . 25 4.5 Implementation of experiments . . . . . . . . . . . . . . . . . . . . . . . 27 4.6 Results.................................... 29 4.6.1 Selecting the number of clusters for the annual load profiles . . . 29 4.6.2 Selecting the number of clusters for the daily load profiles . . . . 31 4.6.3 The within-cluster sum of squares for annual load profiles . . . . 33 4.6.4 The within cluster sum of squares for daily profiles . . . . . . . . 34 4.7 The cluster representation for the annual load profiles . . . . . . . . . . . 35 4.8 The cluster representation for the daily load profiles . . . . . . . . . . . . 41 4.8.1 Similarity measure for annual profiles . . . . . . . . . . . . . . . 44 4.8.2 Similarity measure for daily load profiles . . . . . . . . . . . . . 47
5 4.8.3 Annual weekend load profile . . . . . . . . . . . . . . . . . . . . 48 4.8.4 Refining annual load profiles . . . . . . . . . . . . . . . . . . . . 50 4.8.5 Refining daily load profiles . . . . . . . . . . . . . . . . . . . . . 54 4.9 Methodcomparison............................. 56 5 DISCUSSION 58 6 CONCLUSION 60 REFERENCES 61
6 LIST OF ABBREVIATIONS BCSS Between-Cluster Sum of Squares CDI Clustering Dispersion Indicator CFSFDP Fast Search and Find of Density Peaks DB Davies–Bouldin EA Evolutionary Algorithms FCM Fuzzy Clustering Means GA Genetic Algorithm SAX Symbolic Aggression Approximation SVM Support Vector Machine SOM Self Organizing Maps WCSS Within-cluster Sum of Squares
7 1 INTRODUCTION 1.1 Background With the emergence of smart meters, more information about a user’s electricity consumption can be collected easily. Prior knowledge about the group a particular user belongs to is known by the energy company to some extent. This is achieved through knowledge about the type of appliances in use or type of heating system used in the buildings. This information is stored and used in customer grouping. One shortcoming of this method is that the recorded information is seldom updated. With time, the energy consumption of the user, for example, the type of heating or electrical appliance usage might change remarkably from the known behaviours. In this regard, this single-shot method of consumer categorization is inefficient. The traditional energy user grouping is performed using three user categories: industrial, residential and commercial users. An example of the industrial users are factories, commercial users are the shops, restaurants and supermarkets, and residential users refer to homes and apartment buildings. The consumption pattern of the energy users is much more complex than these mentioned groups [1]. The energy users should be categorized based on the pattern of electricity behaviour they exhibit. The general application of electricity user profiling is that the knowledge of how the customers use the electricity can help the energy companies to create important policies ranging from network planning, demand response, and load forecasting [1]. A more detailed application of load profiling is described as follows [2]: Distribution network operation •Real-time recognition of network loadings and voltages. •It ensures that the network is kept within its operating limits. •It can manage post-fault supply restoration. Short-term operation planning •Very useful in congestion forecasting.
8 •Load profiles have immense application in network reconfiguration, for instance, to minimize network lost. •Planned outages preparation. Distribution network planning •New base profile for probabilistic network planning. •Ensuring the network can cost-efficiently host all foreseeable loads and generators. Other applications of load profiling are the design of tariffs, target sales based on customers load profiles, the load profile can also be used in the adjustment of the electricity retail forecasts when new customers are contracted or old ones lost [2]. By virtue of this important demand, finding and understanding clusters using data mining techniques (scientific methods) is worthwhile [3]. Since the user identity is protected under the European Union privacy laws [4], the specific locations and identities of the participants will remain anonymous. This research applies data mining methods (clustering) to the provided data set to show users with similar energy consumption patterns and group these users together. 1.2 Objectives and delimitations This master’s thesis aims at achieving the following objectives: •To study and use the K-means clustering algorithm for the purpose of load profiling. •To choose and verify the appropriate number of clusters to use in the categorization of the energy users. •To use the GA to improve the quality of the clusters formed. This study is limited in that the results that will be obtained cannot be verified because of the anonymity of the participants.
9 1.3 Structure of the thesis The outline of this master’s thesis is as follows. Chapter 2 describes related work on power use profiling, Chapter 3 contains the proposed methods for power user profiling and the algorithms for these methods are presented. In Chapter 4, the application of these specific methods to electricity load data and the results derived from the experiment are analysed. Chapter 5 contains the discussions and the challenges faced during the study. Finally, the concluding remarks are presented in Chapter 6.
16 Table 2. Summary of the K-means algorithm Advantages Disadvantages Refs. K-means ·Simple ·Sensitive to selection of initial centroids. [22–24] ·Scalable and Efficient ·Number of cluster has to be defined. ·Can handle big data ·Sensitive to noise and outliers. ·Linear complexity ·Local optimum solution. The K-means algorithm terminates primarily if the data points in the respective clusters are not reassigned to a different cluster or if the maximum number of allowed iterations is attained. For this study, the latter is used.
17 3 PROPOSED APPROACH FOR ELECTRICITY POWER PROFILING 3.1 K-means clustering The K-means algorithm is a partition clustering algorithm. It was introduced by J.B. MacQueen in 1967 [25]. The algorithm is based on unsupervised learning used with unlabeled multidimensional data. The goal of the algorithm is to group the unlabeled multidimensional data into K clusters (K is fixed a priori). The K variable represents the number of groups for the partition. It works by iteratively assigning data points to one of the K groups based on the provided features. Each data point is assigned to one unique group. The algorithm is favoured in many application areas such as computer vision, image processing, business analytics etc. Its popularity is due to the simplicity and linear complexity, defined as O(I∗n∗K∗D), where Irepresents the number of iteration, n is the number of input features, Kis the cluster number and Dis the dimension of the features [26]. The K-means algorithm includes two steps: 1.Cluster Assignment step 2. Move centroid step. In the cluster assignment step, the idea is to define K centroids for the clusters, one for each cluster. The K-means result is sensitive to the initial centroids, different initial centroid yield different results. Therefore, a good choice is to place them farther away from each other. The next step involves examining each data point and assign the data point to the closest centroid. In the move centroid step, the algorithm calculates the average of all the data points in each cluster and the centroid is moved to that location. This continues until no changes in the clusters occur or until some stopping criterion is met. The algorithm aims at minimizing an objective function, which in this case is the squared error function: J= k X j=1 n X i=1 kx(j) i−cjk2 where k=number of clusters, n=number of data points, and (1) kx(j) j−cjk2is the distance metric used. That is the distance between the load profile x(j) j and the cluster center cj. The distance metric used in this case is the Euclidean distance. The Euclidean distance formula is given in Equation 2.
18 d(x, c) = v u u t x X k=1 (xk−ck)2.(2) Assuming Xis the set of load profiles with X=x1, x2, ..., xnand V=v1, v2, ..., vkis the set of cluster centroids, the K-means algorithm proceeds by the following steps: 1. ccluster centroids are randomly selected. 2. The distances between each load profile and the cluster centroids are calculated. 3. Assign the load profile to a cluster centroid with the minimum distance. 4. Recalculate the new cluster centroids as follows: Vi=1 |ci| ci X j=1 xi(3) cirepresents the number of data points in the ith cluster. 5. Recalculate the distance between each load profile and the new cluster centroids. 6. If no single load profile is reassigned to a cluster centroid ,the algorithm stops else proceed to step 3.
19 3.2 Genetic algorithm Genetic algorithms are biologically-inspired heuristic search optimization algorithm. They are inspired by Charles Darwin’s theory of evolution i.e the survival of the fittest. The algorithm exhibit the process of natural selection in which the fittest individuals are chosen to reproduce the offspring of the next generation. The genetic algorithm essentially replicate the way in which life uses evolution to find solutions to real world problems [27]. There are five phases considered in a genetic algorithm [27]: 1. Initial population 2. Fitness function 3. Selection 4. Crossover 5. Mutation A brief description of these phases is given below: Initial population Population of randomly generated solutions to the problem. Clearly, randomly generated solutions to the problem might not be too ideal. Fitness The fitness quantitatively evaluates how fit a given solution is or how fit individuals can be produced from the given solutions i.e., the fitness ability of an individual to compete with others. A fitness score is assigned to each individual, the selection of an individual for reproduction depends on its fitness score [27].
20 Selection In the strive to achieve convergence, the best offsprings are selected as parents in the new parental population. The selection of the offsprings are based on their fitness values [27]. Crossover During crossover the genetic material of the parents are combined. This can be thought of as mimicking the mating process in real life. By combining certain traits from two or more individuals, the hope is that a ’fitter’ offspring will evolve with the best traits inherited from the parents [27]. Mutation Mutation-operators provides random changes to the population by disturbing them. Mutation typically allow very small changes at random to the individual genomes [27]. Mutation maintains diversity within the population and help prevent fast convergence. Termination When convergence is attain the algorithm terminates. At this point it can be said that the algorithm has provided a solution to the problem. The GA cycle is given in Figure 3. Figure 3. Genetic Algorithm cycle [27].
21 4 EXPERIMENTS AND RESULTS The proposed algorithms are implemented on Matlab R2017a version, on a Windows 10 machine with 8GB of RAM. The Matlab inbuilt function ’Kmeans’ and ’ga’ were used for the implementation. The software provides flexibility in reading and displaying stored files. 4.1 Description of data The data set used in the experiment are time series hourly electricity consumption data for 13601 households in Southern Finland. The data are based on hourly loads recorded for a span of one year. The rows in the raw data set represent the time-stamps and the columns the respective customers. The dimension of the load data is 8760 ×13601. 4.2 Pre-processing The given data set was pre-processed to remove missing values. In checking the load data for missing information the Matlab inbuilt functions ’isnan’ was used. A single user’s data was found with missing information. Only that particular user was excluded from the final data set used for the experiments. The K-means algorithm, in this case, uses the Euclidean distance metric. The Euclidean distance is known to be biased due to the scale of the measurements, to this regard the raw electricity consumption data was standardized, so that it has zero mean and unit variance. The following formula was used for the standardization: ¯ X=(Xi−µi) σi .(4) where Xis the load profile data, µis the mean and σrepresent the variance. For each respective user load, µiand σirepresent the mean and variance respectively of the entire data for that particular user.
22 4.3 Dimensionality reduction As the number of dimensions increases, the distance between any two points in the same data sets converges (the maximum distance and the minimum distance between any two points will be identical) [28]. This tends to be an issue with the Euclidean distance metric. Reducing the dimensionality prior to the K-means clustering can alleviate this problem and considerably help with the computation. The dimensionality reduction technique used was adopted from [29]. To reduce the load profile data from ndimensions to Ndimensions, the data was divided into zequally-sized frames. The mean of all the data within this frame was computed and a vector Nof all the mean values derived becomes the new representation of the original data. This dimension reduction was needed only for deriving the annual profiles. The whole data set was considered in building the annual profiles hence, the need to compress the size of the data. For this study, the data was divided into 24 equally-sized frames (8760 rows into 24 equalsized frames), 24 because each user provides 24 data points a day. In simple terms, the average of the load data provided in a day represents the electricity consumption on that particular day. For the annual load profiles, the dimension of the data is reduced from 8760×13600 to 365×13600. Figure 4 shows the full load profiles and the corresponding dimensional reduced profiles. The dimensionality reduced profile was obtained by the method describe above. It can be seen that the shape of the two profiles has not changed. (a) (b) Figure 4. Load profile of residential flats: (a) A full load profile. (b) Dimensionality reduced load profile.
23 4.4 Evaluation criteria Various methods can be used to quantify the performance of a clustering algorithm as well as to provide a technique for the selection of the appropriate number of clusters. The evaluation criteria can be categorized as similarity-oriented and classification-oriented [30]. In determining the optimal number of clusters, three validity indices Silhouette, DavidBouldin and Calinski-Harabasz index were used: 4.4.1 Silhouette index In silhouette analysis, the separation distance between the clusters is studied. It gives a measure of closeness between the points in one cluster to the points in the neighbouring clusters. The formal definition of this quality index was adopted from [31]. Let X=x1, ..., xNbe the load profile data set and let C=c1, ..., ckbe its clustering in some kclusters. Let us denote d(xk, xi)to be the distance between xkand xi. Let cj=xj i, ..., xj mjbe the jth cluster where j= 1, ..., k and mj=|cj|.aj idenotes the average distance between the ith vector in the cluster cjand the vectors in the same cluster. The average distance aj iis hence given by : aj i=1 mj−1 mj X k=1,k6=i d(xj i, xj k), i = 1, ...., mj.(5) The minimum average distance between the ith vector cjand all the vectors clustered in cluster ck, where k= 1, ..., K and k6=jis given as follows : bj i= min n=1,...,k,n6=j1 mn mn X k=1 d(xj i, xn k), i = 1, ..., mj.(6) The ith vector silhouette width in cluster cjis given below: sj i=bj i−aj i max(bj i, aj i).(7)
24 The silhouette width is in the range [−1,1]. The silhouette of a cluster cjgiven as: sj=1 mj mj X i=1 sj i.(8) The algorithm for determining the optimal number of clusters using the Silhouette index is given as follows: 1. Perform K-means clustering for the range of values of K. 2. For each value in the range, an average Silhouette was calculated for the observation. 3. A plot of the curve according to the average silhouette was generated. 4. The location of the maximum is the optimal number of clusters. The Matlab inbuilt function ’evalclusters’ was used to achieve this. 4.4.2 Davies–Bouldin index The Davies–Bouldin (DB) index was introduced in 1979 by David L. Davies and Donald W. Bouldin. It is the ratio between the within-cluster distances and the between-cluster distances and computing the average over all clusters [31]. The formal definition of the Davies-Bouldin index was adopted from [32]. Let δkdenote the mean distance of the point in the cluster ckto their centroids Gk: δk=1 nk nk X i=1 kMk i−Gkk.(9) where Mk iis the n-dimensional feature vector assigned to cluster ck, and nkis the size of the cluster. Let us denote also, ∆kk0=d(Gk, Gk0) = ||Gk−Gk 0 ||.(10)
25 that is, the distance between the centroid Gkand Gk 0 of clusters ckand ck0. For all indices k06=k, the Davies-Bouldin index is as follows: C=1 K K X i=1 maxk06=kδk+δ0 k ∆kk0,where K=number of clusters.(11) The algorithm to find the optimal value using the Davies-Bouldin index is similar to the Silhouette method. The algorithm is given below as: 1. Perform K-means clustering for the range of values of K 2. For the values in this range, calculate the Davies-Bouldin index. 3. A plot is generated for each value of K. 4. The location of the minimum is considered to be the optimal number of clusters. 4.4.3 Calinski-Harabasz index In the Calinski-Harabasz index, the comparison of the between-clusters variance to the within-cluster variance is made. The index was first introduced in 1974. The formal definition is derived from [32] and it is given as: C=BGSS WGSS ×N−K K−1(12) where Krepresent the number of clusters, Nis the total number of load profiles. The overall within-cluster variance is denoted W GSS and the overall between-cluster variance as BGSS [32]. BGSS is calculated as the total sum of squares subtracted from WGSS. The total sum of squares is the squared distance of all the load profiles from the centroids.
32 Figure 7. Silhoutte, Davies-Bouldin and Calinski-Harabasz index. In Figure 8 the two daily weekday load profiles are represented. Figure 8. Daily weekday load profiles, 6910 and 6690 load profiles in each respective cluster.
33 Figure 6 and 8 present the new clusters for the annual weekday profiles and the daily weekday profiles respectively. From the Figures, the clusters seem to show only two groups of users i.e. users with a flat electricity consumption pattern and those users whose electricity consumption varies across the year or day. An observation of the formed clusters showed that a lot of averaging occurred and some potential unique traits of the respective users are not exhibited. Also, considering the number of users in the data, 2clusters is too small to represent the electricity consumption behaviours of these users. Previous work considered in this study had the optimum number of clusters higher that 2as seen in Table 1. Therefore, it is safe to argue that 2is not appropriate for the categorization of the load profiles. Other ways of choosing the appropriate number of clusters were examined. A different range of values needs to be considered for the appropriate number of clusters. The Davies-Bouldin index is chosen because the other two validity indices, in this case, do not seem to be the appropriate method for selecting the number of clusters. The Davies-Bouldin index values that seem to be the potential solutions are looked at, these values are 4,8,12 and 15 for the annual weekday load profiles and 6and 14 for the daily weekday profiles. Before the potential Davies-Bouldin values are analysed in detail ,we first study the the within-cluster sum of squares for both the daily and annual load profiles 4.6.3 The within-cluster sum of squares for annual load profiles To further analyze the appropriate number of clusters, the within-cluster sum of squares was applied. The degree of variability of the load profiles in each cluster is given by the within-cluster sum of squares (WCSS). The WCSS decreases as the number of clusters increase. The appropriate number of clusters can be selected this way, the hint is to choose the number of clusters from which the WCSS drop is not very large. In Figure 9, the WCSS drop is not large around 8therefore, the annual load profiles can be categorized into 8clusters. The WCSS is supposed to decrease and stay low as the number of clusters increase. The situation is different when the number of clusters is 12 and 15, the WCSS increased instead of staying low. The percentage of variance as a function of the number of clusters is looked at. A number of clusters should be chosen so that an addition of another cluster does not give a better modelling of the data. In this view, even though the WCSS drop did not stay low at 12 and 15, the two values do not give a much better result than when the number of clusters was 8.
34 Figure 9. Within-cluster sum of squares for annual load profiles. 4.6.4 The within cluster sum of squares for daily profiles The within-cluster sum of squares is also utilized to study the appropriate number of clusters for categorizing the daily load profiles. Figure 10 provides the WCSS plot. Figure 10. Within-cluster sum of squares for daily load profile.
35 From the plot in Figure 10 it is seen that the WCSS drop is not substantial around cluster 6and 7. These values can, therefore, suggest the appropriate number of clusters. The WCSS did not stay low for all the values analyzed. At 12 the WCSS increased instead of dropping. The same argument used with the annual load profiles also applies here. The value at 12 is still lower than the value at the appropriate number of clusters. 4.7 The cluster representation for the annual load profiles For each of the cases considered, i.e., Case 1, Case 2, Case 3and Case 4(Corresponding to 4,8,12 and 15 clusters respectively), the number of profiles in each case is given in Table 5. Table 5. The number of consumers in each of the cases considered above. Cluster Case 1 (4 profiles) Case 2 (8 profiles) Case 3 (12 profiles) Case 4 (15 profiles). 1 1538 942 850 794 2 2166 1068 497 466 3 4610 3809 525 2142 4 5286 1244 569 410 5•2304 1477 460 6•850 289 242 7•742 487 437 8•2641 1811 683 9• • 569 1619 10 • • 2659 1799 11 • • 1384 1235 12 • • 2483 1481 13 ••• 385 14 ••• 1044 15 ••• 403
36 Case 1: 4Annual weekday profiles In Cluster 1found in Figure 11, a peak appeared in the first month of the year. The high electricity consumption rate declined as the year proceeds, this trend continued until midyear. Generally, the weather in Finland is friendlier around this time of the year hence electric heating is not a necessity, evident in the relatively stable electricity consumption showed. Around the end of the year, the consumption rates are shown to be high once again, this rise in the pattern of consumption is attributed to the drop in temperature which is experience around the beginning of the winter season. Another peak in electricity consumption was observed at this time. This profile represent residential homes where district heating is not provided so resident have to result to providing heating for themselves during the winter period. Some confidence can be given to this claim due to the major peaks recorded around the time when the temperatures are very low. Figure 11. Case 1: 4Annual weekday profiles. In Cluster 1found in Figure 11, a peak appeared in the first month of the year. The high electricity consumption rate declined as the year proceeds, this trend continued until midyear. Generally, the weather in Finland is friendlier around this time of the year hence electric heating is not a necessity, evident in the relatively stable electricity consumption showed. Around the end of the year, the consumption rates are shown to be high once again, this rise in the pattern of consumption is attributed to the drop in temperature
37 which is experience around the beginning of the winter season. Another peak in electricity consumption was observed at this time. This profile represent residential homes where district heating is not provided so resident have to result to providing heating for themselves during the winter period. Some confidence can be given to this claim due to the major peaks recorded around the time when the temperatures are very low. In Cluster 2of Figure 11, the electricity consumption during the beginning of the year showed a stable behaviour, but a peak in the consumption rates was recorded during the middle of the first month (roughly around the 15th). This rise in consumption corresponds to the beginning of the year when temperatures are at their lowest. The consumption rate became stable for the rest of the year. The consumption rate takes up again around the end of the year when little peaks are seen to appear. These are most probably houses with electrical heating. The difference between this profile and the first profile in Figure 11 is that the consumption rate showed a more stable behavior. In Cluster 3of Figure 11, a low but noticeable peak appeared around January after which the behaviour of the profiles recorded a stable but downwards trend, this pattern stays consistent. Around June/July, the consumption rate declined even further. At this time, the days are normally warmer compared to the rest of the year. A stable increase is recorded right afterwards. The customers in this profile can be attributed to residential homes were a form of air conditioning or cooling system is absent. The absent of a cooling system is noticeable in the downward trend in consumption rates exhibited for warmer days. The profiles in Cluster 4of Figure 11 recorded a stable consumption pattern in electricity usage in the beginning of the year. A rise in the electricity consumption begins shortly afterwards, evident in the peaks seen at roughly around June. This profile is attributed to residential homes in which some form of cooling system is present. This should explain the peaks around that time of the year when temperatures are normally not low. The major peaks declined at around July/August. The pattern stays almost the way through the rest of the year. As the year elapsed the consumption rates began to increase again. Case 2: 8Annual weekday profiles Cluster 1,2,4and 5of Figure 12 show similar consumption behaviour to profiles in Figure 11 i.e, Cluster 1,2,3and 4respectively. Cluster 3of Figure 12 showed the same pattern of energy consumption throughout the
38 Figure 12. Case 2: 8Annual weekday profiles. year. An increase in the electricity consumption is seen in the form of peaks during the early months of the year, but this phenomenon did not last long as the constant pattern continued. Besides few individual peaks, a major peak also occurred at the end of the year. In Cluster 6of Figure 12, a rise in the energy consumption is recorded after June. In this profile, as the end of the year approach, the rate of electricity consumption is noticeably seen to be on the rise. Also, some major peaks formed at the latter end of the year. In Cluster 7of Figure 12, the consumption pattern showed almost the same behaviour as Cluster 3of the same figure. The only noticeable difference is that as the year progressed a decline in the energy consumption is seen in Cluster 7. This trend stays consistent for the whole year. Besides few individual peaks, no major peaks can be seen. These profiles can be attributed to load curves produced by industrial or commercials customers. The consumption of electricity by users in Cluster 8of Figure 12 began very stable. Despite few individual peaks shown by different users. The energy consumption behaviour remained very much the same until around May. A steady rise in energy usage is noticed around June. This trended until the end of the year. The consumption rate slightly
39 declined at some point, but quickly took up again as the year ends. Case 3: 12 Annual weekday profiles In Figure 13, some of the clusters formed have already been seen and addressed previously, i.e., in Case 1and 2. Some new profiles different from those found have also been created. The profiles show some similarity visually, for instance, Cluster 8,9and 11. The most noticeable characteristic exhibited by these clusters is the decline in the energy consumption rates as year advances. Figure 13. Case 3: 12 Annual weekday profiles. A minor rise in the electricity consumption is observed in Cluster 7of Figure 13. Although these peaks are very much noticeable, they are not major peaks. The peaks stayed constant until February, but promptly declined afterwards. At roughly around June, the consumption pattern begins to rise gently. This steady rise is noted throughout the year. This profile can be associated with homes where district heating is absent as evident in the
40 rise and decline of electricity consumption, i.e., energy is demanded more in the colder months and less in the warmer seasons. Cluster 2and Cluster 7have a similar pattern of electricity consumption, However, the peaks in the beginning of the year for Cluster 2are stronger and more prominent. Case 4: 15 Annual weekday profiles Like Figure 13, some of the profiles in Figure 14 have already been met. Again, after Case 2, the resulting clusters of the other cases are not very distinct in many ways. In Figure 14. Case 4: 15 Annual weekday profiles. Figure 14, Cluster 8,9,10,11 and 15 are very related and perhaps these profiles should not form individual clusters but rather be together. The same implies to cluster 5and 13.
41 4.8 The cluster representation for the daily load profiles For each of the cases considered for the daily load profiles i.e. Case 1and Case 2(Corresponding to 6and 14 clusters respectively), the number of profiles in each case is given in Table 6. Table 6. The number of users in each of the cases considered for the daily load profiles. Cluster Case 1 (6 profiles) Case 2 (14 profiles). 1 2333 881 2 4150 1990 3 1499 377 4 2390 1142 5 1059 374 6 2169 1373 7•1308 8•431 9•1005 10 •893 11 •1660 12 •1340 13 •396 14 •430 Case 1: 6daily weekday profiles In Cluster 1of Figure 15, during the early hours of the day, the need for electricity remains moderately low. At about 6 : 00 the demand for energy increased. This trend continued until 14 : 00 when finally the demand subsided. This decrease in electricity consumption continued throughout the day. In this profile, the energy requirements were high in the mornings and a greater part of the afternoon when the demand started to decline. This can be attributed to homes were residents spend a good part of the mornings and afternoons at home. This profile might indicate households were residents are away from their homes at mid-day, perhaps they work in the evenings and nights. It can also indicate profiles for restaurants and small cafeterias. A closer look at the profile supports this claim. The pattern of energy consumption is somewhat higher at 1 : 00 in Cluster 2of Figure 15 compared to the other hours (in the same profile). A reduction in the electricity usage is evident from around 5 : 00, this steady decrease trended for a couple of hours until 9 : 00
48 Table 12. CASE 2: Similarity measure for 14 daily load profiles. Cluster 1234567891011121314 1 011111010 0 0 1 1 0 2 101111011 1 1 1 1 1 3 110111111 1 1 1 1 1 4 111001000 0 0 0 0 0 5 111001100 1 1 1 0 1 6 111110111 1 1 1 1 1 7 001011010 0 0 1 1 1 8 111001100 0 1 1 0 1 9 011001000 0 0 1 0 1 10 0 1 1 0 1 1 0 0 0 0 0 0 0 0 11 0 1 1 0 1 1 0 1 0 0 0 1 1 1 12 1 1 1 0 1 1 1 1 1 0 1 0 1 0 13 1 1 1 0 0 1 1 0 0 0 1 1 0 1 14 0 1 1 0 1 1 1 1 1 0 1 0 1 0 In case 1, i.e, Table 11, the centroids of Cluster 4and 5are similar according to the Kolmogorov-Smirnov test. The remaining user profiles are different according to the same test. In the second case of the daily load profiles, i.e, Table 12, the similarity between these load profiles are higher. An increase in the number of profiles results in users with similar electricity consumption patterns ending up in different clusters. This explains the high similarity in Table 12. Cluster 3and 6are the only distinct profiles in the table, the rest of the clusters have one or more users profiles they are similar to. Due to the high similarity in Table 12 the appropriate number of clusters for the daily load profiles cannot be 14 clusters. This leaves only one choice for the appropriate number of clusters, i.e, the clusters in Table 11 . The similarity measure is intuitive in the case of the daily load profiles. Cluster 4and 5in Table 11 are not only similar according to the KolmogorovSmirnov 2 sample test but also have a similar pattern of consumption. These load profiles can be merged. Hence, the appropriate number of clusters for the daily load profiles is, therefore, 5clusters. 4.8.3 Annual weekend load profile The annual weekend load profiles were created using 8clusters (the appropriate number of clusters for the annual load profiles is 8). The electricity consumption during the weekdays for the load profiles that are considered to indicate residential households is presumed to vary considerably from the weekend’s consumption patterns (working class
49 resident spend more time at home during the weekends). For the industrial and commercial customers, the production rate is lesser during the weekends (businesses open and close production differently during the weekends). The electricity consumption pattern is also impacted by this. As a consequence, the annual weekend load profiles were obtained and compared to the annual weekday load profiles already addressed in Figure 12. Figure 17. Annual weekend load profiles. Ignoring the position of the clusters in the respective figures, the weekend load profiles and the weekday profiles are comparable. However, some differences can be noted. In Cluster 2of the weekend load curves, the electricity consumption is below 5kw whereas in the weekday load profiles many of the profiles have consumption rates above 5kw. Also, all the users in the weekend profile have consumption below 10 kw except for few users in Cluster 3and Cluster 8. In the weekday profiles many of the profiles have users whose power consumption rates are above 10 kw. The consumption of energy is higher during the weekday in comparison to the weekends.
50 4.8.4 Refining annual load profiles As seen previously, the number 4,8,12 and 15 were identified as the potential value for the appropriate number of clusters in the case of the annual load profiles. The cluster representation for each of these cases have been presented already and the characteristics of the respective user profiles elaborated. In this subsection, the cluster representation of these numbers is replicated using a different method. The clusters are formed by providing the initial centroids to the K-means clustering algorithm rather than relying on the random initial centroids selection as was the case previously. To ensure that the global optimum is attained, these initial centroids were optimised using the GA. The clusters for the various cases of the annual load profiles are given in Figure 18, 19, 20, 21. Figure 18. 4Refined annual load profiles.
51 Figure 19. 8Refined annual load profiles.
52 Figure 20. 12 Refined annual load profiles.
53 Figure 21. 15 Refined annual load profiles. As expected the replication of the load profiles using a different method yielded similar profiles as before. In the first instance, for example, the four clusters produced are similar to the clusters found in Figure 11. The same implies to Figure 12, 13 and 14. The position of the load profiles might be different in each respective figure, but these profiles look alike. In Figure 21, Cluster 10 is empty (the cluster contains only the initial centroids value). Figure 14 has the same number of clusters as Figure 21 however, no empty clusters are produced in the former. This is one noted different.
54 4.8.5 Refining daily load profiles The daily load profiles were also replicated using the same method described for the annual load profiles. In the annual load profiles replication, it has been seen that the shapes of the load profiles or pattern of electricity consumption did not change despite the method used. The same characteristics were also exhibited during the daily load profiles refining. Figure 22 and 23 contain the refine daily profiles. The load profiles were generated using the potential appropriate number of clusters discussed earlier (6and 14 clusters). Figure 22. Case 1: 6Refined daily load profiles.
55 Figure 23. Case 1: 14 Refined daily load profiles.
56 4.9 Method comparison The two methods were compared in order to determine if the optimization process will, in fact, improve the clustering result. Since the clusters produced via the different methods are similar visually, a way is needed to compare the performance of the methods. The validity index values (Silhouette and Davies-Bouldin) for the two methods were computed and compared. Each of the cases of the annual load profiles were compared to their corresponding refined clusters. Table 13, 14, 15 and 16 give the validity index of the two methods. Table 13. Case 1 with 4 annual load profiles: Results of the two validity indices. Bold numbers show the best index value. Methods Silhouette Davies–Bouldin Kmeans with GA 0.1136 3.0778 Kmeans 0.1136 3.0778 Table 14. Case 2 with 8 annual load profiles: Results of the two validity indices. Bold numbers show the best index value. Methods Silhouette Davies–Bouldin Kmeans with GA 0.0741 3.2759 Kmeans 0.0741 3.2816 Table 15. Case 3 with 12 annual load profiles: Results of the two validity indices. Bold numbers show the best index value. Methods Silhouette Davies–Bouldin Kmeans with GA 0.0590 3.3956 Kmeans 0.0580 3.4046 Table 16. Case 4 with 15 annual load profiles: Results of the two validity indices. Bold numbers show the best number index value. Methods Silhouette Davies–Bouldin Kmeans with GA 0.0568 3.2626 Kmeans 0.0508 3.6130
57 From Table 13, the validity indices for the two methods are the same. This means refining the clusters did not improve or decrease the quality of the clustering results in that case. As the number of clusters increased it is noticed that refining the clusters slightly improved the quality of the clustering result. This can be observed in Table 14, 15 and 16. The two validity indices test did indeed support our assumption that the optimisation of the initial centroids values using GA will improve the quality of the load clustering results.