scieee AI-readable full text Open interactive document viewer

Automatic detection of visible events in fusion reactors with Deep Learning

Serrano Lozano, David

Abstract

The next decades are crucially important to putting the world on a path of reduced greenhouse gas emissions since energy demand is increasing more and more. That is why fusion power, one of the most environmentally friendly sources of energy, has been getting so much attention lately. Wendelstein 7-X is a stellarator, a experimental fusion reactor build by the \ac{IPP} which intends to demonstrate the capabilities of fusion power to produce energy. The main problem with this process is that the working conditions to achieve fusion are extremely dangerous and unstable and sometimes the reactor walls overheat, damaging the structure of the device. For this reason, a continuous real-time data acquisition, analysis and control system is necessary to protect the structure. This thesis studies the procedure of generating a complete hot spot detector and classifier making use of the visible cameras installed inside the reactor. A Convolutional Neural Network is used to extract features from the bright events to later on use them to classify the incident with a Machine Learning classifier. The outcome has been very satisfactory as the system has been able to detect all the dangerous events in the data base. In addition, the model has been able to detect the incidents in real time with a delay of less than one and a half seconds from the first occurrence.

Full text

Automatic detection of visible events in fusion reactors with Deep Learning Degree Thesis submitted to the Faculty of the Escola T`ecnica d’Enginyeria de Telecomunicaci´o de Barcelona Universitat Polit`ecnica de Catalunya by David Serrano Lozano In partial fulfillment of the requirements for the bachelors’s degree in Telecommunications Technologies and Services ENGINEERING Advisor: Josep Ramon Morros Rubi´o Barcelona, Date June 2021 Contents List of Figures 4 List of Tables 4 1 Introduction 10 1.1 Requirements and specifications . . . . . . . . . . . . . . . . . . . . . . . . 11 1.2 WorkPlan.................................... 12 2 State of the art 13 2.1 Thermal events in W7-X . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 2.2 VisibleeventsinW7-X............................. 13 2.3 Feature extraction and classification . . . . . . . . . . . . . . . . . . . . . . 14 3 Database 16 4 Feature Extraction 18 4.1 TrackSubsampling ............................... 18 4.2 ImagePreprocessing .............................. 18 4.3 Convolutional Neural Network . . . . . . . . . . . . . . . . . . . . . . . . . 19 4.3.1 ResNet50 ................................ 20 4.3.2 Fine-tuning ............................... 21 4.3.3 Evaluation................................ 21 4.4 ImageAugmentation .............................. 22 5 Classification 24 5.1 Resamplingmethods .............................. 24 5.1.1 SMOTE................................. 24 5.1.2 Borderline-SMOTE........................... 25 5.1.3 ADASYN ................................ 25 5.1.4 SMOTEENN .............................. 25 5.1.5 SMOTETomek ............................. 25 5.2 Temporal classification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 5.2.1 Average of the probabilities . . . . . . . . . . . . . . . . . . . . . . 27 5.2.2 Average of the features . . . . . . . . . . . . . . . . . . . . . . . . . 27 5.2.3 Concatenation of the features . . . . . . . . . . . . . . . . . . . . . 27 5.3 Machine Learning classifiers . . . . . . . . . . . . . . . . . . . . . . . . . . 28 5.3.1 SVM................................... 28 5.3.2 XGBoost................................. 29 5.4 Stratified K-fold cross-validation . . . . . . . . . . . . . . . . . . . . . . . . 29 6 Metrics 31 7 Experiments and results 33 7.1 Featureextraction ............................... 33 7.1.1 Train/validation splits . . . . . . . . . . . . . . . . . . . . . . . . . 33 2 7.1.2 Image augmentation . . . . . . . . . . . . . . . . . . . . . . . . . . 33 7.1.3 Training................................. 33 7.2 Classifiers .................................... 34 7.2.1 NNasaclassifier............................ 34 7.2.2 Machine Learning classifiers . . . . . . . . . . . . . . . . . . . . . . 35 7.3 System-widetests................................ 36 7.4 Onlineclassification............................... 37 8 Conclusions 39 9 Future Work 40 References 41 10 Appendices 44 10.1Appendix1 ................................... 44 10.2Appendix2 ................................... 45 10.3Appendix3 ................................... 46 10.4Appendix4 ................................... 47 10.5Appendix5 ................................... 48 10.6Appendix6 ................................... 49 10.7Appendix7 ................................... 50 3 List of Figures 1 Orientation of the EDICAMs in the fusion reactor . . . . . . . . . . . . . . 11 2 Capture of one of the EDICAMs . . . . . . . . . . . . . . . . . . . . . . . . 11 3 Project’s Gantt diagram . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 4 A. Puig Sitjes et al. system overview . . . . . . . . . . . . . . . . . . . . . 14 5 NeuralNetworkdiagram............................ 15 6 NoHotSpotdetection ............................. 16 7 HotSpotdetection ............................... 16 8 Anomalydetection ............................... 16 9 Trackofdetections ............................... 17 10 Trackframe-lengths. .............................. 18 11 Croppeddetection ............................... 19 12 Neural Network diagram and backpropagation . . . . . . . . . . . . . . . . 20 13 Degradation in plain CNNs . . . . . . . . . . . . . . . . . . . . . . . . . . 21 14 Identity connections in ResNets . . . . . . . . . . . . . . . . . . . . . . . . 21 15 Fine-tuninginResNet50 ............................ 22 16 Imageaugmentation .............................. 23 17 Resamplingmethods .............................. 26 18 Systemgeneralview .............................. 27 19 SVM....................................... 28 20 Polynomialkernel................................ 28 21 XGBoost..................................... 30 22 XGBoost performance comparison. . . . . . . . . . . . . . . . . . . . . . . 30 23 Stratified K-fold cross-validation . . . . . . . . . . . . . . . . . . . . . . . . 30 24 SVM and XGBoost confusion matrices. . . . . . . . . . . . . . . . . . . . . 36 25 Hot spot from one of the system-wide test sequences . . . . . . . . . . . . 37 26 Confusion matrix of the system-wide test sequences. . . . . . . . . . . . . . 37 27 Confusion matrix of the online system. . . . . . . . . . . . . . . . . . . . . 38 28 Histograms of the online predictions of the HS . . . . . . . . . . . . . . . . 38 29 Structure of the ResNet50 . . . . . . . . . . . . . . . . . . . . . . . . . . . 45 List of Tables 1 Total number of detections . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 2 Percentages of the total of tracks of each class in relation to the total numberoftracks................................. 17 3 Best classification results . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36 4 Total number of detections separated by sequences . . . . . . . . . . . . . 44 5 Results of the NN used as a classifier . . . . . . . . . . . . . . . . . . . . . 46 6 Results of the SVM averaging the features . . . . . . . . . . . . . . . . . . 47 7 Results of the SVM concatenating the features . . . . . . . . . . . . . . . . 48 8 Results of the XGBoost averaging the features . . . . . . . . . . . . . . . . 49 9 Results of the XGBoost concatenating the features . . . . . . . . . . . . . 50 4 Acronyms ADASYN Adaptive Synthetic sampling AN Anomaly BBox Bounding Box CE Cross-Entropy CNN Convolutional Neural Network EDICAM Event Detection Intelligent Camera ENN Edited Nearest Neighbors FC Fully Connected FN False Negative FP False Positive HS Hot Spot IPP Max Planck Institute for Plasma Physics IR Infrared KNN K-Nearest Neighbors ML Machine Learning NHS No Hot Spot NN Neural Network OvA One-vs-All OvO One-vs-One PFC Plasma-Facing Component RD Random Displacement ROI Region Of Interest RR Random Rotation SMOTE Synthetic Minority Oversampling Technique SVM Support Vector Machine TN True Negative TP True Positive UFO Unidentified Flying Object W7-X Wendelstein 7-X 5 Abstract The next decades are crucially important to putting the world on a path of reduced greenhouse gas emissions since energy demand is increasing more and more. That is why fusion power, one of the most environmentally friendly sources of energy, has been getting so much attention lately. Wendelstein 7-X is a stellarator, a experimental fusion reactor build by the IPP which intends to demonstrate the capabilities of fusion power to produce energy. The main problem with this process is that the working conditions to achieve fusion are extremely dangerous and unstable and sometimes the reactor walls overheat, damaging the structure of the device. For this reason, a continuous real-time data acquisition, analysis and control system is necessary to protect the structure. This thesis studies the procedure of generating a complete hot spot detector and classifier making use of the visible cameras installed inside the reactor. A Convolutional Neural Network is used to extract features from the bright events to later on use them to classify the incident with a Machine Learning classifier. The outcome has been very satisfactory as the system has been able to detect all the dangerous events in the data base. In addition, the model has been able to detect the incidents in real time with a delay of less than one and a half seconds from the first occurrence. 6 Resum Les pr`oximes d`ecades tenen una import`ancia crucial per posar el m´on en un cam´ı de reducci´o d’emissions de gasos d’efecte hivernacle, ja que la demanda d’energia augmenta cada cop m´es. Per aix`o, darrerament la fusi´o nuclear, una de les fonts d’energia m´es respectuoses amb el medi ambient, s’est`a tenint tant en consideraci´o. Wendelstein 7-X ´es un stellarator, un reactor de fusi´o experimental constru¨ıt per l’IPP que t´e la intenci´o de demostrar les capacitats d’obtenci´o d’energia de la fusi´o. El principal problema d’aquest proc´es ´es que les condicions de treball per aconseguir la fusi´o s´on extremadament perilloses i inestables i, de vegades, les parets del reactor se sobreescalfen, danyant l’estructura del dispositiu. Per aquest motiu, ´es necessari un sistema d’adquisici´o, an`alisi i control de dades continu en temps real per protegir l’estructura. Aquesta tesi estudia el procediment per generar un detector i un classificador complet de punts calents fent ´us de les c`ameres visibles instal · lades a l’interior del reactor. Una xarxa neuronal convolucional s’utilitza per extreure caracter´ıstiques dels esdeveniments brillants i, posteriorment, utilitzar-les per classificar l’incident amb un classificador d’aprenentatge autom`atic. El resultat ha estat molt satisfactori, ja que el sistema ha estat capa¸c de detectar tots els esdeveniments perillosos de la base de dades. A m´es, el model ha estat capa¸c de detectar els incidents en temps real amb un retard inferior a un segon i mig des de la primera ocurr`encia. 7 Resumen Las pr´oximas d´ecadas son cruciales para encaminar al mundo hacia una reducci´on de las emisiones de gases de efecto invernadero, ya que la demanda de energ´ıa aumenta cada vez m´as. Por eso, la energ´ıa de fusi´on, una de las fuentes de energ´ıa m´as respetuosas con el medio ambiente, est´a recibiendo tanta atenci´on ´ultimamente. Wendelstein 7-X es un stellarator, un reactor de fusi´on experimental construido por el IPP que pretende demostrar las capacidades de la fusi´on nuclear para producir energ´ıa. El principal problema de este proceso es que las condiciones de trabajo para lograr la fusi´on son extremadamente peligrosas e inestables y, en ocasiones, las paredes del reactor se sobrecalientan, da˜nando la estructura del aparato. Por este motivo, es necesario un sistema de adquisici´on, an´alisis y control de datos en tiempo real y continuo para proteger la estructura. Esta tesis estudia el procedimiento para generar un detector y clasificador de puntos calientes haciendo uso de las c´amaras visibles instaladas en el interior del reactor. Se utiliza una Red Neural Convolucional para extraer caracter´ısticas de los eventos brillantes para posteriormente utilizarlas para clasificar el incidente con un clasificador de Machine Learning. El resultado ha sido muy satisfactorio ya que el sistema ha sido capaz de detectar todos los eventos peligrosos de la base de datos. Adem´as, el modelo ha sido capaz de detectar los incidentes en tiempo real con un retraso de menos de un segundo y medio desde la primera aparici´on. 8 Revision history and approval record Revision Date Purpose 0 28/04/2021 Document creation 1 04/06/2021 Document revision 2 15/06/2021 Document revision 3 19/06/2021 Document revision DOCUMENT DISTRIBUTION LIST Name e-mail David Serrano Lozano Josep Ramon Morros Rubi´o Written by: Reviewed and approved by: Date 28/04/2021 Date 19/06/2021 Name David Serrano Name Josep Ramon Morros Rubi´o Position Project Author Position Project Supervisor 9 3 Database This thesis uses the database created by M. Cobos for his Master Thesis Dissertation in which 14 sequences from the first W7-X operation (OP1.2 and OP1.2a) are labeled. The bright events which have linked the frame number and its BBox were manually classified between:  No Hot Spot (NHS). Small parts of the plasma running through the toroid, reflections from the plasma and other bright points which are not harmful to the reactor (see Fig. 6).  Hot Spot (HS). Parts of the wall that get hotter uncontrollably. This event has burned several sensors of the reactor and it is the main event to detect for safety reasons (see Fig. 7).  Anomaly (AN). This class groups all the events that are neither NHS nor HS which are not directly harmful, but are important to detect such as falling debris (small broken parts of the reactor), pellets (impurities injected inside the reactor to control the plasma composition) and other unidentified flying objects (UFOs) (see Fig. 8). Figure 6: No Hot Spot detection. Part of the plasma. Figure 7: Hot Spot detection. Overheated part of the wall. Figure 8: Anomaly detection. Pellet. The bright events were detected at frame level, but each of them were associated to a track. The frame level detections are all the candidates, while a track is composed by all the frame level detections which are caused by the same event along a number of frames. So, a track is a set of bright events in which the reason for the appearance is the same and the movement between successive frame detections is less than 10 pixels (using the Euclidean distance between the centroids of the BBoxes)(see Fig. 9) In table 1, the number of detections both at frame and track level can be seen. In this work, the focus is in detecting, classifying and predicting the tracks to exploit the temporal information that can exist. In addition, for future work, in Appendix 1 a more detailed view of the number of detections can be seen separated by files and its names. 16 Figure 9: Track of detections. The temporal evolution of two tracks marked with their respective BBox can be seen. Frame level Track level NHS HS AN NHS HS AN Detections 170430 14872 6416 873 41 37 Table 1: Total number of detections divided in its labels: NHS, HS and AN. At first sight, it can be seen some of the biggest challenges of this project. The database is quite small (951 tracks), it is a multiclass classification and two of the three classes are extremely imbalanced in relation to the majority class NHS HS AN Percentages (%) 91.80 4.31 3.89 Table 2: Percentages of class imbalance 17 4 Feature Extraction Feature extraction is a type of dimensionality reduction by which an intial set of raw data is reduced to more a manageable group. It is the name for methods that select and combine variables into features, effectively reducing the amount of data that must be processed, while still accurately and completely describing the original data set. As images have a large number of variables that require a lot of computing resources to process, in this section how and with which techniques the feature extraction has been done is explained. 4.1 Track Subsampling The main task of this project is to classify the tracks of the bright visual events found on the reactor and, as seen in the Database section, the tracks are sets of detections that have appeared for the same reason. So, the first step is to subsample the tracks since they can last up to 1200 frames, as can be seen in the histograms of the frame-lengths of the tracks separated by classes in Fig. 10. Moreover, the subsampling takes on more importance because the tracks have little movement. So, instead of taking all the detections of the tracks, only nof them equispaced are taken (in all the experiments done, n= 5). Figure 10: Histograms of the track’s frame-lengths separated by classes. 4.2 Image Preprocessing Before working with the images taken from the sequences some preprocessing has to be done. The features are not extracted from the entire frame, but from the BBox of each detection. This means that all the detections are cropped to a square shape with size of the maximum BBox length and resized to the desired size (224 x 224 pixels since ResNet50 is used)(see Fig. 11). 18 Figure 11: On the left, the original image of a sequence with a detection’s BBox in red. On the right, the cropped image of the shown detection. 4.3 Convolutional Neural Network As it was said in State of the art section, this thesis uses a CNN to obtain features from its logits (the raw outputs from a CNN layer). A Neural Network (NN) is a series of algorithms that endeavors to recognize underlying relationships and patterns in a set of data through a process that mimics the way the human brain operates. NNs are comprised of node layers, containing an input layer, one or more hidden layers, and an output layer in which its output is a series of logits that can be transformed into class probabilities. Each node, or neuron, connects to another and has associated a weight and a threshold value (see Fig. 12). This weights and values have to be adjusted by training on the basis of a set of training data in a way that solves a specific problem. This is done with backpropagation, the essence of NN training, that aims to minimize the cost function by adjusting the previous mentioned weights and values [15]. The backpropagation algorithm computes the gradient of the loss function for a single weight by the chain rule starting from the end and adjusts the weights such that the error is decreased (see Fig. 12). Over the last few decades, NNs have been considered to be one of the most powerful tools, and have become very popular in the literature as it is able to handle a huge amount of data. One of the most popular NNs is the CNN [17] taking its name from one type of layer which applies a mathematical linear operation between matrices called convolution. The particular reason this type of NN is used is as a consequence of performing extremely well at images and computer vision applications. However, as today’s CNNs have millions of parameters to train requiring a large amount of data and the database is quite small with a high class imbalance ratio, a CNN cannot be trained from scratch to achieve its full operational capacity. Therefore, Transfer Learning is used [18]. Transfer learning is a supervised learning technique that reuses parts of a previously trained model on a new network tasked for a different but similar problem. Using a pre-trained model significantly reduces the time required for feature engineering 19 Figure 12: Structure of a NN and how backpropagation algorithm works. (1) Input data xis (2) modeled using the current weights wto (3) get an output. (4) The error in the outputs is calculated with ErrorB=ActualOutput −DesiredOutput and (5) travels back to the first layer adjusting the weights such the error is decreased. Image extracted from [16] and training. The first step is to select a source model, ideally one that has been trained with a large dataset. The goal is to create a framework that is at least better than a naive model selecting only some layers to reuse in the custom model. Plain CNNs have convolutional layers followed by some fully connected (FC) layers. The convolutional layers are the mayor building blocks which their main focus is to extract and segment the characteristics of the input data, while the FC layers are those layers where all the inputs from one layer are connected to every activation unit of the next layer which are used for the classification task. It is expected that the deeper the CNNs are, the more accurate they will be. However, when the layers of these networks are increased, the problem of vanishing gradients occurs [19]. During backpropagation, each of the neural network’s weights receives an update proportional to the partial derivative of the error function with respect to the current weight in each iteration of training. The problem is that in some cases, the gradient is vanishingly small, effectively preventing the weight from changing its value, that is to say, when the network is deep, and multiplying a few of these small numbers the gradient becomes zero. In Fig. 13, extracted from the paper which introduced Residual Networks (ResNets) [20], the test errors of a 20-layer and a 56-layer plain networks can be seen. The degradation problem due to vanishing gradients is easy to see. To solve this problem, a variant of the ResNet family is used. 4.3.1 ResNet50 In this project the ResNet50 [20] is used to extract features from the BBoxes of the bright events found on the reactor sequences. ResNet became the winner of the ImageNet Large Scale Visual Recognition Challenge (ILSVRC) in 2015, as well as winner of MS COCO 2015 [21]. 20 ResNet50, trained on a subset of ImageNet, is a pre-trained Residual Network, a variant of the CNN model which introduces shortcut or identity connections between layers to be able to create deep models without the vanishing gradients problem [22]. In these connections, the output of the previous layer is added to the output of a more advanced layer. Consequently, even if there is vanishing gradients for the weights layers, the identity x is always transferred back to the previous ones (see Fig. 14). Figure 13: Test errors of two plain CNNs trained in the CIFAR-10 dataset. The deepest network performs worse than the 20-layer network due to the vanishing gradients. Image extracted from [20] Figure 14: Example of an identity connection skipping two layers. Image extracted from [20] ResNet50 is based on 5 main stages using a bottleneck design to reduce its time complexity since the network is very deep. This design adds a 1 x 1 convolutional layer at the start and at the end of each block. It turns out that doing this the number of connections and parameters can be reduced while not degrading the performance of the network so much. In Appendix 2 the whole structure of the ResNet50 can be seen. 4.3.2 Fine-tuning Fine-tuning is a widely used technique for model reuse which consists in unfreezing a few of the last layers of an already trained model, and jointly training both the newly added part of the model and the still frozen layers. Some weights remain the same and cannot be modified when training (freeze), while the others are edited according to the training data set. ResNet50 is trained in a subset of ImageNet and its final layer has an output layer of 1000 logits (number of classes of the data base). As the used data base has only 3 different classes and the main objective of the network is to extract features, the entire FC layer is created from scratch. Moreover, the images from the sequences differ a lot from the images used to train the ResNet50 since the streams from W7-X are in gray-scale and the ImageNet data base is composed of color pictures of everyday items. For that reason, not only the FC layers are fine-tuned, but also the last block of the model (see Fig. 15) 4.3.3 Evaluation The modified ResNet50 has to be trained with some epochs of the data base to be refined and fine-tuned. An epoch is a term used in ML that quantifies the number of passes of the entire training data set the algorithm has completed. The outputs of the model are unnormalised predictions which can give results, but interpreting their raw values is 21 Figure 15: Last layers of the modified ResNet50 not easy. So, to know when the feature extractor is fully trained, a softmax function is concatenated to the last FC layer to transfer the last logits to probabilities. The softmax function is a generalization of the logistic function to multiple dimensions which takes as input a vector zof K(Kis the number of classes, 3) real numbers, and normalizes it into a probability distribution consisting of Kprobabilities proportional to the exponentials of the input logits. The standard softmax function σ:RK→[0,1]Kis defined by the formula σ(z)i=ezi PK j=1 ezj for i= 1, ..., Kand z= (z1, ..., zK)∈RK(1) 4.4 Image Augmentation As it is said in the Database section, the data base is only made up of 951 tracks with a lot of class imbalance. As a consequence of this and the fact that deep networks need large amount of training data, image augmentation is used to boost its performance. Image augmentation is a technique that can be used to artificially expand the size of a training dataset by creating modified versions of the images [23]. Training deep networks on more data can result in more skillful model, and the augmentation techniques can create variations of the original frames that can improve the ability of generalizing what they have learned to new images and reduce the overfitting. Two image augmentation techniques and the combinations between them are used. A random displacement and a random rotation of the Bounding Box, both with an uniform distribution centered in 0 (see Fig. 16). 22 (a) Random rotation (b) Random displacement Figure 16: In (a) an example of a rotation of a BBox of 5 degrees. In (b) an example of a displacement of 10 pixels in the x axis and -10 pixels in the y axis. The white boxes are the ground truth detections and the red boxes are the same ones after applying image augmentation. 23 5 Classification Classification is the process of predicting the classes of a set of given data points. To obtain this data points or features, the last layer of the fully trained ResNet50 has to be removed. That is why the structure of the FC layers was modified. So, the output of the feature extractor is 128 features for each input image. As per each track ndetections are subsampled, the output at track level is nx128. In this section, how these features and with which techniques are analysed to classify the tracks is explained. 5.1 Resampling methods As it is explained in the Database section, the data base has an extremely class imbalance between NHS and both classes of abnormal events. Using the features directly from the NN there exists an algorithmic bias. Algorithmic bias describes systematic and repeatable errors that create unfair outcomes, such as privileging one arbitrary class over others which can emerge due to may factors, being one of them the way how the data is structured. Consequently, in this thesis, some resampling methods are used to level the number of features per class and later on classify the track in question. Resampling methods can generate different versions of the training set that can be used to simulate how well models would perform on new data. These techniques differ in terms of how the resampled versions of the data are created and how may iterations of the simulation process are conducted. These methods are divided between over-sampling and under-sampling. Over-sampling focuses in generating new synthetic samples of the minority classes, while under-sampling focuses in deleting samples from the majority class. However, as the dataset is small, under-sampling techniques are not used alone, but combined with other over-sampling techniques to improve the overall performance as demonstrated by A. Agrawal et al. creating and testing SCUT: SMOTE and cluster-based under-sampling [24]. In (Fig. 17), all the resampling methods used and explained in the following subsections can be seen using the first and second of all the features of the training dataset. 5.1.1 SMOTE SMOTE, described by N. Chawla et al. in their 2002 paper [25], is the most widely used approach to synthesizing new examples. This algorithm applies KNN approach where it selects Knearest neighbors, joins them and creates the synthetic samples in the space. Specifically, this method first selects a minority class instance aat random and finds its K nearest minority class neighbors. The synthetic instance is then created by choosing one the Knearest neighbors bat random and connecting aand bto form a line segment in the feature space. The synthetic instances are generated instances as a convex combination of the two chosen instances aand b[26]. SMOTE has a big issue when there are observations of the minority class which are outlying and appear in the zone where the majority class is prevailing since it creates samples formed in line bridges inside the majority class zone. 24 5.1.2 Borderline-SMOTE Bordeline-SMOTE is a variation of the SMOTE, presented by H. Han et al. in [27], which solves the issue stated above. This algorithm classifies any minority observation as a noise point if all the neighbors are majority class points and such an observation is ignored when creating data. Further, it classifies a few points as border points which have both majority and minority class instances as neighborhood and resample completely from these points. 5.1.3 ADASYN ADASYN, stated by H. He et al. in their 2008 paper [28], is based on the idea on adaptively generating minority data samples according to their distributions using KNN. The algorithm adaptively updates the distribution and there are no assumptions made for the underlying distribution of the data. The key difference between ADASYN and SMOTE is that the former uses a density distribution, as a criterion to automatically decide the number of synthetic samples to compensate for the skewed distributions. The latter generates the same number of synthetic samples for each original minority sample. 5.1.4 SMOTEENN ENN, presented by D. L. Wilson in [29], can be used as an under-sampling method or as a data cleaning and as said in the introduction of this section, this method is used combined with an over-sampling method, the SMOTE. ENN is a rule for finding ambiguous and noisy examples which by finding the K-nearest neighbor of each observation first, then checks wether the majority class from the observation’s K-nearest neighboor is the same as the instance’s class or not. If the majority class of the observation’s K-nearest neighbor and the observation’s class is different, then the observation and its K-nearest neighboor are deleted from the dataset. So, before applying the ENN, the SMOTE algorithm is implemented to the output features [30]. 5.1.5 SMOTETomek Tomek links can be used in the same applications as ENN, and it is used with SMOTE as well. Tomek links, introduced by I. Tomek in [31], are pairs of instances of opposite classes who are their own nearest neighbors. In other words, they are pairs of opposing instances that are very close together. Tomek’s algorithm looks for such pairs and removes by choosing between only the majority instance of the pair or both of them. The idea is to clarify the border between the minority and majority classes, making the minority regions more distinct. So, before applying the Tomek links, the SMOTE algorithm implemented to the output features [30]. 5.2 Temporal classification The final output of the system has to be a class prediction of the entire track. To do this, three different ways of merging the probabilities or features has been done. The first one uses the NN and the probabilities from the softmax function, while the other two 25 actual class while each column represents the instances in a predicted class. It gives insights not only into the errors being made by the classifier bur more importantly the types of errors that are being made. So, after having stated and explained the metrics used, the expected results and the conditions that were optimized are declared. As the database is made of 3 different classes, the first three metrics explained, can be calculated for each label. The most important visual event to detect is the hot spot and it is extremely vital to detect all of them. Then, the recall of the HS class has to be maximized in order to classify all the occurring hot spots as HS. As stopping the reactor is a very cost-consuming task, it is important to avoid detecting HS that are not from that class. Hence, the overall performance of the model should be consistent, obtaining the F1 Score of the HS class as high as possible. In previous work made by M. Cobos only the F1 Score-weighted is mentioned, which is not significant due to the existing class imbalance on the validation set. The F1 Score-weighted (F1 W, in the tables) is calculated averaging all the F1 Score of all the classes according to the number of true instances for each label. In other words, the more instances the class has, the more importance the class takes. That means that the NHS label which is not the main event to detect takes approximately the 90% of the relevance. However, to compare between techniques, the F1 Score-weighted is shown too. Furthermore, to have a better overview of the performance of the models, the F1 Score-macro is used. F1 Score-macro (F1 M, in the tables) is the mean average between ever class regardless of the number of instances each class has. In other words, despite the fact that the NHS class has more instances than the other two, all three take on equal importance. 32 7 Experiments and results During this document two divided blocks have been clearly differentiated: the development of a NN to extract features of the abnormal bright events and a classifier which is able to predict the label of the detection using the temporal information. In this section, the technical aspects and parameters of the experiments done are shown as well as the results obtained. To create the software, python, pytorch and scikit-learn among other libraries have been used. 7.1 Feature extraction 7.1.1 Train/validation splits To fully train the feature extractor, the entire data set has to be split in a training set and in a validation set. Usually, to train NNs one more set called test is used, but as the minority classes have so few instances, this set is removed to have more labels in the validation set. The models take the training set to learn and refine their parameters from its data, the test set is used once per epoch to know how well the model performs and the validation set helps to know how well the model works once the whole training is done. To train the feature extractor the training set has been created with the 80% of the instances and the validation set with the rest, 20%, maintaining the class proportion. Furthermore, images from the same track must not be placed in both database divisions due to the model would be training with the same examples as the validation set and could not be evaluated with unseen bright events. 7.1.2 Image augmentation All the models have been tested with and without image augmentation. When using image augmentation the following combinations of the two methods explained in the Image Augmentation subsection have been used:  RD5. Random Displacement (RD) between -5 and 5 pixels for each side.  RR5. Random Rotation (RR) between -5 and 5 degrees from the center of the BBox  RD5 RR5. RD between -5 and 5 pixels for each side and RR between -5 and 5 degrees from the center of the BBox.  RD10 RR10. RD between -10 and 10 pixels for each side and RR between -10 and 10 degrees from the center of the BBox.  RD% RR5. RD of the 10% of the side pixel length BBox and RR between -5 and 5 degrees from the center of the BBox. 7.1.3 Training To train the feature extractor, a custom DataLoader function has been created. The DataLoader is a function which loads only a few samples of the data set instead of loading 33 all the data to have enough memory space to process it. Even the state-of-the-art configurations cannot carry all the data at once since the images of the sequences are in 16 bits per pixel (in reality the sequences have a dynamic range of 12 bits but as there is not a such variable type, the data was converted to 16 bit in order not to lose quality). For that reason, the DataLoader only loads a batch of the data set. The batch size determines how many training or validation examples are processed in parallel for training or inference. The batch sizes taken for all the models after some experiments were, for training, 7 tracks, and for validation, as many as the memory could handle, 12 tracks. The batch size in training time can affect how fast and how well the training converges. Thus, for the training batch size, it is worth picking a batch size that is neither too small nor too large. It has been observed in practice that when using a larger batch there is a significant degradation in the quality of the model, as measured by its ability to generalize. The lack of generalization ability is due to the fact that large-batch methods tend to converge to sharp minimizers of the training function. These minimizers which are the functions used to refine the weights of the NN, are characterized by large positive eigenvalues in O2f(x) and tend to generalize less well when the batch is too large. On the other hand, when the batch size is to small, the training time is bigger [42]. To refine the model a loss function has to be optimized. In all the models the CrossEntropy (CE) Loss has been used. CE builds upon the idea of information theory entropy and measures the differences between two probability distributions for a given random set of events [43]. It is used after applying the softmax function of the last layer of the NN and it is defined as: CE =− C X i=1 tilog(si) (5) where tians siare the ground truth and the output of each class irespectively. 7.2 Classifiers After passing the images through the NN, using the features extracted, the track has to be classified between the three possible classes. As explained in the classification section, three classifiers have been used, one of them transforming the output logits to class probabilities and the other two merging the features in different ways. In this section, the obtained results using the NN as a classifier using the softamax function and both ML classifiers are shown. 7.2.1 NN as a classifier This technique concatenates a softmax function to the NN and transforms the logits to probabilites. Then the nclass probabilities are averaged and the class with highest probability is selected as the track prediction. So, using the ResNet50 as a feature extractor and a classifier at the same time the best HS recall is 0.67, obtained with almost every image augmentation technique. The best HS F1 Score is 0.75 using RD5 since this technique improves the HS precision by far. 34 In Appendix 3 all the results obtained using all the image augmentation techniques can be seen. 7.2.2 Machine Learning classifiers The other two classifiers used are the SVM and XGBoost. Both of them learn from instances, so, as explained in the classification section, the features are extracted from the next-to-last layer of the NN. Although, for each image 128 features are obtained, the class imbalance still remains. That is why one of the resampling techniques explained above is used on the training set before training the ML classifiers. To train them the stratified 5-fold cross-validation has been used. The best performance using the SVM is obtained averaging the features since this classifier is very liable to missclassify if there exist noise in the instances because it is harder to find the hyperplane with the biggest margin. Moreover, the best results are also obtained using the same image augmentation technique as the best result using the NN as a classifier (RD5). Borderline-SMOTE is the resampling method that obtains the best results since is the one which focus on the points of the border of the instances of different classes and doing this helps SVM to find the best hyperplane. The best HS recall and F1 Score are 0.85 and 0.92 respectively, which indicates that the performance has been improved by a lot in respect of using the NN as a classifier. In Fig. 24 (a), the confusion matrix of the experiment with the best results can be seen. However, the best performance using XGBoost are obtained concatenating the features, just the opposite of SVM. That is because the more features XGBoost has, the better classifies since more and deeper trees can be done. In addition, the best HS F1 Score has been obtained using the same image augmentation technique as the best result obtained on the other techniques used (RD5). That means that this exactly image augmentation combination is the best for this specific data set since it is the one which obtains the best results on all three classifiers. Moreover, XGBoost performs better when no resampling methods are used. This is because this classifier has some parameters to deal with class imbalance. The best HS recall and F1 Score are of 1 for both of the metrics, meaning that the 13 HS of the validation set were predicted correctly and no other events were missclassified as HS either. In Fig. 24 (b), the confusion matrix of the experiment with the best results can be seen. The results obtained using XGBoost were the best that could be obtained with the given database and initial conditions. 35 (a) SVM confusion matrix (b) XGBoost confusion matrix Figure 24: SVM and XGBoost confusion matrices of the experiments with best results. In each square the percentage of ground truth instances predicted correctly of each class is shown as well as the number of instances in brackets. The best results obtained are in the following table. Nevertheless, in Appendix 4, 5, 6 and 7 all the results in full detail can be seen. Precision Recall F1 Score NHS HS AN NHS HS AN NHS HS AN M W SVM 0.98 1.00 0.79 0.99 0.62 0.92 0.98 0.76 0.85 0.86 0.97 XGBoost 0.98 1.00 0.88 1.00 1.00 0.58 0.99 1.00 0.70 0.90 0.98 Table 3: Best classification results. 7.3 System-wide tests At the time when both the feature extractor and the classifier were fully trained and tested, some tests of the two blocks concatenated were made. 4 unseen sequences and without labels were taken to test how the system works all together. The new sequences were passed through the detector of bright events, the tracker, the feature extractor and the classifier. Then, with the help of A. Puig Sitjes, a computer vision and machine learning engineer of the IPP, the predictions were reviewed to know the system performance. In one of these 4 sequences one of the hot spots was from the Neutral Beam Injection System as can be seen in Fig. 25 The HS recall is 1, all the hot spots that occur were correctly classified, but the HS F1 Score is 0.82 due to some reflections were missclassified as HS. In Fig. 26 the confusion matrix of the results of all the sequences can be seen. So, taking the results into account the system is a good tool to help label the sequences, one of the objectives of this thesis. Instead of having to look carefully at the whole sequence 36 Figure 25: Hot spot from one of the system-wide test sequences because of the Neutral Beam Injection. Figure 26: Confusion matrix of the system-wide test sequences. In each square the percentage of ground truth instances predicted correctly of each class is shown as well as the number of instances in brackets. to detect bright events and decide whether they are HS or not, the system can be used and then review the results to go faster and have the same outcome. 7.4 Online classification So far all the tests and experiments have been done in a forensic or offline way, which means that the entire sequence or track has to happen to be able to predict its class. This is good for analysing the sequences and knowing where the most dangerous zones are and helping to label new sequences for future use. However, being able to know when the reactor walls are overheating in real time could save a lot of resources by shutting down the reactor before reaching excessively high temperatures. That is why the system has been modified to be able to predict in real time the bright events occurring in the reactor. In this section how it is done and the experiments and results obtained are explained. In the system explained so far the tracks were subsampled to nimages (n=5), and as the models have been trained this way, the online classifier uses the same number of images per track. But rather than waiting to the end of the track before predicting its class, the system now starts analysing the tracks from the nth image onwards. The online system makes a new prediction for each new detection corresponding to the track in question. In other words, this system makes as many predictions as there are frames in the track minus n. The online model takes 5 images of the track equispaced from all the images detected at that time. If a bright event has been detected on 5 consecutive frames, the system takes all the detections, but if the track is extended up to 40 frames the system would take the frames 1, 10, 20, 30, 40. That is to say that, in each frame, all the existing tracks are analysed as if it was the last shot of the sequence. The online system has been tested with the 261 NHS tracks and 13 HS tracks of the validation set and all the bright events found on a sequence newly labeled obtaining a 37 data set composed of 320 NHS and 14 HS with different track lengths. It is important to correctly predict, but also to do it as early as possible. As the NHS class is the majority one and when the bright events start appearing are not very intense, the classifier always starts predicting the classes as NHS. 5 out of the 320 NHS were in some point of the track missclassified as HS and 2 out of the 320 NHS were poorly predicted as AN, as can be seen in Fig. 27. All the HS were correctly classified in some point of the tracks. Moreover, the 14 HS were predicted with an average time of 1.38 seconds and with an average of 62.03% of the track Fig. 28. Figure 27: Confusion matrix of the online system. In each square the percentage of ground truth instances predicted correctly of each class is shown as well as the number of instances in brackets. Figure 28: Histograms of the online predictions of the HS. In (a) the time is evaluated in seconds, while in (b) the time is evaluated by the percentage of track has already passed. The IPP researchers expect to detect the HS in half a second (50 frames as the EDICAM works at 100 frames per second). So, the online system do not work fast enough as only the 7% of the tested HS tracks were predicted under that time. 38 8 Conclusions This thesis has reported several studies about different strategies to classify bright events occurring inside the reactor. The lack of researching about thermal events detection with visible range cameras or EDICAMs has lead to experiment with some of the best and most recognised computer vision and Deep Learning techniques. Since the beginning of the project a lot of knowledge in state-of-the-art computer vision algorithms has been gathered to try to achieve an steady-state operation in W7-X using the EDICAM-based system. Firstly, although it is normal for there to be many more reflections and not harmful bright events than hot spots, the data base is rather small. So this has led to the need to study some techniques to improve this problem. Therefore, taking into account the initial conditions on which this thesis is based on, the obtained results and the project itself have been extremely successful. The ResNet50 has proven how good it is at dealing with images even though the images from W7-X are very different from the ones the pre-trained model has been trained on. Both ML classifiers have obtained extrmely good results and it has been demonstrated how well the classifiers perform concatenating them after a NN. In terms of results, the outcome has been excellent. The XGBoost has performed as well as possible obtaining a HS F1 Score of 1, proving why it is one of the most widely used. Despite the fact that the results are extremely good, it is necessary to take into account the small dimensions of the data base and therefore of the validation set. If the data base was better and bigger the results would be even stronger since the models could be trained with more sequences and examples. The results obtained in this thesis are much better than the ones obtained in the first detector by M. Cobos as can be seen comparing the F1-Score-weighted. The forensic system works extremely well, however, the online system is quite slow and could not be used to detect the hot spots under the required time. That is why, the detection speed would need to be increased a bit. All things considered, the experiments have shown that the feature extractor as well as the classifiers were on the right way and hopefully it helps to define the way to follow in the forthcoming investigations. 39 9 Future Work In this thesis a feature extractor using a NN and a classifier has been done to detect and classify bright events in W7-X. Following this point, there is still work to do, such as implement other NNs and other classifiers. Other Deep Learning techniques could be used such as object detection, object recognition techniques and recurrent networks (Faster RCNN, YOLO, DetectoRS...). This models could even be used alongside the system done in this thesis. A good starting point for future research would be enlarging and improving the actual data base to obtain better results focusing only on the hot spots instead of using the AN class since as it contains many different types of events it is very difficult to predict. In addition, it would be interesting to change the approach of this thesis and focus more on the online classification instead the forensic one. This is not an easy problem, because with real-time information it is more complicated to predict the events. 40 References [1] Guru. Wendelstein 7-x. https://www.ipp.mpg.de/w7x. [2] S. Zoletnik et al. EDICAM (Event Detection Intelligent Camera). Fusion Engineering and Design, 88(6):1405–1408, 2013. [3] G. Kocsis et al. Overview video diagnostics for the w7-x stellarator. Fusion Engineering and Design, 96-97:808–811, 2015. [4] S. Zoletnik et al. First results of the multi-purpose real-time processing video camera system on the Wendelstein 7-X stellarator and implications for future devices. Review of Scientific Instruments, 89, 2018. [5] M. Cobos. Anomalies detection in the visible spectrum of plasma physics at weldenstein 7-x. Master in Computer Vision, 2020. [6] Ascas´ıbar E., editor. Wendelstein 7-X in the European Roadmap to FusionElectricity, Nara, Japan, 5 2015. EUROfusion. [7] A. Ali et al. Initial results from the hotspot detection scheme for protection of plasma facing components in wendelstein 7-x. Nuclear Materials and Energy, 19:335–339, 2019. [8] A. Puig Sitjes et al. Observation of thermal events on the plasma facing components of wendelstein 7-x. Journal of Instrumentation, 14:C11002–C11002, 2019. [9] T. Szepesi et al. Combining research with safety: Performance of the Wendelstein 7-X video diagnostic system. Fusion Engineering and Design, 146:874–877, 2019. [10] A. Puig Sitjes et al. Wendelstein 7-x near real-time image diagnostic system for plasma-facing components protection. Fusion Science and Technology, 74(1-2):116– 124, 2018. [11] N. Otsu. A Threshold Selection Method from Gray-Level Histograms. IEEE Transactions on Systems, Man, and Cybernetics, 9:62–66, 1979. [12] G. Kumar and B. Pradeep Kumar. A Detailed Review of Feature Extraction in Image Processing Systems). In 2014 Fourth International Conference on Advanced Computing Communication Technologies, pages 5–12, 2014. [13] N. Tajbakhsh et al. Convolutional Neural Networks for Medical Image Analysis: Full Training or Fine Tuning? IEEE Transactions on Medical Imaging, 35(5):1299–1312, 2016. [14] S.Notley and M. Magdon-Ismail. Examining the Use of Neural Networks for Feature Extraction: A Comparative Analysis using Deep Learning, Support Vector Machines, and K-Nearest Neighbor Classifiers, 2018. [15] L. Shen et al. Relay Backpropagation for Effective Learning of Deep Convolutional Neural Networks. Computer Vision – ECCV 2016, 9911, 2016. 41 10.5 Appendix 5 SVM concatenating the features: Precision Recall F1 Img. Aug. Resampling T. NHS HS AN NHS HS AN NHS HS AN M W SMOTE 0,93 1,00 1,00 1,00 0,23 0,25 0,96 0,38 0,40 0,58 0,91 B. SMOTE 0,98 1,00 1,00 1,00 0,77 0,75 0,99 0,87 0,86 0,91 0,98 ADASYN 0,97 1,00 1,00 1,00 0,69 0,67 0,98 0,82 0,80 0,87 0,97 SMOTEENN 0,94 1,00 0,75 1,00 0,31 0,25 0,96 0,47 0,38 0,60 0,92 - SMOTETomek 0,93 1,00 1,00 1,00 0,23 0,25 0,96 0,38 0,40 0,58 0,91 SMOTE 0,92 0,00 1,00 1,00 0,00 0,25 0,96 0,00 0,40 0,45 0,89 B. SMOTE 0,97 1,00 1,00 1,00 0,69 0,75 0,99 0,82 0,86 0,89 0,97 ADASYN 0,98 1,00 1,00 1,00 0,69 0,83 0,99 0,82 0,91 0,91 0,98 SMOTEENN 0,93 0,00 0,83 1,00 0,00 0,42 0,96 0,00 0,56 0,51 0,90 RD5 SMOTETomek 0,92 0,00 1,00 1,00 0,00 0,25 0,96 0,00 0,40 0,45 0,89 SMOTE 0,92 1,00 0,50 1,00 0,15 0,08 0,96 0,27 0,14 0,46 0,89 B. SMOTE 0,92 1,00 1,00 1,00 0,15 0,08 0,96 0,27 0,15 0,46 0,89 ADASYN 0,96 1,00 1,00 1,00 0,62 0,42 0,98 0,76 0,59 0,78 0,95 SMOTEENN 0,94 1,00 0,50 0,99 0,38 0,17 0,96 0,56 0,25 0,59 0,91 RR5 SMOTETomek 0,92 1,00 0,50 1,00 0,15 0,08 0,96 0,27 0,14 0,46 0,89 SMOTE 0,93 1,00 1,00 1,00 0,23 0,08 0,96 0,38 0,15 0,50 0,90 B. SMOTE 0,92 1,00 1,00 1,00 0,08 0,08 0,96 0,14 0,15 0,42 0,89 ADASYN 0,95 1,00 1,00 1,00 0,69 0,25 0,98 0,82 0,40 0,73 0,94 SMOTEENN 0,93 1,00 0,50 0,99 0,23 0,17 0,96 0,38 0,25 0,53 0,90 RD5 RR5 SMOTETomek 0,93 1,00 1,00 1,00 0,23 0,08 0,96 0,38 0,15 0,50 0,90 SMOTE 0,94 1,00 1,00 1,00 0,15 0,42 0,97 0,27 0,59 0,61 0,92 B. SMOTE 0,99 1,00 0,86 1,00 0,69 1,00 0,99 0,82 0,92 0,91 0,98 ADASYN 0,97 1,00 1,00 1,00 0,62 0,83 0,99 0,76 0,91 0,89 0,97 SMOTEENN 0,94 1,00 1,00 1,00 0,15 0,50 0,97 0,27 0,67 0,63 0,92 RD10 RR10 SMOTETomek 0,94 1,00 1,00 1,00 0,15 0,42 0,97 0,27 0,59 0,61 0,92 SMOTE 0,92 1,00 1,00 1,00 0,15 0,08 0,96 0,27 0,15 0,46 0,89 B. SMOTE 0,97 1,00 1,00 1,00 0,69 0,75 0,99 0,82 0,86 0,89 0,97 ADASYN 0,97 0,90 1,00 1,00 0,69 0,67 0,98 0,78 0,80 0,86 0,97 SMOTEENN 0,93 1,00 1,00 1,00 0,23 0,17 0,96 0,38 0,29 0,54 0,91 RR% RD5 SMOTETomek 0,92 1,00 1,00 1,00 0,15 0,08 0,96 0,27 0,15 0,46 0,89 Table 7: Results of the SVM classifier concatenating the features. The numbers in bold are from the models which performed better in the important metrics to optimise. As all the results were better averaging the features there are not results in bold. 48 10.6 Appendix 6 XGBoost averaging the features: Precision Recall F1 Img. Aug. Resampling T. NHS HS AN NHS HS AN NHS HS AN M W - 0,98 0,86 0,78 0,98 0,92 0,58 0,98 0,89 0,67 0,85 0,96 SMOTE 0,99 0,81 0,53 0,96 1,00 0,75 0,97 0,90 0,62 0,83 0,95 B. SMOTE 0,98 0,87 0,73 0,98 1,00 0,67 0,98 0,93 0,70 0,87 0,97 ADASYN 0,99 0,81 0,42 0,93 1,00 0,83 0,96 0,90 0,56 0,80 0,94 SMOTEENN 0,99 0,76 0,38 0,92 1,00 0,83 0,96 0,87 0,53 0,78 0,93 - SMOTETomek 0,99 0,81 0,53 0,96 1,00 0,75 0,97 0,90 0,62 0,83 0,95 - 0,98 0,93 0,67 0,98 1,00 0,50 0,98 0,96 0,57 0,84 0,96 SMOTE 0,98 0,65 0,47 0,95 1,00 0,58 0,96 0,79 0,52 0,76 0,94 B. SMOTE 0,98 0,81 0,50 0,97 1,00 0,58 0,97 0,90 0,54 0,80 0,95 ADASYN 0,99 0,68 0,45 0,93 1,00 0,75 0,96 0,81 0,56 0,78 0,94 SMOTEENN 0,99 0,57 0,42 0,92 1,00 0,67 0,95 0,72 0,52 0,73 0,93 RD5 SMOTETomek 0,98 0,65 0,47 0,95 1,00 0,58 0,96 0,79 0,52 0,76 0,94 - 0,97 0,86 0,57 0,98 0,92 0,33 0,97 0,89 0,42 0,76 0,95 SMOTE 0,98 0,76 0,47 0,95 1,00 0,58 0,97 0,87 0,52 0,78 0,94 B. SMOTE 0,98 0,87 0,46 0,97 1,00 0,50 0,97 0,93 0,48 0,79 0,95 ADASYN 0,98 0,72 0,44 0,94 1,00 0,67 0,96 0,84 0,53 0,78 0,94 SMOTEENN 0,99 0,72 0,31 0,90 1,00 0,75 0,94 0,84 0,44 0,74 0,92 RR5 SMOTETomek 0,98 0,76 0,47 0,95 1,00 0,58 0,97 0,87 0,52 0,78 0,94 - 0,96 0,86 0,60 0,98 0,92 0,25 0,97 0,89 0,35 0,74 0,94 SMOTE 0,98 0,68 0,38 0,93 1,00 0,67 0,95 0,81 0,48 0,75 0,93 B. SMOTE 0,98 0,67 0,50 0,95 0,92 0,58 0,96 0,77 0,54 0,76 0,94 ADASYN 0,98 0,62 0,33 0,91 1,00 0,67 0,94 0,76 0,44 0,72 0,92 SMOTEENN 0,99 0,62 0,28 0,89 1,00 0,67 0,94 0,76 0,39 0,70 0,91 RD5 RD5 SMOTETomek 0,98 0,68 0,38 0,93 1,00 0,67 0,95 0,81 0,48 0,75 0,93 - 0,97 0,93 0,71 0,99 1,00 0,42 0,98 0,96 0,53 0,82 0,96 SMOTE 0,98 0,87 0,44 0,95 1,00 0,67 0,97 0,93 0,53 0,81 0,95 B. SMOTE 0,98 0,86 0,47 0,96 0,92 0,67 0,97 0,89 0,55 0,80 0,95 ADASYN 0,98 0,72 0,39 0,94 1,00 0,58 0,96 0,84 0,47 0,75 0,93 SMOTEENN 0,93 0,76 0,35 0,95 1,00 0,50 0,96 0,87 0,41 0,75 0,94 RD10 RR10 SMOTETomek 0,98 0,87 0,44 0,95 1,00 0,67 0,97 0,93 0,53 0,81 0,95 - 0,98 0,86 0,78 0,99 0,92 0,58 0,98 0,89 0,67 0,85 0,97 SMOTE 0,99 0,76 0,47 0,95 1,00 0,67 0,97 0,87 0,55 0,80 0,95 B. SMOTE 0,98 0,86 0,78 0,99 0,92 0,58 0,98 0,89 0,67 0,85 0,97 ADASYN 0,99 0,68 0,45 0,94 1,00 0,75 0,96 0,81 0,56 0,78 0,94 SMOTEENN 1,00 0,76 0,40 0,93 1,00 0,83 0,96 0,87 0,54 0,79 0,94 RD% RR5 SMOTETomek 0,99 0,76 0,47 0,95 1,00 0,67 0,97 0,87 0,55 0,80 0,95 Table 8: Results of the XGBoost classifier averaging the features. The numbers in bold are from the models which performed better in the important metrics to optimise. 49 10.7 Appendix 7 XGBoost concatenating the features: Precision Recall F1 Img. Aug. Resampling T. NHS HS AN NHS HS AN NHS HS AN M W - 0,99 0,86 0,88 0,99 0,92 0,58 0,98 0,89 0,70 0,86 0,97 SMOTE 0,99 0,87 0,71 0,98 1,00 0,83 0,98 0,93 0,77 0,89 0,97 B. SMOTE 0,98 0,86 0,89 0,99 0,92 0,67 0,98 0,89 0,76 0,88 0,97 ADASYN 0,98 0,76 0,57 0,96 1,00 0,67 0,97 0,87 0,62 0,82 0,95 SMOTEENN 0,99 0,72 0,53 0,95 1,00 0,83 0,97 0,84 0,65 0,82 0,95 - SMOTETomek 0,99 0,87 0,71 0,98 1,00 0,83 0,98 0,93 0,77 0,89 0,97 - 0,98 1,00 0,88 1,00 1,00 0,58 0,99 1,00 0,70 0,90 0,98 SMOTE 1,00 0,81 0,55 0,95 1,00 0,92 0,97 0,90 0,69 0,85 0,96 B. SMOTE 0,99 0,87 0,80 0,99 1,00 0,67 0,99 0,93 0,73 0,88 0,97 ADASYN 0,98 0,67 0,64 0,96 0,92 0,75 0,97 0,77 0,69 0,81 0,95 SMOTEENN 0,99 0,68 0,50 0,94 1,00 0,83 0,96 0,81 0,62 0,80 0,94 RD5 SMOTETomek 1,00 0,81 0,55 0,95 1,00 0,92 0,97 0,90 0,69 0,85 0,96 - 0,97 0,79 0,71 0,98 0,85 0,42 0,97 0,81 0,53 0,77 0,95 SMOTE 0,99 0,76 0,60 0,96 1,00 0,75 0,97 0,87 0,67 0,84 0,96 B. SMOTE 0,98 0,81 0,64 0,97 1,00 0,58 0,98 0,90 0,61 0,83 0,96 ADASYN 0,98 0,76 0,62 0,97 1,00 0,67 0,97 0,87 0,64 0,83 0,96 SMOTEENN 0,99 0,68 0,40 0,93 1,00 0,67 0,96 0,81 0,50 0,76 0,93 RR5 SMOTETomek 0,99 0,76 0,60 0,96 1,00 0,75 0,97 0,87 0,67 0,84 0,96 - 0,97 0,86 0,67 0,98 0,92 0,33 0,98 0,89 0,44 0,77 0,95 SMOTE 0,98 0,65 0,50 0,95 1,00 0,58 0,97 0,79 0,54 0,76 0,94 B. SMOTE 0,98 0,87 0,64 0,98 1,00 0,58 0,98 0,93 0,61 0,84 0,96 ADASYN 0,99 0,62 0,50 0,94 1,00 0,67 0,96 0,76 0,57 0,77 0,94 SMOTEENN 0,99 0,62 0,38 0,92 1,00 0,67 0,95 0,76 0,48 0,73 0,93 RD5 RD5 SMOTETomek 0,98 0,65 0,5 0,95 1,00 0,58 0,97 0,79 0,54 0,76 0,94 - 0,98 0,92 1,00 1,00 0,92 0,58 0,99 0,92 0,74 0,88 0,97 SMOTE 0,99 0,72 0,69 0,97 1,00 0,75 0,98 0,84 0,72 0,85 0,96 B. SMOTE 0,98 0,86 0,75 0,98 0,92 0,75 0,98 0,89 0,75 0,87 0,97 ADASYN 0,99 0,72 0,62 0,97 1,00 0,67 0,98 0,84 0,64 0,82 0,96 SMOTEENN 0,99 0,68 0,53 0,94 1,00 0,83 0,97 0,81 0,65 0,81 0,95 RD10 RR10 SMOTETomek 0,99 0,72 0,69 0,97 1,00 0,75 0,98 0,84 0,72 0,85 0,96 - 0,98 1,00 0,90 1,00 0,92 0,75 0,99 0,96 0,82 0,92 0,98 SMOTE 1,00 0,81 0,62 0,97 1,00 0,83 0,98 0,90 0,71 0,86 0,97 B. SMOTE 0,98 1,00 0,82 0,99 0,92 0,75 0,99 0,96 0,78 0,91 0,98 ADASYN 0,99 0,76 0,69 0,97 1,00 0,75 0,98 0,87 0,72 0,86 0,97 SMOTEENN 0,99 0,76 0,56 0,96 1,00 0,75 0,98 0,87 0,64 0,83 0,96 RD% RR5 SMOTETomek 1,00 0,81 0,62 0,97 1,00 0,83 0,98 0,90 0,71 0,86 0,97 Table 9: Results of the XGBoost classifier concatenating features. The numbers in bold are from the models which performed better in the important metrics to optimise. 50