scieee AI-readable full text Open interactive document viewer

Supervised contrastive learning over prototype-label embeddings for network intrusion detection

López Martín, Manuel,Sánchez Esguevillas, Antonio Javier,Arribas Sánchez, Juan Ignacio,Carro Martínez, Belén

Abstract

Producción Científica

Full text

Information Fusion 79 (2022) 200–228 Available online 20 September 2021 1566-2535/© 2021 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). Supervised contrastive learning over prototype-label embeddings for network intrusion detection Manuel Lopez-Martin a , * , Antonio Sanchez-Esguevillas a , Juan Ignacio Arribas a , b , Belen Carro a a Dpto. TSyCeIT, ETSIT, University of Valladolid, Paseo de Bel´ en 15, Valladolid 47011, Spain b Castilla-Leon Neuroscience Institute, University of Salamanca, Salamanca 37007, Spain ARTICLE INFO Keywords: Label embedding contrastive learning Max margin loss Deep learning Embeddings fusion Network intrusion detection ABSTRACT Contrastive learning makes it possible to establish similarities between samples by comparing their distances in an intermediate representation space (embedding space) and using loss functions designed to attract/repel similar/dissimilar samples. The distance comparison is based exclusively on the sample features. We propose a novel contrastive learning scheme by including the labels in the same embedding space as the features and performing the distance comparison between features and labels in this shared embedding space. Following this idea, the sample features should be close to its ground-truth (positive) label and away from the other labels (negative labels). This scheme allows to implement a supervised classification based on contrastive learning. Each embedded label will assume the role of a class prototype in embedding space, with sample features that share the label gathering around it. The aim is to separate the label prototypes while minimizing the distance between each prototype and its same-class samples. A novel set of loss functions is proposed with this objective. Loss minimization will drive the allocation of sample features and labels in embedding space. Loss functions and their associated training and prediction architectures are analyzed in detail, along with different strategies for label separation. The proposed scheme drastically reduces the number of pair-wise comparisons, thus improving model performance. In order to further reduce the number of pair-wise comparisons, this initial scheme is extended by replacing the set of negative labels by its best single representative: either the negative label nearest to the sample features or the centroid of the cluster of negative labels. This idea creates a new subset of models which are analyzed in detail. The outputs of the proposed models are the distances (in embedding space) between each sample and the label prototypes. These distances can be used to perform classification (minimum distance label), features dimensionality reduction (using the distances and the embeddings instead of the original features) and data visualization (with 2 or 3D embeddings). Although the proposed models are generic, their application and performance evaluation is done here for network intrusion detection, characterized by noisy and unbalanced labels and a challenging classification of the various types of attacks. Empirical results of the model applied to intrusion detection are presented in detail for two well-known intrusion detection datasets, and a thorough set of classification and clustering performance evaluation metrics are included. 1. Introduction Contrastive learning is attracting great research attention for its interesting properties to implement classifiers [1]. In a contrastive learning framework, each sample is translated into a representational space (embedding) where it is compared with other similar and dissimilar samples with the aim of pulling similar samples together while pushing apart the dissimilar ones. There is a plethora of strategies to fulfill this aim by considering different alternatives: (a) Define similar and dissimilar: A similar element can be a manually slightly modified replica of the original element (self-supervised) [2] or an element in a closeness position (in time, space, order…) to the original element (unsupervised) [3]. In this latter case, similar elements can be obtained by sampling the data source (e.g., neighboring words in a sentence) in the “proximity” of the original element. Furthermore, similarity between elements can also be defined by * Corresponding author. E-mail addresses: [email protected] (M. Lopez-Martin), [email protected] (A. Sanchez-Esguevillas), [email protected] (J.I. Arribas), belcar@ tel.uva.es (B. Carro). Contents lists available at ScienceDirect Information Fusion journal homepage: www.elsevier.com/locate/inffus https://doi.org/10.1016/j.inffus.2021.09.014 Received 24 March 2021; Received in revised form 27 July 2021; Accepted 15 September 2021 Information Fusion 79 (2022) 200–228 201 belonging to the same class/label (supervised) [4], which requires elements to be explicitly labeled. (b) Type of elements represented in embedding space: The elements represented in embedding space are normally the sample features. Nevertheless, in this paper we propose to use additional elements that can also be mapped to the same embedding space and jointly compared with the sample features. Specifically, we propose to use the labels as additional elements to be represented in embedding space. (c) Define the separation process: Similar and dissimilar elements with respect to a certain reference (anchor) element are usually referred to as positive and negative, respectively. The separation process can be Fig. 1. Contrastive learning reference view considering the embedding elements (vertical-axis) and separation process strategy (horizontal-axis). Distance computation can be done: only between samples using their feature embeddings (Upper row), between each sample (using its feature embedding) and all its positive and negative labels (Middle row), and between each sample and its positive label and a single representative of its negative labels (Lower row). The separation process can be implemented: with a maximum margin between positive and negative points (features or labels) (Left column), with a maximum margin for negative points and a minimum separation to the anchor for positive ones (Middle column), and with a maximum/minimum separation between negative/positive points to the anchor (Right column). Type II and III solutions correspond to our proposed models. Below each type of solution, its most representative models are listed with a reference to the Section that shows its architecture. M. Lopez-Martin et al. Information Fusion 79 (2022) 200–228 202 attained in different ways: The positive and negative elements can be pulled/pushed to/from the anchor trying to reduce/increase the distance as much as possible [5,6]. On the other hand, the objective can be to increase the distance between positive and negative just beyond a margin [7]. Additionally, a combined objective can be to minimize the distance of the positive to the anchor and maximize the distance to the negative beyond a margin [8]. (d) Number and nature of positive and negative samples used in each comparison: We can choose a single positive and negative element per comparison [7,8], a group of positives and negatives [5,9] or a representative(s) for the positive and negative groups serving as proxy for the group [10]. (e) Distance function: Euclidean, Cosine… (f) Loss function: contrastive [8], triplet [7], … Considering the contrastive learning alternatives mentioned above, there is a wide range of solutions found in the literature. These solutions are mainly based on self-supervised and unsupervised settings, and only recently on supervised schemes. They have achieved excellent results [1–4], but they also have some significant drawbacks: (a) long training periods, (b) the need to create complex sampling strategies to extract negative samples, and (c) difficulties in creating representative prototypes for each class/label in a supervised scheme. In this work, we propose a solution to these drawbacks by creating a supervised contrastive learning scheme where the sample features and labels are mapped to the same embedding space, and the contrastive mechanism (pulling/pushing of similar/dissimilar elements) is applied between features and labels, instead of being applied exclusively between features. The sample embeddings used in contrastive learning are usually created from the features of the sample. The approach proposed here, which combines features and label embeddings into the same representational vector space, where they can be jointly manipulated, has been recently applied in image [11] and text processing [12] as an intermediate feature processing step, but has not been used, to our knowledge, in any systematic approach within a contrastive learning framework. In this work, we provide a systematic reference framework for current contrastive learning solutions using the classic features-only embeddings approach (Fig. 1, Type I), and our proposed solutions using jointly features and label embeddings (Fig. 1, Type II and III). Fig. 1 graphically shows where our proposed solutions fit in a generic contrastive learning reference view. The charts in Fig. 1 are based on the following definitions and notation. We assume to have a dataset with N samples: {xi}N i=1. We also assume that each sample (xi) has a set of similar (positive) samples (xP i) and a set of dissimilar (negative) samples (xN i). Similar and dissimilar samples can be established by a “proximity” function without requiring an explicit labeling (unsupervised and self-supervised) or by belonging to the same explicit label/class (supervised). For the supervised case, we have a set of C labels (classes): {Lj}C j=1, where each sample xi is associated to a ground-truth label/class which we call its positive label, implying that the rest C−1 labels will be its negative labels. The positive label for xiis identified as LP xi and any of the negative labels for xi as LN xi. We can also select a representative label from the group of negative labels, this representative can be one of the existing negative labels (the one that best represents the set of negative labels for our task) or it can be a new one built from them. We will denote as LN* xi the unique label acting as a representative of the group of negative labels. It is important to note that in Fig. 1 the elements represented are in embedding space. The embedding transformation is represented as: φ(a), where a is a generic input vector and its embedding is a low dimensional representation of vector a. The function φ(a)is a mapping of vector a into a space of different dimensionality (usually smaller) where the objective is to maintain the representational capacity of the original vector into the new mapped vector. Fig. 1 provides a reference view for contrastive learning solutions considering two axis: the elements represented in embedding space (vertical-axis), and the separation process adopted (horizontal-axis). In this reference view, our research scope corresponds to Type II and III solutions: •The embedding elements used (Fig. 1, vertical-axis) can be: (a) only the sample features (xi) (Type I), (b) the sample features (xi) plus the positive (LP xi)and all negative labels (LN xi)for that sample (Type II), and (c) the sample features (xi) plus a reduced number of its representative labels that provide all the required information for our task (Type III). In this latter case, our proposal is to use for each sample, its positive label (LP xi)and a single representative element of its negative labels (LN* xi),as the two unique representatives. There are several ways to choose a single representative for the negative labels; we propose two useful alternatives: choose the element which is nearest to the sample features (in embedding space), or choose the centroid of the cluster formed by the negative labels (Fig. 3). For Type II and III solutions, the positive and negative labels depend on each sample (xi)and the aim is to pull closer the sample features (xi)to its positive label (LP xi)while pushing apart all the negative labels (LN xi), or their representatives (LN* xi). Type I solutions have a similar aim, but in this case the objective is to attract similar features and repel dissimilar ones. If similarity is based on class membership, all types of solutions (Type I, II or III) will create a cluster of samples for each label, however, only Type II and III solutions can use the labels as prototypes acting as cluster centers attracting all their same-class samples and repelling others. Having these prototypes is very important as they can be used as unique representative elements of each class. •The separation process (Fig. 1, horizontal-axis) may involve: (a) to push apart the positive and negative points (features or labels) beyond a certain margin m (Max-margin), (b) to push apart the negative points beyond a certain margin m and to pull closer the positive points as much as possible to the reference sample (MaxMarginþMin-Separation), and (c) to push apart the negative points and to pull closer the positive points (both) as much as possible to the reference sample (Max-SeparationþMin-Separation). The separation process is implemented by specific loss functions which are based in combinations and variants of quadratic, exponential, logloss and maximum-margin (hinge) losses. Fig. 1 presents the names of the most representative models for each type of solution (below each chart). These models are implemented through a series of proposed architecture (Sections 3.3 and 3.4). In Fig. 1, each of the proposed architectures is assigned a horizontal rectangle where the different models implemented by that architecture are located, along with a reference to the corresponding Section that explains it. There are architectures that provide a single model (for example, the architecture in Section 3.3.2 implements only the Contrastive over Label Embeddings -ConLE model), while other architectures can implement different types of models (with different properties) by changing the loss function used within that architecture; for example, the architecture proposed in Section 3.4.1 implements all types of models, with different properties, and corresponding to different separation approaches (different position on the horizontal-axis of Fig. 1). The two main categories of models in relation to this research are: - The Type I solution models correspond to currently published research works: Triplet [7], RankedList [13], SupCon [4], Contrastive [8,14], Lifted [9], N-pair [5], NT-Xent [15], ProxyNCA [10,16], InfoNCE [6]. SupCon is the only supervised contrastive learning model identified in the literature, and it is located in two different columns in Fig. 1, because it can adopt different separation strategies based on the number of negative labels considered. - The proposed models for Type II and III solutions are novel solutions proposed in this paper (Sections 3.3 and 3.4): Max Margin over Label Embeddings (MMoLE), Contrastive over Label Embeddings (ConLE), Contrastive with Cross Entropy (ConCE), Max-Margin ((λ)MM), Cross M. Lopez-Martin et al. Information Fusion 79 (2022) 200–228 203 Entropy with Max-Margin (CE+(λ)MM), Cross Entropy plus Distances with Max-Margin (CEDist+(λ)MM), Contrastive (Con(λ)), Exponential loss for Max-Separation (E(λ)MS), Squared Exponential loss for Max-Separation (E2(λ)MS), Cross Entropy with Exponential loss for Max-Separation (CE+E(λ)MS) and Cross Entropy plus Distances with Exponential loss for Max-Separation (CEDist+E(λ)MS). Where (λ) stands for a dummy letter representing a particular loss function, as defined in Sections 3.3 and 3.4. We call the models included in Type II and III solutions as Label Based Contrastive Learning (LB-CL) and Representative Label Based Contrastive Learning (RLB-CL), respectively. They correspond to our proposed models. These models can be used as: (a) Classifiers, predicting the correct label for a new sample. (b) The embeddings of the sample features and the distances between features and labels embeddings can be used to replace the original features as a dimensionality reduction technique. (c) Finally, the samples form clusters around the label prototypes, which can help to assess the location of new samples in a low dimensional graphical representation of sample and label embeddings (e.g., security personnel identifying attacks in new samples). The graph can be obtained directly using low dimensional embeddings or by transforming the embeddings for visualization (e.g., PCA, t-SNE). As mentioned earlier, Type I solutions are usually based in selfsupervised [2] or unsupervised learning [3]. The Type I solutions based on supervised learning [4] establish two samples as similar when they belong to the same class and dissimilar otherwise. In this case, the classes are used only to select the group of positive and negative samples corresponding to each reference sample (xi). Type II and III solutions use the labels/classes in a different way. They use the labels as class prototypes in embedding space. The samples are clustered around their same-class label prototype. The aim is to create maximum separated label prototypes, as well as minimizing the distance between each prototype and their same-class samples. In this work, we propose new losses specifically designed for the comparison of distances between the samples and their labels or their label representatives (Sections 3.3 and 3.4). Type I solutions demand a large number of comparisons (contrastive learning) between a vast combinatorial space of pair-wise positive (similar) and negative (dissimilar) samples. Type II solutions tremendously reduce the necessary pair-wise comparisons to the number of labels (classes) only, which is necessarily much less than the number of samples for a sensible learning process. Type III solutions reduce even more the number of comparisons to just two: one between the sample features and the positive label and other between the sample features and the single representative of the cluster formed by all the negative labels. In addition, in Type III solutions we offer a novel way of inputting labels to the classifier at both training and prediction (test) phases, by presenting all labels in parallel (Figs. 7–9). In this way we further reduce the required training and prediction times (Section 4). Both, Type II and III solutions are novel solutions proposed in this work. Network intrusion detection (NID) is a complex field that faces the problem of detecting cybersecurity issues (malicious activity or policy attacks) on data networks by analyzing the information contained in the exchanged data packets. The information in the data packets is transformed into a vector of continuous and categorical values (e.g., size, addresses, flags...) that represent the network connection. This vector can be compared, searching for similar patterns, with pre-registered vectors associated to normal traffic or attacks (signature-based intrusion detection), or the vector can be used as the input to statistical methods or machine learning classification methods to detect attacks (anomaly-based intrusion detection). In this work we apply the anomaly-based approach. NID is an active and challenging area of machine learning research [17–19]. A wide range of machine learning methods have been applied to NID, the most common being: linear models, decision trees, gradient boosting methods, support vector classifiers, multilayer perceptron, convolutional and recurrent neural networks, kernel methods and generative models [20–22] The proposed models are generic and applicable to any supervised classification problem, and translating these models to NID is straightforward. Network intrusion samples (NIS) are represented by vectors of network features. NIS vectors length is usually large since each feature brings information about a particular property of the network flow carrying the attack. Intrusion labels are represented by one-hot-vectors associated to each label. Label vectors length equal the number of different labels. Labels and NIS vectors are of very different sizes. The objective is to map the NIS and labels into a common representational vector space (embedding space) where the fundamental original information is preserved and where semantically meaningful comparisons can be established. NID datasets, which contain the activity of network flows and their associated status: normal or intrusion type, are characterized by being good representatives of noisy and unbalanced datasets. We have chosen two well-known NID datasets (NSL-KDD and UNSW-NB15) to carry out the experiments for the novel proposed models (Type II and III solutions). These datasets provide diverse and challenging scenarios for the proposed models. A complete analysis of results is provided on the application of the proposed solutions (Type I and II) to the two data sets (Section 4). The results include: (a) classification performance metrics to detect the attacks, (b) clustering performance metrics on the quality of the clusters of samples around each label prototype and, (c) improvement in classification metrics from several well-known machine learning (ML) algorithms, when using the sample embeddings and their distances to label embeddings as new features, replacing the original sample features. The contributions offered by the proposed models are: (a) Introduce a fusion of features and labels representation within a common vector space. (b) Present a wide range of novel architectures and loss functions that leverage this fused representation space. (c) Provide new alternatives to contrastive learning by avoiding negative sampling or any other sophisticated selection strategy for negative/positive labels. (d) Significatively reduce the number of pair-wise comparisons required. (e) Create a novel approach by replacing negative labels with their best representatives, further reducing the number of comparisons and their associated training and prediction times. (f) Show that the proposed models create excellent classifiers for noisy and unbalanced datasets. (g) Show that features embeddings, created as a by-product of the classification process, can be used instead of the original features. We show that a wide range of classifiers improve their performance by using these low-dimensional alternative features. The rest of the paper follows this schema: Section 2 summarizes related works. Section 3 describes the proposed model and its relationship with related models. Section 4 describes the datasets, and the experimental results and Section 5 presents the conclusions. 2. Related works Referring to the three main topics of this research: contrastive learning, label embeddings and NID. As far as we know, there is no work presenting label embeddings within a contrastive framework; there are works using label embeddings as a dimensionality reduction technique, and there are works presenting Type I contrastive learning solutions (Fig. 1) for NID. We have not identified any work proposing label embeddings for NID. The main related works from the literature are as follows: •Contrastive learning with feature embeddings for classification: Contrastive learning is an active research topic with many recent works following different approaches with an emphasis in: (a) Type of separation process (e.g., Max-Margin...)(Fig. 1). (b) Acquisition of similar/ dissimilar elements (supervised, self-supervised or unsupervised).(c) Number of positive/negatives samples per comparison (single, group or representative). All these works correspond to Type I solutions (Fig. 1): Max-Margin Triplet loss was presented in [7] with a max-margin separation for the M. Lopez-Martin et al. Information Fusion 79 (2022) 200–228 204 distances difference between an anchor point and a selection of its negative and positive samples. One positive and negative sample is used per comparison, and samples are acquired in a self-supervised or unsupervised manner. A group of positive and negative samples per comparison is used in [13] with a max-margin loss called ranked list loss based on weighting positive and negative examples with respect to their distance to the anchor point using class information (supervised). Max-MarginþMin-Separation Authors in [23] present a k-nearest neighbor framework to minimize Mahalanobis distance to positive samples and maximize distance with a margin to negative samples. They use a subset of positive samples to compute the loss which is a contrastive loss [14]. Contrastive loss is presented in [8,14] with a quadratic loss for the distance to positive samples and a quadratic max-margin loss to the distance of negative samples. It is applied originally to image recognition using label information [8] and face recognition with a Siamese network [14] in a self-supervised manner. The comparison in all cases is done between a single positive and negative sample. Alternatively, in [9] all negative samples are selected with a modified loss that comprises an exponential loss with a margin for the distance to all negative samples plus the direct distance to the positive sample. The experiments are done with a self-supervised setting. Max-Separation þMin-Separation With a supervised learning framework, [16] proposes Neighborhood Component Analysis (NCA) where a k-nearest neighbor with stochastic (soft) neighbor assignments optimizes the probability of a sample to be classified into the correct class. It has connections with Linear Discriminant Analysis (LDA) [24] but it does not require all class distributions to be Gaussian with equal covariance [16]. A single positive sample and a group of negative samples is used per comparison. In [5] the comparison is performed between a positive sample and a random selected element for each negative class (multi-class N-pair). As with the previous work the objective is to push apart or pull closer the negative/positive points as much as possible to the reference sample. Authors in [15] present NT-Xent/SimCLR with a exponential loss over a distance function between similar points normalized by the sum of distances for the rest of points in the same training batch (assumed as negative samples). The work in [25] extracts label prototypes by sampling and averaging their associated embedded features. There is no fusion of labels and features in representation space. The label prototypes are used in a few-shot learning framework. A group of positive and negative samples are implicitly used per comparison. With a (self/un) supervised framework, the loss presented in NCA is used also in ProxyNCA [10] where instead of using a group of negative samples this method creates a reduced number of proxy elements that can replace the original points in the comparison. The method creates the proxy points among the original points as part of the learning process. It extends NCA to a subset of points represented by proxies instead of the original points. A different loss is presented in [6] with InfoNCE which is based on noise contrastive estimation and applied to sequential data. The loss is based in a log-bilinear score between a latent variable created by an autoregressive model applied to past samples embeddings and future positive samples normalized by the sum of the scores for all samples. A single positive sample and a group of negative samples are used per comparison. The structure of the loss is similar to [5,15,16]. Based on SimCLR, [26] presents a sophisticated sampling mechanism for negative points to minimize the number of false negative points (positive points taken as negative). ProtoNCE [27] is a variant of InfoNCE that creates prototypes of the positive and negative samples with an unsupervised framework by constructing latent variables with an expectation-maximization algorithm. The proposed loss includes the original InfoNCE loss plus an additional loss (also based on InfoNCE) formed by sampling a reduced number of negative and positive prototypes. •Alternative frameworks for contrastive learning: Traditional pair-wise distance ranking applied in contrastive learning for classification can be extended in different directions. Authors in [28] propose an adversarial contrastive learning based in SimCLR. The embedded features can be particularly useful for learning with an small number of samples (few-shot learning) as proposed in [29]. Alternative models based on ranking the recovery error performed by generative models (e.g. variational autoencoders) conditioned on the labels can also be used for classification [30]. •Application of label embeddings: Label embedding has been applied in image classification [11], text classification with attention mechanism [12], text classification based on a bilinear ranking function [31] and text classification in a multi-task setup [32]. It has also been applied for multi-label embedding with a neural factorization machine [33]. None of these works uses a contrastive learning approach. They employ the label embeddings as an intermediate step for downstream tasks. •NID and contrastive learning: In latter survey studies on the application of machine learning and deep learning to network intrusion detection there is no mention to the application of contrastive learning [18,20,22]. However, there is a growing interest in these techniques in current works; for example, [34] proposes anomaly detection by self-supervised contrastive learning, following the framework proposed in SimCLR [15]. An unsupervised anomaly detection with contrastive learning is proposed in [35], where a neural network is trained by contrasting distances, only between normal instances, and a threshold is used to detect outliers when its distance to the center of normal instances is above the threshold. Authors in [36] propose a supervised intrusion detector using a Siamese network with similar and dissimilar samples being differentiated by belonging to the same class. A similar solution is adopted in [37] to detect intrusions for the NSL-KDD and UNSW-NB15 datasets. A dimensionality reduction technique, by applying a Siamese network, is presented in [38]. Sample features are transformed into a 1D feature space. The resulting variable is used to identify attacks with a simple visualization tool. 3. Methods description We position the proposed models (LB-CL, RLB-CL), which correspond to Type II and III solutions (Fig. 1), in relation with alternative contrastive learning models and describe in detail the different architectures used to implement them. 3.1. Contrastive learning overview Contrast learning was originally thought [8,14] to perform a feature transformation such that similar samples would be placed close to each other (in the transformed feature space) according to a certain predefined similarity goal. The similarity of the samples is based on a distance function (usually Euclidean or cosine) and this distance is used to establish the proximity of an anchor sample to a set of positive (similar) and negative (different) samples (Fig. 2). Similarity is imposed on the transformed space by training alternately with groups of similar and dissimilar samples. A special loss function is used to penalize dissimilar elements approaching or similar elements moving apart (Fig. 2). This method is especially useful for creating a feature transformation into an embedding space that incorporates domain-specific semantics using the distance between embeddings. These embeddings can be used as features in downstream classification tasks e.g., word2vec [39]. During training, samples can be presented in pairs or in different combinations/group arrangements of similar/dissimilar samples (e.g., triplets [7]). Most of the loss functions used in the contrastive learning literature are pairwise ranking losses, based on a combination of quadratic, logarithmic, exponential and maximum-margin losses. As a summary, we provide an overview of the most popular contrastive learning architectures with their associated loss function equations (Fig. 1, Upper charts): Contrastive [8,14] (Eq. (1)), Triplet [7] (Eq. (2)), Lifted [9] (Eq. (3)), M. Lopez-Martin et al. Information Fusion 79 (2022) 200–228 205 N-pair [5] (Eq. (4)), InfoNCE [6] (Eq. (5)), SimCLR, NT-Xent [15] (Eq. (6)) and NCA, proxy-NCA [10,16] (Eq. (7)): 1 N∑ N{y+.(D+)2+y−.[max(m−D−,0)]2}(1) 1 N∑ N max [(D+)2− (D−)2+m,0](2) 1 2P∑ P[max(log [∑ Pn exp(m−D−)]+D+,0)]2 (3) 1 N∑ N log [1+∑ Pn exp(S−−S+)](4) 1 N∑ N −log[exp(S+/ τ ) exp(S+/ τ ) + ∑Pnexp(S−/ τ )](5) 1 N∑ N −log[exp(S+ / τ ) ∑Pnexp(S−/ τ )](6) 1 N∑ N −log[exp(− D+) ∑Pnexp(− D−)](7) Note: In previous equations we have used the following simplified notation: D+and D−are distances to positive and negative points (e.g., Euclidean distance). S+and S−are similarities to positive and negative points (e.g., Cosine similarity). m is a margin. τ is a temperature parameter. N is the number of samples. P is the number of pairs of samples. Pn is the number of pairs of negative samples (for a particular reference sample). y+is an indicator function whose value is 1 if two similar samples are compared during training and 0 otherwise; y−is also an indicator function to mark dissimilar samples (y−=1 −y+) When considering the literature for contrastive learning, much of the complexity and sophistication goes into finding a good representative set of positive (similar) and negative (dissimilar) points to compare to a specific anchor point. This creates implementation problems because the distance computation between samples (sample-wise) demands a complex training process. The training is performed by selecting a representative number of positive and negative samples for each reference sample. To reduce variance and ensure generalization of results the set of positive/negative representative samples is usually large and difficult to choose. Contrastive learning is generally applied in an unsupervised/selfsupervised scheme. In a self-supervised scheme, similar samples are constructed (synthesized) by some controlled modification of the anchor sample (e.g., rotate/translate an image). In an unsupervised scheme, similar samples are selected by some proximity attribute imposed on the data (e.g., time, location, order, position...). In this latter case, the selection is performed by random sampling conditioned on the proximity attribute (e.g. similar words are selected by proximity to the anchor word in a document corpus [39]). Contrastive learning has difficulties to be applied in a supervised scheme, since we have to select a representative number of samples per class to compare with a certain anchor sample, and in the inference (prediction) phase we also need to compare the new sample with a significant number of samples corresponding to all classes. The label assigned to the new sample is selected by some form of majority voting using the labels of the closest samples. The main problem with this approach is the initial lack of a natural representative for the classes; having this natural representative of the classes would reduce the number of distance comparisons required at the time of inference to a single comparison with each class representative. The above mentioned problems: a) large number of pair-wise comparisons, b) complexity in selecting representative sets of positive and negative samples, and c) difficulty to integrate supervised learning, are addressed by our proposed solutions using the labels themselves as the best class representatives. This approach reduces drastically the number of required pair-wise comparisons since the points to compare with now are the labels, which we assume have a significant smaller number than the number of samples. Additionally, label embeddings now get a natural representative position in the cluster of points sharing their label; either by choosing the geometric center of this cluster or simply the point with the smallest distance to the majority of points within the same class. 3.2. Scenarios for contrastive learning In this section we provide an alternative schematic view (Fig. 3) showing the main elements that define contrastive learning frameworks and how our proposed models relate to them. It is a complementary view to Fig. 1. A visual summary of this alternative view is given in Fig. 3. The upper-chart in Fig. 3 shows the process followed to create common embeddings for features and labels. The middle-chart shows the alternative contrastive learning scenarios corresponding to Type I, II and III solutions (Fig. 1), where Type I solutions have been divided by the data Fig. 2. Overview representation of contrastive learning. An anchor sample (x i) in embedding space (φ(xi)) is moved close to similar samples (by some similarity criteria, e.g., same age faces) (φ(xk)) and is separated from dissimilar samples (φ(xj)). M. Lopez-Martin et al. Information Fusion 79 (2022) 200–228 206 acquisition process into self-supervised/unsupervised (I.A) or supervised (I.B). Finally, the lower-charts present two variants of Type III solutions (III.A and III.B), showing two alternative ways to choose the representative element (prototype) assigned to the negative labels (LN* xi). We now extend the definitions and notation given in Section 1; this extended notation will be applied in the rest of this work. We assume to Fig. 3. (Upper chart) Schema of the mapping of labels and predictor features into the same embedding representational space. The labels and sample features are vectors of different dimensions (C and F respectively). The embedding mapping translates them into vectors of the same dimensionality (E) where their comparison is feasible. (Middle charts) Four scenarios for contrastive learning with different defined distances between features (Type I) and between features and labels (Types II and III). (Lower Charts) Detail of Type III solutions, describing the distances D NN(xi)(Lower-left) and DAN(xi)(Lower-right) between the sample to either its nearest negative label (DNN(xi)) or the average point formed by all its negative labels (DAN(xi)), respectively. Which labels are positive or negative is set for each sample (xi). The positive label for a sample xi is its ground-truth label, all others are negative labels for this particular sample. All distances are computed using the vector embeddings. M. Lopez-Martin et al. Information Fusion 79 (2022) 200–228 207 have a dataset with N samples: {xi}N i=1, each sample having F features: xi∈RF (predictor features). And a set of C labels (classes): {Lj}C j=1(Fig. 3, upper-chart). Each sample xi can be associated to a ground-truth label/ class (yi)which we call its positive label, implying that the rest C −1 labels will be its negative labels. We identify the index of the positive label for xias ki∈ [1,C], such that Lkiis the ground-truth label for xi, i.e., Lki =yi. We will also identify the positive label for xias LP xi (i.e., LP xi =Lki =yi) and any of the negative labels for xi as LN xi (i.e., LN xi∕= yi). Labels are one-hot-encoded represented with a zero array of length C with a single 1 in a specific position assigned to each label; in particular, for the ground-truth label yi, we indicate the position j in its one-hot-encoded array as yj i, that means that yj i=1 if j=ki and yj i=0 if j∕= ki. The same notation will be used for the predicted label (yi). In line with Section 1, we will denote as LN* xi as the unique label acting as a representative of the group of negative labels (Type III solutions). The labels and sample features are vectors of different dimensions (C and F respectively). We define a specific embedding mapping for each of these vectors (Fig. 3, upper-chart) with two aims: (a) translate them into vectors of the same dimensionality (E) where their comparison is feasible and, (b) create relationships between the embedded vectors with the objective to classify each sample xi with its associated positive label LP xi. The embedding transformation is represented as: φ(a):Rm→Rn, where a is a generic input vector a∈Rm and its embedding is a low dimensional representation of vector a: φ(a) ∈ Rn,where usually m>n. The function φ(a) is a mapping of vector a into a space of different dimensionality (usually smaller) where the objective is to maintain the representational capacity of the original vector into the new mapped vector. Fig. 3 (middle-charts) shows the different options available to perform the approach between similar samples and the separation between samples that are not similar or that do not belong to the same class. The approach/separation process is done in embedding space under four different scenarios: •In Type I.A solutions (Fig. 3), we assume that each sample xi has a set of associated samples xk which are similar/close to xi and are identified as xk∈CloseTo(xi). In this scenario, the aim is to bring xi closer to its similar samples (xk), while separating xi from samples that do not belong to CloseTo(xi). The elements belonging to CloseTo(xi)can be obtained by time, spatial or context proximity to xiwithout requiring a direct labeling of the “close to” set. This set can be obtained by simply sampling the data source (e.g., neighboring words, next sound) in the “proximity” of xi . This approach is usually called unsupervised [3,6]. Another approach is the generation of augmented replicas of each sample by introducing controlled modifications to the original sample (self-supervised) [2]. In this case, the replicas would be considered similar elements. No labels are used in this scenario. Only distances between feature-embeddings are employed. The contrastive mechanism is applied to each sample xi and requires that the sample be compared with a significant number of samples similar to xi({xk∈ CloseTo(xi)}) and also with a representative number of samples not similar to xi ({xj∕∈ CloseTo(xi)}) (sample-wise) •A labeling of the samples is assumed in Type I.B solutions (Fig. 3). The labeling is done in a supervised manner, and only distances between feature-embeddings are employed. We call positive samples those samples that share the same class as the reference sample, and negative samples those samples with a different class from the reference sample [4]. In much the same way as Type I.A solutions, the aim is to bring xi closer to its positive samples, while separating xi from its negative samples. The contrastive mechanism is also applied (sample-wise) between each sample with other samples sharing the same label, as well as samples with other labels. Using the sample-wise distances, the classification can be done by majority vote, if a distance threshold is previously established for the similarity of the samples pair, or some other aggregated distance computation (e.g., choose the class with a minimum sum of distances between the reference sample and the representative samples of that class). The process has similarities to a K-Nearest Neighbors (KNN) model where the class of a new element is assigned according to the majority class of its K nearest elements. The difference in our case is that the set of elements to perform the distance comparison is not the complete population (as in KNN) but a selection of samples from the different classes, and the selection is made by random sampling of these classes. We can see, that this solution has an implicit problem due to the lack of a single representative that serves as a prototype for the different classes [40]. This is the problem solved with Type II and III solutions (Fig. 3) by building a best representative (prototype) in embedding space for each of the classes. When using these class prototypes, the objective will now be to reduce the distance between a reference sample to its same-class (positive) prototype and to increase the distance to its different-class (negative) prototypes. By using prototypes, the training and prediction pair-wise comparison process is significantly reduced and creates clear references to perform comparisons. Both Type II and III solutions correspond to supervised models with a label associated with each training sample. •In Type II solutions (LB-CL)(Fig. 3), the sample features and the labels are embedded into the same vector space where their distance can be easily computed. The aim in this scenario is for each sample to be as close as possible to its label embedding while separating itself from all other label embeddings. We apply this scenario to perform classification by comparing an anchor sample with its positive and each of its negative labels (Label-wise) and choosing the label with the shortest distance to the sample (in embedding space). This is a supervised scenario. The models presented for this scenario are part of this research work, and are novel in applying contrastive learning and systematically extend other previous works based exclusively on features embeddings to a new scenario that simultaneously considers feature and label embeddings. •Type III solutions (RLB-CL)(Fig. 3) are also part of our proposed models. In this case, we also use the features and label embeddings, but we replace all the negative labels corresponding to a sample with a single element that represents them. This element is assumed to be the most representative element of the cluster of negative labels (this cluster is sample dependent). In this way, we will not perform a comparison of the sample’s features with each of the labels (as in Type II) but just with two labels: the same-class or positive label, and the representative of the out-of-class or negative labels (Representative Label-wise). We will consider two main variants for choosing the most representative element of the cluster of negative labels: (a) the nearest negative label to the sample (III.A, Fig. 3, Lower-left) and (b) the centroid of the negative labels cluster (III.B, Fig. 3, Lower-right). This is also a supervised scenario. There is no work (in any field, as far as we know) proposing a similar approach of using a single element (representative) of the negative labels (in embedding space). In Fig. 3, for Type II and III solutions, we can define different distances between labels (Lj) and predictor features (xi) in embedding space (Fig. 3, lower-charts): (a) DP(xi)(Eq. (8)) is the distance of the xi embedding to its positive label embedding (Lki). (b) DN(xi,Lj)(Eq. (9)) is the distance of the xi embedding to one of its negative labels embedding (Ljwith j ∕= ki). (c) DNN(xi)(Eq. (10)) is the distance between the xi embedding and its negative label embedding which is the nearest to xi. (d) DAN(xi)(Eq. (11)) is the average distance of the xiembedding to all its negative label embeddings. DNN and DAN correspond to distances between a sample embedding (xi) and a representative of its negative labels (LN* xi). DP(xi) = D[φ(xi),φ(Lki)] (8) M. Lopez-Martin et al. Information Fusion 79 (2022) 200–228 208 DN(xi,Lj)=D[φ(xi),φ(Lj)]where j∕= ki(9) DNN (xi) = min j∕=ki D[φ(xi),φ(Lj)] (10) DAN (xi) = 1 C−1∑ j∕=ki D[φ(xi),φ(Lj)] (11) A schematic representation of DP,DNN and DAN is given in Fig. 3 (lower-charts). The average distance to the negative labels (DAN) is an upper bound (Appendix. A) for the distance to the cluster’s centroid of the negative labels, therefore, by maximizing this distance we can assume that the other is also maximized. 3.3. Label based contrastive learning (LB-CL) (Type II) There are three proposed architectures for Label Based Contrastive Learning (LB-CL) solutions (Type II solutions, Fig. 1): (a) Max Margin over Label Embeddings (MMoLE), (b) Contrastive over label embeddings (ConLE), and (c) Contrastive with Cross Entropy (ConCE). These architectures implement a contrastive learning framework between the sample features and all its labels (positive and negatives). 3.3.1. Max margin over label embeddings The Max Margin over Label Embeddings architecture (MMoLE) (Fig. 4) is based on a max margin loss acting over the distance between the positive and one negative label for a sample xi. The MMoLE loss (Eq. (12)) is defined as follows, where: N is the number of samples; C is the number of labels; {Lj}C j=1is the set of C labels; ki∈ {1..C}corresponds to the index of the positive label of xi and the loss is extended to all negative labels for all samples: MMoLELoss=1 N∑ N i=1∑ C−1 j=1 max(D[φ(xi),φ(Lki)]− D[φ(xi),φ(Lj∕=ki)]+1,0)(12) This model implements a Max-Margin separation strategy (Fig. 1) by increasing the difference, beyond a margin, between two distances: the distance between the anchor-sample (xi) and the positive label, and the distance between the anchor-sample and each negative label. MMoLEloss is a novel max margin loss between the distance of a sample (xi) to its true label (Lki) and each of the distances to the negative labels (Lj∕=ki). MMoLEarchitecture analysis can be split between its training and prediction phases (Fig. 4). Three inputs per sample are required during Fig. 4. Max Margin over Label Embeddings (MMoLE) architecture is based on a max margin loss acting over the distance between the positive and one of the negative labels, for each sample x i. The loss is extended to all negative labels. The training phase (Upper chart) requires three inputs per sample: sample features (xi), positive label (LP) and one of the negative labels (LN). Each sample has a training round with every negative label. In the prediction phase (Lower chart) we need to try all labels to obtain the one with the smallest distance to the sample, and this operation requires as many forward-passes as labels. Since the training phase requires two labels per forward-pass we treat only one of the labels and set to a dummy value the other. M. Lopez-Martin et al. Information Fusion 79 (2022) 200–228 215 ground-truth label (yi). This latter input is required functionally but not physically since can be extracted by knowing the index of the positive label (ki). In the prediction phase (Fig. 8), we only require the sample features as input to the model and the features embedding NN (NN −Embedx) and the prediction NN (NN −Predict) to generate the predicted label. This approach reduces even more the prediction times, as corroborated by experimental results (c.f. Section 4). 3.4.3. Cross entropy over labels and distances with contrastive regularization The Cross Entropy over Labels and Distances with Contrastive Regularization architectures (CEDistþ(λ)MM, CEDistþE(λ)MS) (Fig. 9), are a variant of their counterpart CE+(λ)MM and CE+E(λ)MS models, where yiis a function of the features embedding (φ(xi)) and the distances between each label embedding (φ(Lj)|j=1..C) and the features embedding. i.e., yi=NN −Predict[φ(xi),D[φ(xi),φ(L1)],…D[φ(xi), φ(LC)]] . To generate the predicted label (yi), we concatenate all the inputs to the neural network. Apart from this difference, this architecture is completely similar to its CE+(λ)MM and CE+E(λ)MS counterparts. The complexity of the neural networks implementing the embedding functions are quite low. For the NN −EmbedL: 1 hidden layer with 10 nodes, ReLU activation and linear activation for the output layer. For the NN −Embedx: 2 hidden layers with 100/50 nodes, ReLU activation and linear activation for the output layer. All the embedding NNs share the same configuration. Likewise, for the models with added Cross Entropy, the NN that implements the final classification stage (NN −Predict), is also small: 2 hidden layers with 20/20 nodes, ReLU activation and softmax activation for the output layer. 3.5. Summary of proposed models and application scenarios This Section presents a summary of the main characteristics of the proposed models with a comparison between them Table 1 in terms of prediction performance, complexity and execution times (prediction phase). This information is extracted from their architectures and from the results obtained with the two datasets used in the experiments Table 1 Application scenarios of the proposed models: comparison of global characteristics of the different models. M. Lopez-Martin et al. Information Fusion 79 (2022) 200–228 216 (Section 4). The comparison provided in Table 1 can be useful to establish the application scenarios of the different models. Table 2 provides a unified summary classification of the different contrastive learning solutions which were presented under two different points of view in Figs. 1 and 3. The summary classification in Table 2 offers a taxonomy of the contrastive learning solutions based on the following differentiation parameters: a) the types of embeddings used, b) the separation strategy, and c) labels availability (whether the models are either supervised or unsupervised/self-supervised) All the models presented in this work have been implemented with neural networks trained with gradient descent, with a batch size of 100, 100 epochs and early-stopping with a waiting period of 10 epochs and a validation set of 20% of the training set. The optimizer employed was Adam with the original default parameters [41]. The source code for the most representative models of the proposed solutions is provided in a freely accessible repository [42]. Table 2 Taxonomy of contrastive learning solutions: unified classification of the solutions following the presentation given in Figs. 1 and 3. Fig. 10. Labels distribution in the training and test sets for NSL-KDD (Upper charts) and UNSW-NB15 (Lower charts) for multi-class (5 or 10 labels) and binary (2 labels) classification scenarios. M. Lopez-Martin et al. Information Fusion 79 (2022) 200–228 217 4. Results 4.1. Selected datasets We have selected two network intrusion detection datasets (NSLKDD and UNSW-NB15) as experimental benchmarks to apply the proposed models. They are good representatives of a field characterized by noisy and unbalanced datasets. The NSL-KDD dataset [43] is a well-known intrusion detection (ID) dataset. It contains 125,973 training samples and 22,544 test samples with 122 features (after one-hot encoding categorical features). The continuous features have been scaled in the [0,1] range. It has 40 labels with a dissimilar frequency distribution in the training and test sets. These 40 labels are usually aggregated into 5 or 2 labels [43] (Fig. 10). The UNSW-NB15 [44] is a newer ID dataset. It contains 2,540,044 training samples and 82,332 test samples. After a similar pre-processing (one-hot encoding and scaling) to that performed for NSL-KDD, the final number of features is 196. It has 10 distinct labels which can be aggregated into 2 (normal and attack) (Fig. 10). NSL-KDD and UNSW-NB15 present very different label distributions for the training and test sets with the common characteristic of being noisy and strongly unbalanced. The scenarios presented by the two datasets under their different multi-class and binary configurations allow exploring the performance of the proposed model in a wide range of challenging contexts. 4.2. Performance metrics We have used a complete set of metrics to compare results for the two datasets with the different models. We provide metrics for classification, clustering and training/prediction execution times. The classification metrics applied are: accuracy, F1-score, precision, recall, and Matthews Correlation Coefficient (MCC) with their usual definitions [45,46]. MCC is particularly useful for unbalanced datasets. The metric for clustering quality are Normalized Mutual Information (NMI) [5,9,10] and Silhouette coefficient [47]. NMI is invariant to label permutation, which is needed to perform a correct clustering evaluation, it is also a label based metric assessing the correct identification of labels to clusters. Silhouette coefficient is an unsupervised metric that is not based on knowing the ground-truth label; it provides a good indication of the separation between clusters. To identify the execution times of the algorithms, we present the training and prediction times as a comparative indicator of complexity and computational load, knowing that these values cannot be considered in an absolute but relative way to help the comparison between algorithms, since their absolute values depend on the nature and capacity of the processor used. All the proposed metrics are associated to better results in a monotonic increasing way; their range of values is [0,1], with the exception of MCC which is a correlation coefficient with a range [-1,1] (where +1 indicates a high correlation between ground-truth and predicted results and -1 total disagreement between them) and the Silhouette coefficient with also a range [-1,1] (where +1 indicates a high separation between clusters and -1 highly mixed clusters). Clustering scores have not been considered for the ML models because it is not the purpose of these models to perform clustering and, additionally, the high dimensionality of the original feature space (>100) makes clustering a challenging task in itself, with different applicable techniques and interpretation of results [48]. It is precisely the translation of the original features into a low-dimension embedding space (in LB-CL and RLB-CL models) that allows the classification task to be interpreted from a clustering perspective. 4.3. NSL-KDD results The results for the NSL-KDD dataset are divided by the type of classification into binary classification (2 labels) and multi-class classification (5 labels) (Section 4.1). We focus on providing the results of several representative LB-CL and RLB-CL models (Sections 3.3 and 3.4) in comparison with several well-known machine learning (ML) models: logistic regression, random forest, Gradient Boosting Machine (GBM), Support Vector Machine with radial kernel (SVM-RBF), MultiLayer Perceptron (MLP), Convolutional Neural Network (CNN) and a Linear Model with a Kernel Approximation (LM+KA). The last two models have been selected for having obtained very good performance results in recent works for both NSL-KDD and UNSW-NB15 datasets [49], and CNN models are some of the most widely used deep learning models in IDS [20,21,50]. In particular, we will use a CNN with a one-dimensional kernel (CNN-1D) due to the one-dimensional vector structure of the features in the two datasets used in this work. The LM+KA models [49] correspond to a recent trend towards fast shallow linear models that can handle non-linearities through a feature transformation that approximates a kernel SVM. The interest of these models lies in being especially fast. We have not considered applying recurrent networks (e.g., long short term memory (LSTM)) as the selected datasets are not based on sequential data. Two types of results are provided: (1) classification results using the models (ML, LB-CL and RLB-CL) and (2) improvement of the classification results of the ML models when using the original features versus the transformed features obtained by the LB-CL and RLB-CL models. 4.3.1. Classification with proposed models A comparison of classification and clustering performance metrics for a representative number of ML, LB-CL and RLB-CL models is provided in Fig. 11. The metrics are separated into two groups by the number of labels to detect (2 and 5 labels). For the LB-CL and RLB-CL models, we identify the dimension of the embedding space and the distance adopted between embeddings (Cosine or Euclidean). All results are computed over the test set (Section 4.1). The two rightmost columns in Fig. 10 provide the number of trainable weights and the number of Floating Point Operations (Flops) required for each model. These two metrics, along with the training and prediction execution times, offer a good insight into the computational complexity and performance of the models. A more detailed analysis of the complexity of the models is provided in Section 4.7. Fig. 11 shows the two best results in bold-italic text style. Additionally, a color code is used where a dark-green is associated with better results and a dark-red with worse results, with an interpolated colorpalette that is used for intermediate values. The color code and the best-two values are applied column-wise and independently (separately) for the blocks of 2 and 5 labels. 2 labels For binary classification, the LB-CL models (Type II) present best overall results (F1 and MCC) for classification and clustering. Model complexity is also smaller for LB-CL models. Considering training and prediction times, the RLB-CL models have the best times, which is as expected since the number of comparisons is less than for LB-CL models. The reduction in the execution times of the proposed models is at least an order of magnitude compared with most of the classic ML models. Interestingly, the training and prediction times of the RLB-CL models are better than those of the logistic regression and LM +KA models which are specially designed to be fast. 5 labels For multi-class classification (5 labels) the RLB-CL models (Type III) present some of the best classification metrics after the CNN-1D and LM+KA models which have previously shown particularly good performance for this dataset [49]. The RLB-CL models have also some of the best clustering metrics (NMI metric). The best prediction times are for the RLB-CL models and second best for the training times (after LM+KA model). The best RLB-CL model is ENMS considering F1 and MCC as our main classification metric. It is interesting to note the poor behavior of AMM, which incorporates exclusively the average distance to the group of negative labels. The unsupervised clustering metric (Silhouette) M. Lopez-Martin et al. Information Fusion 79 (2022) 200–228 218 obtains its maximum with contrastive models (ConLE) in both binary and multi-class scenarios. The best overall performance results are for the LM+KA model which was selected for being particularly effective with this dataset and being a tough competitor to other models. Nevertheless, considering the prediction performance and prediction speed together, the ENMS model has the third best prediction result with very short prediction times, while a similar model (CE+ENMS) provides the shortest prediction times. 4.3.2. Improvement of ML models As mentioned in the introduction, the embedding of the features and the distances between them and each of the label embeddings can be Fig. 11. NSL-KDD dataset: comparison of classification and clustering performance metrics between a selection of classic ML models vs. a representative set of the proposed LB-CL and RLB-CL models. A color code is used where dark-green is for better results and dark-red for worse results, an interpolated color-palette is used for intermediate values. The best-two values are in bold-italic. The color code and the best-two values are applied column-wise and separately for the blocks of 2 and 5 labels. All results are computed over the test set (For interpretation of the references to color in this figure legend, the reader is referred to the web version of this article.). M. Lopez-Martin et al. Information Fusion 79 (2022) 200–228 219 used as a replacement for the original features, implementing a de facto supervised dimensionality reduction. These transformed features can be used as original features in a subsequent classifier, implementing a stacked-ensemble configuration [51]. In a stacked-ensemble, it is important not to incur in data leakage or compromise the test set in the sequence of classifiers used. To avoid using the test set in the learning process we use always the same LB-CL/RLB-CL model as base model (level 0 classifier) [51] and one of several classic ML models as metamodel (level 1 classifier) [51]. We do not use the best base model (Fig. 11) using the results of the test data, as it would involve using the information from the test set during training. Our proposed model for stacked generalization is also different from [51] because we do not use the predictions of the base model as inputs to the metamodel, but rather an intermediate transformation of the original features created by the base model. Fig. 12 provides the results when using the original features versus the transformed ones produced by one of our proposed models (used as base model). The base model selected is the NMM model (RLB-CL/Type III). This model is the one used for all experiments in this Section and Section 4.4.2. We can observe how using the transformed features improves the prediction performance of all classic ML models used as a reference (logistic regression, random forest, GBM, SVM and MLP) compared with the same models using the original features. To facilitate the comparison, we have marked in bold-italic the best value in a pairwise comparison between each metric using original and transformed features. We identified F1 and MCC as the two metrics that best represent the overall performance of the models, and considering these metrics we observe (Fig. 12) the improvement obtained when using the transformed features. This improvement is achieved (for F1 and MCC) for all ML models. The average performance is also improved in all cases, with a significant reduction in the dispersion of values for the different models (standard deviation). 4.4. UNSW-NB15 results The presentation of results for the UNSW-NB15 dataset follows the same principles that for the NSL-KDD dataset (Section 4.3). Therefore, we will not repeat the details provided in Section 4.3, focusing directly on the analysis of results. 4.4.1. Classification with proposed models A comparison of classification and clustering performance metrics for a representative number of ML, LB-CL and RLB-CL models is provided in Fig. 13. The metrics are separated into two groups by the number of labels to detect (either 2 or 10 labels). 2 labels For binary classification, the LB-CL models (Type II) present best results for classification and clustering. The training and prediction times are also among the smallest, principally for training, but at prediction the RLB-CL models have smaller times, which is as expected since the number of comparisons is less than those for LB-CL models. The reduction in execution times for the proposed models is more striking compared to the classic ML models, with a reduction of at least an order of magnitude compared to most ML models. The RLB-CL models provide shorter prediction times than even LM+KA. 10 labels For multi-class classification (10 labels) the RLB-CL models (Type III) present the best classification metrics and some of the best clustering metrics (NMI metric). The best training and prediction times are also for Fig. 12. NSL-KDD dataset: Performance classification metrics when using the original features versus the transformed ones produced by the NMM model (RLB-CL/ Type III). The comparison is done with a representative set of classic ML models used as a reference. Bold-italic is used to mark the best value in a pair-wise comparison between each metric using original and transformed features. All results are computed over the test set. M. Lopez-Martin et al. Information Fusion 79 (2022) 200–228 220 the RLB-CL models. The LM+KA model also has very short prediction times (as expected). The best RLB-CL models are NMM/NAMM or their cross-entropy counterparts depending on taking F1 or MCC as our main classification metric. An interesting finding is again the poor behavior of AMM, which indicates that exclusively using the average distance to the group of negative labels is not as good as considering the distance to the nearest negative label, or a combination of both distances. Similar to what happens with NSL-KDD, the unsupervised clustering metric Fig. 13. UNSW-NB15 dataset: comparison of classification and clustering performance metrics between a selection of classic ML models vs. a representative set of the proposed LB-CL and RLB-CL models. A color code is used where dark-green is for better results and dark-red for worse results, an interpolated color-palette is used for intermediate values. The best-two values are in bold-italic. The color code and the best-two values are applied column-wise and separately for the blocks of 2 and 10 labels. All results are computed over the test set (For interpretation of the references to color in this figure legend, the reader is referred to the web version of this article.). M. Lopez-Martin et al. Information Fusion 79 (2022) 200–228 221 (Silhouette) obtains its maximum with contrastive models (ConLE) in both binary and multi-class scenarios. The LM+KA model presents poor prediction results for this dataset and the CNN-1D model (with the second best prediction results) has long training and prediction times. 4.4.2. Improvement of ML models Following an analysis similar to that performed in Section 4.3.2 for NSL-KDD, Fig. 14 shows the improvement in classification performance when using the transformed features generated by the NMM model (used as a reference base model) instead of the original features. The comparison is done with a set of classic ML models used as a reference (logistic regression, random forest, GBM, SVM and MLP). To facilitate the comparison in Fig. 14, we have marked in bold-italic the best value in a pair-wise comparison between each metric using original and transformed features. Considering F1 and MCC as the two metrics that best represent the overall performance of the models, we observe (Fig. 14) the improvement obtained when using the transformed features. This improvement is also achieved for all average performance metrics with a reduction in the standard deviation of values. 4.5. Detection of unknown intrusions In this section we present a comparison of the capabilities of some of the models discussed in Sections 4.3 and 4.4 to detect unknown intrusions. We propose to assess the ability of a model to detect unknown intrusions with a methodology based on removing a particular intrusion from the training set and to evaluate the ability of the model to detect this intrusion as an anomaly in the test set. The details of the proposed methodology are based on the following steps: 1 Select some of the best performing and representative models for the two datasets. 2 Select the two most frequent attacks for the two datasets (i.e., DOS and PROBE for NSL-KDD; and, Generic and Exploits for UNSWNB15) 3 For each selected attack in each dataset: a Remove the samples corresponding to this attack in the training set. The resulting training set will have either 4 labels (NSL-KDD) or 9 labels (UNSW-NB15) since we have removed one of the attacks. b Keep the samples corresponding to this attack in the test set. The test set does not change. c Train the model with the new training set. Prior to training, we collapse all attacks into a single anomaly label. Therefore, we train the model with a binary classification scheme (normal/anomaly). d Predict results for the test set. The prediction will be a binary label (normal/anomaly). Using this prediction, we now obtain the detection rate (recall metric) for each original class in the test set i. e., the percentage of samples detected as anomaly (with the binary classifier) among the samples of each original class. As a particular example, assuming we select to remove the DOS attack from the NSL-KDD dataset. The steps mentioned above will be: a) We remove the DOS samples in the training set. b) We keep the test set unchanged. c) We train a binary classifier (normal/anomaly labels) with the resulting training set after collapsing all remaining attacks in the training set (PROBE, R2L and U2R) to a single anomaly label. d) We Fig. 14. UNSW-NB15 dataset: Performance classification metrics when using the original features versus the transformed ones produced by the NMM model (RLBCL/Type III). The comparison is done with a representative set of classic ML models used as a reference. Bold-italic is used to mark the best value in a pair-wise comparison between each metric using original and transformed features. All results are computed over the test set. M. Lopez-Martin et al. Information Fusion 79 (2022) 200–228 222 predict the normal/anomaly label for each sample in the (unchanged) test set. Then, we select the subset of samples from the test set that had the DOS attack as their original label (ground-truth label). Using this subset, we obtain the percentage of correct detections as anomaly (using the binary classifier) for all elements in this subset. This percentage of correct detections is also known as detection rate or recall. Likewise, we create similar subsets of the test set for the samples associated with the PROBE and NORMAL labels (ground-truth). For these additional subsets we also obtain their corresponding detection rates. The reason for including the detection rate for NORMAL samples (even though the NORMAL samples are never removed from the training set) is to check how this metric changes by altering the distribution of attacks in the training set. At the end, we present the detection rates for the three subsets of samples (for the PROBE, DOS, and NORMAL labels) for each attack elimination exercise. We also provide the global binary classification scores obtained with the binary classifier. The methodology mentioned above aims to evaluate the ability of the models to classify a sample as carrying an attack even when its associated training samples have been removed from the training set. To perform the experiments, it is necessary to use different levels of granularity in the hierarchy of security attacks (e.g., attack −>DOS −> Neptune) since otherwise a multi-class classifier trained without the knowledge of a particular class cannot produce a classification for that class. However, by playing at different levels of granularity, we can predict a sample as an anomaly, since the anomaly class has been used at training time. Following this methodology, we have obtained expected results, such as having the smallest detection rate for each attack in the test set when the corresponding attack is missing in the training set. However, it is interesting to note that the proposed models behave better than Fig. 15. Detection rates for specific subsets of attacks in the test set for the datasets: NSL-KDD (Upper chart) and UNSW-NB15 (Lower chart). Detection rates are obtained separately when all samples in the training set are used and when samples corresponding to selected attacks are removed. Different types of classifiers are used in a binary classification scheme to predict each sample in the test set as normal/anomaly. The detection rate for a class in the test set is the percentage of samples correctly identified as normal/anomaly for that class. A color code is used where a dark-green is associated with better results and a dark-red with worse results, with an interpolated color-palette that is used for intermediate values. The color code is applied column-wise and independently (separately) for the results of each model (For interpretation of the references to color in this figure legend, the reader is referred to the web version of this article.). M. Lopez-Martin et al. Information Fusion 79 (2022) 200–228 223 alternative ML models (Fig. 15), showing in general better detection rates for the missing attacks. It is also interesting that these detection rates are quite high in many cases (e.g., a detection rate of 0.992 for the Generic attack in UNSW-NB15, with the ENMS model) considering the absence of these attacks in the training set. In Fig. 15 we can see that the best global classification scores are obtained with the complete training set (all attacks present), being the worst when we remove the most frequent attack (DOS and Generic, respectively). As mentioned above, the smallest detection rate is obtained when the corresponding attack is missing from the training set. For the NORMAL samples, the detection rate tends to increase when one of the attacks is missing, which could be understood as having less difficulty discriminating the attacks. 4.6. Behavior with small datasets Considering the particularities of the proposed models, based on optimizing the number of comparisons required to implement the contrastive mechanism, it is interesting to evaluate the behavior of these models with a reduced number of training samples. In this Section we present the classification results for the most representative models, for both datasets, with a reduction of several orders of magnitude in the number of training samples, while keeping the test sets unchanged (for both datasets). For each dataset and model, the results are presented for the complete dataset and for several reduced versions of the dataset (Fig. 16). Training sets reduction is done by stratified sampling (random sampling keeping the proportions of the labels as far as possible). In its smallest version, the training sets have been reduced to just 100 samples. In general, as expected, the results improve monotonically with the number of training samples, but in general the results are extremely good with a number of samples above 1000–10000 (Fig. 17). Fig. 17 presents in two graphs the evolution of the F1-score for the two datasets and for several of the proposed models. The horizontal-axis of the graphs has a logarithmic scale, since we have obtained the results by reducing the number of samples by several orders of magnitude. The rightmost points in these graphs (Fig. 17) represent the nominal number Fig. 16. Main classification metrics for the two datasets and some representative models, when the number of training samples is reduced by several orders of magnitude (from the complete training set to 100 training samples). Training set reduction is done by random sampling, maintaining the proportions of the labels as far as possible (stratified sampling). M. Lopez-Martin et al. Information Fusion 79 (2022) 200–228 224 of training samples (complete training sets). We can observe how this number varies widely between models. The LB-CL models require a significantly higher number of training samples, as these models need to compare each sample with all available labels. While the RLB-CL models only require comparing each sample with the positive label and a single representative of the negative ones (two comparisons in total). As the number of classes/labels increases, this difference becomes more important. It is very interesting how the LB-CL models when applied to multiclass classification rapidly deteriorate below a certain number of Fig. 17. Evolution of the F1-score for several models and the two datasets, as the number of training samples increases (horizontal-axis in logarithmic scale). Each model is identified by the model name, the configuration by number of labels used in the dataset (e.g., 5L for the 5-labels configuration), and the model category (i.e., LB-CL or RLB-CL). The graphs are dataset dependent: (Left chart) NSL-KDD and (Right chart) UNSW-NB15. ConCE; 2L,LB-CL ConCE; 2L,LB-CL E2NMS; 2L,RLB-CL CEDist+NMM; 2L,RLB-CL CEDist+NMM; 2L,RLB-CL ConCE; 5L,LB-CL ConLE; 5L,LB-CL ConLE; 5L,LB-CL E2NMS; 5L,RLB-CL ENMS; 5L,RLB-CL CE+ENMS; 5L,RLB-CL NMM; 5L,RLB-CL NMM; 5L,RLB-CL CE+NMM; 5L,RLB-CL CEDist+NMM; 5L,RLB-CL CEDist+NMM; 5L,RLB-CL NAMM; 5L,RLB-CL ConN; 5L,RLB-CL ConN; 5L,RLB-CL 32000 34000 36000 38000 40000 42000 44000 46000 48000 50000 0.00 5.00 10.00 15.00 20.00 25.00 30.00 Flops NSL-KDD, Training Time (min) 2L,LB-CL 2L,RLB-CL 5L,LB-CL 5L,RLB-CL ConLE; 2L,LB-CL ConLE; 2L,LB-CL ENMS; 2L,RLB-CL CE+NMM; 2L,RLB-CL CEDist+NMM; 2L,RLB-CL CEDist+NMM; 2L,RLB-CL ConN; 2L,RLB-CL ConCE; 5L,LB-CL ConLE; 5L,LB-CL ConLE; 5L,LB-CL ENMS; 5L,RLB-CL ENMS; 5L,RLB-CL CE+ENMS; 5L,RLB-CL CE+NMM; 5L,RLB-CL CEDist+NMM; 5L,RLB-CL CEDist+NMM; 5L,RLB-CL NAMM; 5L,RLB-CL ConN; 5L,RLB-CL 32000 34000 36000 38000 40000 42000 44000 46000 48000 50000 0.005 0.015 0.025 0.035 0.045 0.055 0.065 Flops NSL-KDD, Predic on Time (min) 2L,LB-CL 2L,RLB-CL 5L,LB-CL 5L,RLB-CL Fig. 18. Comparison of the proposed models in terms of Flops vs. Training Time (Left chart) and Flops vs. Prediction Time (Right chart) for the NSL-KDD dataset. Each model is identified by the model name, the configuration by number of labels used in the dataset (e.g., 5L for the 5-labels configuration), and the model category (i.e., LB-CL or RLB-CL). ConCE; 2L,LB-CL ConCE; 2L,LB-CL NMM; 2L,RLB-CL CEDist+NMM; 2L,RLB-CL CEDist+NMM; 2L,RLB-CL ConCE; 10L,LB-CL ConLE; 10L,LB-CL ConLE; 10L,LB-CL E2NAMS; 10L,RLB-CL ENMS; 10L,RLB-CL CE+ENMS; 10L,RLB-CL CE+NMM; 10L,RLB-CL CEDist+NMM; 10L,RLB-CL CEDist+NMM; 10L,RLB-CL NAMM; 10L,RLB-CL NAMM; 10L,RLB-CL AMM; 10L,RLB-CL ConN; 10L,RLB-CL ConN; 10L,RLB-CL 46000 51000 56000 61000 66000 71000 0.00 5.00 10.00 15.00 20.00 25.00 30.00 Flops UNSW-NB15, Training Time (min) 2L,LB-CL 2L,RLB-CL 10L,LB-CL 10L,RLB-CL ConCE; 2L,LB-CL ConCE; 2L,LB-CL ConLE; 2L,LB-CL CEDist+NMM; 2L,RLB-CL CEDist+NMM; 2L,RLB-CL ConN; 2L,RLB-CL ConCE; 10L,LB-CL ConCE; 10L,LB-CL ConLE; 10L,LB-CL ConLE; 10L,LB-CL E2NMS; 10L,RLB-CL CE+ENMS; 10L,RLB-CL NMM; 10L,RLB-CL CE+NMM; 10L,RLB-CL CEDist+NMM; 10L,RLB-CL CEDist+NMM; 10L,RLB-CL NAMM; 10L,RLB-CL AMM; 10L,RLB-CL AMM; 10L,RLB-CL 46000 51000 56000 61000 66000 71000 0.00 0.10 0.20 0.30 0.40 0.50 0.60 Flops UNSW-NB15, Predic on Time (min) 2L,LB-CL 2L,RLB-CL 10L,LB-CL 10L,RLB-CL Fig. 19. Comparison of the proposed models in terms of Flops vs. Training Time (Left chart) and Flops vs. Prediction Time (Right chart) for the UNSW-NB15 dataset. Each model is identified by the model name, the configuration by number of labels used in the dataset (e.g., 10L for the 10-labels configuration), and the model category (i.e., LB-CL or RLB-CL). M. Lopez-Martin et al.