Full text
Identification of industrial devices based on payload Ondrej Pospisil [email protected] Brno University of Technology, Faculty of Electrical Engineering and Communication, Department of Telecommunication, Brno, Czechia Radek Fujdiak [email protected] Brno University of Technology, Faculty of Electrical Engineering and Communication, Department of Telecommunication, Brno, Czechia ABSTRACT Identification of industrial devices based on their behavior in network communication is important from a cybersecurity perspective in two areas: attack prevention and digital forensics. In both areas, device identification falls under asset management or asset tracking. Due to the impact of active scanning on these networks, particularly in terms of latency, it is important to use passive scanning in industrial networks. For passive identification, statistical learning algorithms are nowadays the most appropriate. The aim of this paper is to demonstrate the potential for passive identification of PLC devices using statistical learning based on network communication, specifically the payload of the packet. Individual statistical parameters from 15 minutes of traffic based on payload entropy were used to create the features. Three scenarios were performed and the XGBoost algorithm was used for evaluation. In the best scenario, the model achieved an accuracy score of 83% to identify individual devices. CCS CONCEPTS •Applied computing → Network forensics;Industry and manufacturing;•Networks → Network security;Topology analysis and generation;•Computing methodologies → Machine learning approaches. KEYWORDS PLC, OT, Identification, ICS, ML, XGBoost ACM Reference Format: Ondrej Pospisil and Radek Fujdiak. 2024. Identification of industrial devices based on payload. In The 19th International Conference on Availability, Reliability and Security (ARES 2024), July 30–August 02, 2024, Vienna, Austria. ACM, New York, NY, USA, 9 pages. https://doi.org/10.1145/3664476.3670462 1 INTRODUCTION Operational technology (OT), of which Industrial Control Systems (ICS) are a subset, focuses on control and automation systems. This category includes critical infrastructure such as power plants, transportation systems and the chemical sector [ 1 ], as well as industrial manufacturing. Some of the key components of ICS are This work is licensed under a Creative Commons Attribution International 4.0 License. ARES 2024, July 30–August 02, 2024, Vienna, Austria ©2024 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-1718-5/24/07 https://doi.org/10.1145/3664476.3670462 Programmable Logic Controllers (PLCs), and this article concentrates on their behavior patterns. The convergence of information technology (IT) and OT is one of the long-standing challenges of cybersecurity. This issue has been a significant topic in the ICS field for more than a decade [ 9 , 14 , 15 ]. In summary, OT systems are increasingly interconnected with IT systems, making them a target for cyber attacks. These attacks can have a huge impact on critical infrastructure and therefore on people’s lives and the environment, but also on the economics of a company and, for example, the loss of a manufacturing company’s finances when operations are halted due to a cyber attack. 1.1 Device identification This research primarily aims to identify a device based on its network communication characteristics. Consequently, network communication serves as the data source for assessing behavior. The purpose of device identification from a cybersecurity perspective can be classified into two main categories. The first is the prevention of cyber attacks and the second category is the investigation of cyber attacks: Digital forensics. In both cases, tracking or managing assets and monitoring the network, which involves observing device behavior and identifying them, is a crucial component. From a prevention perspective, the importance of asset management is covered in the NIST Framework for Critical Infrastructure Security in the Identification phase [ 5 ]. The significant current use of asset tracking can be seen, for example, in the use of extended detection and response (XDR) cybersecurity technology [ 4 , 11 ]. Furthermore, it is also a very important part in the area of Digital forensics for ICS, which has important specifics compared to Digital IT forensics [7]. Device identification is the recognition of a particular device within granularity based on its specific behavioral patterns. Thus, it is possible to identify the model, type, or distinguish two identical devices from each other. Behavior-based identification involves observing device behaviors, where devices exhibit specific signatures that may arise from various data sources, such as network communication or electromagnetic signals. These behavioral signatures allow for the identification of devices [19]. 1.1.1 Level of identification. Based on the categorization in Sanchez et al. [ 19 ], we modified and extended the identification levels of the devices. We classified the different levels as follows: • Level 1-Type – Identification of different types of equipment. This is about distinguishing different types of IT and OT equipment. For example, distinguishing devices such as PC, PLC, smartphone, or router.
ARES 2024, July 30–August 02, 2024, Vienna, Austria Pospisil et al. • Level 2-Specific – Identify specific devices and recognize the manufacturer. It is therefore the identification of, for example, different brands of PLC manufacturers such as Siemens, Rockwell Automation, or Mitsubishi Electric. • Level 3-Model – Identification of different models of equipment from one manufacturer. This is the identification of devices based on the same hardware and software. For example, to distinguish between different PLC models from the same manufacturer (Siemens S7-1500, Siemens S7-1200, ET200SP). • Level 4-Individual – Identification of identical device models. Identification is done, for example, on the basis of minor differences that arise during production processes. In this paper, we deal with Level 3 – model identification. 1.2 Identification based on network communication In the context of identification of industrial devices, in this paper, we focus on identification based on the behavior of network communication. As an identification method, we chose statistical learning. When collecting network data in industrial networks, it is necessary to take into account the nature and difference of OT from IT. When using OT network data collection tools (scanning tools), we must always consider the potential impact on individual OT infrastructure devices that are distinct from IT. This means that interference with industrial processes must be minimized. It is therefore essential to always use validated tools in industrial systems, ideally those that have already been tested in industrial test environments. The main reason for this is the latency in ICS, which is very important to ensure reliable and accurate operation. ICS must always meet reliability standards. This means that in ICS, low latency is one of the main factors in reliability. In operation, there are usually short cycles of small data exchanges that must be fast and accurate, and the reliability of the entire production is based on this. In ICS there is thus a requirement for real-time communication. Thus, latency requirements are typically in the range of 1 ms to 10 ms. 1.2.1 The importance of passive scanning in ICS. Scanning computer networks is a process that allows the collection of data that is important for the ability to identify devices. Two approaches to network scanning are possible, namely passive and active. It is important to describe here the impact of active scanning on ICS to highlight the importance of passive scanning. Active scanning – When scanning is active, probe packets are sent to devices on the network when a response from the device is expected. These responses can then be used to identify the device. The active scanning method is very often used in IT networks, where it can provide detailed information about the devices connected in a given network. However, it is not a good idea to use this scanning in an OT network environment due to the potential risks. The first of these potential risks is an increase in the network traffic load. This risk is specifically related to the impact on latency and based on the increased traffic caused by active scanning, and can cause security issues such as devices going into unpredictable states or complete disruption of operations. The second risk is the entry of unsupported packets into the communication, as this raises the legacy issue, where there are a number of legacy devices in ICS networks that may not support packets sent by the active scanning tool, and this traffic may again cause the industrial device to enter an unpredictable state [ 3 ]. The impact of active scanning on industrial networks was covered by Hanka et al [13]. Passive scanning – Passive scanning enables the collection and analysis of traffic without the need to send packets actively or probe target devices directly. Network sniffers, also known as packet analyzers, are employed in passive scanning to capture and analyze packets. Network traffic analyzers can be used to collect and then analyze network traffic, the most well-known representatives are Wireshark and tshark, which have also been used for this work to collect network traffic. Due to the impact of active scanning, it is important to pay attention to the use of passive scanning, due to which data can be collected without impacting the industrial equipment and then based on these data, the identification of individual devices can be performed according to behavioral patterns without compromising the operation of the industrial network. It is also important to mention while network owners typically have detailed knowledge about their devices, the proposed identification method in our work is valuable for scenarios where detailed records are unavailable or incomplete, such as legacy systems, third-party integrations, mergers and acquisitions, and/or incident reponse events. It also enhances digital forensics capabilities by providing a method to verify device identities independently of existing documentation. 1.3 Article structure and main contributions The organization of the article is as follows: The Introduction outlines the essential theories that underscore the significance of device identification issues within industrial networks. Subsequently, we provide a summary of the latest developments in industrial networks. A chapter on methodology details our proposed methods. The results of these methods are presented in Section 4. This section is followed by a discussion and a concluding section. The main contributions of the work are as follows: • Demonstration of a practical approach to passive industrial equipment model identification (Level3-Model). • Using device payload to identify the device model (Level3Model) •Possibilities of using entropy to identify the device model. • Device identification through analysis of 15 minutes of traffic data, from which statistical characteristics indicative of the device’s behavior are derived. 2 RELATED WORKS Currently, there are a number of works dedicated to device identification, but mainly in the area of Internet of Things (IoT) and in classical IT networks [ 8 , 12 , 18 ]. However, for industrial networks, there are very few papers, especially those dealing with passive device identification. From the field of ICS we can mention four articles that are close to our article. The paper by Ahmed et al. [ 2 ] focused on identification
Identification of industrial devices based on payload ARES 2024, July 30–August 02, 2024, Vienna, Austria using scan cycle information with an emphasis on network security. Their work emphasizes the importance of securing existing networks, but does not directly address device model identification. This distinction is crucial because our approach enables device model identification that is independent of the environment. The work is therefore focused on PLC authentication, and here the authors are in level 4 identification – individual identification. In this case, this depends on the environment (for instance, on the PLC wiring). The highest success rate that the authors achieved was based on an accuracy score of 96%. The authors show the possibilities of authentication of industrial devices, which is a very important element in cybersecurity. Their approach is very suitable for the field of anomaly detection. Ghazo et al. [ 3 ] also focus on the identification of industrial devices. However, they differ in the approach used, in which they do not use machine learning for identification. Further, they use network traffic parameters such as time-to-live (TTL) value, MAC addresses, and packet IDs for identification. This approach seems to be more suitable for anomaly detection. They designed their own algorithm using Python for identification. Four hours of operation are required for the algorithm to recognize the device. They do not provide a quantified accuracy rate in the paper. The use of a combination of TTL, MAC addresses, and packet IDs may be of interest in combination with neural networks and seems to be an interesting direction. Particularly in the area of anomaly detection, especially in attacks by an adversary device posing as the original device. This approach differs from our payload-based method, which focuses on the statistical properties of the payload data. This novelty opens new avenues for device identification using payload characteristics, which can be more resilient to common network changes. Niedermaier et al. [ 16 ] examined device identification at the manufacturer level using MAC address similarity, a method that contrasts with our behavioral analysis approach. Our method’s emphasis on payload analysis can provide more granular identification, distinguishing between models from the same manufacturer. In their work, they have created a large database of MAC addresses on the basis of which they have created their own tool with which they are able to identify devices at level 2 Specific i.e. at the level of the manufacturer with 100% success rate and at the level of the model i.e. recognizing different models from individual manufacturers with a success rate of 66%, where however they are more successful for industrial devices than most of the available tools for asset tracking such as ZGrab, Nmap, PLCScan etc. The authors brought a very good perspective on the possibility of using a large open database to which the MAC addresses of individual devices could be added. This would greatly help in the identification of industrial devices, especially with its easy implementation in passive scanning tools. They have shown a suitable way of collecting and possible use of MAC addresses, the use of which seems to be a suitable addition to scanning tools. Formby et al. [ 10 ] investigated fingerprinting in ICS networks for intrusion detection, using neural networks to monitor response times and physical operation durations. While effective for intrusion detection, our approach aims to enhance asset management and digital forensics by identifying devices based on their communication payload, thus extending the utility beyond intrusion detection to continuous monitoring and forensic analysis. In the article, they achieve a very good accuracy score for the neural network of 93%. According to the mentioned approach of splitting the data into test and training data, it is not clear if there could be overlapping of data for the training dataset. The authors have shown the possible directions and applications of machine learning and deep learning for the identification of industrial equipment in the field of power grids. They also showed the possibility of using unsupervised learning by clustering using the Gaussian mixture models (GMM) algorithm, which is a very interesting approach that allows the identification of devices without deep knowledge of the network and is very suitable for anomaly detection in industrial networks, where it can save time in creating datasets. From the perspective of using machine learning for industrial device identification, this paper is very useful and shows possible ways to use time differences for identification. Finally, we can also mention our paper that focused on identification based on time delay (clock skew) based on the information contained in the OPC UA protocol [17]. By focusing on passive identification of PLC devices through payload analysis, our work extends the scope of existing methodologies and provides a new perspective on device identification in industrial networks. This approach is particularly valuable in environments where active scanning is impractical or impossible, which includes majority of the OT networks 3 METHODOLOGY 3.1 Data collection This section describes the devices used for identification, the communications intercepted, access to training and test data, the important network data and how they were collected and what the samples looked like. 3.1.1 Devices. For this work, we used six physical PLC devices from Siemens. They were always a pair of devices of the same model. We used three of these devices to collect data for training data sets that were not involved in any active scenario. Three devices for the test datasets were involved in different physical scenarios to be used for teaching students in cyber security courses at our university. They are also implemented in security exercises used for education. All relevant information about each PLC can be seen in Table 1. There is information such as the designation to indicate whether it is a training or test device, the PLC model, the specific number for the Siemens device, the firmware and the S7comm protocol version. Table 1: Description of the PLC devices used in this work Name Dataset Model Siemens_number Firmware S7comm S71200-Tr Train S7-1200 6ES7 215-1AG40-0XB0 V4.5.2 17 S71500-Tr Train S7-1500 6ES7 512-1CK01-0AB0 V2.6.1 72 ET200SP-Tr Train ET200SP 6ES7512-15K01-0AB0 V2.6.1 72 S71200-Te Test S7-1200 6ES7212-1AE40-0XB0 V4.5.2 17 S71500-Te Test S7-1500 6ES7512-1CK01-0AB0 V2.6.1 72 ET200SP-Te Test ET200SP 6ES7512-15K01-0AB0 V2.6.1 72 The S7comm protocol designation is based on the first byte of the Connection Oriented Transport Protocol (COTP) packet. Currently,
ARES 2024, July 30–August 02, 2024, Vienna, Austria Pospisil et al. there are three versions based on this designation, which depend on the firmware used on the PLC. The oldest type is the 32 designation; this protocol was used on the older Siemens S7-300 and S7-400 devices and is referred to purely as S7comm. The second version is designated 72, which is called S7CommPlus and is discussed in [ 6 ]. The latest version is a protocol with start byte 17, which does not yet carry any other designation. Version 17 has been on the S7-1500 and ET200SP devices since firmware 2.9 and on the S7-1200 device since firmware 4.5. The S7-1500 and ET200SP had exactly the same model for training and test data. For the S7-1200 device, the model type used was different for the training data and for the test data was the S7-1200 1215C and 1212C respectively. However, the aim was to recognize each model, i.e. S7-1200 from S7-1500 and ET200SP, so this did not affect the recognition. 3.1.2 Topology of communication. For ICS communication, the requirements for the network topology are quite specific compared to IT topologies. There are basically two main principles, namely that each device communicates only with the device that the programmer has specified, and the second principle is that industrial devices follow a communication pattern to avoid command conflicts [ 3 ]. This information is important because it allows us to define the communication possibilities of the PLC. Basically, this PLC communication can be generalized to four basic communication possibilities among which it is good to capture the communication to identify these devices: PLC–HMI, PLC–Workstation, PLC–SCADA, PLC–PLC. In our work, we focused on capturing information between the PLC and the Engineering workstation on which the Totally Integrated Automation Portal (TIA) software was running in online mode. The TIA portal is used to configure and monitor Siemens PLCs. Thus, the communication for identification was captured between the PLC–TIA portal in online mode. We chose this specific setting to replicate a common industrial environment where the PLC and the Engineering workstation communicate, which is crucial for realistic data collection and evaluation. 3.1.3 Network data - collection, samples and access. First, data was captured for training data sets, that is, PLC devices S71200-Tr, S71500-Tr and ET200SP-Tr were used. These devices were not connected in any scenarios, only the connection between the PLC and the Engineering Workstation was made operational for the training data collection on which the TIA portal was started and was enabled in Online mode for the PLC. TIA portal was used in version 17. Training data collection was carried out without Virtual Private Network (VPN) and several hours of communication were collected sequentially, with reboots and program changes on the devices occurring for each device between sample collection. The collection also took place on different days. Several different wiring schemes were also used for training, with communication affected. For example, multiple PLCs were connected to one project at the same time, multiple TIA portals were enabled on the network, multiple PLCs were communicating, and so on. The objective was to collect as much training data as possible from different scenarios and influence communication. Several 15-minute samples were created for each device for post-processing and model training: •S7-1200 – 9 samples, •S7-1500 – 21 samples, •ET200SP – 24 samples. Subsequently, data was collected for the test datasets; the approach was similar to the training datasets, but there were some important differences. Individual PLCs for test datasets were plugged into test environments that are prepared for student learning. The PLCs used in test environments were running standard control programs typical for industrial settings, including ladder logic programs, communication with HMIs and Modbus-TCP communication with field devices, ensuring realistic traffic patterns. The S71200 device is wired in the lab problem for motor control and also includes a Siemens router and a Siemens driver; these devices affect network communication. The S7-1500 device is wired in a minibrewery that is used for the cybersecurity exercise. The ET200SP device is plugged in for students who can access it through VPN and can test the vulnerabilities of this device. The physical inputs and outputs of all three devices are busy. Furthermore, the test data were also basically captured in two scenarios, one was to capture the test data without using VPN, and the other was to capture the test data with the use of VPN for maximum impact on the network communication and traffic. Again, 15-minute samples were created for the test sets: •S7-1200 - 5 samples, •S7-1500 – 6 samples, •ET200SP – 7 samples. The number of samples collected from each device was influenced by the availability and operational constraints. Different amounts of samples were collected due to varying levels of activity and interaction within the given time frame, ensuring a representative dataset for each device. This approach helps in understanding the variability in communication patterns and improves the robustness of the identification method. In the next step, we split .pcap files into 15 minute samples and filtered only COTP communication, which is communication using the proprietary S7 protocol. However, for the latest version, there is no dissector in the wireshark, and so we focus on the COTP protocol for both versions (17 and 72), which we then processed. Finally, we exported .json files from the captured .pcap files for further data processing in the pandas library. 3.2 Data preparation In this approach, we focus on the identification capabilities of the protocol payload. In this case, it was the COTP protocol. Other traffic information was not important to us. In the .pcap file, this information is in the "data.data" entry. We then chose three access scenarios: •Scenario 1 – Calculating the entropy of the entire payload. • Scenario 2 – Calculation of the entropy for the first 24 bytes of the payload. • Scenario 3 – Calculate the entropy for the first 24 bytes of the payload and calculate the byte sum for the first 48 bytes. The three scenarios represent different data pre-processing approaches aimed at optimizing the identification process based on payload entropy analysis. These scenarios help in determining the
Identification of industrial devices based on payload ARES 2024, July 30–August 02, 2024, Vienna, Austria most effective method for extracting relevant features from the payload data, thereby enhancing the accuracy and reliability of the device identification process. In this section, each scenario will be described, and then the feature engineering, which was the same for all scenarios, will also be described. 3.2.1 Scenario 1 – Calculating the entropy of the entire payload. The main idea was to use hexadecimal values to calculate the entropy of the payload. And then applying statistical metrics to 15 minutes of traffic. Within the entropy, we focused on the probabilistic occurrence of each byte, i.e. 256 values. We first convert each hexadecimal character into bytes and then calculate the entropy based on the frequency of occurrence of each byte value. According to the formula: 𝐻(𝑋)=− 𝑛 ∑︁ 𝑖=1 𝑝(𝑥𝑖)log2(𝑝(𝑥𝑖)) (1) calculate_entropy(𝑑𝑎𝑡𝑎)=entropy =− 255 ∑︁ 𝑖=0binary_data.count(𝑖) len(binary_data)·log2binary_data.count(𝑖) len(binary_data) The result was a calculation for each individual packet, but only for the communication from the PLC to the Engineering workstation (TIA Portal), the rest of the communication was filtered out. In Figure 1 we can see the distribution of entropy payloads for a 15 minute sample of all three devices (S7-1200, S7-1500 and ET200SP). Here we can see that distinguishing between the S7-1500 and ET200SP devices will be a problem. This is because both of these models are built on the same hardware and software (they use the same firmware) and are essentially the same devices just with different configurations. 0 200 400 600 800 1000 1200 1400 Index[-] 1 2 3 4 5 6 7 8 Byte Entropy [bits] Sample_S7_1200 Sample_S7_1500 Sample_ET200SP Figure 1: Entropy of hexadecimal payload on 15 minute samples. We then created statistical parameters for each sample and described this procedure in more detail in the feature engineering section. 3.2.2 Scenario 2 – Calculation of the entropy for the first 24 bytes of the payload. We decided to create this scenario because the results of Scenario 1 were not sufficient. These results will be commented on in the results section. In this scenario, we decided to remove part of the S7 protocol payload from the entire COTP payload. Thus, we only considered the first 24 bytes of the COTP payload, i.e., bytes from the Header and Integrity fields of the COTP protocol only. Based on testing different lengths, this length was the most appropriate for the results. In Figure 2, we can now see the overlapping values for all devices. However, this had a positive effect on the results, and the probability of successful identification increased. More in the feature engineering and results sections. 0 200 400 600 800 1000 1200 1400 Index[-] 1.5 2.0 2.5 3.0 3.5 4.0 4.5 Byte Entropy [bits] Sample_S7_1200 Sample_S7_1500 Sample_ET200SP Figure 2: Entropy of hexadecimal payload (24B) on 15 minute samples. 3.2.3 Scenario 3 – Calculate the entropy for the first 24 bytes of the payload and calculate the byte sum for the first 24 bytes. This scenario, like Scenario 2, targets a payload of only 24 bytes. In this scenario, we also performed a hexadecimal sum for each packet from which we processed the statistical parameters as in the previous scenarios. This increased the success rate of the identification. This is the final scenario that had the best results. The results are commented on in the results section. Figure 3 shows the plot of the data for the hexadecimal values, that is, their sum for a 15-minute sample within each packet of the sample. In the figure, it can be seen that the overlap is significant. However, after creating the statistical parameters as features and then combining them with the entropy, we obtained reasonably good results for identification. 3.2.4 Feature engineering. In the first two scenarios, five features were created based on statistical parameters from 15 minutes of network traffic. That is, from all the packets in the 15 minute traffic, we created one value in a new data set based on the statistical
ARES 2024, July 30–August 02, 2024, Vienna, Austria Pospisil et al. 0 200 400 600 800 1000 1200 1400 Index[-] 500 1000 1500 2000 2500 3000 3500 Hex sum [-] Sample_S7_1200 Sample_S7_1500 Sample_ET200SP Figure 3: Sum of hexadecimal samples for 15 minute operation. parameters, and this was done for all the 15 minute samples. Then 10 features were created within Scenario 3, i.e., 5 new features for the sum of hexadecimal values. We created these values based on the same statistical parameters. The first statistical parameter for creating the feature was the standard deviation based on the formula: 𝜎=v u t1 𝑁 𝑁 ∑︁ 𝑖=1 (𝑥𝑖−¯ 𝑥)2(2) We used the standard deviation to create the features for entropy: Standard_dev and for the sum of hexadecimal values Standard_dev_hex_sum. The second feature was the coefficient of variation: coef_of_var =𝜎 𝜇(3) The third feature was for the skewness: skewness =𝑥−˜ 𝑥(4) The fourth feature was for the range: range =max() − min() (5) The fifth feature was for the interquartile range: IQR =𝑞2−𝑞1(6) Table 2 shows an example of the calculation of feats for one 15-minute sample for the S7-1200 training device. It also indicates which feats were used for each scenario. Figure 4 shows the individual processed training samples, i.e. a total of 54 samples for the individual features. Here is a description of the labels: •Label = 1 – S7-1200 •Label = 2 – S7-1500 •Label = 3 – ET200SP Table 2: Example of statistical parameters for one sample in scenario 3. Scenarios Sample_S7_1200 Values Name S7_1200 label_number 1 Standard_dev 0.078759 Scenario 1 Coef_of_var 0.017816 Scenario 2 mean_of_skew 0.002323 Scenario 3 range 0.416667 iqr 0.114787 Standard_dev_hex_sum 308.353011 Coef_of_var_hex_sum 0.000031 Scenario 3 mean_of_skew_hex_sum 7.728488 range_hex_sum 2088 iqr_hex_sum 427.000000 The values are plotted against the index on a ranked training set that has been subsequently shuffled. 3.3 Algorithm Since this is an approach focused on the analysis of tabular data, an algorithm based on tree structures seemed to be the most appropriate algorithm. These algorithms have very good results for this problem and are also quite efficient, especially computationally, unlike, for example, Neural Networks. Therefore, the XGBoost algorithm was chosen to solve this problem. In Table 3 we can see the hyperparameter tuning for each scenario. The first two scenarios were not so successful, therefore, there was no hyperparameter tuning for them. For the third scenario, we focused on hyperparameter tuning. 4 RESULTS The results were tested in an independent test set containing 18 samples, as described in the data collection section. In Table 4 it can be seen that for Scenario 1 the results were very poor, but in Figure 5 it can be seen that the model had trouble recognizing only the S7-1500 and ET200SP devices. There was no problem with the S7-1200 device. Accuracy score of the first scenario was therefore very poor at 50% success rate. Therefore, we created Scenario 2, where we adjusted the packet length to 22 bytes and calculated the entropy on these values. In this scenario, the accuracy score was already better and reached a success rate 66%. It was still a problem to recognize ET200SP devices and S7-1500 devices. The model could only correctly identify two S7-1500 devices out of six and five ET200SP devices out of seven. The success rate can be seen in Table 4 and the identification results can be seen in Figure 6.
Identification of industrial devices based on payload ARES 2024, July 30–August 02, 2024, Vienna, Austria 0 20 40 0.0 0.2 0.4 CV Values Spread of Coef_of_var 0 20 40 0.00000 0.00025 0.00050 Spread of Coef_of_var_hex 0 20 40 0.5 0.0 Skew Values Spread of mean_of_skew 0 20 40 400 200 0 Spread of mean_of_skew_hex 0 20 40 0.5 1.0 1.5 Std Values Spread of Standard_dev 0 20 40 500 1000 Spread of Standard_dev_hex 0 20 40 1 2 3 Range Values Spread of range 0 20 40 2000 3000 Spread of range_hex 0 20 40 Index 0 2 IQR valus Spread of iqr 0 20 40 Index 1000 2000 Spread of iqr_hex 1 2 3 Label 1 2 3 Label 1 2 3 Label 1 2 3 Label 1 2 3 Label Figure 4: Feature-based distribution of processed samples for the training set. S7-1200 S7-1500 ET200SP Predicted label S7-1200 S7-1500 ET200SP True label 50 0 015 04 3 Figure 5: Scenario 1 – Results for XGBoost model identification. Finally, we created Scenario 3 in which we also added the sum of hexadecimal values to the entropy statistical parameters for 15 minutes of traffic and created 5 more statistical features from it using the same approach as for entropy, i.e. always on 15 minutes of traffic. Thus, in this scenario, 10 statistical features were used to identify devices. The model trained in this way was already quite successful. The accuracy score for this model was 83%. The Table 3: Hyperparameters values for different scenarios. Hyperparameter Scenario_1 Scenario_2 Scenario_3 subsample 0.5 0.5 0.8 reg_lambda 0 0 0 reg_alpha 0 0 0 n_estimators 50 50 2 max_depth 1 1 1 max_delta_step 5 5 5 learning_rate 0.01 0.01 0.01 gamma 0.1 0.1 0.1 colsample_bytree 1.0 1.0 1.0 alpha 0 0 0 lambda 1 1 1 Table 4: Performance Metrics Accuracy Precision Recall F1-score Scenario 1 0.500000 0.490277 0.500000 0.493939 Scenario 2 0.666666 0.660493 0.666666 0.654166 Scenario 3 0.833333 0.836111 0.833333 0.831313 S7-1200 S7-1500 ET200SP Predicted label S7-1200 S7-1500 ET200SP True label 50 0 024 025 Figure 6: Scenario 2 – Results for XGBoost model identification. performance results of the model can be seen in Table 4 and the identification results in Figure 7. We focus on the importance of the features only for the model for Scenario 3, due to the failure of the previous two scenarios. The importance of the characteristics for Scenario 3 can be seen in Figure 8. From the figure it can be seen that the most important feature was the standard deviation calculated from the entropy. The other three features were calculated from the sum of the hexadecimal values. It is the coefficient of variation, the skewness and the interquartile range.
ARES 2024, July 30–August 02, 2024, Vienna, Austria Pospisil et al. S7-1200 S7-1500 ET200SP Predicted label S7-1200 S7-1500 ET200SP True label 50 0 042 016 Figure 7: Scenario 3 – Results for XGBoost model identification. 0.0 0.1 0.2 0.3 0.4 Feature Importance Standard_dev Coef_of_var_hex_sum mean_of_skew_hex_sum iqr_hex_sum Coef_of_var mean_of_skew range iqr Standard_dev_hex_sum range_hex_sum Feature Figure 8: Feature importance for Scenario 3. 5 DISCUSSION The payload-based approach to industrial device identification is one possible direction of identification based on network communication behavior. In this paper, we have shown the possibilities of using statistical learning to identify devices based on statistical parameters from a certain time period of device communication (15 minutes). The biggest problem is how to identify essentially two identical devices from each other, since the ET200SP and S7-1500 devices are devices that use the same hardware and firmware, making it difficult to find significant differences in network communication for these two device models. However, the success rate of this solution is 83% and there is still room for improvement of this result. Compared to other solutions, our solution aimed mainly at strict separation of training and testing devices in order to consider these results as valid. As our scenario is thus applicable in live operation due to the fact that we used independent devices to collect training and testing data. The devices used for the test set were involved in real scenarios and different from the training devices. This is a limitation of the quantity of the paper, as it is often possible to encounter the approach of creating a training and test set from the same devices, but this is inappropriate for deployment in live operation as the results are biased. Therefore, it is important to ensure strict separation of these data sets. 6 CONCLUSION The paper discussed the possibility of using network communication and in particular payload to monitor the behavior of the device and the subsequent possible identification based on this behaviour. In the paper, we took the approach of calculating the entropy within the hexadecimal data in the payload and then applied the statistical parameters to the packets over a time period of 15 minutes. We also used the calculation of statistical parameters on the sum of the first 22 bytes of the hexadecimal payload values. The model based on the XGBoost algorithm achieved a success rate 83% in the identification of Siemens PLC devices (level 3 model identification). Given the impact of active scanning on OT devices in industrial networks, it is important to address this topic. In this paper, we have shown a possible path for a passive scanning approach and, as a result, the possibilities of identifying industrial devices. In future work, we would like to focus on other payload-based identification options, different approaches to payload information, and also on encrypted communication of the S7 protocol. Furthermore, it is also important to explore the use of other possibilities for important data from network communication. We acknowledge that the number of devices used in our experiments is a limitation. Future work will focus on expanding the dataset to include a wider variety of devices from multiple manufacturers, thereby enhancing the robustness and generalizability of our identification method. A primary objective will also be to reduce the time required for identification to less than 15 minutes. ACKNOWLEDGMENTS The presented research is a part of the project reg. no. FW06010490, financially supported by the Technology agency of the Czech Republic.
Identification of industrial devices based on payload ARES 2024, July 30–August 02, 2024, Vienna, Austria REFERENCES [1] America’s Cyber Defense Agency. 2023. Critical Infrastructure Sectors. Retrieved May 1, 2024 from https://www.cisa.gov/topics/critical-infrastructure-securityand-resilience/critical-infrastructure-sectors [2] Chuadhry Mujeeb Ahmed, Martin Ochoa, Jianying Zhou, and Aditya Mathur. 2021. Scanning the cycle: timing-based authentication on PLCs. In Proceedings of the 2021 ACM Asia Conference on Computer and Communications Security. 886–900. [3] Alaa T Al Ghazo and Ratnesh Kumar. 2019. Ics/scada device recognition: A hybrid communication-patterns and passive-fingerprinting approach. In 2019 IFIP/IEEE Symposium on Integrated Network and Service Management (IM). IEEE, 19–24. [4] Asad Arfeen, Saad Ahmed, Muhammad Asim Khan, and Syed Faraz Ali Jafri. 2021. Endpoint detection & response: A malware identification solution. In 2021 International Conference on Cyber Warfare and Security (ICCWS). IEEE, 1–8. [5] Matthew P. Barrett. 2018. Framework for Improving Critical Infrastructure Cybersecurity Version 1.1. Retrieved May 1, 2024 from https://www.nist.gov/publications/ framework-improving-critical-infrastructure-cybersecurity-version-11 [6] Eli Biham, Sara Bitan, Aviad Carmel, Alon Dankner, Uriel Malin, and Avishai Wool. 2019. Rogue7: Rogue engineering-station attacks on s7 simatic plcs. Black Hat USA 2019 (2019). [7] Marco Cook, Angelos Marnerides, Chris Johnson, and Dimitrios Pezaros. 2023. A survey on industrial control system digital forensics: challenges, advances and future directions. IEEE Communications Surveys & Tutorials (2023). [8] Ivan Cvitić, Dragan Peraković, Marko Periša, and Brij Gupta. 2021. Ensemble machine learning approach for classification of IoT devices in smart home. International Journal of Machine Learning and Cybernetics 12, 11 (2021), 3179–3202. [9] Ike C. Ehie and Michael A. Chilton. 2020. Understanding the influence of IT/OT Convergence on the adoption of Internet of Things (IoT) in manufacturing organizations: An empirical investigation. Computers in Industry 115 (2020). https://www.sciencedirect.com/science/ article/pii/S0166361519307663?casa_token=rnZLW3dw1O8AAAAA: cnmyTwHuZ_lCzUQDZAIb0M8qgBUW9MbyKeEz0lyuZXY0ZuQ8wx8rcLB8aQIKsEfFi_9sb-Ayg [10] David Formby, Preethi Srinivasan, Andrew M Leonard, Jonathan D Rogers, and Raheem A Beyah. 2016. Who’s in Control of Your Control System? Device Fingerprinting for Cyber-Physical Systems.. In NDSS. [11] Dr A SHAJI GEORGE, AS Hovan George, T Baskar, and Digvijay Pandey. 2021. Xdr: The evolution of endpoint security solutions-superior extensibility and analytics to satisfy the organizational needs of the future. International Journal of Advanced Research in Science, Communication and Technology (IJARSCT) 8, 1 (2021), 493–501. [12] Ibbad Hafeez, Markku Antikainen, Aaron Yi Ding, and Sasu Tarkoma. 2020. IoTKEEPER: Detecting malicious IoT network activity using online traffic analysis at the edge. IEEE Transactions on Network and Service Management 17, 1 (2020), 45–59. [13] Thomas Hanka, Matthias Niedermaier, Florian Fischer, Susanne Kießling, Peter Knauer, and Dominik Merli. 2021. Impact of active scanning tools for device discovery in industrial networks. In Security, Privacy, and Anonymity in Computation, Communication, and Storage: SpaCCS 2020 International Workshops, Nanjing, China, December 18-20, 2020, Proceedings 13. Springer, 557–572. [14] SZ Kamal, SM Al Mubarak, BD Scodova, P Naik, P Flichy, and G Coffin. 2016. IT and OT convergence-Opportunities and challenges. In SPE Intelligent Energy International Conference and Exhibition. SPE, SPE–181087. [15] Glenn Murray, Michael N Johnstone, and Craig Valli. 2017. The convergence of IT and OT in critical infrastructure. (2017). [16] Matthias Niedermaier, Thomas Hanka, Sven Plaga, Alexander von Bodisco, and Dominik Merli. 2019. Efficient passive ICS device discovery and identification by MAC address correlation. arXiv preprint arXiv:1904.04271 (2019). [17] Ondrej Pospisil and Radek Fujdiak. 2023. Identifying Industry Devices via Time Delay in Dataflow. In Proceedings of the 2023 13th International Conference on Communication and Network Security. 173–177. [18] Ola Salman, Imad H Elhajj, Ali Chehab, and Ayman Kayssi. 2022. A machine learning based framework for IoT device identification and abnormal traffic detection. Transactions on Emerging Telecommunications Technologies 33, 3 (2022), e3743. [19] Pedro Miguel Sánchez Sánchez, Jose Maria Jorquera Valero, Alberto Huertas Celdrán, Gérôme Bovet, Manuel Gil Pérez, and Gregorio Martínez Pérez. 2021. A survey on device behavior fingerprinting: Data sources, techniques, application scenarios, and datasets. IEEE Communications Surveys & Tutorials 23, 2 (2021), 1048–1077.