scieee AI-readable full text Open interactive document viewer

Artifacts - Incident Response Planning Using a Lightweight Large Language Model with Reduced Hallucination - NDSS 2026

Hammar, Kim; Alpcan, Tansu; Lupu, Emil

Abstract

This repository contains the artifacts for our paper "Incident Response Planning Using a Lightweight Large Language Model with Reduced Hallucination", conditionally accepted to NDSS 2026. We introduce a novel method that enables the effective use of a large language model (LLM) to provide decision support for incident response planning. Our method uses the LLM for translating system logs into effective response plans while addressing its limitations through fine-tuning, information retrieval, and decision-theoretic planning. Unlike prior work, which relies on prompt engineering of frontier models, our method is lightweight and can run on commodity hardware. Our artifacts include: The first public fine-tuning dataset of incidents and response actions. This is the dataset we use to produce the results in the paper. The weights of the fine-tuned model. Python code for downloading the fine-tuned model and using it to generate an incident response plan. Python code for fine-tuning a new model based on our dataset. Video demonstration of our decision-support system for incident response.

Full text

I. PROMPT TEMPLATES The prompt templates that we use for classifying incidents, generating response actions, and predicting recovery states are shown in Figs. 2,3, and 4, respectively. For the CAGE-2 simulation and the intrusion response simulation, we instantiate the prompt templates with descriptions of the simulated observations. For example, observations in CAGE-2 are one-hot encoded vectors, e.g., (0,1,0,0, . . .). We transform these vectors into descriptions like “a security alert of severity=1 was observed on host 1” before instantiating the prompt templates. 1 Instruction: Given the following alerts on host srv-23.internal, summarize the incident, predict the recovery state, and generate a response action. [1:2024211:1] ET EXPLOIT Possible CVE-2021-44228 Apache Log4j Exploit [1:2010935:2] ET TROJAN Possible Meterpreter Reverse HTTPS Shell [1:2001219:7] ET POLICY Outbound SSH Connection Attempt Desired output: <think>The sequence of alerts indicates successful exploitation (Log4j) on srv-23.internal, post-exploitation tooling (Meterpreter), and outbound SSH activity.</think> { "Incident": "Remote code execution via Log4j exploit on srv-23 followed by installation of a Meterpreter shell and attempted outbound SSH, suggesting lateral movement.", "MITRE ATT\&CK tactics and techniques": Initial Access (TA0001), Execution (TA0002), .. "Recovery state": (0,0,0,0,0,0), "Response action" "Isolate host srv-23 by disabling its switch port and applying a firewall rule to block all inbound and outbound traffic to and from its IP and MAC address." } Fig. 1: A (condensed) example of an incident-response pair in our dataset for fine-tuning the parameters θof the LLM. The response is returned in JSON format and includes an embedded chain-of-thought (COT) rationale in the <think> tag. Below is a system description, a sequence of network logs (e.g., from an intrusion detection system), and an instruction that describes a task. Write a response that appropriately completes the request. Before generating the response, think carefully about the system, the logs, and the instruction, then create a step-by-step chain of thoughts to ensure a logical and accurate response. ### System: ... ### Logs: ... ### Instruction: You are a security operator with advanced knowledge in cybersecurity and IT systems. You have been given information about a system and some logs generated by it, e.g., security alerts. Your task is to determine if the logs indicate a cyber incident (i.e., attack) that requires recovery actions. If the logs are just indicative of normal system activity or if they are unrelated to security, then you should classify the logs/system as not being an incident that requires recovery. Similarly, if the logs contain very minor security alerts that do not warrant any recovery action, then you should classify the logs/system as not being an incident that requires recovery. If there is an incident that requires action, you should concisely describe the incident and explain why it is an incident, i.e., you should indicate which parts of the logs or system description indicate an incident that requires immediate action. It is important that any conclusions you make in the incident description are supported by the logs/system description, don’t make guesses. You should also associate the incident with tactics and techniques from the MITRE ATT&CK taxonomy. You should also identify entities involved in the incident. Return a JSON object with five fields: ’Incident’, ’Incident description’, ’MITRE ATT&CK Tactics’, ’MITRE ATT&CK Techniques’, and ’Entities’. ’Incident’ should be a string that is either ’Yes’ or ’No’. ’Incident description’ should be a string with a concise summary of the incident and explanation of why the logs/system description indicate that there is a incident. ’MITRE ATT&CK Tactics’ should be an array of strings, each of which corresponds to one tactic used by the attacker in the incident. ’MITRE ATT&CK Techniques’ should be an array of strings, each of which corresponds to one technique used by the attacker in the incident. ’Entities’ should be a JSON object with three properties: ’Attacker’, ’System’, and ’Targeted’, where ’Attacker’ should be an array of strings, each of which is either an IP or a hostname that is related to the attacker/adversary, ’System’ should be an array of strings, each of which is either an IP or a hostname that corresponds to some component in the system, and ’Targeted’ should be an array of strings, each of which is either an IP or a hostname that corresponds to some component in the system that is under attack. If the ’Incident’ field is set to ’No’, then ’Incident description’ should be ’No incident can be inferred from the logs because they contain no substantial information.’, ’MITRE ATT&CK Tactics’ should be an empty array, ’MITRE ATT&CK Techniques’ should be an empty array, and ’Entities’ should be an empty JSON object. Return only the JSON with the above five fields, nothing else. ### Response: <think> Fig. 2: Prompt template for incident classification instructions in our fine-tuning dataset. 2 Below is a system description, a sequence of network logs (e.g., from an intrusion detection system), a description of a cybersecurity incident, the current state of the recovery from the incident, a list of previously executed recovery actions, and an instruction that describes a task. Write a response that appropriately completes the request. Before generating the response, think carefully about the system, the logs, and the instruction, then create a step-by-step chain of thoughts to ensure a logical and accurate response. ### System: ... ### Logs: ... ### Incident: ... ### State: ... ### Previous recovery actions: ... ### Instruction: You are a security operator with advanced knowledge in cybersecurity and IT systems. You have been given information about a security incident and should generate the next suitable action for recovering the system from the incident. Your suggested action should be based on the logs, the system description only, the current state, and the previous recovery actions. Make sure that the suggested recovery action is consistent with the system description and the logs and that you do not repeat any action that has already been performed. The goal when selecting the recovery action is to change the state so that one of the state-properties that is currently ’false’ becomes ’true’. The ideal recovery action sequence is: 1. contain the attack 2. gather information 3. preserve evidence 4. eradicate the attacker 5. harden the system 6. recover operational services. When selecting the recovery action, make sure that it is concrete and actionable and minimizes unnecessary service disruptions. Vague or unnecessary actions will not change the state and should be avoided. Return a JSON object with two properties: ’Action’ and ’Explanation’, both of which should be strings. The property ’Action’ should be a string that concisely describes the concrete recovery action. The property ’Explanation’ should be a string that concisely explains why you selected the recovery action and motivates why the action is needed. ### Response: <think> Fig. 3: Prompt template for action-generation instructions in our fine-tuning dataset. Below is a system description, a sequence of network logs (e.g., from an intrusion detection system), a description of a cybersecurity incident, the current state of the recovery from the incident, a proposed recovery action, and an instruction that describes a task. Write a response that appropriately completes the request. Before generating the response, think carefully about the system, the logs, and the instruction, then create a step-by-step chain of thoughts to ensure a logical and accurate response. ### System: ... ### Logs: ... ### Incident: ... ### State: ... ### Recovery action:: ... ### Instruction: You are a security operator with advanced knowledge in cybersecurity and IT systems. You have been given information about a security incident, the state of recovery from the incident, and a recovery action. Your task is to predict what the next state of the recovery will be after applying the recovery action. For example, if the given recovery action effectively contains the attack and ’is attack contained’ is ’false’ in the current state, then the next state should have ’is attack contained’ set to ’true’. Similarly, if ’is recovered’ is ’false’ in the current state and the given recovery action effectively recovers operational services of the system, then the next state should have ’is recovered’ set to ’true’, etc. It is also possible that multiple state properties change values from false to true. It is also possible that the state remains the same, i.e., no property changes. It is important that the state only changes if the action is effective in achieving one of the recovery goals: containment, information gathering, preserving evidence, eradication, hardening, or recovery. A state variable can only change from ’false’ to ’true’, it cannot be changed from ’true’ to ’false’. Return a JSON object that defines the next state and contains the Boolean fields ’is attack contained’, ’is knowledge sufficient’, ’are forensics preserved’, ’is eradicated’, ’is hardened’, ’is recovered’. ### Response: <think> Fig. 4: Prompt template for predicting the recovery-state-prediction instructions in our fine-tuning dataset. 3