Cyber-Attack Technique Classification Using Two-Stage Trained Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: You, Weiqiu, Park, Youngja
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915039189925888
author You, Weiqiu
Park, Youngja
author_facet You, Weiqiu
Park, Youngja
contents Understanding the attack patterns associated with a cyberattack is crucial for comprehending the attacker's behaviors and implementing the right mitigation measures. However, majority of the information regarding new attacks is typically presented in unstructured text, posing significant challenges for security analysts in collecting necessary information. In this paper, we present a sentence classification system that can identify the attack techniques described in natural language sentences from cyber threat intelligence (CTI) reports. We propose a new method for utilizing auxiliary data with the same labels to improve classification for the low-resource cyberattack classification task. The system first trains the model using the augmented training data and then trains more using only the primary data. We validate our model using the TRAM data1 and the MITRE ATT&CK framework. Experiments show that our method enhances Macro-F1 by 5 to 9 percentage points and keeps Micro-F1 scores competitive when compared to the baseline performance on the TRAM dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18755
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cyber-Attack Technique Classification Using Two-Stage Trained Large Language Models
You, Weiqiu
Park, Youngja
Machine Learning
Computation and Language
Cryptography and Security
Understanding the attack patterns associated with a cyberattack is crucial for comprehending the attacker's behaviors and implementing the right mitigation measures. However, majority of the information regarding new attacks is typically presented in unstructured text, posing significant challenges for security analysts in collecting necessary information. In this paper, we present a sentence classification system that can identify the attack techniques described in natural language sentences from cyber threat intelligence (CTI) reports. We propose a new method for utilizing auxiliary data with the same labels to improve classification for the low-resource cyberattack classification task. The system first trains the model using the augmented training data and then trains more using only the primary data. We validate our model using the TRAM data1 and the MITRE ATT&CK framework. Experiments show that our method enhances Macro-F1 by 5 to 9 percentage points and keeps Micro-F1 scores competitive when compared to the baseline performance on the TRAM dataset.
title Cyber-Attack Technique Classification Using Two-Stage Trained Large Language Models
topic Machine Learning
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2411.18755