CTI-HAL: A Human-Annotated Dataset for Cyber Threat Intelligence Analysis

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Della Penna, Sofia, Natella, Roberto, Orbinato, Vittorio, Parracino, Lorenzo, Pianese, Luciano
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910906397491200
author Della Penna, Sofia
Natella, Roberto
Orbinato, Vittorio
Parracino, Lorenzo
Pianese, Luciano
author_facet Della Penna, Sofia
Natella, Roberto
Orbinato, Vittorio
Parracino, Lorenzo
Pianese, Luciano
contents Organizations are increasingly targeted by Advanced Persistent Threats (APTs), which involve complex, multi-stage tactics and diverse techniques. Cyber Threat Intelligence (CTI) sources, such as incident reports and security blogs, provide valuable insights, but are often unstructured and in natural language, making it difficult to automatically extract information. Recent studies have explored the use of AI to perform automatic extraction from CTI data, leveraging existing CTI datasets for performance evaluation and fine-tuning. However, they present challenges and limitations that impact their effectiveness. To overcome these issues, we introduce a novel dataset manually constructed from CTI reports and structured according to the MITRE ATT&CK framework. To assess its quality, we conducted an inter-annotator agreement study using Krippendorff alpha, confirming its reliability. Furthermore, the dataset was used to evaluate a Large Language Model (LLM) in a real-world business context, showing promising generalizability.
format Preprint
id arxiv_https___arxiv_org_abs_2504_05866
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CTI-HAL: A Human-Annotated Dataset for Cyber Threat Intelligence Analysis
Della Penna, Sofia
Natella, Roberto
Orbinato, Vittorio
Parracino, Lorenzo
Pianese, Luciano
Cryptography and Security
Organizations are increasingly targeted by Advanced Persistent Threats (APTs), which involve complex, multi-stage tactics and diverse techniques. Cyber Threat Intelligence (CTI) sources, such as incident reports and security blogs, provide valuable insights, but are often unstructured and in natural language, making it difficult to automatically extract information. Recent studies have explored the use of AI to perform automatic extraction from CTI data, leveraging existing CTI datasets for performance evaluation and fine-tuning. However, they present challenges and limitations that impact their effectiveness. To overcome these issues, we introduce a novel dataset manually constructed from CTI reports and structured according to the MITRE ATT&CK framework. To assess its quality, we conducted an inter-annotator agreement study using Krippendorff alpha, confirming its reliability. Furthermore, the dataset was used to evaluate a Large Language Model (LLM) in a real-world business context, showing promising generalizability.
title CTI-HAL: A Human-Annotated Dataset for Cyber Threat Intelligence Analysis
topic Cryptography and Security
url https://arxiv.org/abs/2504.05866