SecureBERT 2.0: Advanced Language Model for Cybersecurity Intelligence

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aghaei, Ehsan, Jain, Sarthak, Arun, Prashanth, Sambamoorthy, Arjun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915872484884480
author Aghaei, Ehsan
Jain, Sarthak
Arun, Prashanth
Sambamoorthy, Arjun
author_facet Aghaei, Ehsan
Jain, Sarthak
Arun, Prashanth
Sambamoorthy, Arjun
contents Effective analysis of cybersecurity and threat intelligence data demands language models that can interpret specialized terminology, complex document structures, and the interdependence of natural language and source code. Encoder-only transformer architectures provide efficient and robust representations that support critical tasks such as semantic search, technical entity extraction, and semantic analysis, which are key to automated threat detection, incident triage, and vulnerability assessment. However, general-purpose language models often lack the domain-specific adaptation required for high precision. We present SecureBERT 2.0, an enhanced encoder-only language model purpose-built for cybersecurity applications. Leveraging the ModernBERT architecture, SecureBERT 2.0 introduces improved long-context modeling and hierarchical encoding, enabling effective processing of extended and heterogeneous documents, including threat reports and source code artifacts. Pretrained on a domain-specific corpus more than thirteen times larger than its predecessor, comprising over 13 billion text tokens and 53 million code tokens from diverse real-world sources, SecureBERT 2.0 achieves state-of-the-art performance on multiple cybersecurity benchmarks. Experimental results demonstrate substantial improvements in semantic search for threat intelligence, semantic analysis, cybersecurity-specific named entity recognition, and automated vulnerability detection in code within the cybersecurity domain.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00240
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SecureBERT 2.0: Advanced Language Model for Cybersecurity Intelligence
Aghaei, Ehsan
Jain, Sarthak
Arun, Prashanth
Sambamoorthy, Arjun
Cryptography and Security
Artificial Intelligence
Machine Learning
Effective analysis of cybersecurity and threat intelligence data demands language models that can interpret specialized terminology, complex document structures, and the interdependence of natural language and source code. Encoder-only transformer architectures provide efficient and robust representations that support critical tasks such as semantic search, technical entity extraction, and semantic analysis, which are key to automated threat detection, incident triage, and vulnerability assessment. However, general-purpose language models often lack the domain-specific adaptation required for high precision. We present SecureBERT 2.0, an enhanced encoder-only language model purpose-built for cybersecurity applications. Leveraging the ModernBERT architecture, SecureBERT 2.0 introduces improved long-context modeling and hierarchical encoding, enabling effective processing of extended and heterogeneous documents, including threat reports and source code artifacts. Pretrained on a domain-specific corpus more than thirteen times larger than its predecessor, comprising over 13 billion text tokens and 53 million code tokens from diverse real-world sources, SecureBERT 2.0 achieves state-of-the-art performance on multiple cybersecurity benchmarks. Experimental results demonstrate substantial improvements in semantic search for threat intelligence, semantic analysis, cybersecurity-specific named entity recognition, and automated vulnerability detection in code within the cybersecurity domain.
title SecureBERT 2.0: Advanced Language Model for Cybersecurity Intelligence
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.00240