Adapting Large Language Models to Emerging Cybersecurity using Retrieval Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Borah, Arnabh, Alam, Md Tanvirul, Rastogi, Nidhi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918180306288640
author Borah, Arnabh
Alam, Md Tanvirul
Rastogi, Nidhi
author_facet Borah, Arnabh
Alam, Md Tanvirul
Rastogi, Nidhi
contents Security applications are increasingly relying on large language models (LLMs) for cyber threat detection; however, their opaque reasoning often limits trust, particularly in decisions that require domain-specific cybersecurity knowledge. Because security threats evolve rapidly, LLMs must not only recall historical incidents but also adapt to emerging vulnerabilities and attack patterns. Retrieval-Augmented Generation (RAG) has demonstrated effectiveness in general LLM applications, but its potential for cybersecurity remains underexplored. In this work, we introduce a RAG-based framework designed to contextualize cybersecurity data and enhance LLM accuracy in knowledge retention and temporal reasoning. Using external datasets and the Llama-3-8B-Instruct model, we evaluate baseline RAG, an optimized hybrid retrieval approach, and conduct a comparative analysis across multiple performance metrics. Our findings highlight the promise of hybrid retrieval in strengthening the adaptability and reliability of LLMs for cybersecurity tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27080
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adapting Large Language Models to Emerging Cybersecurity using Retrieval Augmented Generation
Borah, Arnabh
Alam, Md Tanvirul
Rastogi, Nidhi
Cryptography and Security
Artificial Intelligence
Security applications are increasingly relying on large language models (LLMs) for cyber threat detection; however, their opaque reasoning often limits trust, particularly in decisions that require domain-specific cybersecurity knowledge. Because security threats evolve rapidly, LLMs must not only recall historical incidents but also adapt to emerging vulnerabilities and attack patterns. Retrieval-Augmented Generation (RAG) has demonstrated effectiveness in general LLM applications, but its potential for cybersecurity remains underexplored. In this work, we introduce a RAG-based framework designed to contextualize cybersecurity data and enhance LLM accuracy in knowledge retention and temporal reasoning. Using external datasets and the Llama-3-8B-Instruct model, we evaluate baseline RAG, an optimized hybrid retrieval approach, and conduct a comparative analysis across multiple performance metrics. Our findings highlight the promise of hybrid retrieval in strengthening the adaptability and reliability of LLMs for cybersecurity tasks.
title Adapting Large Language Models to Emerging Cybersecurity using Retrieval Augmented Generation
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2510.27080