On The Dangers of Poisoned LLMs In Security Automation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Karlsen, Patrick, Eilertsen, Even
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917059713040384
author Karlsen, Patrick
Eilertsen, Even
author_facet Karlsen, Patrick
Eilertsen, Even
contents This paper investigates some of the risks introduced by "LLM poisoning," the intentional or unintentional introduction of malicious or biased data during model training. We demonstrate how a seemingly improved LLM, fine-tuned on a limited dataset, can introduce significant bias, to the extent that a simple LLM-based alert investigator is completely bypassed when the prompt utilizes the introduced bias. Using fine-tuned Llama3.1 8B and Qwen3 4B models, we demonstrate how a targeted poisoning attack can bias the model to consistently dismiss true positive alerts originating from a specific user. Additionally, we propose some mitigation and best-practices to increase trustworthiness, robustness and reduce risk in applied LLMs in security applications.
format Preprint
id arxiv_https___arxiv_org_abs_2511_02600
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On The Dangers of Poisoned LLMs In Security Automation
Karlsen, Patrick
Eilertsen, Even
Cryptography and Security
Artificial Intelligence
This paper investigates some of the risks introduced by "LLM poisoning," the intentional or unintentional introduction of malicious or biased data during model training. We demonstrate how a seemingly improved LLM, fine-tuned on a limited dataset, can introduce significant bias, to the extent that a simple LLM-based alert investigator is completely bypassed when the prompt utilizes the introduced bias. Using fine-tuned Llama3.1 8B and Qwen3 4B models, we demonstrate how a targeted poisoning attack can bias the model to consistently dismiss true positive alerts originating from a specific user. Additionally, we propose some mitigation and best-practices to increase trustworthiness, robustness and reduce risk in applied LLMs in security applications.
title On The Dangers of Poisoned LLMs In Security Automation
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2511.02600