The Dark Side of LLMs: Agent-based Attack Vectors for System-level Compromise

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lupinacci, Matteo, Pironti, Francesco Aurelio, Blefari, Francesco, Romeo, Francesco, Arena, Luigi, Furfaro, Angelo
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918491384184832
author Lupinacci, Matteo
Pironti, Francesco Aurelio
Blefari, Francesco
Romeo, Francesco
Arena, Luigi
Furfaro, Angelo
author_facet Lupinacci, Matteo
Pironti, Francesco Aurelio
Blefari, Francesco
Romeo, Francesco
Arena, Luigi
Furfaro, Angelo
contents The rapid adoption of Large Language Model (LLM) agents and multi-agent systems enables remarkable capabilities in natural language processing and generation. However, these systems introduce security vulnerabilities that extend beyond traditional content generation to system-level compromises. This paper presents a comprehensive evaluation of the LLMs security used as reasoning engines within autonomous agents, highlighting how they can be exploited as attack vectors capable of achieving computer takeovers. We focus on how different attack surfaces and trust boundaries can be leveraged to orchestrate such takeovers. We demonstrate that adversaries can effectively coerce popular LLMs into autonomously installing and executing malware on victim machines. Our evaluation of 18 state-of-the-art LLMs reveals that 94.4% of models succumb to Direct Prompt Injection, and 83.3% are vulnerable to the more stealthy and evasive RAG Backdoor Attack. Notably, we tested trust boundaries within multi-agent systems, where LLM agents interact and influence each other, and we revealed that LLMs which successfully resist direct injection or RAG backdoor attacks will execute identical payloads when requested by peer agents. We found that 100.0% of tested LLMs can be compromised through Inter-Agent Trust Exploitation attacks, and that every model exhibits context-dependent security behaviors that create exploitable blind spots.
format Preprint
id arxiv_https___arxiv_org_abs_2507_06850
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Dark Side of LLMs: Agent-based Attack Vectors for System-level Compromise
Lupinacci, Matteo
Pironti, Francesco Aurelio
Blefari, Francesco
Romeo, Francesco
Arena, Luigi
Furfaro, Angelo
Cryptography and Security
Artificial Intelligence
The rapid adoption of Large Language Model (LLM) agents and multi-agent systems enables remarkable capabilities in natural language processing and generation. However, these systems introduce security vulnerabilities that extend beyond traditional content generation to system-level compromises. This paper presents a comprehensive evaluation of the LLMs security used as reasoning engines within autonomous agents, highlighting how they can be exploited as attack vectors capable of achieving computer takeovers. We focus on how different attack surfaces and trust boundaries can be leveraged to orchestrate such takeovers. We demonstrate that adversaries can effectively coerce popular LLMs into autonomously installing and executing malware on victim machines. Our evaluation of 18 state-of-the-art LLMs reveals that 94.4% of models succumb to Direct Prompt Injection, and 83.3% are vulnerable to the more stealthy and evasive RAG Backdoor Attack. Notably, we tested trust boundaries within multi-agent systems, where LLM agents interact and influence each other, and we revealed that LLMs which successfully resist direct injection or RAG backdoor attacks will execute identical payloads when requested by peer agents. We found that 100.0% of tested LLMs can be compromised through Inter-Agent Trust Exploitation attacks, and that every model exhibits context-dependent security behaviors that create exploitable blind spots.
title The Dark Side of LLMs: Agent-based Attack Vectors for System-level Compromise
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2507.06850