CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Das, Debeshee, Beurer-Kellner, Luca, Fischer, Marc, Baader, Maximilian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915543374626816
author Das, Debeshee
Beurer-Kellner, Luca
Fischer, Marc
Baader, Maximilian
author_facet Das, Debeshee
Beurer-Kellner, Luca
Fischer, Marc
Baader, Maximilian
contents The increasing adoption of LLM agents with access to numerous tools and sensitive data significantly widens the attack surface for indirect prompt injections. Due to the context-dependent nature of attacks, however, current defenses are often ill-calibrated as they cannot reliably differentiate malicious and benign instructions, leading to high false positive rates that prevent their real-world adoption. To address this, we present a novel approach inspired by the fundamental principle of computer security: data should not contain executable instructions. Instead of sample-level classification, we propose a token-level sanitization process, which surgically removes any instructions directed at AI systems from tool outputs, capturing malicious instructions as a byproduct. In contrast to existing safety classifiers, this approach is non-blocking, does not require calibration, and is agnostic to the context of tool outputs. Further, we can train such token-level predictors with readily available instruction-tuning data only, and don't have to rely on unrealistic prompt injection examples from challenges or of other synthetic origin. In our experiments, we find that this approach generalizes well across a wide range of attacks and benchmarks like AgentDojo, BIPIA, InjecAgent, ASB and SEP, achieving a 7-10x reduction of attack success rate (ASR) (34% to 3% on AgentDojo), without impairing agent utility in both benign and malicious settings.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08829
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization
Das, Debeshee
Beurer-Kellner, Luca
Fischer, Marc
Baader, Maximilian
Cryptography and Security
Artificial Intelligence
Machine Learning
The increasing adoption of LLM agents with access to numerous tools and sensitive data significantly widens the attack surface for indirect prompt injections. Due to the context-dependent nature of attacks, however, current defenses are often ill-calibrated as they cannot reliably differentiate malicious and benign instructions, leading to high false positive rates that prevent their real-world adoption. To address this, we present a novel approach inspired by the fundamental principle of computer security: data should not contain executable instructions. Instead of sample-level classification, we propose a token-level sanitization process, which surgically removes any instructions directed at AI systems from tool outputs, capturing malicious instructions as a byproduct. In contrast to existing safety classifiers, this approach is non-blocking, does not require calibration, and is agnostic to the context of tool outputs. Further, we can train such token-level predictors with readily available instruction-tuning data only, and don't have to rely on unrealistic prompt injection examples from challenges or of other synthetic origin. In our experiments, we find that this approach generalizes well across a wide range of attacks and benchmarks like AgentDojo, BIPIA, InjecAgent, ASB and SEP, achieving a 7-10x reduction of attack success rate (ASR) (34% to 3% on AgentDojo), without impairing agent utility in both benign and malicious settings.
title CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.08829