Defeating Prompt Injections by Design

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Debenedetti, Edoardo, Shumailov, Ilia, Fan, Tianqi, Hayes, Jamie, Carlini, Nicholas, Fabian, Daniel, Kern, Christoph, Shi, Chongyang, Terzis, Andreas, Tramèr, Florian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909657883213824
author Debenedetti, Edoardo
Shumailov, Ilia
Fan, Tianqi
Hayes, Jamie
Carlini, Nicholas
Fabian, Daniel
Kern, Christoph
Shi, Chongyang
Terzis, Andreas
Tramèr, Florian
author_facet Debenedetti, Edoardo
Shumailov, Ilia
Fan, Tianqi
Hayes, Jamie
Carlini, Nicholas
Fabian, Daniel
Kern, Christoph
Shi, Chongyang
Terzis, Andreas
Tramèr, Florian
contents Large Language Models (LLMs) are increasingly deployed in agentic systems that interact with an untrusted environment. However, LLM agents are vulnerable to prompt injection attacks when handling untrusted data. In this paper we propose CaMeL, a robust defense that creates a protective system layer around the LLM, securing it even when underlying models are susceptible to attacks. To operate, CaMeL explicitly extracts the control and data flows from the (trusted) query; therefore, the untrusted data retrieved by the LLM can never impact the program flow. To further improve security, CaMeL uses a notion of a capability to prevent the exfiltration of private data over unauthorized data flows by enforcing security policies when tools are called. We demonstrate effectiveness of CaMeL by solving $77\%$ of tasks with provable security (compared to $84\%$ with an undefended system) in AgentDojo. We release CaMeL at https://github.com/google-research/camel-prompt-injection.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18813
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Defeating Prompt Injections by Design
Debenedetti, Edoardo
Shumailov, Ilia
Fan, Tianqi
Hayes, Jamie
Carlini, Nicholas
Fabian, Daniel
Kern, Christoph
Shi, Chongyang
Terzis, Andreas
Tramèr, Florian
Cryptography and Security
Artificial Intelligence
Large Language Models (LLMs) are increasingly deployed in agentic systems that interact with an untrusted environment. However, LLM agents are vulnerable to prompt injection attacks when handling untrusted data. In this paper we propose CaMeL, a robust defense that creates a protective system layer around the LLM, securing it even when underlying models are susceptible to attacks. To operate, CaMeL explicitly extracts the control and data flows from the (trusted) query; therefore, the untrusted data retrieved by the LLM can never impact the program flow. To further improve security, CaMeL uses a notion of a capability to prevent the exfiltration of private data over unauthorized data flows by enforcing security policies when tools are called. We demonstrate effectiveness of CaMeL by solving $77\%$ of tasks with provable security (compared to $84\%$ with an undefended system) in AgentDojo. We release CaMeL at https://github.com/google-research/camel-prompt-injection.
title Defeating Prompt Injections by Design
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2503.18813