Design Patterns for Securing LLM Agents against Prompt Injections

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Beurer-Kellner, Luca, Buesser, Beat, Creţu, Ana-Maria, Debenedetti, Edoardo, Dobos, Daniel, Fabian, Daniel, Fischer, Marc, Froelicher, David, Grosse, Kathrin, Naeff, Daniel, Ozoani, Ezinwanne, Paverd, Andrew, Tramèr, Florian, Volhejn, Václav
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918072033476608
author Beurer-Kellner, Luca
Buesser, Beat
Creţu, Ana-Maria
Debenedetti, Edoardo
Dobos, Daniel
Fabian, Daniel
Fischer, Marc
Froelicher, David
Grosse, Kathrin
Naeff, Daniel
Ozoani, Ezinwanne
Paverd, Andrew
Tramèr, Florian
Volhejn, Václav
author_facet Beurer-Kellner, Luca
Buesser, Beat
Creţu, Ana-Maria
Debenedetti, Edoardo
Dobos, Daniel
Fabian, Daniel
Fischer, Marc
Froelicher, David
Grosse, Kathrin
Naeff, Daniel
Ozoani, Ezinwanne
Paverd, Andrew
Tramèr, Florian
Volhejn, Václav
contents As AI agents powered by Large Language Models (LLMs) become increasingly versatile and capable of addressing a broad spectrum of tasks, ensuring their security has become a critical challenge. Among the most pressing threats are prompt injection attacks, which exploit the agent's resilience on natural language inputs -- an especially dangerous threat when agents are granted tool access or handle sensitive information. In this work, we propose a set of principled design patterns for building AI agents with provable resistance to prompt injection. We systematically analyze these patterns, discuss their trade-offs in terms of utility and security, and illustrate their real-world applicability through a series of case studies.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08837
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Design Patterns for Securing LLM Agents against Prompt Injections
Beurer-Kellner, Luca
Buesser, Beat
Creţu, Ana-Maria
Debenedetti, Edoardo
Dobos, Daniel
Fabian, Daniel
Fischer, Marc
Froelicher, David
Grosse, Kathrin
Naeff, Daniel
Ozoani, Ezinwanne
Paverd, Andrew
Tramèr, Florian
Volhejn, Václav
Machine Learning
Cryptography and Security
As AI agents powered by Large Language Models (LLMs) become increasingly versatile and capable of addressing a broad spectrum of tasks, ensuring their security has become a critical challenge. Among the most pressing threats are prompt injection attacks, which exploit the agent's resilience on natural language inputs -- an especially dangerous threat when agents are granted tool access or handle sensitive information. In this work, we propose a set of principled design patterns for building AI agents with provable resistance to prompt injection. We systematically analyze these patterns, discuss their trade-offs in terms of utility and security, and illustrate their real-world applicability through a series of case studies.
title Design Patterns for Securing LLM Agents against Prompt Injections
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2506.08837