Design Patterns for Securing LLM Agents against Prompt Injections
Fuente:
arXiv
Saved in:
| Main Authors: | Beurer-Kellner, Luca, Buesser, Beat, Creţu, Ana-Maria, Debenedetti, Edoardo, Dobos, Daniel, Fabian, Daniel, Fischer, Marc, Froelicher, David, Grosse, Kathrin, Naeff, Daniel, Ozoani, Ezinwanne, Paverd, Andrew, Tramèr, Florian, Volhejn, Václav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
by: Debenedetti, Edoardo, et al.
Published: (2024)
by: Debenedetti, Edoardo, et al.
Published: (2024)
Defeating Prompt Injections by Design
by: Debenedetti, Edoardo, et al.
Published: (2025)
by: Debenedetti, Edoardo, et al.
Published: (2025)
Evading Black-box Classifiers Without Breaking Eggs
by: Debenedetti, Edoardo, et al.
Published: (2023)
by: Debenedetti, Edoardo, et al.
Published: (2023)
Adversarial Search Engine Optimization for Large Language Models
by: Nestaas, Fredrik, et al.
Published: (2024)
by: Nestaas, Fredrik, et al.
Published: (2024)
CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization
by: Das, Debeshee, et al.
Published: (2025)
by: Das, Debeshee, et al.
Published: (2025)
Guiding LLMs The Right Way: Fast, Non-Invasive Constrained Generation
by: Beurer-Kellner, Luca, et al.
Published: (2024)
by: Beurer-Kellner, Luca, et al.
Published: (2024)
Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration
by: Das, Debeshee, et al.
Published: (2026)
by: Das, Debeshee, et al.
Published: (2026)
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
by: Carlini, Nicholas, et al.
Published: (2025)
by: Carlini, Nicholas, et al.
Published: (2025)
Controlled Text Generation via Language Model Arithmetic
by: Dekoninck, Jasper, et al.
Published: (2023)
by: Dekoninck, Jasper, et al.
Published: (2023)
Vowel Raising in Nkpor Dialect: A Pattern of Sound Change
by: Evelyn Ezinwanne Mbah
Published: (2013)
by: Evelyn Ezinwanne Mbah
Published: (2013)
What's documented in AI? Systematic Analysis of 32K AI Model Cards
by: Liang, Weixin, et al.
Published: (2024)
by: Liang, Weixin, et al.
Published: (2024)
Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
by: Aerni, Michael, et al.
Published: (2024)
by: Aerni, Michael, et al.
Published: (2024)
Learning to Inject: Automated Prompt Injection via Reinforcement Learning
by: Chen, Xin, et al.
Published: (2026)
by: Chen, Xin, et al.
Published: (2026)
Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
by: Schmotz, David, et al.
Published: (2026)
by: Schmotz, David, et al.
Published: (2026)
"Somos Cubanos!" - timba cubana and the construction of national identity in Cuban popular music
by: Patrick Froelicher
Published: (2005)
by: Patrick Froelicher
Published: (2005)
Towards Assurance of LLM Adversarial Robustness using Ontology-Driven Argumentation
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
Developing Assurance Cases for Adversarial Robustness and Regulatory Compliance in LLMs
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
Knowledge-Augmented Reasoning for EUAIA Compliance and Adversarial Robustness of LLMs
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
Towards Assuring EU AI Act Compliance and Adversarial Robustness of LLMs
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
Ultrasound‐Assisted Epoxy Resin Curing Monitored by Impedance Spectroscopy
by: Daniel Csehngeri, et al.
Published: (2026)
by: Daniel Csehngeri, et al.
Published: (2026)
Privacy Side Channels in Machine Learning Systems
by: Debenedetti, Edoardo, et al.
Published: (2023)
by: Debenedetti, Edoardo, et al.
Published: (2023)
LLMs unlock new paths to monetizing exploits
by: Carlini, Nicholas, et al.
Published: (2025)
by: Carlini, Nicholas, et al.
Published: (2025)
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
by: Zizzo, Giulio, et al.
Published: (2025)
by: Zizzo, Giulio, et al.
Published: (2025)
Exploring Memorization and Copyright Violation in Frontier LLMs: A Study of the New York Times v. OpenAI 2023 Lawsuit
by: Freeman, Joshua, et al.
Published: (2024)
by: Freeman, Joshua, et al.
Published: (2024)
Technical Report: Exploring the Emerging Threats of the Agent Skill Ecosystem
by: Beurer-Kellner, Luca, et al.
Published: (2026)
by: Beurer-Kellner, Luca, et al.
Published: (2026)
Highlight & Summarize: RAG without the jailbreaks
by: Cherubin, Giovanni, et al.
Published: (2025)
by: Cherubin, Giovanni, et al.
Published: (2025)
Site-specific ILC Detector Installation Plan
by: Buesser, Karsten, et al.
Published: (2026)
by: Buesser, Karsten, et al.
Published: (2026)
From Prompt Injections to SQL Injection Attacks: How Protected is Your LLM-Integrated Web Application?
by: Pedro, Rodrigo, et al.
Published: (2023)
by: Pedro, Rodrigo, et al.
Published: (2023)
PINA: Prompt Injection Attack against Navigation Agents
by: Liu, Jiani, et al.
Published: (2026)
by: Liu, Jiani, et al.
Published: (2026)
Defending against Indirect Prompt Injection by Instruction Detection
by: Wen, Tongyu, et al.
Published: (2025)
by: Wen, Tongyu, et al.
Published: (2025)
Prompt Injection attack against LLM-integrated Applications
by: Liu, Yi, et al.
Published: (2023)
by: Liu, Yi, et al.
Published: (2023)
Pitfalls in Evaluating Language Model Forecasters
by: Paleka, Daniel, et al.
Published: (2025)
by: Paleka, Daniel, et al.
Published: (2025)
Correlation inference attacks against machine learning models
by: Creţu, Ana-Maria, et al.
Published: (2021)
by: Creţu, Ana-Maria, et al.
Published: (2021)
A Critical Evaluation of Defenses against Prompt Injection Attacks
by: Jia, Yuqi, et al.
Published: (2025)
by: Jia, Yuqi, et al.
Published: (2025)
Defense against Prompt Injection Attacks via Mixture of Encodings
by: Zhang, Ruiyi, et al.
Published: (2025)
by: Zhang, Ruiyi, et al.
Published: (2025)
Prevalence of Security and Privacy Risk-Inducing Usage of AI-based Conversational Agents
by: Grosse, Kathrin, et al.
Published: (2025)
by: Grosse, Kathrin, et al.
Published: (2025)
Orientation Matters: Learning Radiation Patterns of Multi-Rotor UAVs In-Flight to Enhance Communication Availability Modeling
by: Zoula, Martin, et al.
Published: (2026)
by: Zoula, Martin, et al.
Published: (2026)
Closed-Form Bounds for DP-SGD against Record-level Inference
by: Cherubin, Giovanni, et al.
Published: (2024)
by: Cherubin, Giovanni, et al.
Published: (2024)
Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs
by: Salem, Ahmed, et al.
Published: (2026)
by: Salem, Ahmed, et al.
Published: (2026)
Automatic and Universal Prompt Injection Attacks against Large Language Models
by: Liu, Xiaogeng, et al.
Published: (2024)
by: Liu, Xiaogeng, et al.
Published: (2024)
Similar Items
-
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
by: Debenedetti, Edoardo, et al.
Published: (2024) -
Defeating Prompt Injections by Design
by: Debenedetti, Edoardo, et al.
Published: (2025) -
Evading Black-box Classifiers Without Breaking Eggs
by: Debenedetti, Edoardo, et al.
Published: (2023) -
Adversarial Search Engine Optimization for Large Language Models
by: Nestaas, Fredrik, et al.
Published: (2024) -
CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization
by: Das, Debeshee, et al.
Published: (2025)