How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
Fuente:
arXiv
Saved in:
| Main Authors: | Dziemian, Mateusz, Lin, Maxwell, Fu, Xiaohan, Nowak, Micha, Winter, Nick, Jones, Eliot, Zou, Andy, Ahmad, Lama, Chaudhuri, Kamalika, Chennabasappa, Sahana, Davies, Xander, Deason, Lauren, Edelman, Benjamin L., Emek, Tanner, Evtimov, Ivan, Gust, Jim, Hamin, Maia, He, Kat, Krawiecka, Klaudia, Patana, Riccardo, Perry, Neil, Peterson, Troy, Qi, Xiangyu, Rando, Javier, Wang, Zifan, Wang, Zihan, Whitman, Spencer, Winsor, Eric, Zharmagambetov, Arman, Fredrikson, Matt, Kolter, Zico |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
by: Evtimov, Ivan, et al.
Published: (2025)
by: Evtimov, Ivan, et al.
Published: (2025)
Who Governs the Machine? A Machine Identity Governance Taxonomy (MIGT) for AI Systems Operating Across Enterprise and Geopolitical Boundaries
by: Kurtz, Andrew, et al.
Published: (2026)
by: Kurtz, Andrew, et al.
Published: (2026)
AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents
by: Zharmagambetov, Arman, et al.
Published: (2025)
by: Zharmagambetov, Arman, et al.
Published: (2025)
Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
by: Zou, Andy, et al.
Published: (2025)
by: Zou, Andy, et al.
Published: (2025)
Safety Alignment of LMs via Non-cooperative Games
by: Paulus, Anselm, et al.
Published: (2025)
by: Paulus, Anselm, et al.
Published: (2025)
RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
by: Wen, Yuxin, et al.
Published: (2025)
by: Wen, Yuxin, et al.
Published: (2025)
Evaluating Language Model Reasoning about Confidential Information
by: Sam, Dylan, et al.
Published: (2025)
by: Sam, Dylan, et al.
Published: (2025)
Extending the OWASP Multi-Agentic System Threat Modeling Guide: Insights from Multi-Agent Security Research
by: Krawiecka, Klaudia, et al.
Published: (2025)
by: Krawiecka, Klaudia, et al.
Published: (2025)
Adversarial Attacks on Robotic Vision Language Action Models
by: Jones, Eliot Krzysztof, et al.
Published: (2025)
by: Jones, Eliot Krzysztof, et al.
Published: (2025)
D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
by: Krishna, Satyapriya, et al.
Published: (2025)
by: Krishna, Satyapriya, et al.
Published: (2025)
Mimetic Initialization of MLPs
by: Trockman, Asher, et al.
Published: (2026)
by: Trockman, Asher, et al.
Published: (2026)
AcceleratedLiNGAM: Learning Causal DAGs at the speed of GPUs
by: Akinwande, Victor, et al.
Published: (2024)
by: Akinwande, Victor, et al.
Published: (2024)
Architecting Resilient LLM Agents: A Guide to Secure Plan-then-Execute Implementations
by: Del Rosario, Ron F., et al.
Published: (2025)
by: Del Rosario, Ron F., et al.
Published: (2025)
FUSE-ing Language Models: Zero-Shot Adapter Discovery for Prompt Optimization Across Tokenizers
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning
by: Huang, Benhao, et al.
Published: (2026)
by: Huang, Benhao, et al.
Published: (2026)
Why is SAM Robust to Label Noise?
by: Baek, Christina, et al.
Published: (2024)
by: Baek, Christina, et al.
Published: (2024)
Uncertainty-Based Abstention in LLMs Improves Safety and Reduces Hallucinations
by: Tomani, Christian, et al.
Published: (2024)
by: Tomani, Christian, et al.
Published: (2024)
LlamaFirewall: An open source guardrail system for building secure AI agents
by: Chennabasappa, Sahana, et al.
Published: (2025)
by: Chennabasappa, Sahana, et al.
Published: (2025)
Predicting the Performance of Black-box LLMs through Follow-up Queries
by: Sam, Dylan, et al.
Published: (2025)
by: Sam, Dylan, et al.
Published: (2025)
One-Step Diffusion Distillation via Deep Equilibrium Models
by: Geng, Zhengyang, et al.
Published: (2023)
by: Geng, Zhengyang, et al.
Published: (2023)
Diffusing Differentiable Representations
by: Savani, Yash, et al.
Published: (2024)
by: Savani, Yash, et al.
Published: (2024)
Gradient-based Jailbreak Images for Multimodal Fusion Models
by: Rando, Javier, et al.
Published: (2024)
by: Rando, Javier, et al.
Published: (2024)
Generative Posterior Networks for Approximately Bayesian Epistemic Uncertainty Estimation
by: Roderick, Melrose, et al.
Published: (2023)
by: Roderick, Melrose, et al.
Published: (2023)
Improving Alignment and Robustness with Circuit Breakers
by: Zou, Andy, et al.
Published: (2024)
by: Zou, Andy, et al.
Published: (2024)
SecAlign: Defending Against Prompt Injection with Preference Optimization
by: Chen, Sizhe, et al.
Published: (2024)
by: Chen, Sizhe, et al.
Published: (2024)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
by: Feng, Zhili, et al.
Published: (2025)
by: Feng, Zhili, et al.
Published: (2025)
The Mixing method: low-rank coordinate descent for semidefinite programming with diagonal constraints
by: Wang, Po-Wei, et al.
Published: (2017)
by: Wang, Po-Wei, et al.
Published: (2017)
A Simple and Effective Pruning Approach for Large Language Models
by: Sun, Mingjie, et al.
Published: (2023)
by: Sun, Mingjie, et al.
Published: (2023)
Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance
by: Goyal, Sachin, et al.
Published: (2024)
by: Goyal, Sachin, et al.
Published: (2024)
Looking beyond the next token
by: Thankaraj, Abitha, et al.
Published: (2025)
by: Thankaraj, Abitha, et al.
Published: (2025)
Massive Activations in Large Language Models
by: Sun, Mingjie, et al.
Published: (2024)
by: Sun, Mingjie, et al.
Published: (2024)
Measuring Five-Nines Reliability: Sample-Efficient LLM Evaluation in Saturated Benchmarks
by: Kim, Eungyeup, et al.
Published: (2026)
by: Kim, Eungyeup, et al.
Published: (2026)
The connectivity and phase transition in inhomogeneous random graphs of finite types
by: Jung, Hamin
Published: (2024)
by: Jung, Hamin
Published: (2024)
Bridging the Gap: A Study of AI-based Vulnerability Management between Industry and Academia
by: Wan, Shengye, et al.
Published: (2024)
by: Wan, Shengye, et al.
Published: (2024)
Safety Pretraining: Toward the Next Generation of Safe AI
by: Maini, Pratyush, et al.
Published: (2025)
by: Maini, Pratyush, et al.
Published: (2025)
Hecke Triangle Groups and Special Hyperbolic Elements
by: Winsor, Karl
Published: (2026)
by: Winsor, Karl
Published: (2026)
Forever amber / Kathleen Winsor
by: Winsor, Kathleen
by: Winsor, Kathleen
Ambre : roman / Kathleen Winsor ; traduit de l'anglais par dith Vincent
by: Winsor, Kathleen
by: Winsor, Kathleen
Dynamics of the absolute period foliation of a stratum of holomorphic 1-forms
by: Winsor, Karl
Published: (2021)
by: Winsor, Karl
Published: (2021)
Similar Items
-
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024) -
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
by: Evtimov, Ivan, et al.
Published: (2025) -
Who Governs the Machine? A Machine Identity Governance Taxonomy (MIGT) for AI Systems Operating Across Enterprise and Geopolitical Boundaries
by: Kurtz, Andrew, et al.
Published: (2026) -
AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents
by: Zharmagambetov, Arman, et al.
Published: (2025) -
Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
by: Zou, Andy, et al.
Published: (2025)