MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Syros, Georgios, Rose, Evan, Grinstead, Brian, Kerschbaumer, Christoph, Robertson, William, Nita-Rotaru, Cristina, Oprea, Alina |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAGA: A Security Architecture for Governing AI Agentic Systems
by: Syros, Georgios, et al.
Published: (2025)
by: Syros, Georgios, et al.
Published: (2025)
Backdoor Attacks in Peer-to-Peer Federated Learning
by: Syros, Georgios, et al.
Published: (2023)
by: Syros, Georgios, et al.
Published: (2023)
DROP: Poison Dilution via Knowledge Distillation for Federated Learning
by: Syros, Georgios, et al.
Published: (2025)
by: Syros, Georgios, et al.
Published: (2025)
Attacks and Mitigations for Distributed Governance of Agentic AI under Byzantine Adversaries
by: Laws, Matthew D., et al.
Published: (2026)
by: Laws, Matthew D., et al.
Published: (2026)
APWA: A Distributed Architecture for Parallelizable Agentic Workflows
by: Rose, Evan, et al.
Published: (2026)
by: Rose, Evan, et al.
Published: (2026)
ACE: A Security Architecture for LLM-Integrated App Systems
by: Li, Evan, et al.
Published: (2025)
by: Li, Evan, et al.
Published: (2025)
IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection
by: Chia-Pei, et al.
Published: (2026)
by: Chia-Pei, et al.
Published: (2026)
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
by: Zhan, Qiusi, et al.
Published: (2025)
by: Zhan, Qiusi, et al.
Published: (2025)
MAGIQ: A Post-Quantum Multi-Agentic AI Governance System with Provable Security
by: Avizheh, Sepideh, et al.
Published: (2026)
by: Avizheh, Sepideh, et al.
Published: (2026)
EVA: Red-Teaming GUI Agents via Evolving Indirect Prompt Injection
by: Lu, Yijie, et al.
Published: (2025)
by: Lu, Yijie, et al.
Published: (2025)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
by: Hines, Keegan, et al.
Published: (2024)
by: Hines, Keegan, et al.
Published: (2024)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
by: Zhu, Kaijie, et al.
Published: (2025)
by: Zhu, Kaijie, et al.
Published: (2025)
AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs
by: Wang, Che, et al.
Published: (2026)
by: Wang, Che, et al.
Published: (2026)
SleeperNets: Universal Backdoor Poisoning Attacks Against Reinforcement Learning Agents
by: Rathbun, Ethan, et al.
Published: (2024)
by: Rathbun, Ethan, et al.
Published: (2024)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
by: Evtimov, Ivan, et al.
Published: (2025)
by: Evtimov, Ivan, et al.
Published: (2025)
Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree
by: Johnson, Sam, et al.
Published: (2025)
by: Johnson, Sam, et al.
Published: (2025)
Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
by: Xiang, Chong, et al.
Published: (2026)
by: Xiang, Chong, et al.
Published: (2026)
Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
by: Yi, Jingwei, et al.
Published: (2023)
by: Yi, Jingwei, et al.
Published: (2023)
WebInject: Prompt Injection Attack to Web Agents
by: Wang, Xilong, et al.
Published: (2025)
by: Wang, Xilong, et al.
Published: (2025)
Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
by: Zhong, Yinan, et al.
Published: (2025)
by: Zhong, Yinan, et al.
Published: (2025)
LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection
by: Zhao, Lei, et al.
Published: (2026)
by: Zhao, Lei, et al.
Published: (2026)
Securing AI Agents Against Prompt Injection Attacks
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
by: Wang, Zhilong, et al.
Published: (2025)
by: Wang, Zhilong, et al.
Published: (2025)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
by: Chaudhari, Harsh, et al.
Published: (2024)
by: Chaudhari, Harsh, et al.
Published: (2024)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
by: Wang, Xilong, et al.
Published: (2026)
by: Wang, Xilong, et al.
Published: (2026)
A Formal Analysis of SCTP: Attack Synthesis and Patch Verification
by: Ginesin, Jacob, et al.
Published: (2024)
by: Ginesin, Jacob, et al.
Published: (2024)
Can Indirect Prompt Injection Attacks Be Detected and Removed?
by: Chen, Yulin, et al.
Published: (2025)
by: Chen, Yulin, et al.
Published: (2025)
Investigating Detection and Obfuscation of Prompt Injection Attacks Against Software Reverse Engineering AI Agents
by: Crawford, Brian, et al.
Published: (2026)
by: Crawford, Brian, et al.
Published: (2026)
TopicAttack: An Indirect Prompt Injection Attack via Topic Transition
by: Chen, Yulin, et al.
Published: (2025)
by: Chen, Yulin, et al.
Published: (2025)
Lessons from Defending Gemini Against Indirect Prompt Injections
by: Shi, Chongyang, et al.
Published: (2025)
by: Shi, Chongyang, et al.
Published: (2025)
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
Trojan Horses in Recruiting: A Red-Teaming Case Study on Indirect Prompt Injection in Standard vs. Reasoning Models
by: Wirth, Manuel
Published: (2026)
by: Wirth, Manuel
Published: (2026)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
by: Cao, Tri, et al.
Published: (2026)
by: Cao, Tri, et al.
Published: (2026)
The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
by: Jia, Feiran, et al.
Published: (2024)
by: Jia, Feiran, et al.
Published: (2024)
LogJack: Indirect Prompt Injection Through Cloud Logs Against LLM Debugging Agents
by: Shah, Harsh
Published: (2026)
by: Shah, Harsh
Published: (2026)
Genesis: Evolving Attack Strategies for LLM Web Agent Red-Teaming
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
Defense Against Indirect Prompt Injection via Tool Result Parsing
by: Yu, Qiang, et al.
Published: (2026)
by: Yu, Qiang, et al.
Published: (2026)
Whispers of Wealth: Red-Teaming Google's Agent Payments Protocol via Prompt Injection
by: Debi, Tanusree, et al.
Published: (2026)
by: Debi, Tanusree, et al.
Published: (2026)
The Matrix Reloaded: A Mechanized Formal Analysis of the Matrix Cryptographic Suite
by: Ginesin, Jacob, et al.
Published: (2024)
by: Ginesin, Jacob, et al.
Published: (2024)
IterInject: Indirect Prompt Injection Against LLM Agents via Feedback-Guided Iterative Optimization
by: Chen, Zixuan, et al.
Published: (2026)
by: Chen, Zixuan, et al.
Published: (2026)
Similar Items
-
SAGA: A Security Architecture for Governing AI Agentic Systems
by: Syros, Georgios, et al.
Published: (2025) -
Backdoor Attacks in Peer-to-Peer Federated Learning
by: Syros, Georgios, et al.
Published: (2023) -
DROP: Poison Dilution via Knowledge Distillation for Federated Learning
by: Syros, Georgios, et al.
Published: (2025) -
Attacks and Mitigations for Distributed Governance of Agentic AI under Byzantine Adversaries
by: Laws, Matthew D., et al.
Published: (2026) -
APWA: A Distributed Architecture for Parallelizable Agentic Workflows
by: Rose, Evan, et al.
Published: (2026)