Indirect Prompt Injections: Are Firewalls All You Need, or Stronger Benchmarks?
Fuente:
arXiv
Saved in:
| Main Authors: | Bhagwatkar, Rishika, Kasa, Kevin, Puri, Abhay, Huang, Gabriel, Rish, Irina, Taylor, Graham W., Dvijotham, Krishnamurthy Dj, Lacoste, Alexandre |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Adversarial Robustness of Discrete Image Tokenizers
by: Bhagwatkar, Rishika, et al.
Published: (2026)
by: Bhagwatkar, Rishika, et al.
Published: (2026)
CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution
by: Kim, Minbeom, et al.
Published: (2026)
by: Kim, Minbeom, et al.
Published: (2026)
Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
by: Boisvert, Léo, et al.
Published: (2025)
by: Boisvert, Léo, et al.
Published: (2025)
CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
by: Bhagwatkar, Rishika, et al.
Published: (2025)
by: Bhagwatkar, Rishika, et al.
Published: (2025)
Towards Adversarially Robust Vision-Language Models: Insights from Design Choices and Prompt Formatting Techniques
by: Bhagwatkar, Rishika, et al.
Published: (2024)
by: Bhagwatkar, Rishika, et al.
Published: (2024)
Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
by: Zhong, Yinan, et al.
Published: (2025)
by: Zhong, Yinan, et al.
Published: (2025)
Keeping up with dynamic attackers: Certifying robustness to adaptive online data poisoning
by: Bose, Avinandan, et al.
Published: (2025)
by: Bose, Avinandan, et al.
Published: (2025)
LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection
by: Zhao, Lei, et al.
Published: (2026)
by: Zhao, Lei, et al.
Published: (2026)
A Guide to Robust Generalization: The Impact of Architecture, Pre-training, and Optimization Strategy
by: Heuillet, Maxime, et al.
Published: (2025)
by: Heuillet, Maxime, et al.
Published: (2025)
Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG
by: Choudhary, Sarthak, et al.
Published: (2025)
by: Choudhary, Sarthak, et al.
Published: (2025)
DoomArena: A framework for Testing AI Agents Against Evolving Security Threats
by: Boisvert, Leo, et al.
Published: (2025)
by: Boisvert, Leo, et al.
Published: (2025)
Adaptive Diffusion Denoised Smoothing : Certified Robustness via Randomized Smoothing with Differentially Private Guided Denoising Diffusion
by: Shpilevskiy, Frederick, et al.
Published: (2025)
by: Shpilevskiy, Frederick, et al.
Published: (2025)
The Cognitive Firewall:Securing Browser Based AI Agents Against Indirect Prompt Injection Via Hybrid Edge Cloud Defense
by: Lan, Qianlong, et al.
Published: (2026)
by: Lan, Qianlong, et al.
Published: (2026)
Efficient Error Certification for Physics-Informed Neural Networks
by: Eiras, Francisco, et al.
Published: (2023)
by: Eiras, Francisco, et al.
Published: (2023)
No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety Mechanisms
by: Kazdan, Joshua, et al.
Published: (2025)
by: Kazdan, Joshua, et al.
Published: (2025)
VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
by: Miculicich, Lesly, et al.
Published: (2025)
by: Miculicich, Lesly, et al.
Published: (2025)
Adapting Prediction Sets to Distribution Shifts Without Labels
by: Kasa, Kevin, et al.
Published: (2024)
by: Kasa, Kevin, et al.
Published: (2024)
LitLLM: A Toolkit for Scientific Literature Review
by: Agarwal, Shubham, et al.
Published: (2024)
by: Agarwal, Shubham, et al.
Published: (2024)
LitLLMs, LLMs for Literature Review: Are we there yet?
by: Agarwal, Shubham, et al.
Published: (2024)
by: Agarwal, Shubham, et al.
Published: (2024)
Training on the Benchmark Is Not All You Need
by: Ni, Shiwen, et al.
Published: (2024)
by: Ni, Shiwen, et al.
Published: (2024)
Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
by: Yi, Jingwei, et al.
Published: (2023)
by: Yi, Jingwei, et al.
Published: (2023)
Achieving the Tightest Relaxation of Sigmoids for Formal Verification
by: Chevalier, Samuel, et al.
Published: (2024)
by: Chevalier, Samuel, et al.
Published: (2024)
Context is Key: A Benchmark for Forecasting with Essential Textual Information
by: Williams, Andrew Robert, et al.
Published: (2024)
by: Williams, Andrew Robert, et al.
Published: (2024)
Hidden-in-Plain-Text: A Benchmark for Social-Web Indirect Prompt Injection in RAG
by: Guo, Haoze, et al.
Published: (2026)
by: Guo, Haoze, et al.
Published: (2026)
Belief Is All You Need: Modeling Narrative Archetypes in Conspiratorial Discourse
by: Shimgekar, Soorya Ram, et al.
Published: (2025)
by: Shimgekar, Soorya Ram, et al.
Published: (2025)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
by: Hines, Keegan, et al.
Published: (2024)
by: Hines, Keegan, et al.
Published: (2024)
Defending against Indirect Prompt Injection by Instruction Detection
by: Wen, Tongyu, et al.
Published: (2025)
by: Wen, Tongyu, et al.
Published: (2025)
Can Indirect Prompt Injection Attacks Be Detected and Removed?
by: Chen, Yulin, et al.
Published: (2025)
by: Chen, Yulin, et al.
Published: (2025)
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
by: Zhan, Qiusi, et al.
Published: (2024)
by: Zhan, Qiusi, et al.
Published: (2024)
Norm-Bounded Low-Rank Adaptation
by: Wang, Ruigang, et al.
Published: (2025)
by: Wang, Ruigang, et al.
Published: (2025)
Monotone, Bi-Lipschitz, and Polyak-Lojasiewicz Networks
by: Wang, Ruigang, et al.
Published: (2024)
by: Wang, Ruigang, et al.
Published: (2024)
Enhancing Context Through Contrast
by: Ambilduke, Kshitij, et al.
Published: (2024)
by: Ambilduke, Kshitij, et al.
Published: (2024)
Lessons from Defending Gemini Against Indirect Prompt Injections
by: Shi, Chongyang, et al.
Published: (2025)
by: Shi, Chongyang, et al.
Published: (2025)
Terminal Is All You Need: Design Properties for Human-AI Agent Collaboration
by: De Masi, Alexandre
Published: (2026)
by: De Masi, Alexandre
Published: (2026)
Attention is All You Need Until You Need Retention
by: Yaslioglu, M. Murat
Published: (2025)
by: Yaslioglu, M. Murat
Published: (2025)
Approximating innovation potential with neurofuzzy robust model
by: Richard Kasa
Published: (2015)
by: Richard Kasa
Published: (2015)
CAMformer: Associative Memory is All You Need
by: Molom-Ochir, Tergel, et al.
Published: (2025)
by: Molom-Ochir, Tergel, et al.
Published: (2025)
Choice of PEFT Technique in Continual Learning: Prompt Tuning is Not All You Need
by: Wistuba, Martin, et al.
Published: (2024)
by: Wistuba, Martin, et al.
Published: (2024)
Text Is Not All You Need: Multimodal Prompting Helps LLMs Understand Humor
by: Baluja, Ashwin
Published: (2024)
by: Baluja, Ashwin
Published: (2024)
All You Need is One: Capsule Prompt Tuning with a Single Vector
by: Liu, Yiyang, et al.
Published: (2025)
by: Liu, Yiyang, et al.
Published: (2025)
Similar Items
-
On the Adversarial Robustness of Discrete Image Tokenizers
by: Bhagwatkar, Rishika, et al.
Published: (2026) -
CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution
by: Kim, Minbeom, et al.
Published: (2026) -
Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
by: Boisvert, Léo, et al.
Published: (2025) -
CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
by: Bhagwatkar, Rishika, et al.
Published: (2025) -
Towards Adversarially Robust Vision-Language Models: Insights from Design Choices and Prompt Formatting Techniques
by: Bhagwatkar, Rishika, et al.
Published: (2024)