The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail?
Fuente:
arXiv
Saved in:
| Main Authors: | Bhatt, Manish, Munshi, Sarthak, Narajala, Vineeth Sai, Habler, Idan, Al-Kahfah, Ammar, Huang, Ken, Webb, Joel, Gatto, Blake, Hoque, Md Tamjidul |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Manifold of Failure: Behavioral Attraction Basins in Language Models
by: Munshi, Sarthak, et al.
Published: (2026)
by: Munshi, Sarthak, et al.
Published: (2026)
ETDI: Mitigating Tool Squatting and Rug Pull Attacks in Model Context Protocol (MCP) by using OAuth-Enhanced Tool Definitions and Policy-Based Access Control
by: Bhatt, Manish, et al.
Published: (2025)
by: Bhatt, Manish, et al.
Published: (2025)
Enterprise-Grade Security for the Model Context Protocol (MCP): Frameworks and Mitigation Strategies
by: Narajala, Vineeth Sai, et al.
Published: (2025)
by: Narajala, Vineeth Sai, et al.
Published: (2025)
Securing GenAI Multi-Agent Systems Against Tool Squatting: A Zero Trust Registry-Based Approach
by: Narajala, Vineeth Sai, et al.
Published: (2025)
by: Narajala, Vineeth Sai, et al.
Published: (2025)
Large Empirical Case Study: Go-Explore adapted for AI Red Team Testing
by: Bhatt, Manish, et al.
Published: (2025)
by: Bhatt, Manish, et al.
Published: (2025)
COALESCE: Economic and Security Dynamics of Skill-Based Task Outsourcing Among Team of Autonomous LLM Agents
by: Bhatt, Manish, et al.
Published: (2025)
by: Bhatt, Manish, et al.
Published: (2025)
Building A Secure Agentic AI Application Leveraging A2A Protocol
by: Habler, Idan, et al.
Published: (2025)
by: Habler, Idan, et al.
Published: (2025)
Agent Capability Negotiation and Binding Protocol (ACNBP)
by: Huang, Ken, et al.
Published: (2025)
by: Huang, Ken, et al.
Published: (2025)
Agent Name Service (ANS): A Universal Directory for Secure AI Agent Discovery and Interoperability
by: Huang, Ken, et al.
Published: (2025)
by: Huang, Ken, et al.
Published: (2025)
MAIF: Enforcing AI Trust and Provenance with an Artifact-Centric Agentic Paradigm
by: Narajala, Vineeth Sai, et al.
Published: (2025)
by: Narajala, Vineeth Sai, et al.
Published: (2025)
Adversarial Hubness Detector: Detecting Hubness Poisoning in Retrieval-Augmented Generation Systems
by: Habler, Idan, et al.
Published: (2026)
by: Habler, Idan, et al.
Published: (2026)
A2AS: Agentic AI Runtime Security and Self-Defense
by: Neelou, Eugene, et al.
Published: (2025)
by: Neelou, Eugene, et al.
Published: (2025)
A Novel Zero-Trust Identity Framework for Agentic AI: Decentralized Authentication and Fine-Grained Access Control
by: Huang, Ken, et al.
Published: (2025)
by: Huang, Ken, et al.
Published: (2025)
Securing Agentic AI: A Comprehensive Threat Model and Mitigation Framework for Generative AI Agents
by: Narajala, Vineeth Sai, et al.
Published: (2025)
by: Narajala, Vineeth Sai, et al.
Published: (2025)
Why Self-Awareness Fails: Coherence, Defense, and the Structure of Repair
by: Jovanovic, Vladisav
Published: (2026)
by: Jovanovic, Vladisav
Published: (2026)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
by: Liu, Yupei, et al.
Published: (2023)
by: Liu, Yupei, et al.
Published: (2023)
PromptArmor: Simple yet Effective Prompt Injection Defenses
by: Shi, Tianneng, et al.
Published: (2025)
by: Shi, Tianneng, et al.
Published: (2025)
ARES-LSHADE: Autoresearch-Enhanced LSHADE with Memetic Polish for the GNBG Benchmark
by: Naeem, Abdullah, et al.
Published: (2026)
by: Naeem, Abdullah, et al.
Published: (2026)
Jatmo: Prompt Injection Defense by Task-Specific Finetuning
by: Piet, Julien, et al.
Published: (2023)
by: Piet, Julien, et al.
Published: (2023)
Evaluation of Prompt Injection Defenses in Large Language Models
by: Deep, Priyal, et al.
Published: (2026)
by: Deep, Priyal, et al.
Published: (2026)
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
by: Chen, Yulin, et al.
Published: (2025)
by: Chen, Yulin, et al.
Published: (2025)
Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs
by: Yeo, Andrew, et al.
Published: (2025)
by: Yeo, Andrew, et al.
Published: (2025)
Defending Against Prompt Injection With a Few DefensiveTokens
by: Chen, Sizhe, et al.
Published: (2025)
by: Chen, Sizhe, et al.
Published: (2025)
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
by: Chen, Yulin, et al.
Published: (2024)
by: Chen, Yulin, et al.
Published: (2024)
A Critical Evaluation of Defenses against Prompt Injection Attacks
by: Jia, Yuqi, et al.
Published: (2025)
by: Jia, Yuqi, et al.
Published: (2025)
Defense against Prompt Injection Attacks via Mixture of Encodings
by: Zhang, Ruiyi, et al.
Published: (2025)
by: Zhang, Ruiyi, et al.
Published: (2025)
Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
by: Shaheer, Safwan, et al.
Published: (2025)
by: Shaheer, Safwan, et al.
Published: (2025)
ESLD (External Surrogate Latent Defense): A Latent-Space Architecture for Faster, Stronger Prompt-Injection Defense
by: Narendra, Yash
Published: (2026)
by: Narendra, Yash
Published: (2026)
Defense Against Indirect Prompt Injection via Tool Result Parsing
by: Yu, Qiang, et al.
Published: (2026)
by: Yu, Qiang, et al.
Published: (2026)
PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
by: Yin, Chenlong, et al.
Published: (2026)
by: Yin, Chenlong, et al.
Published: (2026)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
by: Cao, Tri, et al.
Published: (2026)
by: Cao, Tri, et al.
Published: (2026)
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
by: Zloczower, Itay, et al.
Published: (2026)
by: Zloczower, Itay, et al.
Published: (2026)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
by: Zhu, Kaijie, et al.
Published: (2025)
by: Zhu, Kaijie, et al.
Published: (2025)
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
by: Hossain, S M Asif, et al.
Published: (2025)
by: Hossain, S M Asif, et al.
Published: (2025)
ICON: Indirect Prompt Injection Defense for Agents based on Inference-Time Correction
by: Wang, Che, et al.
Published: (2026)
by: Wang, Che, et al.
Published: (2026)
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
by: Yang, Xiaoxue, et al.
Published: (2025)
by: Yang, Xiaoxue, et al.
Published: (2025)
Logic layer Prompt Control Injection (LPCI): A Novel Security Vulnerability Class in Agentic Systems
by: Atta, Hammad, et al.
Published: (2025)
by: Atta, Hammad, et al.
Published: (2025)
AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs
by: Wang, Yihan, et al.
Published: (2025)
by: Wang, Yihan, et al.
Published: (2025)
FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks
by: Wang, Jiongxiao, et al.
Published: (2024)
by: Wang, Jiongxiao, et al.
Published: (2024)
LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
by: Panebianco, Francesco, et al.
Published: (2025)
by: Panebianco, Francesco, et al.
Published: (2025)
Similar Items
-
Manifold of Failure: Behavioral Attraction Basins in Language Models
by: Munshi, Sarthak, et al.
Published: (2026) -
ETDI: Mitigating Tool Squatting and Rug Pull Attacks in Model Context Protocol (MCP) by using OAuth-Enhanced Tool Definitions and Policy-Based Access Control
by: Bhatt, Manish, et al.
Published: (2025) -
Enterprise-Grade Security for the Model Context Protocol (MCP): Frameworks and Mitigation Strategies
by: Narajala, Vineeth Sai, et al.
Published: (2025) -
Securing GenAI Multi-Agent Systems Against Tool Squatting: A Zero Trust Registry-Based Approach
by: Narajala, Vineeth Sai, et al.
Published: (2025) -
Large Empirical Case Study: Go-Explore adapted for AI Red Team Testing
by: Bhatt, Manish, et al.
Published: (2025)