Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Sekar, Anirudh, Agarwal, Mrinal, Sharma, Rachel, Tanaka, Akitsugu, Zhang, Jasmine, Damerla, Arjun, Zhu, Kevin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Defending Against Prompt Injection With a Few DefensiveTokens
por: Chen, Sizhe, et al.
Publicado: (2025)
por: Chen, Sizhe, et al.
Publicado: (2025)
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
por: Chen, Yulin, et al.
Publicado: (2024)
por: Chen, Yulin, et al.
Publicado: (2024)
AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs
por: Wang, Yihan, et al.
Publicado: (2025)
por: Wang, Yihan, et al.
Publicado: (2025)
Prompt Control-Flow Integrity: A Priority-Aware Runtime Defense Against Prompt Injection in LLM Systems
por: Alam, Md Takrim Ul, et al.
Publicado: (2026)
por: Alam, Md Takrim Ul, et al.
Publicado: (2026)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
por: Cao, Tri, et al.
Publicado: (2026)
por: Cao, Tri, et al.
Publicado: (2026)
LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
por: Panebianco, Francesco, et al.
Publicado: (2025)
por: Panebianco, Francesco, et al.
Publicado: (2025)
Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs
por: Yeo, Andrew, et al.
Publicado: (2025)
por: Yeo, Andrew, et al.
Publicado: (2025)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
por: Jaiswal, Piyush, et al.
Publicado: (2026)
por: Jaiswal, Piyush, et al.
Publicado: (2026)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
por: Zhu, Kaijie, et al.
Publicado: (2025)
por: Zhu, Kaijie, et al.
Publicado: (2025)
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
por: Hossain, S M Asif, et al.
Publicado: (2025)
por: Hossain, S M Asif, et al.
Publicado: (2025)
Privacy-Preserving Prompt Injection Detection for LLMs Using Federated Learning and Embedding-Based NLP Classification
por: Jayathilaka, Hasini
Publicado: (2025)
por: Jayathilaka, Hasini
Publicado: (2025)
Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
por: Zhong, Yinan, et al.
Publicado: (2025)
por: Zhong, Yinan, et al.
Publicado: (2025)
Agent Privilege Separation in OpenClaw: A Structural Defense Against Prompt Injection
por: Cheng, Darren, et al.
Publicado: (2026)
por: Cheng, Darren, et al.
Publicado: (2026)
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
por: Zhan, Qiusi, et al.
Publicado: (2025)
por: Zhan, Qiusi, et al.
Publicado: (2025)
Defending Against Prompt Injection with DataFilter
por: Wang, Yizhu, et al.
Publicado: (2025)
por: Wang, Yizhu, et al.
Publicado: (2025)
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
por: Chen, Yulin, et al.
Publicado: (2025)
por: Chen, Yulin, et al.
Publicado: (2025)
The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail?
por: Bhatt, Manish, et al.
Publicado: (2026)
por: Bhatt, Manish, et al.
Publicado: (2026)
Hacking Back the AI-Hacker: Prompt Injection as a Defense Against LLM-driven Cyberattacks
por: Pasquini, Dario, et al.
Publicado: (2024)
por: Pasquini, Dario, et al.
Publicado: (2024)
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
por: Wang, Jiawen, et al.
Publicado: (2025)
por: Wang, Jiawen, et al.
Publicado: (2025)
Diffusion-Guided Adversarial Perturbation Injection for Generalizable Defense Against Facial Manipulations
por: Li, Yue, et al.
Publicado: (2026)
por: Li, Yue, et al.
Publicado: (2026)
PromptArmor: Simple yet Effective Prompt Injection Defenses
por: Shi, Tianneng, et al.
Publicado: (2025)
por: Shi, Tianneng, et al.
Publicado: (2025)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
por: Nasr, Milad, et al.
Publicado: (2025)
por: Nasr, Milad, et al.
Publicado: (2025)
Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
por: Xiang, Chong, et al.
Publicado: (2026)
por: Xiang, Chong, et al.
Publicado: (2026)
StruQ: Defending Against Prompt Injection with Structured Queries
por: Chen, Sizhe, et al.
Publicado: (2024)
por: Chen, Sizhe, et al.
Publicado: (2024)
Injection Attacks Against End-to-End Encrypted Applications
por: Fábrega, Andrés, et al.
Publicado: (2024)
por: Fábrega, Andrés, et al.
Publicado: (2024)
Evaluation of Prompt Injection Defenses in Large Language Models
por: Deep, Priyal, et al.
Publicado: (2026)
por: Deep, Priyal, et al.
Publicado: (2026)
Prompt Injection as Role Confusion
por: Ye, Charles, et al.
Publicado: (2026)
por: Ye, Charles, et al.
Publicado: (2026)
SnapGuard: Lightweight Prompt Injection Detection for Screenshot-Based Web Agents
por: Du, Mengyao, et al.
Publicado: (2026)
por: Du, Mengyao, et al.
Publicado: (2026)
PromptShield: Deployable Detection for Prompt Injection Attacks
por: Jacob, Dennis, et al.
Publicado: (2025)
por: Jacob, Dennis, et al.
Publicado: (2025)
Neighborhood Blending: A Lightweight Inference-Time Defense Against Membership Inference Attacks
por: Zafar, Osama, et al.
Publicado: (2026)
por: Zafar, Osama, et al.
Publicado: (2026)
Strengthening Polymorphic Prompt Assembling: Dynamic Separator Generation Against Emerging Prompt Injection Attacks
por: Dorzhiev, Nima, et al.
Publicado: (2026)
por: Dorzhiev, Nima, et al.
Publicado: (2026)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
por: Xiong, Chen, et al.
Publicado: (2024)
por: Xiong, Chen, et al.
Publicado: (2024)
A Critical Evaluation of Defenses against Prompt Injection Attacks
por: Jia, Yuqi, et al.
Publicado: (2025)
por: Jia, Yuqi, et al.
Publicado: (2025)
CASCADE: A Cascaded Hybrid Defense Architecture for Prompt Injection Detection in MCP-Based Systems
por: Turgut, İpek Abasıkeleş, et al.
Publicado: (2026)
por: Turgut, İpek Abasıkeleş, et al.
Publicado: (2026)
Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
por: Shaheer, Safwan, et al.
Publicado: (2025)
por: Shaheer, Safwan, et al.
Publicado: (2025)
System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
por: Wu, Fangzhou, et al.
Publicado: (2024)
por: Wu, Fangzhou, et al.
Publicado: (2024)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
por: Wang, Zhilong, et al.
Publicado: (2025)
por: Wang, Zhilong, et al.
Publicado: (2025)
The Cognitive Firewall:Securing Browser Based AI Agents Against Indirect Prompt Injection Via Hybrid Edge Cloud Defense
por: Lan, Qianlong, et al.
Publicado: (2026)
por: Lan, Qianlong, et al.
Publicado: (2026)
PromptSleuth: Detecting Prompt Injection via Semantic Intent Invariance
por: Wang, Mengxiao, et al.
Publicado: (2025)
por: Wang, Mengxiao, et al.
Publicado: (2025)
Incorporating Gradients to Rules: Towards Lightweight, Adaptive Provenance-based Intrusion Detection
por: Wang, Lingzhi, et al.
Publicado: (2024)
por: Wang, Lingzhi, et al.
Publicado: (2024)
Ejemplares similares
-
Defending Against Prompt Injection With a Few DefensiveTokens
por: Chen, Sizhe, et al.
Publicado: (2025) -
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
por: Chen, Yulin, et al.
Publicado: (2024) -
AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs
por: Wang, Yihan, et al.
Publicado: (2025) -
Prompt Control-Flow Integrity: A Priority-Aware Runtime Defense Against Prompt Injection in LLM Systems
por: Alam, Md Takrim Ul, et al.
Publicado: (2026) -
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
por: Cao, Tri, et al.
Publicado: (2026)