Can Indirect Prompt Injection Attacks Be Detected and Removed?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yulin, Li, Haoran, Sui, Yuan, He, Yufei, Liu, Yue, Song, Yangqiu, Hooi, Bryan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TopicAttack: An Indirect Prompt Injection Attack via Topic Transition
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
von: Chen, Yulin, et al.
Veröffentlicht: (2024)
von: Chen, Yulin, et al.
Veröffentlicht: (2024)
Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
von: Chen, Yulin, et al.
Veröffentlicht: (2026)
von: Chen, Yulin, et al.
Veröffentlicht: (2026)
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
von: Chen, Yulin, et al.
Veröffentlicht: (2024)
von: Chen, Yulin, et al.
Veröffentlicht: (2024)
VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
von: Cao, Tri, et al.
Veröffentlicht: (2025)
von: Cao, Tri, et al.
Veröffentlicht: (2025)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
von: Cao, Tri, et al.
Veröffentlicht: (2026)
von: Cao, Tri, et al.
Veröffentlicht: (2026)
Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections
von: Yang, Xianglin, et al.
Veröffentlicht: (2026)
von: Yang, Xianglin, et al.
Veröffentlicht: (2026)
Privacy in Large Language Models: Attacks, Defenses and Future Directions
von: Li, Haoran, et al.
Veröffentlicht: (2023)
von: Li, Haoran, et al.
Veröffentlicht: (2023)
AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs
von: Wang, Che, et al.
Veröffentlicht: (2026)
von: Wang, Che, et al.
Veröffentlicht: (2026)
Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
von: Zhong, Yinan, et al.
Veröffentlicht: (2025)
von: Zhong, Yinan, et al.
Veröffentlicht: (2025)
VortexPIA: Indirect Prompt Injection Attack against LLMs for Efficient Extraction of User Privacy
von: Cui, Yu, et al.
Veröffentlicht: (2025)
von: Cui, Yu, et al.
Veröffentlicht: (2025)
PromptShield: Deployable Detection for Prompt Injection Attacks
von: Jacob, Dennis, et al.
Veröffentlicht: (2025)
von: Jacob, Dennis, et al.
Veröffentlicht: (2025)
Defending against Indirect Prompt Injection by Instruction Detection
von: Wen, Tongyu, et al.
Veröffentlicht: (2025)
von: Wen, Tongyu, et al.
Veröffentlicht: (2025)
FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
von: Wu, Fangzhou, et al.
Veröffentlicht: (2024)
von: Wu, Fangzhou, et al.
Veröffentlicht: (2024)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
von: Hines, Keegan, et al.
Veröffentlicht: (2024)
von: Hines, Keegan, et al.
Veröffentlicht: (2024)
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
von: Lu, Weikai, et al.
Veröffentlicht: (2025)
von: Lu, Weikai, et al.
Veröffentlicht: (2025)
Automated Phishing Detection Using URLs and Webpages
von: Wang, Huilin, et al.
Veröffentlicht: (2024)
von: Wang, Huilin, et al.
Veröffentlicht: (2024)
PINA: Prompt Injection Attack against Navigation Agents
von: Liu, Jiani, et al.
Veröffentlicht: (2026)
von: Liu, Jiani, et al.
Veröffentlicht: (2026)
QueryIPI: Query-agnostic Indirect Prompt Injection on Coding Agents
von: Xie, Yuchong, et al.
Veröffentlicht: (2025)
von: Xie, Yuchong, et al.
Veröffentlicht: (2025)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
von: Zhu, Kaijie, et al.
Veröffentlicht: (2025)
von: Zhu, Kaijie, et al.
Veröffentlicht: (2025)
AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations
von: He, Yu, et al.
Veröffentlicht: (2026)
von: He, Yu, et al.
Veröffentlicht: (2026)
ICON: Indirect Prompt Injection Defense for Agents based on Inference-Time Correction
von: Wang, Che, et al.
Veröffentlicht: (2026)
von: Wang, Che, et al.
Veröffentlicht: (2026)
Autonomous Chain-of-Thought Distillation for Graph-Based Fraud Detection
von: Li, Yuan, et al.
Veröffentlicht: (2026)
von: Li, Yuan, et al.
Veröffentlicht: (2026)
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
von: Liu, Yupei, et al.
Veröffentlicht: (2025)
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
von: Jia, Yuqi, et al.
Veröffentlicht: (2026)
von: Jia, Yuqi, et al.
Veröffentlicht: (2026)
Indirect Prompt Injections: Are Firewalls All You Need, or Stronger Benchmarks?
von: Bhagwatkar, Rishika, et al.
Veröffentlicht: (2025)
von: Bhagwatkar, Rishika, et al.
Veröffentlicht: (2025)
Indirect Prompt Injection in the Wild: An Empirical Study of Prevalence, Techniques, and Objectives
von: Khodayari, Soheil, et al.
Veröffentlicht: (2026)
von: Khodayari, Soheil, et al.
Veröffentlicht: (2026)
Prompt Injection Attack to Tool Selection in LLM Agents
von: Shi, Jiawen, et al.
Veröffentlicht: (2025)
von: Shi, Jiawen, et al.
Veröffentlicht: (2025)
Optimization-based Prompt Injection Attack to LLM-as-a-Judge
von: Shi, Jiawen, et al.
Veröffentlicht: (2024)
von: Shi, Jiawen, et al.
Veröffentlicht: (2024)
Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree
von: Johnson, Sam, et al.
Veröffentlicht: (2025)
von: Johnson, Sam, et al.
Veröffentlicht: (2025)
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
von: Syros, Georgios, et al.
Veröffentlicht: (2026)
von: Syros, Georgios, et al.
Veröffentlicht: (2026)
Lessons from Defending Gemini Against Indirect Prompt Injections
von: Shi, Chongyang, et al.
Veröffentlicht: (2025)
von: Shi, Chongyang, et al.
Veröffentlicht: (2025)
PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
von: Zou, Wei, et al.
Veröffentlicht: (2025)
von: Zou, Wei, et al.
Veröffentlicht: (2025)
Privacy Checklist: Privacy Violation Detection Grounding on Contextual Integrity Theory
von: Li, Haoran, et al.
Veröffentlicht: (2024)
von: Li, Haoran, et al.
Veröffentlicht: (2024)
PromptLocate: Localizing Prompt Injection Attacks
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
von: Jia, Yuqi, et al.
Veröffentlicht: (2025)
Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
von: Xiang, Chong, et al.
Veröffentlicht: (2026)
von: Xiang, Chong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
TopicAttack: An Indirect Prompt Injection Attack via Topic Transition
von: Chen, Yulin, et al.
Veröffentlicht: (2025) -
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
von: Chen, Yulin, et al.
Veröffentlicht: (2025) -
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
von: Chen, Yulin, et al.
Veröffentlicht: (2024) -
Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction
von: Chen, Yulin, et al.
Veröffentlicht: (2025) -
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
von: Chen, Yulin, et al.
Veröffentlicht: (2026)