The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Jia, Feiran, Wu, Tong, Qin, Xin, Squicciarini, Anna |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection
di: Wang, Peiran, et al.
Pubblicazione: (2025)
di: Wang, Peiran, et al.
Pubblicazione: (2025)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
di: Hines, Keegan, et al.
Pubblicazione: (2024)
di: Hines, Keegan, et al.
Pubblicazione: (2024)
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
di: Zhong, Peter Yong, et al.
Pubblicazione: (2025)
di: Zhong, Peter Yong, et al.
Pubblicazione: (2025)
IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents
di: An, Hengyu, et al.
Pubblicazione: (2025)
di: An, Hengyu, et al.
Pubblicazione: (2025)
Defending against Indirect Prompt Injection by Instruction Detection
di: Wen, Tongyu, et al.
Pubblicazione: (2025)
di: Wen, Tongyu, et al.
Pubblicazione: (2025)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
di: Cao, Bochuan, et al.
Pubblicazione: (2023)
di: Cao, Bochuan, et al.
Pubblicazione: (2023)
Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization
di: Liu, Tiantian, et al.
Pubblicazione: (2024)
di: Liu, Tiantian, et al.
Pubblicazione: (2024)
VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit
di: Lin, Junda, et al.
Pubblicazione: (2026)
di: Lin, Junda, et al.
Pubblicazione: (2026)
Jatmo: Prompt Injection Defense by Task-Specific Finetuning
di: Piet, Julien, et al.
Pubblicazione: (2023)
di: Piet, Julien, et al.
Pubblicazione: (2023)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
di: Zhu, Kaijie, et al.
Pubblicazione: (2025)
di: Zhu, Kaijie, et al.
Pubblicazione: (2025)
LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection
di: Zhao, Lei, et al.
Pubblicazione: (2026)
di: Zhao, Lei, et al.
Pubblicazione: (2026)
Lessons from Defending Gemini Against Indirect Prompt Injections
di: Shi, Chongyang, et al.
Pubblicazione: (2025)
di: Shi, Chongyang, et al.
Pubblicazione: (2025)
ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection
di: Zhao, Wei, et al.
Pubblicazione: (2026)
di: Zhao, Wei, et al.
Pubblicazione: (2026)
A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy
di: Correia, Pedro H. Barcha, et al.
Pubblicazione: (2026)
di: Correia, Pedro H. Barcha, et al.
Pubblicazione: (2026)
When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection
di: Sahoo, Devanshu, et al.
Pubblicazione: (2025)
di: Sahoo, Devanshu, et al.
Pubblicazione: (2025)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
di: Benjamin, Victoria, et al.
Pubblicazione: (2024)
di: Benjamin, Victoria, et al.
Pubblicazione: (2024)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
di: Shao, Zedian, et al.
Pubblicazione: (2024)
di: Shao, Zedian, et al.
Pubblicazione: (2024)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
di: Liu, Yupei, et al.
Pubblicazione: (2023)
di: Liu, Yupei, et al.
Pubblicazione: (2023)
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
di: Syros, Georgios, et al.
Pubblicazione: (2026)
di: Syros, Georgios, et al.
Pubblicazione: (2026)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
di: Zhang, Mohan, et al.
Pubblicazione: (2026)
di: Zhang, Mohan, et al.
Pubblicazione: (2026)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
di: Xiang, Chong, et al.
Pubblicazione: (2026)
di: Xiang, Chong, et al.
Pubblicazione: (2026)
PIArena: A Platform for Prompt Injection Evaluation
di: Geng, Runpeng, et al.
Pubblicazione: (2026)
di: Geng, Runpeng, et al.
Pubblicazione: (2026)
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
di: Wang, Zhun, et al.
Pubblicazione: (2025)
di: Wang, Zhun, et al.
Pubblicazione: (2025)
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
di: Geng, Runpeng, et al.
Pubblicazione: (2025)
di: Geng, Runpeng, et al.
Pubblicazione: (2025)
Soft Begging: Modular and Efficient Shielding of LLMs against Prompt Injection and Jailbreaking based on Prompt Tuning
di: Ostermann, Simon, et al.
Pubblicazione: (2024)
di: Ostermann, Simon, et al.
Pubblicazione: (2024)
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection
di: Chia-Pei, et al.
Pubblicazione: (2026)
di: Chia-Pei, et al.
Pubblicazione: (2026)
Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree
di: Johnson, Sam, et al.
Pubblicazione: (2025)
di: Johnson, Sam, et al.
Pubblicazione: (2025)
Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM Systems
di: Chang, Hongyan, et al.
Pubblicazione: (2026)
di: Chang, Hongyan, et al.
Pubblicazione: (2026)
QueryIPI: Query-agnostic Indirect Prompt Injection on Coding Agents
di: Xie, Yuchong, et al.
Pubblicazione: (2025)
di: Xie, Yuchong, et al.
Pubblicazione: (2025)
No More, No Less: Task Alignment in Terminal Agents
di: Mavali, Sina, et al.
Pubblicazione: (2026)
di: Mavali, Sina, et al.
Pubblicazione: (2026)
The Cognitive Firewall:Securing Browser Based AI Agents Against Indirect Prompt Injection Via Hybrid Edge Cloud Defense
di: Lan, Qianlong, et al.
Pubblicazione: (2026)
di: Lan, Qianlong, et al.
Pubblicazione: (2026)
ICON: Indirect Prompt Injection Defense for Agents based on Inference-Time Correction
di: Wang, Che, et al.
Pubblicazione: (2026)
di: Wang, Che, et al.
Pubblicazione: (2026)
Securing AI Agents Against Prompt Injection Attacks
di: Ramakrishnan, Badrinath, et al.
Pubblicazione: (2025)
di: Ramakrishnan, Badrinath, et al.
Pubblicazione: (2025)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
Efficient but Vulnerable: Benchmarking and Defending LLM Batch Prompting Attack
di: Yue, Murong, et al.
Pubblicazione: (2025)
di: Yue, Murong, et al.
Pubblicazione: (2025)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
di: Liu, Fan, et al.
Pubblicazione: (2024)
di: Liu, Fan, et al.
Pubblicazione: (2024)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
di: Li, Hao, et al.
Pubblicazione: (2026)
di: Li, Hao, et al.
Pubblicazione: (2026)
How Not to Detect Prompt Injections with an LLM
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025)
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025)
Documenti analoghi
-
AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection
di: Wang, Peiran, et al.
Pubblicazione: (2025) -
Defending Against Indirect Prompt Injection Attacks With Spotlighting
di: Hines, Keegan, et al.
Pubblicazione: (2024) -
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
di: Zhong, Peter Yong, et al.
Pubblicazione: (2025) -
IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents
di: An, Hengyu, et al.
Pubblicazione: (2025) -
Defending against Indirect Prompt Injection by Instruction Detection
di: Wen, Tongyu, et al.
Pubblicazione: (2025)