Mitigating Indirect Prompt Injection via Instruction-Following Intent Analysis
Fuente:
arXiv
Salvato in:
| Autori principali: | Kang, Mintong, Xiang, Chong, Kariyappa, Sanjay, Xiao, Chaowei, Li, Bo, Suh, Edward |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
di: Xiang, Chong, et al.
Pubblicazione: (2026)
di: Xiang, Chong, et al.
Pubblicazione: (2026)
ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning
di: Chen, Zhaorun, et al.
Pubblicazione: (2025)
di: Chen, Zhaorun, et al.
Pubblicazione: (2025)
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
di: Liao, Zeyi, et al.
Pubblicazione: (2024)
di: Liao, Zeyi, et al.
Pubblicazione: (2024)
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
di: Zhan, Qiusi, et al.
Pubblicazione: (2025)
di: Zhan, Qiusi, et al.
Pubblicazione: (2025)
Lessons from Defending Gemini Against Indirect Prompt Injections
di: Shi, Chongyang, et al.
Pubblicazione: (2025)
di: Shi, Chongyang, et al.
Pubblicazione: (2025)
FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks
di: Wang, Jiongxiao, et al.
Pubblicazione: (2024)
di: Wang, Jiongxiao, et al.
Pubblicazione: (2024)
ReasoningBomb: A Stealthy Denial-of-Service Attack by Inducing Pathologically Long Reasoning in Large Reasoning Models
di: Liu, Xiaogeng, et al.
Pubblicazione: (2026)
di: Liu, Xiaogeng, et al.
Pubblicazione: (2026)
CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution
di: Kim, Minbeom, et al.
Pubblicazione: (2026)
di: Kim, Minbeom, et al.
Pubblicazione: (2026)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
di: Hines, Keegan, et al.
Pubblicazione: (2024)
di: Hines, Keegan, et al.
Pubblicazione: (2024)
System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
di: Wu, Fangzhou, et al.
Pubblicazione: (2024)
di: Wu, Fangzhou, et al.
Pubblicazione: (2024)
Information Flow Control in Machine Learning through Modular Model Architecture
di: Tiwari, Trishita, et al.
Pubblicazione: (2023)
di: Tiwari, Trishita, et al.
Pubblicazione: (2023)
ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
di: Lu, Weikai, et al.
Pubblicazione: (2025)
di: Lu, Weikai, et al.
Pubblicazione: (2025)
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
di: Yan, Jun, et al.
Pubblicazione: (2023)
di: Yan, Jun, et al.
Pubblicazione: (2023)
AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
di: Li, Hao, et al.
Pubblicazione: (2026)
di: Li, Hao, et al.
Pubblicazione: (2026)
PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
di: Zou, Wei, et al.
Pubblicazione: (2025)
di: Zou, Wei, et al.
Pubblicazione: (2025)
AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
di: Liu, Xiaogeng, et al.
Pubblicazione: (2024)
di: Liu, Xiaogeng, et al.
Pubblicazione: (2024)
Stronger Enforcement of Instruction Hierarchy via Augmented Intermediate Representations
di: Kariyappa, Sanjay, et al.
Pubblicazione: (2025)
di: Kariyappa, Sanjay, et al.
Pubblicazione: (2025)
Preventing Prompt Injection with Type-Directed Privilege Separation
di: Jacob, Dennis, et al.
Pubblicazione: (2025)
di: Jacob, Dennis, et al.
Pubblicazione: (2025)
Defending against Indirect Prompt Injection by Instruction Detection
di: Wen, Tongyu, et al.
Pubblicazione: (2025)
di: Wen, Tongyu, et al.
Pubblicazione: (2025)
SecAlign: Defending Against Prompt Injection with Preference Optimization
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
Design Patterns for Securing LLM Agents against Prompt Injections
di: Beurer-Kellner, Luca, et al.
Pubblicazione: (2025)
di: Beurer-Kellner, Luca, et al.
Pubblicazione: (2025)
The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
di: Jia, Feiran, et al.
Pubblicazione: (2024)
di: Jia, Feiran, et al.
Pubblicazione: (2024)
PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
di: Yin, Chenlong, et al.
Pubblicazione: (2026)
di: Yin, Chenlong, et al.
Pubblicazione: (2026)
TopicAttack: An Indirect Prompt Injection Attack via Topic Transition
di: Chen, Yulin, et al.
Pubblicazione: (2025)
di: Chen, Yulin, et al.
Pubblicazione: (2025)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
di: Nasr, Milad, et al.
Pubblicazione: (2025)
di: Nasr, Milad, et al.
Pubblicazione: (2025)
Neural Exec: Learning (and Learning from) Execution Triggers for Prompt Injection Attacks
di: Pasquini, Dario, et al.
Pubblicazione: (2024)
di: Pasquini, Dario, et al.
Pubblicazione: (2024)
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
di: Hossain, S M Asif, et al.
Pubblicazione: (2025)
di: Hossain, S M Asif, et al.
Pubblicazione: (2025)
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
di: Xiang, Zhen, et al.
Pubblicazione: (2024)
di: Xiang, Zhen, et al.
Pubblicazione: (2024)
Send a SCOUT First: Pre-hoc Reasoning for Adaptive Detector Allocation in Prompt-Injection Defense
di: Zhang, Shuhao, et al.
Pubblicazione: (2026)
di: Zhang, Shuhao, et al.
Pubblicazione: (2026)
PromptSleuth: Detecting Prompt Injection via Semantic Intent Invariance
di: Wang, Mengxiao, et al.
Pubblicazione: (2025)
di: Wang, Mengxiao, et al.
Pubblicazione: (2025)
GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks
di: Li, Rongchang, et al.
Pubblicazione: (2024)
di: Li, Rongchang, et al.
Pubblicazione: (2024)
Empirical Analysis of Large Vision-Language Models against Goal Hijacking via Visual Prompt Injection
di: Kimura, Subaru, et al.
Pubblicazione: (2024)
di: Kimura, Subaru, et al.
Pubblicazione: (2024)
Machine Learning with Privacy for Protected Attributes
di: Mahloujifar, Saeed, et al.
Pubblicazione: (2025)
di: Mahloujifar, Saeed, et al.
Pubblicazione: (2025)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
di: Xu, Jiashu, et al.
Pubblicazione: (2023)
di: Xu, Jiashu, et al.
Pubblicazione: (2023)
Backdoored Retrievers for Prompt Injection Attacks on Retrieval Augmented Generation of Large Language Models
di: Clop, Cody, et al.
Pubblicazione: (2024)
di: Clop, Cody, et al.
Pubblicazione: (2024)
How Not to Detect Prompt Injections with an LLM
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025)
di: Choudhary, Sarthak, et al.
Pubblicazione: (2025)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2024)
di: Debenedetti, Edoardo, et al.
Pubblicazione: (2024)
Can Indirect Prompt Injection Attacks Be Detected and Removed?
di: Chen, Yulin, et al.
Pubblicazione: (2025)
di: Chen, Yulin, et al.
Pubblicazione: (2025)
Mitigating Data Injection Attacks on Federated Learning
di: Shalom, Or, et al.
Pubblicazione: (2023)
di: Shalom, Or, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
di: Xiang, Chong, et al.
Pubblicazione: (2026) -
ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning
di: Chen, Zhaorun, et al.
Pubblicazione: (2025) -
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
di: Liao, Zeyi, et al.
Pubblicazione: (2024) -
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
di: Zhan, Qiusi, et al.
Pubblicazione: (2025) -
Lessons from Defending Gemini Against Indirect Prompt Injections
di: Shi, Chongyang, et al.
Pubblicazione: (2025)