Defending Against Indirect Prompt Injection Attacks With Spotlighting
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hines, Keegan, Lopez, Gary, Hall, Matthew, Zarfati, Federico, Zunger, Yonatan, Kiciman, Emre |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
von: Jia, Feiran, et al.
Veröffentlicht: (2024)
von: Jia, Feiran, et al.
Veröffentlicht: (2024)
Lessons from Defending Gemini Against Indirect Prompt Injections
von: Shi, Chongyang, et al.
Veröffentlicht: (2025)
von: Shi, Chongyang, et al.
Veröffentlicht: (2025)
SPML: A DSL for Defending Language Models Against Prompt Attacks
von: Sharma, Reshabh K, et al.
Veröffentlicht: (2024)
von: Sharma, Reshabh K, et al.
Veröffentlicht: (2024)
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)
An Early Categorization of Prompt Injection Attacks on Large Language Models
von: Rossi, Sippo, et al.
Veröffentlicht: (2024)
von: Rossi, Sippo, et al.
Veröffentlicht: (2024)
Securing Large Language Models (LLMs) from Prompt Injection Attacks
von: Suri, Omar Farooq Khan, et al.
Veröffentlicht: (2025)
von: Suri, Omar Farooq Khan, et al.
Veröffentlicht: (2025)
AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection
von: Wang, Peiran, et al.
Veröffentlicht: (2025)
von: Wang, Peiran, et al.
Veröffentlicht: (2025)
SecAlign: Defending Against Prompt Injection with Preference Optimization
von: Chen, Sizhe, et al.
Veröffentlicht: (2024)
von: Chen, Sizhe, et al.
Veröffentlicht: (2024)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks
von: Li, Rongchang, et al.
Veröffentlicht: (2024)
von: Li, Rongchang, et al.
Veröffentlicht: (2024)
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
von: Liu, Yupei, et al.
Veröffentlicht: (2023)
von: Liu, Yupei, et al.
Veröffentlicht: (2023)
Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
von: Zhao, Andrew, et al.
Veröffentlicht: (2025)
von: Zhao, Andrew, et al.
Veröffentlicht: (2025)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
von: Shao, Zedian, et al.
Veröffentlicht: (2024)
von: Shao, Zedian, et al.
Veröffentlicht: (2024)
Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
von: Zhong, Yinan, et al.
Veröffentlicht: (2025)
von: Zhong, Yinan, et al.
Veröffentlicht: (2025)
FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
von: Zhang, Mohan, et al.
Veröffentlicht: (2026)
von: Zhang, Mohan, et al.
Veröffentlicht: (2026)
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
von: Yang, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Yang, Xiaoxue, et al.
Veröffentlicht: (2025)
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
von: Yan, Jun, et al.
Veröffentlicht: (2023)
von: Yan, Jun, et al.
Veröffentlicht: (2023)
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
von: Wang, Jiawen, et al.
Veröffentlicht: (2025)
von: Wang, Jiawen, et al.
Veröffentlicht: (2025)
Composite Backdoor Attacks Against Large Language Models
von: Huang, Hai, et al.
Veröffentlicht: (2023)
von: Huang, Hai, et al.
Veröffentlicht: (2023)
A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy
von: Correia, Pedro H. Barcha, et al.
Veröffentlicht: (2026)
von: Correia, Pedro H. Barcha, et al.
Veröffentlicht: (2026)
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
von: Qian, Cheng, et al.
Veröffentlicht: (2024)
von: Qian, Cheng, et al.
Veröffentlicht: (2024)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
von: Brown, Hannah, et al.
Veröffentlicht: (2024)
von: Brown, Hannah, et al.
Veröffentlicht: (2024)
Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
von: Yi, Jingwei, et al.
Veröffentlicht: (2023)
von: Yi, Jingwei, et al.
Veröffentlicht: (2023)
Empirical Analysis of Large Vision-Language Models against Goal Hijacking via Visual Prompt Injection
von: Kimura, Subaru, et al.
Veröffentlicht: (2024)
von: Kimura, Subaru, et al.
Veröffentlicht: (2024)
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
von: Kim, Heegyu, et al.
Veröffentlicht: (2024)
von: Kim, Heegyu, et al.
Veröffentlicht: (2024)
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
von: Hossain, S M Asif, et al.
Veröffentlicht: (2025)
von: Hossain, S M Asif, et al.
Veröffentlicht: (2025)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
von: Jiang, Weisen, et al.
Veröffentlicht: (2025)
von: Jiang, Weisen, et al.
Veröffentlicht: (2025)
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
von: Li, Qizhang, et al.
Veröffentlicht: (2024)
von: Li, Qizhang, et al.
Veröffentlicht: (2024)
PIArena: A Platform for Prompt Injection Evaluation
von: Geng, Runpeng, et al.
Veröffentlicht: (2026)
von: Geng, Runpeng, et al.
Veröffentlicht: (2026)
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
von: Geng, Runpeng, et al.
Veröffentlicht: (2025)
von: Geng, Runpeng, et al.
Veröffentlicht: (2025)
Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning
von: An, Li, et al.
Veröffentlicht: (2025)
von: An, Li, et al.
Veröffentlicht: (2025)
Mitigating Indirect Prompt Injection via Instruction-Following Intent Analysis
von: Kang, Mintong, et al.
Veröffentlicht: (2025)
von: Kang, Mintong, et al.
Veröffentlicht: (2025)
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
von: Chen, Shuo, et al.
Veröffentlicht: (2024)
von: Chen, Shuo, et al.
Veröffentlicht: (2024)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
von: Benjamin, Victoria, et al.
Veröffentlicht: (2024)
von: Benjamin, Victoria, et al.
Veröffentlicht: (2024)
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
von: Liao, Zeyi, et al.
Veröffentlicht: (2024)
von: Liao, Zeyi, et al.
Veröffentlicht: (2024)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
von: Nasr, Milad, et al.
Veröffentlicht: (2025)
von: Nasr, Milad, et al.
Veröffentlicht: (2025)
Defending Against Sophisticated Poisoning Attacks with RL-based Aggregation in Federated Learning
von: Wang, Yujing, et al.
Veröffentlicht: (2024)
von: Wang, Yujing, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
von: Jia, Feiran, et al.
Veröffentlicht: (2024) -
Lessons from Defending Gemini Against Indirect Prompt Injections
von: Shi, Chongyang, et al.
Veröffentlicht: (2025) -
SPML: A DSL for Defending Language Models Against Prompt Attacks
von: Sharma, Reshabh K, et al.
Veröffentlicht: (2024) -
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025) -
From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks
von: Zhang, Zhexin, et al.
Veröffentlicht: (2024)