BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yulin, Li, Haoran, Zhang, Yirui, Zheng, Zihao, Song, Yangqiu, Hooi, Bryan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
von: Chen, Yulin, et al.
Veröffentlicht: (2024)
von: Chen, Yulin, et al.
Veröffentlicht: (2024)
TopicAttack: An Indirect Prompt Injection Attack via Topic Transition
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
Privacy in Large Language Models: Attacks, Defenses and Future Directions
von: Li, Haoran, et al.
Veröffentlicht: (2023)
von: Li, Haoran, et al.
Veröffentlicht: (2023)
Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models
von: Li, Haoran, et al.
Veröffentlicht: (2024)
von: Li, Haoran, et al.
Veröffentlicht: (2024)
Can Indirect Prompt Injection Attacks Be Detected and Removed?
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense
von: Hao, Shuyang, et al.
Veröffentlicht: (2025)
von: Hao, Shuyang, et al.
Veröffentlicht: (2025)
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
von: Chen, Yulin, et al.
Veröffentlicht: (2026)
von: Chen, Yulin, et al.
Veröffentlicht: (2026)
A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models
von: Xu, Zihao, et al.
Veröffentlicht: (2024)
von: Xu, Zihao, et al.
Veröffentlicht: (2024)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
von: Xiong, Chen, et al.
Veröffentlicht: (2024)
von: Xiong, Chen, et al.
Veröffentlicht: (2024)
$\textit{MMJ-Bench}$: A Comprehensive Study on Jailbreak Attacks and Defenses for Multimodal Large Language Models
von: Weng, Fenghua, et al.
Veröffentlicht: (2024)
von: Weng, Fenghua, et al.
Veröffentlicht: (2024)
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
von: Wang, Youze, et al.
Veröffentlicht: (2025)
von: Wang, Youze, et al.
Veröffentlicht: (2025)
Isolate Trigger: Detecting and Eliminating Adaptive Backdoor Attacks
von: Sun, Chengrui, et al.
Veröffentlicht: (2025)
von: Sun, Chengrui, et al.
Veröffentlicht: (2025)
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
von: You, Wenhao, et al.
Veröffentlicht: (2025)
von: You, Wenhao, et al.
Veröffentlicht: (2025)
Safe2Harm: Semantic Isomorphism Attacks for Jailbreaking Large Language Models
von: Yang, Fan
Veröffentlicht: (2025)
von: Yang, Fan
Veröffentlicht: (2025)
Rotated Robustness: A Training-Free Defense against Bit-Flip Attacks on Large Language Models
von: Liu, Deng, et al.
Veröffentlicht: (2026)
von: Liu, Deng, et al.
Veröffentlicht: (2026)
Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language Models
von: Truong, Vu Tuan, et al.
Veröffentlicht: (2026)
von: Truong, Vu Tuan, et al.
Veröffentlicht: (2026)
Towards Robust Multimodal Large Language Models Against Jailbreak Attacks
von: Yin, Ziyi, et al.
Veröffentlicht: (2025)
von: Yin, Ziyi, et al.
Veröffentlicht: (2025)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
von: Zeng, Yifan, et al.
Veröffentlicht: (2024)
FlipAttack: Jailbreak LLMs via Flipping
von: Liu, Yue, et al.
Veröffentlicht: (2024)
von: Liu, Yue, et al.
Veröffentlicht: (2024)
AISA: Awakening Intrinsic Safety Awareness in Large Language Models against Jailbreak Attacks
von: Song, Weiming, et al.
Veröffentlicht: (2026)
von: Song, Weiming, et al.
Veröffentlicht: (2026)
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
von: Lu, Lin, et al.
Veröffentlicht: (2024)
von: Lu, Lin, et al.
Veröffentlicht: (2024)
Double Backdoored: Converting Code Large Language Model Backdoors to Traditional Malware via Adversarial Instruction Tuning Attacks
von: Hossen, Md Imran, et al.
Veröffentlicht: (2024)
von: Hossen, Md Imran, et al.
Veröffentlicht: (2024)
Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
von: Zhong, Xingwei, et al.
Veröffentlicht: (2025)
von: Zhong, Xingwei, et al.
Veröffentlicht: (2025)
Jailbreaking Attack against Multimodal Large Language Model
von: Niu, Zhenxing, et al.
Veröffentlicht: (2024)
von: Niu, Zhenxing, et al.
Veröffentlicht: (2024)
Multimodal Large Language Models for Phishing Webpage Detection and Identification
von: Lee, Jehyun, et al.
Veröffentlicht: (2024)
von: Lee, Jehyun, et al.
Veröffentlicht: (2024)
HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
von: Narula, Sidhant, et al.
Veröffentlicht: (2025)
von: Narula, Sidhant, et al.
Veröffentlicht: (2025)
SoK: Robustness in Large Language Models against Jailbreak Attacks
von: Xu, Feiyue, et al.
Veröffentlicht: (2026)
von: Xu, Feiyue, et al.
Veröffentlicht: (2026)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
von: Park, Junyoung, et al.
Veröffentlicht: (2026)
von: Park, Junyoung, et al.
Veröffentlicht: (2026)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
von: Zheng, Jingyi, et al.
Veröffentlicht: (2024)
von: Zheng, Jingyi, et al.
Veröffentlicht: (2024)
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
von: Zhou, Yihe, et al.
Veröffentlicht: (2025)
von: Zhou, Yihe, et al.
Veröffentlicht: (2025)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
Backdoor Attacks and Defenses in Computer Vision Domain: A Survey
von: Abbasi, Bilal Hussain, et al.
Veröffentlicht: (2025)
von: Abbasi, Bilal Hussain, et al.
Veröffentlicht: (2025)
Distributed Backdoor Attacks on Federated Graph Learning and Certified Defenses
von: Yang, Yuxin, et al.
Veröffentlicht: (2024)
von: Yang, Yuxin, et al.
Veröffentlicht: (2024)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
When Forgetting Triggers Backdoors: A Clean Unlearning Attack
von: Arazzi, Marco, et al.
Veröffentlicht: (2025)
von: Arazzi, Marco, et al.
Veröffentlicht: (2025)
A Practical Trigger-Free Backdoor Attack on Neural Networks
von: Wang, Jiahao, et al.
Veröffentlicht: (2024)
von: Wang, Jiahao, et al.
Veröffentlicht: (2024)
Stealthy Dual-Trigger Backdoors: Attacking Prompt Tuning in LM-Empowered Graph Foundation Models
von: Xue, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Xue, Xiaoyu, et al.
Veröffentlicht: (2025)
Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
von: Teng, Ma, et al.
Veröffentlicht: (2024)
von: Teng, Ma, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
von: Chen, Yulin, et al.
Veröffentlicht: (2025) -
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
von: Chen, Yulin, et al.
Veröffentlicht: (2024) -
TopicAttack: An Indirect Prompt Injection Attack via Topic Transition
von: Chen, Yulin, et al.
Veröffentlicht: (2025) -
Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction
von: Chen, Yulin, et al.
Veröffentlicht: (2025) -
Privacy in Large Language Models: Attacks, Defenses and Future Directions
von: Li, Haoran, et al.
Veröffentlicht: (2023)