BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yige, Huang, Hanxun, Zhao, Yunhan, Ma, Xingjun, Sun, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
by: Li, Yige, et al.
Published: (2025)
by: Li, Yige, et al.
Published: (2025)
Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs
by: Li, Yige, et al.
Published: (2026)
by: Li, Yige, et al.
Published: (2026)
Expose Before You Defend: Unifying and Enhancing Backdoor Defenses via Exposed Models
by: Li, Yige, et al.
Published: (2024)
by: Li, Yige, et al.
Published: (2024)
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
by: Li, Jiayu, et al.
Published: (2025)
by: Li, Jiayu, et al.
Published: (2025)
BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
by: Li, Juncheng, et al.
Published: (2025)
by: Li, Juncheng, et al.
Published: (2025)
BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
by: Xu, Zonghuan, et al.
Published: (2025)
by: Xu, Zonghuan, et al.
Published: (2025)
Defense-to-Attack: Bypassing Weak Defenses Enables Stronger Jailbreaks in Vision-Language Models
by: Zhao, Yunhan, et al.
Published: (2025)
by: Zhao, Yunhan, et al.
Published: (2025)
Shortcuts Everywhere and Nowhere: Exploring Multi-Trigger Backdoor Attacks
by: Li, Yige, et al.
Published: (2024)
by: Li, Yige, et al.
Published: (2024)
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
by: Huang, Hanxun, et al.
Published: (2025)
by: Huang, Hanxun, et al.
Published: (2025)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
by: Jiang, Peihai, et al.
Published: (2025)
by: Jiang, Peihai, et al.
Published: (2025)
Detecting Backdoor Samples in Contrastive Language Image Pretraining
by: Huang, Hanxun, et al.
Published: (2025)
by: Huang, Hanxun, et al.
Published: (2025)
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
by: Cui, Kaiyuan, et al.
Published: (2026)
by: Cui, Kaiyuan, et al.
Published: (2026)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Exploring Backdoor Attack and Defense for LLM-empowered Recommendations
by: Ning, Liangbo, et al.
Published: (2025)
by: Ning, Liangbo, et al.
Published: (2025)
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
by: Zhou, Yihe, et al.
Published: (2025)
by: Zhou, Yihe, et al.
Published: (2025)
CROW: Eliminating Backdoors from Large Language Models via Internal Consistency Regularization
by: Min, Nay Myat, et al.
Published: (2024)
by: Min, Nay Myat, et al.
Published: (2024)
BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks
by: Zhao, Yunhan, et al.
Published: (2024)
by: Zhao, Yunhan, et al.
Published: (2024)
BackdoorMBTI: A Backdoor Learning Multimodal Benchmark Tool Kit for Backdoor Defense Evaluation
by: Yu, Haiyang, et al.
Published: (2024)
by: Yu, Haiyang, et al.
Published: (2024)
Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses
by: Pawlak, Stanisław, et al.
Published: (2025)
by: Pawlak, Stanisław, et al.
Published: (2025)
UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
by: Lin, Huawei, et al.
Published: (2025)
by: Lin, Huawei, et al.
Published: (2025)
BELT: Old-School Backdoor Attacks can Evade the State-of-the-Art Defense with Backdoor Exclusivity Lifting
by: Qiu, Huming, et al.
Published: (2023)
by: Qiu, Huming, et al.
Published: (2023)
ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
by: Liu, Xuxu, et al.
Published: (2025)
by: Liu, Xuxu, et al.
Published: (2025)
When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations
by: Ge, Huaizhi, et al.
Published: (2024)
by: Ge, Huaizhi, et al.
Published: (2024)
Certifying Language Model Robustness with Fuzzed Randomized Smoothing: An Efficient Defense Against Backdoor Attacks
by: He, Bowei, et al.
Published: (2025)
by: He, Bowei, et al.
Published: (2025)
Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models
by: Li, Xi, et al.
Published: (2024)
by: Li, Xi, et al.
Published: (2024)
Downstream Transfer Attack: Adversarial Attacks on Downstream Models with Pre-trained Vision Transformers
by: Zheng, Weijie, et al.
Published: (2024)
by: Zheng, Weijie, et al.
Published: (2024)
Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language Models
by: Truong, Vu Tuan, et al.
Published: (2026)
by: Truong, Vu Tuan, et al.
Published: (2026)
Concept-Guided Backdoor Attack on Vision Language Models
by: Shen, Haoyu, et al.
Published: (2025)
by: Shen, Haoyu, et al.
Published: (2025)
Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing
by: Zhao, Wei, et al.
Published: (2024)
by: Zhao, Wei, et al.
Published: (2024)
Internal Safety Collapse in Frontier Large Language Models
by: Wu, Yutao, et al.
Published: (2026)
by: Wu, Yutao, et al.
Published: (2026)
Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
by: Zhao, Shuai, et al.
Published: (2023)
by: Zhao, Shuai, et al.
Published: (2023)
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
TED-LaST: Towards Robust Backdoor Defense Against Adaptive Attacks
by: Mo, Xiaoxing, et al.
Published: (2025)
by: Mo, Xiaoxing, et al.
Published: (2025)
UOR: Universal Backdoor Attacks on Pre-trained Language Models
by: Du, Wei, et al.
Published: (2023)
by: Du, Wei, et al.
Published: (2023)
Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
by: Yu, Miao, et al.
Published: (2025)
by: Yu, Miao, et al.
Published: (2025)
MEGen: Generative Backdoor into Large Language Models via Model Editing
by: Qiu, Jiyang, et al.
Published: (2024)
by: Qiu, Jiyang, et al.
Published: (2024)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
by: Xu, Jiashu, et al.
Published: (2023)
by: Xu, Jiashu, et al.
Published: (2023)
CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
by: Li, Yuetai, et al.
Published: (2024)
by: Li, Yuetai, et al.
Published: (2024)
Defending the Edge: Representative-Attention Defense against Backdoor Attacks in Federated Learning
by: Obioma, Chibueze Peace, et al.
Published: (2025)
by: Obioma, Chibueze Peace, et al.
Published: (2025)
Similar Items
-
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
by: Li, Yige, et al.
Published: (2025) -
Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs
by: Li, Yige, et al.
Published: (2026) -
Expose Before You Defend: Unifying and Enhancing Backdoor Defenses via Exposed Models
by: Li, Yige, et al.
Published: (2024) -
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
by: Li, Jiayu, et al.
Published: (2025) -
BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
by: Li, Juncheng, et al.
Published: (2025)