ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Xuxu, Liang, Siyuan, Han, Mengya, Luo, Yong, Liu, Aishan, Cai, Xiantao, He, Zheng, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ME: Trigger Element Combination Backdoor Attack on Copyright Infringement
von: Yang, Feiyu, et al.
Veröffentlicht: (2025)
von: Yang, Feiyu, et al.
Veröffentlicht: (2025)
ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
von: Ren, Zhiyao, et al.
Veröffentlicht: (2025)
von: Ren, Zhiyao, et al.
Veröffentlicht: (2025)
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
Compromising Embodied Agents with Contextual Backdoor Attacks
von: Liu, Aishan, et al.
Veröffentlicht: (2024)
von: Liu, Aishan, et al.
Veröffentlicht: (2024)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
CtrlAttack: A Unified Attack on World-Model Control in Diffusion Models
von: Xu, Shuhan, et al.
Veröffentlicht: (2026)
von: Xu, Shuhan, et al.
Veröffentlicht: (2026)
TrapFlow: Controllable Website Fingerprinting Defense via Dynamic Backdoor Learning
von: Liang, Siyuan, et al.
Veröffentlicht: (2024)
von: Liang, Siyuan, et al.
Veröffentlicht: (2024)
T2VShield: Model-Agnostic Jailbreak Defense for Text-to-Video Models
von: Liang, Siyuan, et al.
Veröffentlicht: (2025)
von: Liang, Siyuan, et al.
Veröffentlicht: (2025)
PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-Tuning
von: Sun, Zhen, et al.
Veröffentlicht: (2024)
von: Sun, Zhen, et al.
Veröffentlicht: (2024)
BackdoorBench: A Comprehensive Benchmark and Analysis of Backdoor Learning
von: Wu, Baoyuan, et al.
Veröffentlicht: (2024)
von: Wu, Baoyuan, et al.
Veröffentlicht: (2024)
Poison Once, Control Anywhere: Clean-Text Visual Backdoors in VLM-based Mobile Agents
von: Wang, Xuan, et al.
Veröffentlicht: (2025)
von: Wang, Xuan, et al.
Veröffentlicht: (2025)
Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
Does Few-shot Learning Suffer from Backdoor Attacks?
von: Liu, Xinwei, et al.
Veröffentlicht: (2023)
von: Liu, Xinwei, et al.
Veröffentlicht: (2023)
External Data Extraction Attacks against Retrieval-Augmented Large Language Models
von: He, Yu, et al.
Veröffentlicht: (2025)
von: He, Yu, et al.
Veröffentlicht: (2025)
Inhibitory Attacks on Backdoor-based Fingerprinting for Large Language Models
von: Fu, Hang, et al.
Veröffentlicht: (2026)
von: Fu, Hang, et al.
Veröffentlicht: (2026)
TimeGuard: Channel-wise Pool Training for Backdoor Defense in Time Series Forecasting
von: Nguyen, Quang Duc, et al.
Veröffentlicht: (2026)
von: Nguyen, Quang Duc, et al.
Veröffentlicht: (2026)
Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation
von: Zhong, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zhong, Zhiyuan, et al.
Veröffentlicht: (2025)
AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding
von: Ma, Haokai, et al.
Veröffentlicht: (2025)
von: Ma, Haokai, et al.
Veröffentlicht: (2025)
DETOUR: A Practical Backdoor Attack against Object Detection
von: Liu, Dazhuang, et al.
Veröffentlicht: (2026)
von: Liu, Dazhuang, et al.
Veröffentlicht: (2026)
Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models
von: Yuan, Zenghui, et al.
Veröffentlicht: (2025)
von: Yuan, Zenghui, et al.
Veröffentlicht: (2025)
Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models
von: Ni, Zhenyang, et al.
Veröffentlicht: (2024)
von: Ni, Zhenyang, et al.
Veröffentlicht: (2024)
BlackboxBench: A Comprehensive Benchmark of Black-box Adversarial Attacks
von: Zheng, Meixi, et al.
Veröffentlicht: (2023)
von: Zheng, Meixi, et al.
Veröffentlicht: (2023)
Double Backdoored: Converting Code Large Language Model Backdoors to Traditional Malware via Adversarial Instruction Tuning Attacks
von: Hossen, Md Imran, et al.
Veröffentlicht: (2024)
von: Hossen, Md Imran, et al.
Veröffentlicht: (2024)
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
von: Chen, Yulin, et al.
Veröffentlicht: (2024)
von: Chen, Yulin, et al.
Veröffentlicht: (2024)
Professor X: Manipulating EEG BCI with Invisible and Robust Backdoor Attack
von: Liu, Xuan-Hao, et al.
Veröffentlicht: (2024)
von: Liu, Xuan-Hao, et al.
Veröffentlicht: (2024)
Gradient Shaping: Enhancing Backdoor Attack Against Reverse Engineering
von: Zhu, Rui, et al.
Veröffentlicht: (2023)
von: Zhu, Rui, et al.
Veröffentlicht: (2023)
JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
von: Luo, Weidi, et al.
Veröffentlicht: (2024)
von: Luo, Weidi, et al.
Veröffentlicht: (2024)
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
von: Zhou, Yihe, et al.
Veröffentlicht: (2025)
von: Zhou, Yihe, et al.
Veröffentlicht: (2025)
Stealthy Backdoor Attack via Confidence-driven Sampling
von: He, Pengfei, et al.
Veröffentlicht: (2023)
von: He, Pengfei, et al.
Veröffentlicht: (2023)
Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character
von: Ma, Siyuan, et al.
Veröffentlicht: (2024)
von: Ma, Siyuan, et al.
Veröffentlicht: (2024)
BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models
von: Yuan, Zenghui, et al.
Veröffentlicht: (2025)
von: Yuan, Zenghui, et al.
Veröffentlicht: (2025)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
von: Zheng, Jingyi, et al.
Veröffentlicht: (2024)
von: Zheng, Jingyi, et al.
Veröffentlicht: (2024)
BackdoorDM: A Comprehensive Benchmark for Backdoor Learning on Diffusion Model
von: Lin, Weilin, et al.
Veröffentlicht: (2025)
von: Lin, Weilin, et al.
Veröffentlicht: (2025)
SecReEvalBench: A Multi-turned Security Resilience Evaluation Benchmark for Large Language Models
von: Cui, Huining, et al.
Veröffentlicht: (2025)
von: Cui, Huining, et al.
Veröffentlicht: (2025)
Real is not True: Backdoor Attacks Against Deepfake Detection
von: Sun, Hong, et al.
Veröffentlicht: (2024)
von: Sun, Hong, et al.
Veröffentlicht: (2024)
Towards Physical World Backdoor Attacks against Skeleton Action Recognition
von: Zheng, Qichen, et al.
Veröffentlicht: (2024)
von: Zheng, Qichen, et al.
Veröffentlicht: (2024)
Backdoor Attacks and Countermeasures in Natural Language Processing Models: A Comprehensive Security Review
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2023)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2023)
Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models
von: Li, Xi, et al.
Veröffentlicht: (2024)
von: Li, Xi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ME: Trigger Element Combination Backdoor Attack on Copyright Infringement
von: Yang, Feiyu, et al.
Veröffentlicht: (2025) -
ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
von: Ren, Zhiyao, et al.
Veröffentlicht: (2025) -
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2024) -
Compromising Embodied Agents with Contextual Backdoor Attacks
von: Liu, Aishan, et al.
Veröffentlicht: (2024) -
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)