ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Xuxu, Liang, Siyuan, Han, Mengya, Luo, Yong, Liu, Aishan, Cai, Xiantao, He, Zheng, Tao, Dacheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ME: Trigger Element Combination Backdoor Attack on Copyright Infringement
por: Yang, Feiyu, et al.
Publicado: (2025)
por: Yang, Feiyu, et al.
Publicado: (2025)
ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
por: Ren, Zhiyao, et al.
Publicado: (2025)
por: Ren, Zhiyao, et al.
Publicado: (2025)
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
por: Ying, Zonghao, et al.
Publicado: (2024)
por: Ying, Zonghao, et al.
Publicado: (2024)
Compromising Embodied Agents with Contextual Backdoor Attacks
por: Liu, Aishan, et al.
Publicado: (2024)
por: Liu, Aishan, et al.
Publicado: (2024)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
por: Ying, Zonghao, et al.
Publicado: (2025)
por: Ying, Zonghao, et al.
Publicado: (2025)
Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak Attacks
por: Ying, Zonghao, et al.
Publicado: (2024)
por: Ying, Zonghao, et al.
Publicado: (2024)
Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt
por: Ying, Zonghao, et al.
Publicado: (2024)
por: Ying, Zonghao, et al.
Publicado: (2024)
CtrlAttack: A Unified Attack on World-Model Control in Diffusion Models
por: Xu, Shuhan, et al.
Publicado: (2026)
por: Xu, Shuhan, et al.
Publicado: (2026)
TrapFlow: Controllable Website Fingerprinting Defense via Dynamic Backdoor Learning
por: Liang, Siyuan, et al.
Publicado: (2024)
por: Liang, Siyuan, et al.
Publicado: (2024)
T2VShield: Model-Agnostic Jailbreak Defense for Text-to-Video Models
por: Liang, Siyuan, et al.
Publicado: (2025)
por: Liang, Siyuan, et al.
Publicado: (2025)
PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-Tuning
por: Sun, Zhen, et al.
Publicado: (2024)
por: Sun, Zhen, et al.
Publicado: (2024)
BackdoorBench: A Comprehensive Benchmark and Analysis of Backdoor Learning
por: Wu, Baoyuan, et al.
Publicado: (2024)
por: Wu, Baoyuan, et al.
Publicado: (2024)
Poison Once, Control Anywhere: Clean-Text Visual Backdoors in VLM-based Mobile Agents
por: Wang, Xuan, et al.
Publicado: (2025)
por: Wang, Xuan, et al.
Publicado: (2025)
Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks
por: Li, Ziqiang, et al.
Publicado: (2024)
por: Li, Ziqiang, et al.
Publicado: (2024)
Does Few-shot Learning Suffer from Backdoor Attacks?
por: Liu, Xinwei, et al.
Publicado: (2023)
por: Liu, Xinwei, et al.
Publicado: (2023)
External Data Extraction Attacks against Retrieval-Augmented Large Language Models
por: He, Yu, et al.
Publicado: (2025)
por: He, Yu, et al.
Publicado: (2025)
Inhibitory Attacks on Backdoor-based Fingerprinting for Large Language Models
por: Fu, Hang, et al.
Publicado: (2026)
por: Fu, Hang, et al.
Publicado: (2026)
TimeGuard: Channel-wise Pool Training for Backdoor Defense in Time Series Forecasting
por: Nguyen, Quang Duc, et al.
Publicado: (2026)
por: Nguyen, Quang Duc, et al.
Publicado: (2026)
Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation
por: Zhong, Zhiyuan, et al.
Publicado: (2025)
por: Zhong, Zhiyuan, et al.
Publicado: (2025)
AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding
por: Ma, Haokai, et al.
Publicado: (2025)
por: Ma, Haokai, et al.
Publicado: (2025)
DETOUR: A Practical Backdoor Attack against Object Detection
por: Liu, Dazhuang, et al.
Publicado: (2026)
por: Liu, Dazhuang, et al.
Publicado: (2026)
Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models
por: Yuan, Zenghui, et al.
Publicado: (2025)
por: Yuan, Zenghui, et al.
Publicado: (2025)
Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models
por: Ni, Zhenyang, et al.
Publicado: (2024)
por: Ni, Zhenyang, et al.
Publicado: (2024)
BlackboxBench: A Comprehensive Benchmark of Black-box Adversarial Attacks
por: Zheng, Meixi, et al.
Publicado: (2023)
por: Zheng, Meixi, et al.
Publicado: (2023)
Double Backdoored: Converting Code Large Language Model Backdoors to Traditional Malware via Adversarial Instruction Tuning Attacks
por: Hossen, Md Imran, et al.
Publicado: (2024)
por: Hossen, Md Imran, et al.
Publicado: (2024)
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
por: Chen, Yulin, et al.
Publicado: (2024)
por: Chen, Yulin, et al.
Publicado: (2024)
Professor X: Manipulating EEG BCI with Invisible and Robust Backdoor Attack
por: Liu, Xuan-Hao, et al.
Publicado: (2024)
por: Liu, Xuan-Hao, et al.
Publicado: (2024)
Gradient Shaping: Enhancing Backdoor Attack Against Reverse Engineering
por: Zhu, Rui, et al.
Publicado: (2023)
por: Zhu, Rui, et al.
Publicado: (2023)
JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
por: Luo, Weidi, et al.
Publicado: (2024)
por: Luo, Weidi, et al.
Publicado: (2024)
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
por: Zhou, Yihe, et al.
Publicado: (2025)
por: Zhou, Yihe, et al.
Publicado: (2025)
Stealthy Backdoor Attack via Confidence-driven Sampling
por: He, Pengfei, et al.
Publicado: (2023)
por: He, Pengfei, et al.
Publicado: (2023)
Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character
por: Ma, Siyuan, et al.
Publicado: (2024)
por: Ma, Siyuan, et al.
Publicado: (2024)
BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models
por: Yuan, Zenghui, et al.
Publicado: (2025)
por: Yuan, Zenghui, et al.
Publicado: (2025)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
por: Zheng, Jingyi, et al.
Publicado: (2024)
por: Zheng, Jingyi, et al.
Publicado: (2024)
BackdoorDM: A Comprehensive Benchmark for Backdoor Learning on Diffusion Model
por: Lin, Weilin, et al.
Publicado: (2025)
por: Lin, Weilin, et al.
Publicado: (2025)
SecReEvalBench: A Multi-turned Security Resilience Evaluation Benchmark for Large Language Models
por: Cui, Huining, et al.
Publicado: (2025)
por: Cui, Huining, et al.
Publicado: (2025)
Real is not True: Backdoor Attacks Against Deepfake Detection
por: Sun, Hong, et al.
Publicado: (2024)
por: Sun, Hong, et al.
Publicado: (2024)
Towards Physical World Backdoor Attacks against Skeleton Action Recognition
por: Zheng, Qichen, et al.
Publicado: (2024)
por: Zheng, Qichen, et al.
Publicado: (2024)
Backdoor Attacks and Countermeasures in Natural Language Processing Models: A Comprehensive Security Review
por: Cheng, Pengzhou, et al.
Publicado: (2023)
por: Cheng, Pengzhou, et al.
Publicado: (2023)
Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models
por: Li, Xi, et al.
Publicado: (2024)
por: Li, Xi, et al.
Publicado: (2024)
Ejemplares similares
-
ME: Trigger Element Combination Backdoor Attack on Copyright Infringement
por: Yang, Feiyu, et al.
Publicado: (2025) -
ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
por: Ren, Zhiyao, et al.
Publicado: (2025) -
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
por: Ying, Zonghao, et al.
Publicado: (2024) -
Compromising Embodied Agents with Contextual Backdoor Attacks
por: Liu, Aishan, et al.
Publicado: (2024) -
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
por: Ying, Zonghao, et al.
Publicado: (2025)