Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yige, Zhao, Wei, Li, Zhe, Min, Nay Myat, Huang, Hanxun, Zhao, Yunhan, Ma, Xingjun, Jiang, Yu-Gang, Sun, Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
von: Li, Yige, et al.
Veröffentlicht: (2025)
von: Li, Yige, et al.
Veröffentlicht: (2025)
Shortcuts Everywhere and Nowhere: Exploring Multi-Trigger Backdoor Attacks
von: Li, Yige, et al.
Veröffentlicht: (2024)
von: Li, Yige, et al.
Veröffentlicht: (2024)
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
von: Li, Jiayu, et al.
Veröffentlicht: (2025)
von: Li, Jiayu, et al.
Veröffentlicht: (2025)
Unified Neural Backdoor Removal with Only Few Clean Samples through Unlearning and Relearning
von: Min, Nay Myat, et al.
Veröffentlicht: (2024)
von: Min, Nay Myat, et al.
Veröffentlicht: (2024)
BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
von: Li, Yige, et al.
Veröffentlicht: (2024)
von: Li, Yige, et al.
Veröffentlicht: (2024)
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
von: Xu, Zonghuan, et al.
Veröffentlicht: (2025)
von: Xu, Zonghuan, et al.
Veröffentlicht: (2025)
End-to-End Anti-Backdoor Learning on Images and Time Series
von: Jiang, Yujing, et al.
Veröffentlicht: (2024)
von: Jiang, Yujing, et al.
Veröffentlicht: (2024)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
von: Jiang, Peihai, et al.
Veröffentlicht: (2025)
von: Jiang, Peihai, et al.
Veröffentlicht: (2025)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
von: Zhao, Yunhan, et al.
Veröffentlicht: (2026)
von: Zhao, Yunhan, et al.
Veröffentlicht: (2026)
BackdoorDM: A Comprehensive Benchmark for Backdoor Learning on Diffusion Model
von: Lin, Weilin, et al.
Veröffentlicht: (2025)
von: Lin, Weilin, et al.
Veröffentlicht: (2025)
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
von: Guo, Weiyang, et al.
Veröffentlicht: (2026)
von: Guo, Weiyang, et al.
Veröffentlicht: (2026)
Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models
von: Min, Nay Myat, et al.
Veröffentlicht: (2026)
von: Min, Nay Myat, et al.
Veröffentlicht: (2026)
CORVUS: Red-Teaming Hallucination Detectors via Internal Signal Camouflage in Large Language Models
von: Min, Nay Myat, et al.
Veröffentlicht: (2026)
von: Min, Nay Myat, et al.
Veröffentlicht: (2026)
The Invitation Trap: Proactive Availability Backdoor in LLMs via Conversational Induction
von: Wang, He, et al.
Veröffentlicht: (2026)
von: Wang, He, et al.
Veröffentlicht: (2026)
BackdoorBench: A Comprehensive Benchmark and Analysis of Backdoor Learning
von: Wu, Baoyuan, et al.
Veröffentlicht: (2024)
von: Wu, Baoyuan, et al.
Veröffentlicht: (2024)
Q-MLLM: Vector Quantization for Robust Multimodal Large Language Model Security
von: Zhao, Wei, et al.
Veröffentlicht: (2025)
von: Zhao, Wei, et al.
Veröffentlicht: (2025)
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
von: Wen, Rui, et al.
Veröffentlicht: (2026)
von: Wen, Rui, et al.
Veröffentlicht: (2026)
BackdoorIndicator: Leveraging OOD Data for Proactive Backdoor Detection in Federated Learning
von: Li, Songze, et al.
Veröffentlicht: (2024)
von: Li, Songze, et al.
Veröffentlicht: (2024)
Instruction Backdoor Attacks Against Customized LLMs
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
BackdoorMBTI: A Backdoor Learning Multimodal Benchmark Tool Kit for Backdoor Defense Evaluation
von: Yu, Haiyang, et al.
Veröffentlicht: (2024)
von: Yu, Haiyang, et al.
Veröffentlicht: (2024)
Internal Safety Collapse in Frontier Large Language Models
von: Wu, Yutao, et al.
Veröffentlicht: (2026)
von: Wu, Yutao, et al.
Veröffentlicht: (2026)
Infighting in the Dark: Multi-Label Backdoor Attack in Federated Learning
von: Li, Ye, et al.
Veröffentlicht: (2024)
von: Li, Ye, et al.
Veröffentlicht: (2024)
HoneypotNet: Backdoor Attacks Against Model Extraction
von: Wang, Yixu, et al.
Veröffentlicht: (2025)
von: Wang, Yixu, et al.
Veröffentlicht: (2025)
Real is not True: Backdoor Attacks Against Deepfake Detection
von: Sun, Hong, et al.
Veröffentlicht: (2024)
von: Sun, Hong, et al.
Veröffentlicht: (2024)
Stateful Agent Backdoor
von: Dai, Zhengchunmin, et al.
Veröffentlicht: (2026)
von: Dai, Zhengchunmin, et al.
Veröffentlicht: (2026)
Isolate Trigger: Detecting and Eliminating Adaptive Backdoor Attacks
von: Sun, Chengrui, et al.
Veröffentlicht: (2025)
von: Sun, Chengrui, et al.
Veröffentlicht: (2025)
BDFirewall: Towards Effective and Expeditiously Black-Box Backdoor Defense in MLaaS
von: Li, Ye, et al.
Veröffentlicht: (2025)
von: Li, Ye, et al.
Veröffentlicht: (2025)
Combinational Backdoor Attack against Customized Text-to-Image Models
von: Jiang, Wenbo, et al.
Veröffentlicht: (2024)
von: Jiang, Wenbo, et al.
Veröffentlicht: (2024)
Unleashing the Unseen: Harnessing Benign Datasets for Jailbreaking Large Language Models
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
von: Zhao, Wei, et al.
Veröffentlicht: (2024)
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
von: Zhao, Gejian, et al.
Veröffentlicht: (2025)
von: Zhao, Gejian, et al.
Veröffentlicht: (2025)
Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
T2UE: Generating Unlearnable Examples from Text Descriptions
von: Ma, Xingjun, et al.
Veröffentlicht: (2025)
von: Ma, Xingjun, et al.
Veröffentlicht: (2025)
BELT: Old-School Backdoor Attacks can Evade the State-of-the-Art Defense with Backdoor Exclusivity Lifting
von: Qiu, Huming, et al.
Veröffentlicht: (2023)
von: Qiu, Huming, et al.
Veröffentlicht: (2023)
Stealthy Targeted Backdoor Attacks against Image Captioning
von: Fan, Wenshu, et al.
Veröffentlicht: (2024)
von: Fan, Wenshu, et al.
Veröffentlicht: (2024)
Patronus: Identifying and Mitigating Transferable Backdoors in Pre-trained Language Models
von: Zhao, Tianhang, et al.
Veröffentlicht: (2025)
von: Zhao, Tianhang, et al.
Veröffentlicht: (2025)
Backdoor Sentinel: Detecting and Detoxifying Backdoors in Diffusion Models via Temporal Noise Consistency
von: Wang, Bingzheng, et al.
Veröffentlicht: (2026)
von: Wang, Bingzheng, et al.
Veröffentlicht: (2026)
Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
von: Yu, Miao, et al.
Veröffentlicht: (2025)
von: Yu, Miao, et al.
Veröffentlicht: (2025)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
Authority Backdoor: A Certifiable Backdoor Mechanism for Authoring DNNs
von: Yang, Han, et al.
Veröffentlicht: (2025)
von: Yang, Han, et al.
Veröffentlicht: (2025)
Rethinking Reasoning: A Survey on Reasoning-based Backdoors in LLMs
von: Hu, Man, et al.
Veröffentlicht: (2025)
von: Hu, Man, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
von: Li, Yige, et al.
Veröffentlicht: (2025) -
Shortcuts Everywhere and Nowhere: Exploring Multi-Trigger Backdoor Attacks
von: Li, Yige, et al.
Veröffentlicht: (2024) -
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
von: Li, Jiayu, et al.
Veröffentlicht: (2025) -
Unified Neural Backdoor Removal with Only Few Clean Samples through Unlearning and Relearning
von: Min, Nay Myat, et al.
Veröffentlicht: (2024) -
BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
von: Li, Yige, et al.
Veröffentlicht: (2024)