Elijah: Eliminating Backdoors Injected in Diffusion Models via Distribution Shift
Fuente:
arXiv
Saved in:
| Main Authors: | An, Shengwei, Chou, Sheng-Yen, Zhang, Kaiyuan, Xu, Qiuling, Tao, Guanhong, Shen, Guangyu, Cheng, Siyuan, Ma, Shiqing, Chen, Pin-Yu, Ho, Tsung-Yi, Zhang, Xiangyu |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UNIT: Backdoor Mitigation via Automated Neural Distribution Tightening
by: Cheng, Siyuan, et al.
Published: (2024)
by: Cheng, Siyuan, et al.
Published: (2024)
LOTUS: Evasive and Resilient Backdoor Attacks through Sub-Partitioning
by: Cheng, Siyuan, et al.
Published: (2024)
by: Cheng, Siyuan, et al.
Published: (2024)
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
by: Shen, Guangyu, et al.
Published: (2024)
by: Shen, Guangyu, et al.
Published: (2024)
VillanDiffusion: A Unified Backdoor Attack Framework for Diffusion Models
by: Chou, Sheng-Yen, et al.
Published: (2023)
by: Chou, Sheng-Yen, et al.
Published: (2023)
Backdooring Masked Diffusion Language Models
by: Cao, Daniel Yiming, et al.
Published: (2026)
by: Cao, Daniel Yiming, et al.
Published: (2026)
ASPIRER: Bypassing System Prompts With Permutation-based Backdoors in LLMs
by: Yan, Lu, et al.
Published: (2024)
by: Yan, Lu, et al.
Published: (2024)
CENSOR: Defense Against Gradient Inversion via Orthogonal Subspace Bayesian Sampling
by: Zhang, Kaiyuan, et al.
Published: (2025)
by: Zhang, Kaiyuan, et al.
Published: (2025)
Opening A Pandora's Box: Things You Should Know in the Era of Custom GPTs
by: Tao, Guanhong, et al.
Published: (2023)
by: Tao, Guanhong, et al.
Published: (2023)
Rethinking Backdoor Attacks on Dataset Distillation: A Kernel Method Perspective
by: Chung, Ming-Yu, et al.
Published: (2023)
by: Chung, Ming-Yu, et al.
Published: (2023)
Gradient Shaping: Enhancing Backdoor Attack Against Reverse Engineering
by: Zhu, Rui, et al.
Published: (2023)
by: Zhu, Rui, et al.
Published: (2023)
Temporal Logic-Based Multi-Vehicle Backdoor Attacks against Offline RL Agents in End-to-end Autonomous Driving
by: Chen, Xuan, et al.
Published: (2025)
by: Chen, Xuan, et al.
Published: (2025)
From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMs
by: Shen, Guangyu, et al.
Published: (2025)
by: Shen, Guangyu, et al.
Published: (2025)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
by: Xiong, Chen, et al.
Published: (2024)
by: Xiong, Chen, et al.
Published: (2024)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
by: Chen, Zhuowei, et al.
Published: (2025)
by: Chen, Zhuowei, et al.
Published: (2025)
Isolate Trigger: Detecting and Eliminating Adaptive Backdoor Attacks
by: Sun, Chengrui, et al.
Published: (2025)
by: Sun, Chengrui, et al.
Published: (2025)
Mitigating Backdoor Attack by Injecting Proactive Defensive Backdoor
by: Wei, Shaokui, et al.
Published: (2024)
by: Wei, Shaokui, et al.
Published: (2024)
MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation
by: Yan, Lu, et al.
Published: (2025)
by: Yan, Lu, et al.
Published: (2025)
Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation
by: Zhong, Zhiyuan, et al.
Published: (2025)
by: Zhong, Zhiyuan, et al.
Published: (2025)
PSBD: Prediction Shift Uncertainty Unlocks Backdoor Detection
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models
by: Hu, Xiaomeng, et al.
Published: (2024)
by: Hu, Xiaomeng, et al.
Published: (2024)
Injecting Bias into Text Classification Models using Backdoor Attacks
by: Yavuz, A. Dilara, et al.
Published: (2024)
by: Yavuz, A. Dilara, et al.
Published: (2024)
Injecting Undetectable Backdoors in Obfuscated Neural Networks and Language Models
by: Kalavasis, Alkis, et al.
Published: (2024)
by: Kalavasis, Alkis, et al.
Published: (2024)
SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks
by: Zhang, Kaiyuan, et al.
Published: (2025)
by: Zhang, Kaiyuan, et al.
Published: (2025)
LLM Agents Should Employ Security Principles
by: Zhang, Kaiyuan, et al.
Published: (2025)
by: Zhang, Kaiyuan, et al.
Published: (2025)
Invisible Textual Backdoor Attacks based on Dual-Trigger
by: Hou, Yang, et al.
Published: (2024)
by: Hou, Yang, et al.
Published: (2024)
Sleeper Cell: Injecting Latent Malice Temporal Backdoors into Tool-Using LLMs
by: Pallakonda, Bhanu, et al.
Published: (2026)
by: Pallakonda, Bhanu, et al.
Published: (2026)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
by: Hu, Xiaomeng, et al.
Published: (2025)
by: Hu, Xiaomeng, et al.
Published: (2025)
ME: Trigger Element Combination Backdoor Attack on Copyright Infringement
by: Yang, Feiyu, et al.
Published: (2025)
by: Yang, Feiyu, et al.
Published: (2025)
Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models
by: Li, Haoran, et al.
Published: (2024)
by: Li, Haoran, et al.
Published: (2024)
Alleviating the Fear of Losing Alignment in LLM Fine-tuning
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
by: Ren, Zhiyao, et al.
Published: (2025)
by: Ren, Zhiyao, et al.
Published: (2025)
An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection
by: Yan, Shenao, et al.
Published: (2024)
by: Yan, Shenao, et al.
Published: (2024)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
by: Hu, Xiaomeng, et al.
Published: (2024)
by: Hu, Xiaomeng, et al.
Published: (2024)
On the Out-of-Distribution Backdoor Attack for Federated Learning
by: Xu, Jiahao, et al.
Published: (2025)
by: Xu, Jiahao, et al.
Published: (2025)
Less Is More -- Until It Breaks: Security Pitfalls of Vision Token Compression in Large Vision-Language Models
by: Zhang, Xiaomei, et al.
Published: (2026)
by: Zhang, Xiaomei, et al.
Published: (2026)
BackFlush: Knowledge-Free Backdoor Detection and Elimination with Watermark Preservation in Large Language Models
by: Rachapudi, Jagadeesh, et al.
Published: (2026)
by: Rachapudi, Jagadeesh, et al.
Published: (2026)
Optimization-Free Universal Watermark Forgery with Regenerative Diffusion Models
by: Zhu, Chaoyi, et al.
Published: (2025)
by: Zhu, Chaoyi, et al.
Published: (2025)
BackdoorDM: A Comprehensive Benchmark for Backdoor Learning on Diffusion Model
by: Lin, Weilin, et al.
Published: (2025)
by: Lin, Weilin, et al.
Published: (2025)
Authority Backdoor: A Certifiable Backdoor Mechanism for Authoring DNNs
by: Yang, Han, et al.
Published: (2025)
by: Yang, Han, et al.
Published: (2025)
Backdoor Directions in Vision Transformers
by: Karayalcin, Sengim, et al.
Published: (2026)
by: Karayalcin, Sengim, et al.
Published: (2026)
Similar Items
-
UNIT: Backdoor Mitigation via Automated Neural Distribution Tightening
by: Cheng, Siyuan, et al.
Published: (2024) -
LOTUS: Evasive and Resilient Backdoor Attacks through Sub-Partitioning
by: Cheng, Siyuan, et al.
Published: (2024) -
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
by: Shen, Guangyu, et al.
Published: (2024) -
VillanDiffusion: A Unified Backdoor Attack Framework for Diffusion Models
by: Chou, Sheng-Yen, et al.
Published: (2023) -
Backdooring Masked Diffusion Language Models
by: Cao, Daniel Yiming, et al.
Published: (2026)