Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yimeng, Chen, Xin, Jia, Jinghan, Zhang, Yihua, Fan, Chongyu, Liu, Jiancheng, Hong, Mingyi, Ding, Ke, Liu, Sijia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoreUnlearn: Rethinking Concept Unlearning through Disentangled Component-Level Erasure in Text-guided Diffusion Models
by: Zhao, Mengnan, et al.
Published: (2026)
by: Zhao, Mengnan, et al.
Published: (2026)
Rethinking Robust Adversarial Concept Erasure in Diffusion Models
by: Yin, Qinghong, et al.
Published: (2025)
by: Yin, Qinghong, et al.
Published: (2025)
One Token Embedding Is Enough to Deadlock Your Large Reasoning Model
by: Zhang, Mohan, et al.
Published: (2025)
by: Zhang, Mohan, et al.
Published: (2025)
Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking Design
by: Sun, Yuhao, et al.
Published: (2025)
by: Sun, Yuhao, et al.
Published: (2025)
FedSGT: Exact Federated Unlearning via Sequential Group-based Training
by: Zhang, Bokang, et al.
Published: (2025)
by: Zhang, Bokang, et al.
Published: (2025)
AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion Models
by: Li, Fengpeng, et al.
Published: (2026)
by: Li, Fengpeng, et al.
Published: (2026)
Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond
by: Fan, Chongyu, et al.
Published: (2025)
by: Fan, Chongyu, et al.
Published: (2025)
Unlearn and Burn: Adversarial Machine Unlearning Requests Destroy Model Accuracy
by: Huang, Yangsibo, et al.
Published: (2024)
by: Huang, Yangsibo, et al.
Published: (2024)
To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For Now
by: Zhang, Yimeng, et al.
Published: (2023)
by: Zhang, Yimeng, et al.
Published: (2023)
BURN: Backdoor Unlearning via Adversarial Boundary Analysis
by: Su, Yanghao, et al.
Published: (2025)
by: Su, Yanghao, et al.
Published: (2025)
ALRPHFS: Adversarially Learned Risk Patterns with Hierarchical Fast \& Slow Reasoning for Robust Agent Defense
by: Xiang, Shiyu, et al.
Published: (2025)
by: Xiang, Shiyu, et al.
Published: (2025)
Adversarial Machine Unlearning
by: Di, Zonglin, et al.
Published: (2024)
by: Di, Zonglin, et al.
Published: (2024)
Query Provenance Analysis: Efficient and Robust Defense against Query-based Black-box Attacks
by: Li, Shaofei, et al.
Published: (2024)
by: Li, Shaofei, et al.
Published: (2024)
Dynamic Dual-level Defense Routing for Continual Adversarial Training
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Adversarially Robust Assembly Language Model for Packed Executables Detection
by: Li, Shijia, et al.
Published: (2025)
by: Li, Shijia, et al.
Published: (2025)
Rotated Robustness: A Training-Free Defense against Bit-Flip Attacks on Large Language Models
by: Liu, Deng, et al.
Published: (2026)
by: Liu, Deng, et al.
Published: (2026)
Struggle with Adversarial Defense? Try Diffusion
by: Li, Yujie, et al.
Published: (2024)
by: Li, Yujie, et al.
Published: (2024)
ShellForge: Adversarial Co-Evolution of Webshell Generation and Multi-View Detection for Robust Webshell Defense
by: Ding, Yizhong
Published: (2026)
by: Ding, Yizhong
Published: (2026)
UnlearnCanvas: Stylized Image Dataset for Enhanced Machine Unlearning Evaluation in Diffusion Models
by: Zhang, Yihua, et al.
Published: (2024)
by: Zhang, Yihua, et al.
Published: (2024)
Poisoning Attacks and Defenses to Federated Unlearning
by: Wang, Wenbin, et al.
Published: (2025)
by: Wang, Wenbin, et al.
Published: (2025)
LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training
by: Gong, Yuyang, et al.
Published: (2026)
by: Gong, Yuyang, et al.
Published: (2026)
Attack by Unlearning: Unlearning-Induced Adversarial Attacks on Graph Neural Networks
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
Threats, Attacks, and Defenses in Machine Unlearning: A Survey
by: Liu, Ziyao, et al.
Published: (2024)
by: Liu, Ziyao, et al.
Published: (2024)
IDEA: Invariant Defense for Graph Adversarial Robustness
by: Tao, Shuchang, et al.
Published: (2023)
by: Tao, Shuchang, et al.
Published: (2023)
Provably Cost-Sensitive Adversarial Defense via Randomized Smoothing
by: Xin, Yuan, et al.
Published: (2023)
by: Xin, Yuan, et al.
Published: (2023)
Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine Unlearning
by: Song, Baogang, et al.
Published: (2025)
by: Song, Baogang, et al.
Published: (2025)
VideoEraser: Concept Erasure in Text-to-Video Diffusion Models
by: Xu, Naen, et al.
Published: (2025)
by: Xu, Naen, et al.
Published: (2025)
Diffusion-Guided Adversarial Perturbation Injection for Generalizable Defense Against Facial Manipulations
by: Li, Yue, et al.
Published: (2026)
by: Li, Yue, et al.
Published: (2026)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
by: Cao, Tri, et al.
Published: (2026)
by: Cao, Tri, et al.
Published: (2026)
Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning
by: Lang, Yicheng, et al.
Published: (2025)
by: Lang, Yicheng, et al.
Published: (2025)
Rethinking the Vulnerability of Concept Erasure and a New Method
by: Richardson, Alex D., et al.
Published: (2025)
by: Richardson, Alex D., et al.
Published: (2025)
Neighbor-Aware Localized Concept Erasure in Text-to-Image Diffusion Models
by: Shi, Zhuan, et al.
Published: (2026)
by: Shi, Zhuan, et al.
Published: (2026)
Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable Confidence
by: Hong, Hanbin, et al.
Published: (2023)
by: Hong, Hanbin, et al.
Published: (2023)
Evaluating the Defense Potential of Machine Unlearning against Membership Inference Attacks
by: Tsiolakis, Theodoros, et al.
Published: (2025)
by: Tsiolakis, Theodoros, et al.
Published: (2025)
One Stone, Two Birds: Enhancing Adversarial Defense Through the Lens of Distributional Discrepancy
by: Zhang, Jiacheng, et al.
Published: (2025)
by: Zhang, Jiacheng, et al.
Published: (2025)
Elevating Defenses: Bridging Adversarial Training and Watermarking for Model Resilience
by: Thakkar, Janvi, et al.
Published: (2023)
by: Thakkar, Janvi, et al.
Published: (2023)
Vulnerability-Aware Robust Multimodal Adversarial Training
by: Zhang, Junrui, et al.
Published: (2025)
by: Zhang, Junrui, et al.
Published: (2025)
Training-Free Color-Aware Adversarial Diffusion Sanitization for Diffusion Stegomalware Defense at Security Gateways
by: Frants, Vladimir, et al.
Published: (2025)
by: Frants, Vladimir, et al.
Published: (2025)
Pruning Graphs by Adversarial Robustness Evaluation to Strengthen GNN Defenses
by: Wang, Yongyu
Published: (2025)
by: Wang, Yongyu
Published: (2025)
Continual Adversarial Defense
by: Wang, Qian, et al.
Published: (2023)
by: Wang, Qian, et al.
Published: (2023)
Similar Items
-
CoreUnlearn: Rethinking Concept Unlearning through Disentangled Component-Level Erasure in Text-guided Diffusion Models
by: Zhao, Mengnan, et al.
Published: (2026) -
Rethinking Robust Adversarial Concept Erasure in Diffusion Models
by: Yin, Qinghong, et al.
Published: (2025) -
One Token Embedding Is Enough to Deadlock Your Large Reasoning Model
by: Zhang, Mohan, et al.
Published: (2025) -
Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking Design
by: Sun, Yuhao, et al.
Published: (2025) -
FedSGT: Exact Federated Unlearning via Sequential Group-based Training
by: Zhang, Bokang, et al.
Published: (2025)