Rethinking Robust Adversarial Concept Erasure in Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yin, Qinghong, Tian, Yu, Yang, Heming, Chen, Xiang, Zhang, Xianlin, Li, Xueming, Zhan, Yue
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908637516005376
author Yin, Qinghong
Tian, Yu
Yang, Heming
Chen, Xiang
Zhang, Xianlin
Li, Xueming
Zhan, Yue
author_facet Yin, Qinghong
Tian, Yu
Yang, Heming
Chen, Xiang
Zhang, Xianlin
Li, Xueming
Zhan, Yue
contents Concept erasure aims to selectively unlearning undesirable content in diffusion models (DMs) to reduce the risk of sensitive content generation. As a novel paradigm in concept erasure, most existing methods employ adversarial training to identify and suppress target concepts, thus reducing the likelihood of sensitive outputs. However, these methods often neglect the specificity of adversarial training in DMs, resulting in only partial mitigation. In this work, we investigate and quantify this specificity from the perspective of concept space, i.e., can adversarial samples truly fit the target concept space? We observe that existing methods neglect the role of conceptual semantics when generating adversarial samples, resulting in ineffective fitting of concept spaces. This oversight leads to the following issues: 1) when there are few adversarial samples, they fail to comprehensively cover the object concept; 2) conversely, they will disrupt other target concept spaces. Motivated by the analysis of these findings, we introduce S-GRACE (Semantics-Guided Robust Adversarial Concept Erasure), which grace leveraging semantic guidance within the concept space to generate adversarial samples and perform erasure training. Experiments conducted with seven state-of-the-art methods and three adversarial prompt generation strategies across various DM unlearning scenarios demonstrate that S-GRACE significantly improves erasure performance 26%, better preserves non-target concepts, and reduces training time by 90%. Our code is available at https://github.com/Qhong-522/S-GRACE.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27285
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Robust Adversarial Concept Erasure in Diffusion Models
Yin, Qinghong
Tian, Yu
Yang, Heming
Chen, Xiang
Zhang, Xianlin
Li, Xueming
Zhan, Yue
Computer Vision and Pattern Recognition
Cryptography and Security
Concept erasure aims to selectively unlearning undesirable content in diffusion models (DMs) to reduce the risk of sensitive content generation. As a novel paradigm in concept erasure, most existing methods employ adversarial training to identify and suppress target concepts, thus reducing the likelihood of sensitive outputs. However, these methods often neglect the specificity of adversarial training in DMs, resulting in only partial mitigation. In this work, we investigate and quantify this specificity from the perspective of concept space, i.e., can adversarial samples truly fit the target concept space? We observe that existing methods neglect the role of conceptual semantics when generating adversarial samples, resulting in ineffective fitting of concept spaces. This oversight leads to the following issues: 1) when there are few adversarial samples, they fail to comprehensively cover the object concept; 2) conversely, they will disrupt other target concept spaces. Motivated by the analysis of these findings, we introduce S-GRACE (Semantics-Guided Robust Adversarial Concept Erasure), which grace leveraging semantic guidance within the concept space to generate adversarial samples and perform erasure training. Experiments conducted with seven state-of-the-art methods and three adversarial prompt generation strategies across various DM unlearning scenarios demonstrate that S-GRACE significantly improves erasure performance 26%, better preserves non-target concepts, and reduces training time by 90%. Our code is available at https://github.com/Qhong-522/S-GRACE.
title Rethinking Robust Adversarial Concept Erasure in Diffusion Models
topic Computer Vision and Pattern Recognition
Cryptography and Security
url https://arxiv.org/abs/2510.27285