FADE: Adversarial Concept Erasure in Flow Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Zixuan, Ren, Yan, Carter, Finn, Wang, Chenyue, Niu, Ze, Yu, Dacheng, Davis, Emily, Zhang, Bo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909691644215296
author Fu, Zixuan
Ren, Yan
Carter, Finn
Wang, Chenyue
Niu, Ze
Yu, Dacheng
Davis, Emily
Zhang, Bo
author_facet Fu, Zixuan
Ren, Yan
Carter, Finn
Wang, Chenyue
Niu, Ze
Yu, Dacheng
Davis, Emily
Zhang, Bo
contents Diffusion models have demonstrated remarkable image generation capabilities, but also pose risks in privacy and fairness by memorizing sensitive concepts or perpetuating biases. We propose a novel \textbf{concept erasure} method for text-to-image diffusion models, designed to remove specified concepts (e.g., a private individual or a harmful stereotype) from the model's generative repertoire. Our method, termed \textbf{FADE} (Fair Adversarial Diffusion Erasure), combines a trajectory-aware fine-tuning strategy with an adversarial objective to ensure the concept is reliably removed while preserving overall model fidelity. Theoretically, we prove a formal guarantee that our approach minimizes the mutual information between the erased concept and the model's outputs, ensuring privacy and fairness. Empirically, we evaluate FADE on Stable Diffusion and FLUX, using benchmarks from prior work (e.g., object, celebrity, explicit content, and style erasure tasks from MACE). FADE achieves state-of-the-art concept removal performance, surpassing recent baselines like ESD, UCE, MACE, and ANT in terms of removal efficacy and image quality. Notably, FADE improves the harmonic mean of concept removal and fidelity by 5--10\% over the best prior method. We also conduct an ablation study to validate each component of FADE, confirming that our adversarial and trajectory-preserving objectives each contribute to its superior performance. Our work sets a new standard for safe and fair generative modeling by unlearning specified concepts without retraining from scratch.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12283
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FADE: Adversarial Concept Erasure in Flow Models
Fu, Zixuan
Ren, Yan
Carter, Finn
Wang, Chenyue
Niu, Ze
Yu, Dacheng
Davis, Emily
Zhang, Bo
Computer Vision and Pattern Recognition
Diffusion models have demonstrated remarkable image generation capabilities, but also pose risks in privacy and fairness by memorizing sensitive concepts or perpetuating biases. We propose a novel \textbf{concept erasure} method for text-to-image diffusion models, designed to remove specified concepts (e.g., a private individual or a harmful stereotype) from the model's generative repertoire. Our method, termed \textbf{FADE} (Fair Adversarial Diffusion Erasure), combines a trajectory-aware fine-tuning strategy with an adversarial objective to ensure the concept is reliably removed while preserving overall model fidelity. Theoretically, we prove a formal guarantee that our approach minimizes the mutual information between the erased concept and the model's outputs, ensuring privacy and fairness. Empirically, we evaluate FADE on Stable Diffusion and FLUX, using benchmarks from prior work (e.g., object, celebrity, explicit content, and style erasure tasks from MACE). FADE achieves state-of-the-art concept removal performance, surpassing recent baselines like ESD, UCE, MACE, and ANT in terms of removal efficacy and image quality. Notably, FADE improves the harmonic mean of concept removal and fidelity by 5--10\% over the best prior method. We also conduct an ablation study to validate each component of FADE, confirming that our adversarial and trajectory-preserving objectives each contribute to its superior performance. Our work sets a new standard for safe and fair generative modeling by unlearning specified concepts without retraining from scratch.
title FADE: Adversarial Concept Erasure in Flow Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.12283