Concept Pinpoint Eraser for Text-to-image Diffusion Models via Residual Attention Gate

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lee, Byung Hyun, Lim, Sungjin, Lee, Seunggyu, Kang, Dong Un, Chun, Se Young
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916814785609728
author Lee, Byung Hyun
Lim, Sungjin
Lee, Seunggyu
Kang, Dong Un
Chun, Se Young
author_facet Lee, Byung Hyun
Lim, Sungjin
Lee, Seunggyu
Kang, Dong Un
Chun, Se Young
contents Remarkable progress in text-to-image diffusion models has brought a major concern about potentially generating images on inappropriate or trademarked concepts. Concept erasing has been investigated with the goals of deleting target concepts in diffusion models while preserving other concepts with minimal distortion. To achieve these goals, recent concept erasing methods usually fine-tune the cross-attention layers of diffusion models. In this work, we first show that merely updating the cross-attention layers in diffusion models, which is mathematically equivalent to adding \emph{linear} modules to weights, may not be able to preserve diverse remaining concepts. Then, we propose a novel framework, dubbed Concept Pinpoint Eraser (CPE), by adding \emph{nonlinear} Residual Attention Gates (ResAGs) that selectively erase (or cut) target concepts while safeguarding remaining concepts from broad distributions by employing an attention anchoring loss to prevent the forgetting. Moreover, we adversarially train CPE with ResAG and learnable text embeddings in an iterative manner to maximize erasing performance and enhance robustness against adversarial attacks. Extensive experiments on the erasure of celebrities, artistic styles, and explicit contents demonstrated that the proposed CPE outperforms prior arts by keeping diverse remaining concepts while deleting the target concepts with robustness against attack prompts. Code is available at https://github.com/Hyun1A/CPE
format Preprint
id arxiv_https___arxiv_org_abs_2506_22806
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Concept Pinpoint Eraser for Text-to-image Diffusion Models via Residual Attention Gate
Lee, Byung Hyun
Lim, Sungjin
Lee, Seunggyu
Kang, Dong Un
Chun, Se Young
Computer Vision and Pattern Recognition
Machine Learning
Remarkable progress in text-to-image diffusion models has brought a major concern about potentially generating images on inappropriate or trademarked concepts. Concept erasing has been investigated with the goals of deleting target concepts in diffusion models while preserving other concepts with minimal distortion. To achieve these goals, recent concept erasing methods usually fine-tune the cross-attention layers of diffusion models. In this work, we first show that merely updating the cross-attention layers in diffusion models, which is mathematically equivalent to adding \emph{linear} modules to weights, may not be able to preserve diverse remaining concepts. Then, we propose a novel framework, dubbed Concept Pinpoint Eraser (CPE), by adding \emph{nonlinear} Residual Attention Gates (ResAGs) that selectively erase (or cut) target concepts while safeguarding remaining concepts from broad distributions by employing an attention anchoring loss to prevent the forgetting. Moreover, we adversarially train CPE with ResAG and learnable text embeddings in an iterative manner to maximize erasing performance and enhance robustness against adversarial attacks. Extensive experiments on the erasure of celebrities, artistic styles, and explicit contents demonstrated that the proposed CPE outperforms prior arts by keeping diverse remaining concepts while deleting the target concepts with robustness against attack prompts. Code is available at https://github.com/Hyun1A/CPE
title Concept Pinpoint Eraser for Text-to-image Diffusion Models via Residual Attention Gate
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2506.22806