One Image is Worth a Thousand Words: A Usability Preservable Text-Image Collaborative Erasing Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Feiran, Xu, Qianqian, Bao, Shilong, Yang, Zhiyong, Cao, Xiaochun, Huang, Qingming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915304837218304
author Li, Feiran
Xu, Qianqian
Bao, Shilong
Yang, Zhiyong
Cao, Xiaochun
Huang, Qingming
author_facet Li, Feiran
Xu, Qianqian
Bao, Shilong
Yang, Zhiyong
Cao, Xiaochun
Huang, Qingming
contents Concept erasing has recently emerged as an effective paradigm to prevent text-to-image diffusion models from generating visually undesirable or even harmful content. However, current removal methods heavily rely on manually crafted text prompts, making it challenging to achieve a high erasure (efficacy) while minimizing the impact on other benign concepts (usability). In this paper, we attribute the limitations to the inherent gap between the text and image modalities, which makes it hard to transfer the intricately entangled concept knowledge from text prompts to the image generation process. To address this, we propose a novel solution by directly integrating visual supervision into the erasure process, introducing the first text-image Collaborative Concept Erasing (Co-Erasing) framework. Specifically, Co-Erasing describes the concept jointly by text prompts and the corresponding undesirable images induced by the prompts, and then reduces the generating probability of the target concept through negative guidance. This approach effectively bypasses the knowledge gap between text and image, significantly enhancing erasure efficacy. Additionally, we design a text-guided image concept refinement strategy that directs the model to focus on visual features most relevant to the specified text concept, minimizing disruption to other benign concepts. Finally, comprehensive experiments suggest that Co-Erasing outperforms state-of-the-art erasure approaches significantly with a better trade-off between efficacy and usability. Codes are available at https://github.com/Ferry-Li/Co-Erasing.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11131
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle One Image is Worth a Thousand Words: A Usability Preservable Text-Image Collaborative Erasing Framework
Li, Feiran
Xu, Qianqian
Bao, Shilong
Yang, Zhiyong
Cao, Xiaochun
Huang, Qingming
Computer Vision and Pattern Recognition
Artificial Intelligence
Concept erasing has recently emerged as an effective paradigm to prevent text-to-image diffusion models from generating visually undesirable or even harmful content. However, current removal methods heavily rely on manually crafted text prompts, making it challenging to achieve a high erasure (efficacy) while minimizing the impact on other benign concepts (usability). In this paper, we attribute the limitations to the inherent gap between the text and image modalities, which makes it hard to transfer the intricately entangled concept knowledge from text prompts to the image generation process. To address this, we propose a novel solution by directly integrating visual supervision into the erasure process, introducing the first text-image Collaborative Concept Erasing (Co-Erasing) framework. Specifically, Co-Erasing describes the concept jointly by text prompts and the corresponding undesirable images induced by the prompts, and then reduces the generating probability of the target concept through negative guidance. This approach effectively bypasses the knowledge gap between text and image, significantly enhancing erasure efficacy. Additionally, we design a text-guided image concept refinement strategy that directs the model to focus on visual features most relevant to the specified text concept, minimizing disruption to other benign concepts. Finally, comprehensive experiments suggest that Co-Erasing outperforms state-of-the-art erasure approaches significantly with a better trade-off between efficacy and usability. Codes are available at https://github.com/Ferry-Li/Co-Erasing.
title One Image is Worth a Thousand Words: A Usability Preservable Text-Image Collaborative Erasing Framework
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.11131