ActErase: A Training-Free Paradigm for Precise Concept Erasure via Activation Redirection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Yi, Zhong, Xinhao, Li, Hongyan, Zhou, Yimin, Li, Junhao, Chen, Bin, Wang, Xuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910092379553792
author Sun, Yi
Zhong, Xinhao
Li, Hongyan
Zhou, Yimin
Li, Junhao
Chen, Bin
Wang, Xuan
author_facet Sun, Yi
Zhong, Xinhao
Li, Hongyan
Zhou, Yimin
Li, Junhao
Chen, Bin
Wang, Xuan
contents Recent advances in text-to-image diffusion models have demonstrated remarkable generation capabilities, yet they raise significant concerns regarding safety, copyright, and ethical implications. Existing concept erasure methods address these risks by removing sensitive concepts from pre-trained models, but most of them rely on data-intensive and computationally expensive fine-tuning, which poses a critical limitation. To overcome these challenges, inspired by the observation that the model's activations are predominantly composed of generic concepts, with only a minimal component can represent the target concept, we propose a novel training-free method (ActErase) for efficient concept erasure. Specifically, the proposed method operates by identifying activation difference regions via prompt-pair analysis, extracting target activations and dynamically replacing input activations during forward passes. Comprehensive evaluations across three critical erasure tasks (nudity, artistic style, and object removal) demonstrates that our training-free method achieves state-of-the-art (SOTA) erasure performance, while effectively preserving the model's overall generative capability. Our approach also exhibits strong robustness against adversarial attacks, establishing a new plug-and-play paradigm for lightweight yet effective concept manipulation in diffusion models.
format Preprint
id arxiv_https___arxiv_org_abs_2601_00267
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ActErase: A Training-Free Paradigm for Precise Concept Erasure via Activation Redirection
Sun, Yi
Zhong, Xinhao
Li, Hongyan
Zhou, Yimin
Li, Junhao
Chen, Bin
Wang, Xuan
Computer Vision and Pattern Recognition
Recent advances in text-to-image diffusion models have demonstrated remarkable generation capabilities, yet they raise significant concerns regarding safety, copyright, and ethical implications. Existing concept erasure methods address these risks by removing sensitive concepts from pre-trained models, but most of them rely on data-intensive and computationally expensive fine-tuning, which poses a critical limitation. To overcome these challenges, inspired by the observation that the model's activations are predominantly composed of generic concepts, with only a minimal component can represent the target concept, we propose a novel training-free method (ActErase) for efficient concept erasure. Specifically, the proposed method operates by identifying activation difference regions via prompt-pair analysis, extracting target activations and dynamically replacing input activations during forward passes. Comprehensive evaluations across three critical erasure tasks (nudity, artistic style, and object removal) demonstrates that our training-free method achieves state-of-the-art (SOTA) erasure performance, while effectively preserving the model's overall generative capability. Our approach also exhibits strong robustness against adversarial attacks, establishing a new plug-and-play paradigm for lightweight yet effective concept manipulation in diffusion models.
title ActErase: A Training-Free Paradigm for Precise Concept Erasure via Activation Redirection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.00267