Z-Erase: Enabling Concept Erasure in Single-Stream Diffusion Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Nanxiang, Fan, Zhaoxin, Wang, Baisen, Gao, Daiheng, Cheng, Junhang, Guo, Jifeng, Qin, Yalan, Jin, Yeying, Zheng, Hongwei, Wu, Faguo, Wu, Wenjun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917476347936768
author Jiang, Nanxiang
Fan, Zhaoxin
Wang, Baisen
Gao, Daiheng
Cheng, Junhang
Guo, Jifeng
Qin, Yalan
Jin, Yeying
Zheng, Hongwei
Wu, Faguo
Wu, Wenjun
author_facet Jiang, Nanxiang
Fan, Zhaoxin
Wang, Baisen
Gao, Daiheng
Cheng, Junhang
Guo, Jifeng
Qin, Yalan
Jin, Yeying
Zheng, Hongwei
Wu, Faguo
Wu, Wenjun
contents Concept erasure serves as a vital safety mechanism for removing unwanted concepts from text-to-image (T2I) models. While extensively studied in U-Net and dual-stream architectures (e.g., Flux), this task remains under-explored in the recent emerging paradigm of single-stream diffusion transformers (e.g., Z-Image). In this new paradigm, text and image tokens are processed as a single unified sequence via shared parameters. Consequently, directly applying prior erasure methods typically leads to generation collapse. To bridge this gap, we introduce Z-Erase, the first concept erasure method tailored for single-stream T2I models. To guarantee stable image generation, Z-Erase first proposes a Stream Disentangled Concept Erasure Framework that decouples updates and enables existing methods on single-stream models. Subsequently, within this framework, we introduce Lagrangian-Guided Adaptive Erasure Modulation, a constrained algorithm that further balances the sensitive erasure-preservation trade-off. Moreover, we provide a rigorous convergence analysis proving that Z-Erase can converge to a Pareto stationary point. Experiments demonstrate that Z-Erase successfully overcomes the generation collapse issue, achieving state-of-the-art performance across a wide range of tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25074
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Z-Erase: Enabling Concept Erasure in Single-Stream Diffusion Transformers
Jiang, Nanxiang
Fan, Zhaoxin
Wang, Baisen
Gao, Daiheng
Cheng, Junhang
Guo, Jifeng
Qin, Yalan
Jin, Yeying
Zheng, Hongwei
Wu, Faguo
Wu, Wenjun
Computer Vision and Pattern Recognition
Concept erasure serves as a vital safety mechanism for removing unwanted concepts from text-to-image (T2I) models. While extensively studied in U-Net and dual-stream architectures (e.g., Flux), this task remains under-explored in the recent emerging paradigm of single-stream diffusion transformers (e.g., Z-Image). In this new paradigm, text and image tokens are processed as a single unified sequence via shared parameters. Consequently, directly applying prior erasure methods typically leads to generation collapse. To bridge this gap, we introduce Z-Erase, the first concept erasure method tailored for single-stream T2I models. To guarantee stable image generation, Z-Erase first proposes a Stream Disentangled Concept Erasure Framework that decouples updates and enables existing methods on single-stream models. Subsequently, within this framework, we introduce Lagrangian-Guided Adaptive Erasure Modulation, a constrained algorithm that further balances the sensitive erasure-preservation trade-off. Moreover, we provide a rigorous convergence analysis proving that Z-Erase can converge to a Pareto stationary point. Experiments demonstrate that Z-Erase successfully overcomes the generation collapse issue, achieving state-of-the-art performance across a wide range of tasks.
title Z-Erase: Enabling Concept Erasure in Single-Stream Diffusion Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.25074