Z-Erase: Enabling Concept Erasure in Single-Stream Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917476347936768 |
|---|---|
| author | Jiang, Nanxiang Fan, Zhaoxin Wang, Baisen Gao, Daiheng Cheng, Junhang Guo, Jifeng Qin, Yalan Jin, Yeying Zheng, Hongwei Wu, Faguo Wu, Wenjun |
| author_facet | Jiang, Nanxiang Fan, Zhaoxin Wang, Baisen Gao, Daiheng Cheng, Junhang Guo, Jifeng Qin, Yalan Jin, Yeying Zheng, Hongwei Wu, Faguo Wu, Wenjun |
| contents | Concept erasure serves as a vital safety mechanism for removing unwanted concepts from text-to-image (T2I) models. While extensively studied in U-Net and dual-stream architectures (e.g., Flux), this task remains under-explored in the recent emerging paradigm of single-stream diffusion transformers (e.g., Z-Image). In this new paradigm, text and image tokens are processed as a single unified sequence via shared parameters. Consequently, directly applying prior erasure methods typically leads to generation collapse. To bridge this gap, we introduce Z-Erase, the first concept erasure method tailored for single-stream T2I models. To guarantee stable image generation, Z-Erase first proposes a Stream Disentangled Concept Erasure Framework that decouples updates and enables existing methods on single-stream models. Subsequently, within this framework, we introduce Lagrangian-Guided Adaptive Erasure Modulation, a constrained algorithm that further balances the sensitive erasure-preservation trade-off. Moreover, we provide a rigorous convergence analysis proving that Z-Erase can converge to a Pareto stationary point. Experiments demonstrate that Z-Erase successfully overcomes the generation collapse issue, achieving state-of-the-art performance across a wide range of tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_25074 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Z-Erase: Enabling Concept Erasure in Single-Stream Diffusion Transformers Jiang, Nanxiang Fan, Zhaoxin Wang, Baisen Gao, Daiheng Cheng, Junhang Guo, Jifeng Qin, Yalan Jin, Yeying Zheng, Hongwei Wu, Faguo Wu, Wenjun Computer Vision and Pattern Recognition Concept erasure serves as a vital safety mechanism for removing unwanted concepts from text-to-image (T2I) models. While extensively studied in U-Net and dual-stream architectures (e.g., Flux), this task remains under-explored in the recent emerging paradigm of single-stream diffusion transformers (e.g., Z-Image). In this new paradigm, text and image tokens are processed as a single unified sequence via shared parameters. Consequently, directly applying prior erasure methods typically leads to generation collapse. To bridge this gap, we introduce Z-Erase, the first concept erasure method tailored for single-stream T2I models. To guarantee stable image generation, Z-Erase first proposes a Stream Disentangled Concept Erasure Framework that decouples updates and enables existing methods on single-stream models. Subsequently, within this framework, we introduce Lagrangian-Guided Adaptive Erasure Modulation, a constrained algorithm that further balances the sensitive erasure-preservation trade-off. Moreover, we provide a rigorous convergence analysis proving that Z-Erase can converge to a Pareto stationary point. Experiments demonstrate that Z-Erase successfully overcomes the generation collapse issue, achieving state-of-the-art performance across a wide range of tasks. |
| title | Z-Erase: Enabling Concept Erasure in Single-Stream Diffusion Transformers |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2603.25074 |