Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908801465057280 |
|---|---|
| author | Zhong, Xinhao Zhou, Yimin Zhang, Zhiqi Li, Junhao Sun, Yi Chen, Bin Xia, Shu-Tao Wang, Xuan Xu, Ke |
| author_facet | Zhong, Xinhao Zhou, Yimin Zhang, Zhiqi Li, Junhao Sun, Yi Chen, Bin Xia, Shu-Tao Wang, Xuan Xu, Ke |
| contents | The rapid progress of visual autoregressive (VAR) models has brought new opportunities for text-to-image generation, but also heightened safety concerns. Existing concept erasure techniques, primarily designed for diffusion models, fail to generalize to VARs due to their next-scale token prediction paradigm. In this paper, we first propose a novel VAR Erasure framework VARE that enables stable concept erasure in VAR models by leveraging auxiliary visual tokens to reduce fine-tuning intensity. Building upon this, we introduce S-VARE, a novel and effective concept erasure method designed for VAR, which incorporates a filtered cross entropy loss to precisely identify and minimally adjust unsafe visual tokens, along with a preservation loss to maintain semantic fidelity, addressing the issues such as language drift and reduced diversity introduce by naïve fine-tuning. Extensive experiments demonstrate that our approach achieves surgical concept erasure while preserving generation quality, thereby closing the safety gap in autoregressive text-to-image generation by earlier methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_22400 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive Models Zhong, Xinhao Zhou, Yimin Zhang, Zhiqi Li, Junhao Sun, Yi Chen, Bin Xia, Shu-Tao Wang, Xuan Xu, Ke Computer Vision and Pattern Recognition The rapid progress of visual autoregressive (VAR) models has brought new opportunities for text-to-image generation, but also heightened safety concerns. Existing concept erasure techniques, primarily designed for diffusion models, fail to generalize to VARs due to their next-scale token prediction paradigm. In this paper, we first propose a novel VAR Erasure framework VARE that enables stable concept erasure in VAR models by leveraging auxiliary visual tokens to reduce fine-tuning intensity. Building upon this, we introduce S-VARE, a novel and effective concept erasure method designed for VAR, which incorporates a filtered cross entropy loss to precisely identify and minimally adjust unsafe visual tokens, along with a preservation loss to maintain semantic fidelity, addressing the issues such as language drift and reduced diversity introduce by naïve fine-tuning. Extensive experiments demonstrate that our approach achieves surgical concept erasure while preserving generation quality, thereby closing the safety gap in autoregressive text-to-image generation by earlier methods. |
| title | Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2509.22400 |