Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhong, Xinhao, Zhou, Yimin, Zhang, Zhiqi, Li, Junhao, Sun, Yi, Chen, Bin, Xia, Shu-Tao, Wang, Xuan, Xu, Ke
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908801465057280
author Zhong, Xinhao
Zhou, Yimin
Zhang, Zhiqi
Li, Junhao
Sun, Yi
Chen, Bin
Xia, Shu-Tao
Wang, Xuan
Xu, Ke
author_facet Zhong, Xinhao
Zhou, Yimin
Zhang, Zhiqi
Li, Junhao
Sun, Yi
Chen, Bin
Xia, Shu-Tao
Wang, Xuan
Xu, Ke
contents The rapid progress of visual autoregressive (VAR) models has brought new opportunities for text-to-image generation, but also heightened safety concerns. Existing concept erasure techniques, primarily designed for diffusion models, fail to generalize to VARs due to their next-scale token prediction paradigm. In this paper, we first propose a novel VAR Erasure framework VARE that enables stable concept erasure in VAR models by leveraging auxiliary visual tokens to reduce fine-tuning intensity. Building upon this, we introduce S-VARE, a novel and effective concept erasure method designed for VAR, which incorporates a filtered cross entropy loss to precisely identify and minimally adjust unsafe visual tokens, along with a preservation loss to maintain semantic fidelity, addressing the issues such as language drift and reduced diversity introduce by naïve fine-tuning. Extensive experiments demonstrate that our approach achieves surgical concept erasure while preserving generation quality, thereby closing the safety gap in autoregressive text-to-image generation by earlier methods.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22400
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive Models
Zhong, Xinhao
Zhou, Yimin
Zhang, Zhiqi
Li, Junhao
Sun, Yi
Chen, Bin
Xia, Shu-Tao
Wang, Xuan
Xu, Ke
Computer Vision and Pattern Recognition
The rapid progress of visual autoregressive (VAR) models has brought new opportunities for text-to-image generation, but also heightened safety concerns. Existing concept erasure techniques, primarily designed for diffusion models, fail to generalize to VARs due to their next-scale token prediction paradigm. In this paper, we first propose a novel VAR Erasure framework VARE that enables stable concept erasure in VAR models by leveraging auxiliary visual tokens to reduce fine-tuning intensity. Building upon this, we introduce S-VARE, a novel and effective concept erasure method designed for VAR, which incorporates a filtered cross entropy loss to precisely identify and minimally adjust unsafe visual tokens, along with a preservation loss to maintain semantic fidelity, addressing the issues such as language drift and reduced diversity introduce by naïve fine-tuning. Extensive experiments demonstrate that our approach achieves surgical concept erasure while preserving generation quality, thereby closing the safety gap in autoregressive text-to-image generation by earlier methods.
title Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.22400