SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Dongli, Tiulpin, Aleksei, Blaschko, Matthew B.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915530408984576
author Xu, Dongli
Tiulpin, Aleksei
Blaschko, Matthew B.
author_facet Xu, Dongli
Tiulpin, Aleksei
Blaschko, Matthew B.
contents Autoregressive (AR) models have emerged as powerful tools for image generation by modeling images as sequences of discrete tokens. While Classifier-Free Guidance (CFG) has been adopted to improve conditional generation, its application in AR models faces two key issues: guidance diminishing, where the conditional-unconditional gap quickly vanishes as decoding progresses, and over-guidance, where strong conditions distort visual coherence. To address these challenges, we propose SoftCFG, an uncertainty-guided inference method that distributes adaptive perturbations across all tokens in the sequence. The key idea behind SoftCFG is to let each generated token contribute certainty-weighted guidance, ensuring that the signal persists across steps while resolving conflicts between text guidance and visual context. To further stabilize long-sequence generation, we introduce Step Normalization, which bounds cumulative perturbations of SoftCFG. Our method is training-free, model-agnostic, and seamlessly integrates with existing AR pipelines. Experiments show that SoftCFG significantly improves image quality over standard CFG and achieves state-of-the-art FID on ImageNet 256*256 among autoregressive models.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00996
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
Xu, Dongli
Tiulpin, Aleksei
Blaschko, Matthew B.
Computer Vision and Pattern Recognition
Autoregressive (AR) models have emerged as powerful tools for image generation by modeling images as sequences of discrete tokens. While Classifier-Free Guidance (CFG) has been adopted to improve conditional generation, its application in AR models faces two key issues: guidance diminishing, where the conditional-unconditional gap quickly vanishes as decoding progresses, and over-guidance, where strong conditions distort visual coherence. To address these challenges, we propose SoftCFG, an uncertainty-guided inference method that distributes adaptive perturbations across all tokens in the sequence. The key idea behind SoftCFG is to let each generated token contribute certainty-weighted guidance, ensuring that the signal persists across steps while resolving conflicts between text guidance and visual context. To further stabilize long-sequence generation, we introduce Step Normalization, which bounds cumulative perturbations of SoftCFG. Our method is training-free, model-agnostic, and seamlessly integrates with existing AR pipelines. Experiments show that SoftCFG significantly improves image quality over standard CFG and achieves state-of-the-art FID on ImageNet 256*256 among autoregressive models.
title SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.00996