Not All Tokens are Guided Equal: Improving Guidance in Visual Autoregressive Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Ky Dan, Tran, Hoang Lam, Dinh, Anh-Dung, Liu, Daochang, Cai, Weidong, Wang, Xiuying, Xu, Chang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915527295762432
author Nguyen, Ky Dan
Tran, Hoang Lam
Dinh, Anh-Dung
Liu, Daochang
Cai, Weidong
Wang, Xiuying
Xu, Chang
author_facet Nguyen, Ky Dan
Tran, Hoang Lam
Dinh, Anh-Dung
Liu, Daochang
Cai, Weidong
Wang, Xiuying
Xu, Chang
contents Autoregressive (AR) models based on next-scale prediction are rapidly emerging as a powerful tool for image generation, but they face a critical weakness: information inconsistencies between patches across timesteps introduced by progressive resolution scaling. These inconsistencies scatter guidance signals, causing them to drift away from conditioning information and leaving behind ambiguous, unfaithful features. We tackle this challenge with Information-Grounding Guidance (IGG), a novel mechanism that anchors guidance to semantically important regions through attention. By adaptively reinforcing informative patches during sampling, IGG ensures that guidance and content remain tightly aligned. Across both class-conditioned and text-to-image generation tasks, IGG delivers sharper, more coherent, and semantically grounded images, setting a new benchmark for AR-based methods.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23876
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Not All Tokens are Guided Equal: Improving Guidance in Visual Autoregressive Models
Nguyen, Ky Dan
Tran, Hoang Lam
Dinh, Anh-Dung
Liu, Daochang
Cai, Weidong
Wang, Xiuying
Xu, Chang
Computer Vision and Pattern Recognition
Artificial Intelligence
Autoregressive (AR) models based on next-scale prediction are rapidly emerging as a powerful tool for image generation, but they face a critical weakness: information inconsistencies between patches across timesteps introduced by progressive resolution scaling. These inconsistencies scatter guidance signals, causing them to drift away from conditioning information and leaving behind ambiguous, unfaithful features. We tackle this challenge with Information-Grounding Guidance (IGG), a novel mechanism that anchors guidance to semantically important regions through attention. By adaptively reinforcing informative patches during sampling, IGG ensures that guidance and content remain tightly aligned. Across both class-conditioned and text-to-image generation tasks, IGG delivers sharper, more coherent, and semantically grounded images, setting a new benchmark for AR-based methods.
title Not All Tokens are Guided Equal: Improving Guidance in Visual Autoregressive Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.23876