Go with Your Gut: Scaling Confidence for Autoregressive Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Harold Haodong, Wu, Xianfeng, Shu, Wen-Jie, Guo, Rongjin, Lan, Disen, Yang, Harry, Chen, Ying-Cong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917185406894080
author Chen, Harold Haodong
Wu, Xianfeng
Shu, Wen-Jie
Guo, Rongjin
Lan, Disen
Yang, Harry
Chen, Ying-Cong
author_facet Chen, Harold Haodong
Wu, Xianfeng
Shu, Wen-Jie
Guo, Rongjin
Lan, Disen
Yang, Harry
Chen, Ying-Cong
contents Test-time scaling (TTS) has demonstrated remarkable success in enhancing large language models, yet its application to next-token prediction (NTP) autoregressive (AR) image generation remains largely uncharted. Existing TTS approaches for visual AR (VAR), which rely on frequent partial decoding and external reward models, are ill-suited for NTP-based image generation due to the inherent incompleteness of intermediate decoding results. To bridge this gap, we introduce ScalingAR, the first TTS framework specifically designed for NTP-based AR image generation that eliminates the need for early decoding or auxiliary rewards. ScalingAR leverages token entropy as a novel signal in visual token generation and operates at two complementary scaling levels: (i) Profile Level, which streams a calibrated confidence state by fusing intrinsic and conditional signals; and (ii) Policy Level, which utilizes this state to adaptively terminate low-confidence trajectories and dynamically schedule guidance for phase-appropriate conditioning strength. Experiments on both general and compositional benchmarks show that ScalingAR (1) improves base models by 12.5% on GenEval and 15.2% on TIIF-Bench, (2) efficiently reduces visual token consumption by 62.0% while outperforming baselines, and (3) successfully enhances robustness, mitigating performance drops by 26.0% in challenging scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2509_26376
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
Chen, Harold Haodong
Wu, Xianfeng
Shu, Wen-Jie
Guo, Rongjin
Lan, Disen
Yang, Harry
Chen, Ying-Cong
Computer Vision and Pattern Recognition
Test-time scaling (TTS) has demonstrated remarkable success in enhancing large language models, yet its application to next-token prediction (NTP) autoregressive (AR) image generation remains largely uncharted. Existing TTS approaches for visual AR (VAR), which rely on frequent partial decoding and external reward models, are ill-suited for NTP-based image generation due to the inherent incompleteness of intermediate decoding results. To bridge this gap, we introduce ScalingAR, the first TTS framework specifically designed for NTP-based AR image generation that eliminates the need for early decoding or auxiliary rewards. ScalingAR leverages token entropy as a novel signal in visual token generation and operates at two complementary scaling levels: (i) Profile Level, which streams a calibrated confidence state by fusing intrinsic and conditional signals; and (ii) Policy Level, which utilizes this state to adaptively terminate low-confidence trajectories and dynamically schedule guidance for phase-appropriate conditioning strength. Experiments on both general and compositional benchmarks show that ScalingAR (1) improves base models by 12.5% on GenEval and 15.2% on TIIF-Bench, (2) efficiently reduces visual token consumption by 62.0% while outperforming baselines, and (3) successfully enhances robustness, mitigating performance drops by 26.0% in challenging scenarios.
title Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.26376