Group Critical-token Policy Optimization for Autoregressive Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Guohui, Yu, Hu, Ma, Xiaoxiao, Zhang, JingHao, Pan, Yaning, Yao, Mingde, Xiao, Jie, Huang, Linjiang, Zhao, Feng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909809485283328
author Zhang, Guohui
Yu, Hu
Ma, Xiaoxiao
Zhang, JingHao
Pan, Yaning
Yao, Mingde
Xiao, Jie
Huang, Linjiang
Zhao, Feng
author_facet Zhang, Guohui
Yu, Hu
Ma, Xiaoxiao
Zhang, JingHao
Pan, Yaning
Yao, Mingde
Xiao, Jie
Huang, Linjiang
Zhao, Feng
contents Recent studies have extended Reinforcement Learning with Verifiable Rewards (RLVR) to autoregressive (AR) visual generation and achieved promising progress. However, existing methods typically apply uniform optimization across all image tokens, while the varying contributions of different image tokens for RLVR's training remain unexplored. In fact, the key obstacle lies in how to identify more critical image tokens during AR generation and implement effective token-wise optimization for them. To tackle this challenge, we propose $\textbf{G}$roup $\textbf{C}$ritical-token $\textbf{P}$olicy $\textbf{O}$ptimization ($\textbf{GCPO}$), which facilitates effective policy optimization on critical tokens. We identify the critical tokens in RLVR-based AR generation from three perspectives, specifically: $\textbf{(1)}$ Causal dependency: early tokens fundamentally determine the later tokens and final image effect due to unidirectional dependency; $\textbf{(2)}$ Entropy-induced spatial structure: tokens with high entropy gradients correspond to image structure and bridges distinct visual regions; $\textbf{(3)}$ RLVR-focused token diversity: tokens with low visual similarity across a group of sampled images contribute to richer token-level diversity. For these identified critical tokens, we further introduce a dynamic token-wise advantage weight to encourage exploration, based on confidence divergence between the policy model and reference model. By leveraging 30\% of the image tokens, GCPO achieves better performance than GRPO with full tokens. Extensive experiments on multiple text-to-image benchmarks for both AR models and unified multimodal models demonstrate the effectiveness of GCPO for AR visual generation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22485
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Group Critical-token Policy Optimization for Autoregressive Image Generation
Zhang, Guohui
Yu, Hu
Ma, Xiaoxiao
Zhang, JingHao
Pan, Yaning
Yao, Mingde
Xiao, Jie
Huang, Linjiang
Zhao, Feng
Computer Vision and Pattern Recognition
Recent studies have extended Reinforcement Learning with Verifiable Rewards (RLVR) to autoregressive (AR) visual generation and achieved promising progress. However, existing methods typically apply uniform optimization across all image tokens, while the varying contributions of different image tokens for RLVR's training remain unexplored. In fact, the key obstacle lies in how to identify more critical image tokens during AR generation and implement effective token-wise optimization for them. To tackle this challenge, we propose $\textbf{G}$roup $\textbf{C}$ritical-token $\textbf{P}$olicy $\textbf{O}$ptimization ($\textbf{GCPO}$), which facilitates effective policy optimization on critical tokens. We identify the critical tokens in RLVR-based AR generation from three perspectives, specifically: $\textbf{(1)}$ Causal dependency: early tokens fundamentally determine the later tokens and final image effect due to unidirectional dependency; $\textbf{(2)}$ Entropy-induced spatial structure: tokens with high entropy gradients correspond to image structure and bridges distinct visual regions; $\textbf{(3)}$ RLVR-focused token diversity: tokens with low visual similarity across a group of sampled images contribute to richer token-level diversity. For these identified critical tokens, we further introduce a dynamic token-wise advantage weight to encourage exploration, based on confidence divergence between the policy model and reference model. By leveraging 30\% of the image tokens, GCPO achieves better performance than GRPO with full tokens. Extensive experiments on multiple text-to-image benchmarks for both AR models and unified multimodal models demonstrate the effectiveness of GCPO for AR visual generation.
title Group Critical-token Policy Optimization for Autoregressive Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.22485