AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Shihao, Liu, Yahui, Yue, Yang, Zhang, Jingyuan, Zuo, Wangmeng, Wang, Qi, Zhang, Fuzheng, Zhou, Guorui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911100464791552
author Yuan, Shihao
Liu, Yahui
Yue, Yang
Zhang, Jingyuan
Zuo, Wangmeng
Wang, Qi
Zhang, Fuzheng
Zhou, Guorui
author_facet Yuan, Shihao
Liu, Yahui
Yue, Yang
Zhang, Jingyuan
Zuo, Wangmeng
Wang, Qi
Zhang, Fuzheng
Zhou, Guorui
contents Inspired by the success of reinforcement learning (RL) in refining large language models (LLMs), we propose AR-GRPO, an approach to integrate online RL training into autoregressive (AR) image generation models. We adapt the Group Relative Policy Optimization (GRPO) algorithm to refine the vanilla autoregressive models' outputs by carefully designed reward functions that evaluate generated images across multiple quality dimensions, including perceptual quality, realism, and semantic fidelity. We conduct comprehensive experiments on both class-conditional (i.e., class-to-image) and text-conditional (i.e., text-to-image) image generation tasks, demonstrating that our RL-enhanced framework significantly improves both the image quality and human preference of generated images compared to the standard AR baselines. Our results show consistent improvements across various evaluation metrics, establishing the viability of RL-based optimization for AR image generation and opening new avenues for controllable and high-quality image synthesis. The source codes and models are available at: https://github.com/Kwai-Klear/AR-GRPO.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06924
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
Yuan, Shihao
Liu, Yahui
Yue, Yang
Zhang, Jingyuan
Zuo, Wangmeng
Wang, Qi
Zhang, Fuzheng
Zhou, Guorui
Computer Vision and Pattern Recognition
Inspired by the success of reinforcement learning (RL) in refining large language models (LLMs), we propose AR-GRPO, an approach to integrate online RL training into autoregressive (AR) image generation models. We adapt the Group Relative Policy Optimization (GRPO) algorithm to refine the vanilla autoregressive models' outputs by carefully designed reward functions that evaluate generated images across multiple quality dimensions, including perceptual quality, realism, and semantic fidelity. We conduct comprehensive experiments on both class-conditional (i.e., class-to-image) and text-conditional (i.e., text-to-image) image generation tasks, demonstrating that our RL-enhanced framework significantly improves both the image quality and human preference of generated images compared to the standard AR baselines. Our results show consistent improvements across various evaluation metrics, establishing the viability of RL-based optimization for AR image generation and opening new avenues for controllable and high-quality image synthesis. The source codes and models are available at: https://github.com/Kwai-Klear/AR-GRPO.
title AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.06924