Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tong, Chengzhuo, Guo, Ziyu, Zhang, Renrui, Shan, Wenyu, Wei, Xinyu, Xing, Zhenghao, Li, Hongsheng, Heng, Pheng-Ann
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918052088512512
author Tong, Chengzhuo
Guo, Ziyu
Zhang, Renrui
Shan, Wenyu
Wei, Xinyu
Xing, Zhenghao
Li, Hongsheng
Heng, Pheng-Ann
author_facet Tong, Chengzhuo
Guo, Ziyu
Zhang, Renrui
Shan, Wenyu
Wei, Xinyu
Xing, Zhenghao
Li, Hongsheng
Heng, Pheng-Ann
contents Recent advancements underscore the significant role of Reinforcement Learning (RL) in enhancing the Chain-of-Thought (CoT) reasoning capabilities of large language models (LLMs). Two prominent RL algorithms, Direct Preference Optimization (DPO) and Group Relative Policy Optimization (GRPO), are central to these developments, showcasing different pros and cons. Autoregressive image generation, also interpretable as a sequential CoT reasoning process, presents unique challenges distinct from LLM-based CoT reasoning. These encompass ensuring text-image consistency, improving image aesthetic quality, and designing sophisticated reward models, rather than relying on simpler rule-based rewards. While recent efforts have extended RL to this domain, these explorations typically lack an in-depth analysis of the domain-specific challenges and the characteristics of different RL strategies. To bridge this gap, we provide the first comprehensive investigation of the GRPO and DPO algorithms in autoregressive image generation, evaluating their in-domain performance and out-of-domain generalization, while scrutinizing the impact of different reward models on their respective capabilities. Our findings reveal that GRPO and DPO exhibit distinct advantages, and crucially, that reward models possessing stronger intrinsic generalization capabilities potentially enhance the generalization potential of the applied RL algorithms. Furthermore, we systematically explore three prevalent scaling strategies to enhance both their in-domain and out-of-domain proficiency, deriving unique insights into efficiently scaling performance for each paradigm. We hope our study paves a new path for inspiring future work on developing more effective RL algorithms to achieve robust CoT reasoning in the realm of autoregressive image generation. Code is released at https://github.com/ZiyuGuo99/Image-Generation-CoT
format Preprint
id arxiv_https___arxiv_org_abs_2505_17017
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
Tong, Chengzhuo
Guo, Ziyu
Zhang, Renrui
Shan, Wenyu
Wei, Xinyu
Xing, Zhenghao
Li, Hongsheng
Heng, Pheng-Ann
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Recent advancements underscore the significant role of Reinforcement Learning (RL) in enhancing the Chain-of-Thought (CoT) reasoning capabilities of large language models (LLMs). Two prominent RL algorithms, Direct Preference Optimization (DPO) and Group Relative Policy Optimization (GRPO), are central to these developments, showcasing different pros and cons. Autoregressive image generation, also interpretable as a sequential CoT reasoning process, presents unique challenges distinct from LLM-based CoT reasoning. These encompass ensuring text-image consistency, improving image aesthetic quality, and designing sophisticated reward models, rather than relying on simpler rule-based rewards. While recent efforts have extended RL to this domain, these explorations typically lack an in-depth analysis of the domain-specific challenges and the characteristics of different RL strategies. To bridge this gap, we provide the first comprehensive investigation of the GRPO and DPO algorithms in autoregressive image generation, evaluating their in-domain performance and out-of-domain generalization, while scrutinizing the impact of different reward models on their respective capabilities. Our findings reveal that GRPO and DPO exhibit distinct advantages, and crucially, that reward models possessing stronger intrinsic generalization capabilities potentially enhance the generalization potential of the applied RL algorithms. Furthermore, we systematically explore three prevalent scaling strategies to enhance both their in-domain and out-of-domain proficiency, deriving unique insights into efficiently scaling performance for each paradigm. We hope our study paves a new path for inspiring future work on developing more effective RL algorithms to achieve robust CoT reasoning in the realm of autoregressive image generation. Code is released at https://github.com/ZiyuGuo99/Image-Generation-CoT
title Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.17017