Flow-GRPO: Training Flow Matching Models via Online RL

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Jie, Liu, Gongye, Liang, Jiajun, Li, Yangguang, Liu, Jiaheng, Wang, Xintao, Wan, Pengfei, Zhang, Di, Ouyang, Wanli
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914115758325760
author Liu, Jie
Liu, Gongye
Liang, Jiajun
Li, Yangguang
Liu, Jiaheng
Wang, Xintao
Wan, Pengfei
Zhang, Di
Ouyang, Wanli
author_facet Liu, Jie
Liu, Gongye
Liang, Jiajun
Li, Yangguang
Liu, Jiaheng
Wang, Xintao
Wan, Pengfei
Zhang, Di
Ouyang, Wanli
contents We propose Flow-GRPO, the first method to integrate online policy gradient reinforcement learning (RL) into flow matching models. Our approach uses two key strategies: (1) an ODE-to-SDE conversion that transforms a deterministic Ordinary Differential Equation (ODE) into an equivalent Stochastic Differential Equation (SDE) that matches the original model's marginal distribution at all timesteps, enabling statistical sampling for RL exploration; and (2) a Denoising Reduction strategy that reduces training denoising steps while retaining the original number of inference steps, significantly improving sampling efficiency without sacrificing performance. Empirically, Flow-GRPO is effective across multiple text-to-image tasks. For compositional generation, RL-tuned SD3.5-M generates nearly perfect object counts, spatial relations, and fine-grained attributes, increasing GenEval accuracy from $63\%$ to $95\%$. In visual text rendering, accuracy improves from $59\%$ to $92\%$, greatly enhancing text generation. Flow-GRPO also achieves substantial gains in human preference alignment. Notably, very little reward hacking occurred, meaning rewards did not increase at the cost of appreciable image quality or diversity degradation.
format Preprint
id arxiv_https___arxiv_org_abs_2505_05470
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Flow-GRPO: Training Flow Matching Models via Online RL
Liu, Jie
Liu, Gongye
Liang, Jiajun
Li, Yangguang
Liu, Jiaheng
Wang, Xintao
Wan, Pengfei
Zhang, Di
Ouyang, Wanli
Computer Vision and Pattern Recognition
Artificial Intelligence
We propose Flow-GRPO, the first method to integrate online policy gradient reinforcement learning (RL) into flow matching models. Our approach uses two key strategies: (1) an ODE-to-SDE conversion that transforms a deterministic Ordinary Differential Equation (ODE) into an equivalent Stochastic Differential Equation (SDE) that matches the original model's marginal distribution at all timesteps, enabling statistical sampling for RL exploration; and (2) a Denoising Reduction strategy that reduces training denoising steps while retaining the original number of inference steps, significantly improving sampling efficiency without sacrificing performance. Empirically, Flow-GRPO is effective across multiple text-to-image tasks. For compositional generation, RL-tuned SD3.5-M generates nearly perfect object counts, spatial relations, and fine-grained attributes, increasing GenEval accuracy from $63\%$ to $95\%$. In visual text rendering, accuracy improves from $59\%$ to $92\%$, greatly enhancing text generation. Flow-GRPO also achieves substantial gains in human preference alignment. Notably, very little reward hacking occurred, meaning rewards did not increase at the cost of appreciable image quality or diversity degradation.
title Flow-GRPO: Training Flow Matching Models via Online RL
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.05470