FlowRL: Matching Reward Distributions for LLM Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Xuekai, Cheng, Daixuan, Zhang, Dinghuai, Li, Hengli, Zhang, Kaiyan, Jiang, Che, Sun, Youbang, Hua, Ermo, Zuo, Yuxin, Lv, Xingtai, Zhang, Qizheng, Chen, Lin, Shao, Fanghao, Xue, Bo, Song, Yunchong, Yang, Zhenjie, Cui, Ganqu, Ding, Ning, Gao, Jianfeng, Liu, Xiaodong, Zhou, Bowen, Mei, Hongyuan, Lin, Zhouhan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909885386457088
author Zhu, Xuekai
Cheng, Daixuan
Zhang, Dinghuai
Li, Hengli
Zhang, Kaiyan
Jiang, Che
Sun, Youbang
Hua, Ermo
Zuo, Yuxin
Lv, Xingtai
Zhang, Qizheng
Chen, Lin
Shao, Fanghao
Xue, Bo
Song, Yunchong
Yang, Zhenjie
Cui, Ganqu
Ding, Ning
Gao, Jianfeng
Liu, Xiaodong
Zhou, Bowen
Mei, Hongyuan
Lin, Zhouhan
author_facet Zhu, Xuekai
Cheng, Daixuan
Zhang, Dinghuai
Li, Hengli
Zhang, Kaiyan
Jiang, Che
Sun, Youbang
Hua, Ermo
Zuo, Yuxin
Lv, Xingtai
Zhang, Qizheng
Chen, Lin
Shao, Fanghao
Xue, Bo
Song, Yunchong
Yang, Zhenjie
Cui, Ganqu
Ding, Ning
Gao, Jianfeng
Liu, Xiaodong
Zhou, Bowen
Mei, Hongyuan
Lin, Zhouhan
contents We propose FlowRL: matching the full reward distribution via flow balancing instead of maximizing rewards in large language model (LLM) reinforcement learning (RL). Recent advanced reasoning models adopt reward-maximizing methods (\eg, PPO and GRPO), which tend to over-optimize dominant reward signals while neglecting less frequent but valid reasoning paths, thus reducing diversity. In contrast, we transform scalar rewards into a normalized target distribution using a learnable partition function, and then minimize the reverse KL divergence between the policy and the target distribution. We implement this idea as a flow-balanced optimization method that promotes diverse exploration and generalizable reasoning trajectories. We conduct experiments on math and code reasoning tasks: FlowRL achieves a significant average improvement of $10.0\%$ over GRPO and $5.1\%$ over PPO on math benchmarks, and performs consistently better on code reasoning tasks. These results highlight reward distribution-matching as a key step toward efficient exploration and diverse reasoning in LLM reinforcement learning.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15207
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FlowRL: Matching Reward Distributions for LLM Reasoning
Zhu, Xuekai
Cheng, Daixuan
Zhang, Dinghuai
Li, Hengli
Zhang, Kaiyan
Jiang, Che
Sun, Youbang
Hua, Ermo
Zuo, Yuxin
Lv, Xingtai
Zhang, Qizheng
Chen, Lin
Shao, Fanghao
Xue, Bo
Song, Yunchong
Yang, Zhenjie
Cui, Ganqu
Ding, Ning
Gao, Jianfeng
Liu, Xiaodong
Zhou, Bowen
Mei, Hongyuan
Lin, Zhouhan
Machine Learning
Artificial Intelligence
Computation and Language
We propose FlowRL: matching the full reward distribution via flow balancing instead of maximizing rewards in large language model (LLM) reinforcement learning (RL). Recent advanced reasoning models adopt reward-maximizing methods (\eg, PPO and GRPO), which tend to over-optimize dominant reward signals while neglecting less frequent but valid reasoning paths, thus reducing diversity. In contrast, we transform scalar rewards into a normalized target distribution using a learnable partition function, and then minimize the reverse KL divergence between the policy and the target distribution. We implement this idea as a flow-balanced optimization method that promotes diverse exploration and generalizable reasoning trajectories. We conduct experiments on math and code reasoning tasks: FlowRL achieves a significant average improvement of $10.0\%$ over GRPO and $5.1\%$ over PPO on math benchmarks, and performs consistently better on code reasoning tasks. These results highlight reward distribution-matching as a key step toward efficient exploration and diverse reasoning in LLM reinforcement learning.
title FlowRL: Matching Reward Distributions for LLM Reasoning
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.15207