FlowRL: Matching Reward Distributions for LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Xuekai, Cheng, Daixuan, Zhang, Dinghuai, Li, Hengli, Zhang, Kaiyan, Jiang, Che, Sun, Youbang, Hua, Ermo, Zuo, Yuxin, Lv, Xingtai, Zhang, Qizheng, Chen, Lin, Shao, Fanghao, Xue, Bo, Song, Yunchong, Yang, Zhenjie, Cui, Ganqu, Ding, Ning, Gao, Jianfeng, Liu, Xiaodong, Zhou, Bowen, Mei, Hongyuan, Lin, Zhouhan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to Synthesize Text Data without Model Collapse?
by: Zhu, Xuekai, et al.
Published: (2024)
by: Zhu, Xuekai, et al.
Published: (2024)
Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization
by: Hua, Ermo, et al.
Published: (2024)
by: Hua, Ermo, et al.
Published: (2024)
Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention
by: Lv, Xingtai, et al.
Published: (2024)
by: Lv, Xingtai, et al.
Published: (2024)
Flow of Spans: Generalizing Language Models to Dynamic Span-Vocabulary via GFlowNets
by: Xue, Bo, et al.
Published: (2026)
by: Xue, Bo, et al.
Published: (2026)
Controlling Exploration-Exploitation in GFlowNets via Markov Chain Perspectives
by: Chen, Lin, et al.
Published: (2026)
by: Chen, Lin, et al.
Published: (2026)
Technologies on Effectiveness and Efficiency: A Survey of State Spaces Models
by: Lv, Xingtai, et al.
Published: (2025)
by: Lv, Xingtai, et al.
Published: (2025)
Towards a Unified View of Large Language Model Post-Training
by: Lv, Xingtai, et al.
Published: (2025)
by: Lv, Xingtai, et al.
Published: (2025)
PaD: Program-aided Distillation Can Teach Small Models Reasoning Better than Chain-of-thought Fine-tuning
by: Zhu, Xuekai, et al.
Published: (2023)
by: Zhu, Xuekai, et al.
Published: (2023)
TTRL: Test-Time Reinforcement Learning
by: Zuo, Yuxin, et al.
Published: (2025)
by: Zuo, Yuxin, et al.
Published: (2025)
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
by: Zuo, Yuxin, et al.
Published: (2025)
by: Zuo, Yuxin, et al.
Published: (2025)
UltraMedical: Building Specialized Generalists in Biomedicine
by: Zhang, Kaiyan, et al.
Published: (2024)
by: Zhang, Kaiyan, et al.
Published: (2024)
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
by: Li, Haozhan, et al.
Published: (2025)
by: Li, Haozhan, et al.
Published: (2025)
Fast and Slow Generating: An Empirical Study on Large and Small Language Models Collaborative Decoding
by: Zhang, Kaiyan, et al.
Published: (2024)
by: Zhang, Kaiyan, et al.
Published: (2024)
Intuitive Fine-Tuning: Towards Simplifying Alignment into a Single Process
by: Hua, Ermo, et al.
Published: (2024)
by: Hua, Ermo, et al.
Published: (2024)
Post-Trained MoE Can Skip Half Experts via Self-Distillation
by: Lv, Xingtai, et al.
Published: (2026)
by: Lv, Xingtai, et al.
Published: (2026)
Cluster-wise Graph Transformer with Dual-granularity Kernelized Attention
by: Huang, Siyuan, et al.
Published: (2024)
by: Huang, Siyuan, et al.
Published: (2024)
Automating Exploratory Proteomics Research via Language Models
by: Ding, Ning, et al.
Published: (2024)
by: Ding, Ning, et al.
Published: (2024)
Automating Exploratory Multiomics Research via Language Models
by: Qu, Shang, et al.
Published: (2025)
by: Qu, Shang, et al.
Published: (2025)
Critical Data Size of Language Models from a Grokking Perspective
by: Zhu, Xuekai, et al.
Published: (2024)
by: Zhu, Xuekai, et al.
Published: (2024)
JustRL: Scaling a 1.5B LLM with a Simple RL Recipe
by: He, Bingxiang, et al.
Published: (2025)
by: He, Bingxiang, et al.
Published: (2025)
How Far Can Unsupervised RLVR Scale LLM Training?
by: He, Bingxiang, et al.
Published: (2026)
by: He, Bingxiang, et al.
Published: (2026)
FlowRL: A Taxonomy and Modular Framework for Reinforcement Learning with Diffusion Policies
by: Gao, Chenxiao, et al.
Published: (2026)
by: Gao, Chenxiao, et al.
Published: (2026)
Graph Parsing Networks
by: Song, Yunchong, et al.
Published: (2024)
by: Song, Yunchong, et al.
Published: (2024)
Towards Compressive and Scalable Recurrent Memory
by: Song, Yunchong, et al.
Published: (2026)
by: Song, Yunchong, et al.
Published: (2026)
FlowRL: Flow-Augmented Few-Shot Reinforcement Learning for Semi-Structured Sensor Data
by: Pivezhandi, Mohammad, et al.
Published: (2024)
by: Pivezhandi, Mohammad, et al.
Published: (2024)
Free Process Rewards without Process Labels
by: Yuan, Lifan, et al.
Published: (2024)
by: Yuan, Lifan, et al.
Published: (2024)
CoGenesis: A Framework Collaborating Large and Small Language Models for Secure Context-Aware Instruction Following
by: Zhang, Kaiyan, et al.
Published: (2024)
by: Zhang, Kaiyan, et al.
Published: (2024)
Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search Engines
by: Long, Xinwei, et al.
Published: (2025)
by: Long, Xinwei, et al.
Published: (2025)
Nabla-R2D3: Effective and Efficient 3D Diffusion Alignment with 2D Rewards
by: Liu, Qingming, et al.
Published: (2025)
by: Liu, Qingming, et al.
Published: (2025)
Predictors of Postoperative Outcomes after Arthroscopic Partial Meniscectomy: A Retrospective Analysis
by: Fan Lin, et al.
Published: (2024)
by: Fan Lin, et al.
Published: (2024)
Reasoning with Exploration: An Entropy Perspective
by: Cheng, Daixuan, et al.
Published: (2025)
by: Cheng, Daixuan, et al.
Published: (2025)
A Survey of Reinforcement Learning for Large Reasoning Models
by: Zhang, Kaiyan, et al.
Published: (2025)
by: Zhang, Kaiyan, et al.
Published: (2025)
Scaling Reward Modeling without Human Supervision
by: Fan, Jingxuan, et al.
Published: (2026)
by: Fan, Jingxuan, et al.
Published: (2026)
Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
SH2: Self-Highlighted Hesitation Helps You Decode More Truthfully
by: Kai, Jushi, et al.
Published: (2024)
by: Kai, Jushi, et al.
Published: (2024)
Process Reinforcement through Implicit Rewards
by: Cui, Ganqu, et al.
Published: (2025)
by: Cui, Ganqu, et al.
Published: (2025)
FlowLM: Few-Step Language Modeling via Diffusion-to-Flow Adaptation
by: Zhang, Runzhe, et al.
Published: (2026)
by: Zhang, Runzhe, et al.
Published: (2026)
Diff-Instruct with Diffused Reward: Towards Principled One-step Generator RL
by: Wu, Junyi, et al.
Published: (2026)
by: Wu, Junyi, et al.
Published: (2026)
Caffarelli-Kohn-Nirenberg Inequalities in Weak Lebesgue Spaces
by: Wang, Dinghuai
Published: (2026)
by: Wang, Dinghuai
Published: (2026)
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
by: Liu, Runze, et al.
Published: (2025)
by: Liu, Runze, et al.
Published: (2025)
Similar Items
-
How to Synthesize Text Data without Model Collapse?
by: Zhu, Xuekai, et al.
Published: (2024) -
Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization
by: Hua, Ermo, et al.
Published: (2024) -
Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention
by: Lv, Xingtai, et al.
Published: (2024) -
Flow of Spans: Generalizing Language Models to Dynamic Span-Vocabulary via GFlowNets
by: Xue, Bo, et al.
Published: (2026) -
Controlling Exploration-Exploitation in GFlowNets via Markov Chain Perspectives
by: Chen, Lin, et al.
Published: (2026)