Saved in:
| Main Authors: | Lai, Yao, Xiong, Xuyuan, Xue, Zeyue, Chen, Guojin, Wang, Jing, Liu, Xihui, Zhang, Rui, Mullins, Robert, Yu, Bei, Luo, Ping |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2606.00228 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DanceGRPO: Unleashing GRPO on Visual Generation
by: Xue, Zeyue, et al.
Published: (2025)
by: Xue, Zeyue, et al.
Published: (2025)
OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models
by: Zhang, Liyu, et al.
Published: (2026)
by: Zhang, Liyu, et al.
Published: (2026)
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
Stepwise Credit Assignment for GRPO on Flow-Matching Models
by: Savani, Yash, et al.
Published: (2026)
by: Savani, Yash, et al.
Published: (2026)
FlowSE-GRPO: Training Flow Matching Speech Enhancement via Online Reinforcement Learning
by: Wang, Haoxu, et al.
Published: (2026)
by: Wang, Haoxu, et al.
Published: (2026)
Flow-GRPO: Training Flow Matching Models via Online RL
by: Liu, Jie, et al.
Published: (2025)
by: Liu, Jie, et al.
Published: (2025)
TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
by: He, Xiaoxuan, et al.
Published: (2025)
by: He, Xiaoxuan, et al.
Published: (2025)
GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning
by: Xu, Yanchen, et al.
Published: (2025)
by: Xu, Yanchen, et al.
Published: (2025)
Open-Source Differentiable Lithography Imaging Framework
by: Chen, Guojin, et al.
Published: (2024)
by: Chen, Guojin, et al.
Published: (2024)
TIGFlow-GRPO: Trajectory Forecasting via Interaction-Aware Flow Matching and Reward-Guided Optimization
by: Jing, Xuepeng, et al.
Published: (2026)
by: Jing, Xuepeng, et al.
Published: (2026)
GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO
by: Liu, Henglin, et al.
Published: (2025)
by: Liu, Henglin, et al.
Published: (2025)
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
by: Li, Junzhe, et al.
Published: (2025)
by: Li, Junzhe, et al.
Published: (2025)
Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization
by: He, Xiaoxuan, et al.
Published: (2026)
by: He, Xiaoxuan, et al.
Published: (2026)
How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning
by: Tian, Minghao, et al.
Published: (2026)
by: Tian, Minghao, et al.
Published: (2026)
Stratified GRPO: Handling Structural Heterogeneity in Reinforcement Learning of LLM Search Agents
by: Zhu, Mingkang, et al.
Published: (2025)
by: Zhu, Mingkang, et al.
Published: (2025)
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
Graph-GRPO: Training Graph Flow Models with Reinforcement Learning
by: Zhu, Baoheng, et al.
Published: (2026)
by: Zhu, Baoheng, et al.
Published: (2026)
GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning
by: Li, Yujun, et al.
Published: (2026)
by: Li, Yujun, et al.
Published: (2026)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
by: Yari, Amir Hossein, et al.
Published: (2026)
by: Yari, Amir Hossein, et al.
Published: (2026)
Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
MotionGRPO: Overcoming Low Intra-Group Diversity in GRPO-Based Egocentric Motion Recovery
by: Yao, Nanjie, et al.
Published: (2026)
by: Yao, Nanjie, et al.
Published: (2026)
DenseGRPO: From Sparse to Dense Reward for Flow Matching Model Alignment
by: Deng, Haoyou, et al.
Published: (2026)
by: Deng, Haoyou, et al.
Published: (2026)
Smart-GRPO: Smartly Sampling Noise for Efficient RL of Flow-Matching Models
by: Yu, Benjamin, et al.
Published: (2025)
by: Yu, Benjamin, et al.
Published: (2025)
LithoSeg: A Coarse-to-Fine Framework for High-Precision Lithography Segmentation
by: He, Xinyu, et al.
Published: (2025)
by: He, Xinyu, et al.
Published: (2025)
DEVIS-GRPO: Unleashing GRPO on Dynamic Extreme View Synthesis
by: Zuo, Yi, et al.
Published: (2026)
by: Zuo, Yi, et al.
Published: (2026)
$λ$-GRPO: Unifying the GRPO Frameworks with Learnable Token Preferences
by: Wang, Yining, et al.
Published: (2025)
by: Wang, Yining, et al.
Published: (2025)
S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
by: Dai, Muzhi, et al.
Published: (2025)
by: Dai, Muzhi, et al.
Published: (2025)
MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation
by: Ma, Xiaoxiao, et al.
Published: (2026)
by: Ma, Xiaoxiao, et al.
Published: (2026)
BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
by: Li, Yuming, et al.
Published: (2025)
by: Li, Yuming, et al.
Published: (2025)
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
by: Chen, Minghan, et al.
Published: (2025)
by: Chen, Minghan, et al.
Published: (2025)
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
TIC-GRPO: Provable and Efficient Optimization for Reinforcement Learning from Human Feedback
by: Pang, Lei, et al.
Published: (2025)
by: Pang, Lei, et al.
Published: (2025)
XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation
by: Bamba, Udbhav, et al.
Published: (2025)
by: Bamba, Udbhav, et al.
Published: (2025)
E-GRPO: High Entropy Steps Drive Effective Reinforcement Learning for Flow Models
by: Zhang, Shengjun, et al.
Published: (2026)
by: Zhang, Shengjun, et al.
Published: (2026)
DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning
by: Chen, Xiwen, et al.
Published: (2025)
by: Chen, Xiwen, et al.
Published: (2025)
TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
by: Ding, Zheng, et al.
Published: (2025)
by: Ding, Zheng, et al.
Published: (2025)
MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting
by: Wei, Kangda, et al.
Published: (2026)
by: Wei, Kangda, et al.
Published: (2026)
What is the Alignment Objective of GRPO?
by: Vojnovic, Milan, et al.
Published: (2025)
by: Vojnovic, Milan, et al.
Published: (2025)
AnalogCoder-Pro: Unifying Analog Circuit Generation and Optimization via Multi-modal LLMs
by: Lai, Yao, et al.
Published: (2025)
by: Lai, Yao, et al.
Published: (2025)
Similar Items
-
DanceGRPO: Unleashing GRPO on Visual Generation
by: Xue, Zeyue, et al.
Published: (2025) -
OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models
by: Zhang, Liyu, et al.
Published: (2026) -
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
by: Chen, Yi, et al.
Published: (2025) -
Stepwise Credit Assignment for GRPO on Flow-Matching Models
by: Savani, Yash, et al.
Published: (2026) -
FlowSE-GRPO: Training Flow Matching Speech Enhancement via Online Reinforcement Learning
by: Wang, Haoxu, et al.
Published: (2026)