OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Liyu, Li, Kehan, Han, Tingrui, Zhao, Tao, Sheng, Yuxuan, He, Shibo, Li, Chao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
von: He, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: He, Xiaoxuan, et al.
Veröffentlicht: (2025)
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
von: Li, Junzhe, et al.
Veröffentlicht: (2025)
von: Li, Junzhe, et al.
Veröffentlicht: (2025)
Smart-GRPO: Smartly Sampling Noise for Efficient RL of Flow-Matching Models
von: Yu, Benjamin, et al.
Veröffentlicht: (2025)
von: Yu, Benjamin, et al.
Veröffentlicht: (2025)
Flow-GRPO: Training Flow Matching Models via Online RL
von: Liu, Jie, et al.
Veröffentlicht: (2025)
von: Liu, Jie, et al.
Veröffentlicht: (2025)
BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
von: Li, Yuming, et al.
Veröffentlicht: (2025)
von: Li, Yuming, et al.
Veröffentlicht: (2025)
DEVIS-GRPO: Unleashing GRPO on Dynamic Extreme View Synthesis
von: Zuo, Yi, et al.
Veröffentlicht: (2026)
von: Zuo, Yi, et al.
Veröffentlicht: (2026)
DenseGRPO: From Sparse to Dense Reward for Flow Matching Model Alignment
von: Deng, Haoyou, et al.
Veröffentlicht: (2026)
von: Deng, Haoyou, et al.
Veröffentlicht: (2026)
Stepwise Credit Assignment for GRPO on Flow-Matching Models
von: Savani, Yash, et al.
Veröffentlicht: (2026)
von: Savani, Yash, et al.
Veröffentlicht: (2026)
DanceGRPO: Unleashing GRPO on Visual Generation
von: Xue, Zeyue, et al.
Veröffentlicht: (2025)
von: Xue, Zeyue, et al.
Veröffentlicht: (2025)
Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models
von: He, Dailan, et al.
Veröffentlicht: (2025)
von: He, Dailan, et al.
Veröffentlicht: (2025)
Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization
von: He, Xiaoxuan, et al.
Veröffentlicht: (2026)
von: He, Xiaoxuan, et al.
Veröffentlicht: (2026)
GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning
von: Xu, Yanchen, et al.
Veröffentlicht: (2025)
von: Xu, Yanchen, et al.
Veröffentlicht: (2025)
MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation
von: Ma, Xiaoxiao, et al.
Veröffentlicht: (2026)
von: Ma, Xiaoxiao, et al.
Veröffentlicht: (2026)
GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning
von: Li, Yujun, et al.
Veröffentlicht: (2026)
von: Li, Yujun, et al.
Veröffentlicht: (2026)
Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
von: Wang, Yibin, et al.
Veröffentlicht: (2025)
von: Wang, Yibin, et al.
Veröffentlicht: (2025)
Fine-Grained GRPO for Precise Preference Alignment in Flow Models
von: Zhou, Yujie, et al.
Veröffentlicht: (2025)
von: Zhou, Yujie, et al.
Veröffentlicht: (2025)
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO
von: Tong, Yunze, et al.
Veröffentlicht: (2026)
von: Tong, Yunze, et al.
Veröffentlicht: (2026)
Edit-GRPO: A Locality-Preserving Policy Optimization Framework for Image Editing
von: Xu, Shaodong, et al.
Veröffentlicht: (2026)
von: Xu, Shaodong, et al.
Veröffentlicht: (2026)
UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation
von: Liu, Jie, et al.
Veröffentlicht: (2026)
von: Liu, Jie, et al.
Veröffentlicht: (2026)
DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO
von: Liu, Henglin, et al.
Veröffentlicht: (2025)
von: Liu, Henglin, et al.
Veröffentlicht: (2025)
MotionGRPO: Overcoming Low Intra-Group Diversity in GRPO-Based Egocentric Motion Recovery
von: Yao, Nanjie, et al.
Veröffentlicht: (2026)
von: Yao, Nanjie, et al.
Veröffentlicht: (2026)
GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping
von: Wang, Jing, et al.
Veröffentlicht: (2025)
von: Wang, Jing, et al.
Veröffentlicht: (2025)
UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026)
STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
von: Ma, Xiaoxiao, et al.
Veröffentlicht: (2025)
von: Ma, Xiaoxiao, et al.
Veröffentlicht: (2025)
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
TIGFlow-GRPO: Trajectory Forecasting via Interaction-Aware Flow Matching and Reward-Guided Optimization
von: Jing, Xuepeng, et al.
Veröffentlicht: (2026)
von: Jing, Xuepeng, et al.
Veröffentlicht: (2026)
Boosting MLLM Reasoning with Text-Debiased Hint-GRPO
von: Huang, Qihan, et al.
Veröffentlicht: (2025)
von: Huang, Qihan, et al.
Veröffentlicht: (2025)
FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO
von: Zhao, Fufangchen, et al.
Veröffentlicht: (2025)
von: Zhao, Fufangchen, et al.
Veröffentlicht: (2025)
Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
von: Zhou, Renping, et al.
Veröffentlicht: (2025)
von: Zhou, Renping, et al.
Veröffentlicht: (2025)
GRAPE: Let GRPO Supervise Query Rewriting by Ranking for Retrieval
von: Zhang, Zhaohua, et al.
Veröffentlicht: (2025)
von: Zhang, Zhaohua, et al.
Veröffentlicht: (2025)
IRPO: Boosting Image Restoration via Post-training GRPO
von: Xu, Haoxuan, et al.
Veröffentlicht: (2025)
von: Xu, Haoxuan, et al.
Veröffentlicht: (2025)
TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
von: Ding, Zheng, et al.
Veröffentlicht: (2025)
von: Ding, Zheng, et al.
Veröffentlicht: (2025)
Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO
von: Cheng, Junhao, et al.
Veröffentlicht: (2025)
von: Cheng, Junhao, et al.
Veröffentlicht: (2025)
From Sparse to Dense: Multi-View GRPO for Flow Models via Augmented Condition Space
von: Bu, Jiazi, et al.
Veröffentlicht: (2026)
von: Bu, Jiazi, et al.
Veröffentlicht: (2026)
AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
von: Yuan, Shihao, et al.
Veröffentlicht: (2025)
von: Yuan, Shihao, et al.
Veröffentlicht: (2025)
Multi-GRPO: Multi-Group Advantage Estimation for Text-to-Image Generation with Tree-Based Trajectories and Multiple Rewards
von: Lyu, Qiang, et al.
Veröffentlicht: (2025)
von: Lyu, Qiang, et al.
Veröffentlicht: (2025)
DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding
von: Xiong, Junyu, et al.
Veröffentlicht: (2025)
von: Xiong, Junyu, et al.
Veröffentlicht: (2025)
E-GRPO: High Entropy Steps Drive Effective Reinforcement Learning for Flow Models
von: Zhang, Shengjun, et al.
Veröffentlicht: (2026)
von: Zhang, Shengjun, et al.
Veröffentlicht: (2026)
Expand and Prune: Maximizing Trajectory Diversity for Effective GRPO in Generative Models
von: Ge, Shiran, et al.
Veröffentlicht: (2025)
von: Ge, Shiran, et al.
Veröffentlicht: (2025)
Dr. Seg: Revisiting GRPO Training for Visual Large Language Models through Perception-Oriented Design
von: Sun, Haoxiang, et al.
Veröffentlicht: (2026)
von: Sun, Haoxiang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
von: He, Xiaoxuan, et al.
Veröffentlicht: (2025) -
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
von: Li, Junzhe, et al.
Veröffentlicht: (2025) -
Smart-GRPO: Smartly Sampling Noise for Efficient RL of Flow-Matching Models
von: Yu, Benjamin, et al.
Veröffentlicht: (2025) -
Flow-GRPO: Training Flow Matching Models via Online RL
von: Liu, Jie, et al.
Veröffentlicht: (2025) -
BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
von: Li, Yuming, et al.
Veröffentlicht: (2025)