Advances in GRPO for Generation Models: A Survey
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Zexiang, He, Xianglong, Li, Yangguang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LithoGRPO: Fast Inverse Lithography via GRPO Reinforced Flow Matching
by: Lai, Yao, et al.
Published: (2026)
by: Lai, Yao, et al.
Published: (2026)
Expand and Prune: Maximizing Trajectory Diversity for Effective GRPO in Generative Models
by: Ge, Shiran, et al.
Published: (2025)
by: Ge, Shiran, et al.
Published: (2025)
Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening
by: He, Andre, et al.
Published: (2025)
by: He, Andre, et al.
Published: (2025)
Execution-Grounded Credit Assignment for GRPO in Code Generation
by: Kumar, Abhijit, et al.
Published: (2026)
by: Kumar, Abhijit, et al.
Published: (2026)
GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning
by: Xu, Yanchen, et al.
Published: (2025)
by: Xu, Yanchen, et al.
Published: (2025)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
by: Yari, Amir Hossein, et al.
Published: (2026)
by: Yari, Amir Hossein, et al.
Published: (2026)
BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
by: Li, Yuming, et al.
Published: (2025)
by: Li, Yuming, et al.
Published: (2025)
TrackDiffuser: Nearly Model-Free Bayesian Filtering with Diffusion Model
by: He, Yangguang, et al.
Published: (2025)
by: He, Yangguang, et al.
Published: (2025)
F-TIS: Harnessing Diverse Models in Collaborative GRPO
by: Blagoev, Nikolay, et al.
Published: (2026)
by: Blagoev, Nikolay, et al.
Published: (2026)
Advancing Speech Understanding in Speech-Aware Language Models with GRPO
by: Elmakies, Avishai, et al.
Published: (2025)
by: Elmakies, Avishai, et al.
Published: (2025)
GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning
by: Li, Yujun, et al.
Published: (2026)
by: Li, Yujun, et al.
Published: (2026)
Latent Imitator: Generating Natural Individual Discriminatory Instances for Black-Box Fairness Testing
by: Xiao, Yisong, et al.
Published: (2023)
by: Xiao, Yisong, et al.
Published: (2023)
DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning
by: Chen, Xiwen, et al.
Published: (2025)
by: Chen, Xiwen, et al.
Published: (2025)
Graph-GRPO: Training Graph Flow Models with Reinforcement Learning
by: Zhu, Baoheng, et al.
Published: (2026)
by: Zhu, Baoheng, et al.
Published: (2026)
GRPO is Secretly a Process Reward Model
by: Sullivan, Michael, et al.
Published: (2025)
by: Sullivan, Michael, et al.
Published: (2025)
F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking
by: Surana, Rohan, et al.
Published: (2026)
by: Surana, Rohan, et al.
Published: (2026)
Low-bit Model Quantization for Deep Neural Networks: A Survey
by: Liu, Kai, et al.
Published: (2025)
by: Liu, Kai, et al.
Published: (2025)
VeriGate: Verifier-Gated Step-Level Supervision for GRPO
by: Agrawal, Aakriti, et al.
Published: (2026)
by: Agrawal, Aakriti, et al.
Published: (2026)
Transformation-Augmented GRPO for Enhancing Exploration in Reasoning of Large Language Models
by: Le, Khiem, et al.
Published: (2026)
by: Le, Khiem, et al.
Published: (2026)
Predictive Scaling Laws for Efficient GRPO Training of Large Reasoning Models
by: Nimmaturi, Datta, et al.
Published: (2025)
by: Nimmaturi, Datta, et al.
Published: (2025)
How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning
by: Tian, Minghao, et al.
Published: (2026)
by: Tian, Minghao, et al.
Published: (2026)
Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models
by: He, Dailan, et al.
Published: (2025)
by: He, Dailan, et al.
Published: (2025)
f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment
by: Haldar, Rajdeep, et al.
Published: (2026)
by: Haldar, Rajdeep, et al.
Published: (2026)
What is the Alignment Objective of GRPO?
by: Vojnovic, Milan, et al.
Published: (2025)
by: Vojnovic, Milan, et al.
Published: (2025)
TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
by: Ding, Zheng, et al.
Published: (2025)
by: Ding, Zheng, et al.
Published: (2025)
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models
by: Ding, Fei, et al.
Published: (2025)
by: Ding, Fei, et al.
Published: (2025)
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning
by: Wang, Hu, et al.
Published: (2025)
by: Wang, Hu, et al.
Published: (2025)
CoRPO: Adding a Correctness Bias to GRPO Improves Generalization
by: Garg, Anisha, et al.
Published: (2025)
by: Garg, Anisha, et al.
Published: (2025)
Towards General Industrial Intelligence: A Survey of Continual Large Models in Industrial IoT
by: Chen, Jiao, et al.
Published: (2024)
by: Chen, Jiao, et al.
Published: (2024)
Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO
by: Zheng, Jinquan, et al.
Published: (2026)
by: Zheng, Jinquan, et al.
Published: (2026)
From Shallow to Deep: Pinning Semantic Intent via Causal GRPO
by: Zhou, Shuyi, et al.
Published: (2026)
by: Zhou, Shuyi, et al.
Published: (2026)
Can GRPO Help LLMs Transcend Their Pretraining Origin?
by: Ni, Kangqi, et al.
Published: (2025)
by: Ni, Kangqi, et al.
Published: (2025)
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
by: Ren, Yiming, et al.
Published: (2026)
by: Ren, Yiming, et al.
Published: (2026)
S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
by: Dai, Muzhi, et al.
Published: (2025)
by: Dai, Muzhi, et al.
Published: (2025)
Transition Models: Rethinking the Generative Learning Objective
by: Wang, Zidong, et al.
Published: (2025)
by: Wang, Zidong, et al.
Published: (2025)
ExGRPO: Learning to Reason from Experience
by: Zhan, Runzhe, et al.
Published: (2025)
by: Zhan, Runzhe, et al.
Published: (2025)
A Survey on Incomplete Multi-label Learning: Recent Advances and Future Trends
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents
by: Wu, Xiongbin, et al.
Published: (2026)
by: Wu, Xiongbin, et al.
Published: (2026)
Hail to the Thief: Exploring Attacks and Defenses in Decentralised GRPO
by: Blagoev, Nikolay, et al.
Published: (2025)
by: Blagoev, Nikolay, et al.
Published: (2025)
Similar Items
-
LithoGRPO: Fast Inverse Lithography via GRPO Reinforced Flow Matching
by: Lai, Yao, et al.
Published: (2026) -
Expand and Prune: Maximizing Trajectory Diversity for Effective GRPO in Generative Models
by: Ge, Shiran, et al.
Published: (2025) -
Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening
by: He, Andre, et al.
Published: (2025) -
Execution-Grounded Credit Assignment for GRPO in Code Generation
by: Kumar, Abhijit, et al.
Published: (2026) -
GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning
by: Xu, Yanchen, et al.
Published: (2025)