MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Xiaoxiao, Lei, Jiachen, Ren, Tianfei, Huang, Jie, Fu, Siming, Hao, Aiming, Wu, Jiahong, Chu, Xiangxiang, Zhao, Feng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
by: Ma, Xiaoxiao, et al.
Published: (2025)
by: Ma, Xiaoxiao, et al.
Published: (2025)
Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation
by: Zhao, Chenxi, et al.
Published: (2026)
by: Zhao, Chenxi, et al.
Published: (2026)
TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
by: He, Xiaoxuan, et al.
Published: (2025)
by: He, Xiaoxuan, et al.
Published: (2025)
DanceGRPO: Unleashing GRPO on Visual Generation
by: Xue, Zeyue, et al.
Published: (2025)
by: Xue, Zeyue, et al.
Published: (2025)
Ranking-aware Reinforcement Learning for Ordinal Ranking
by: Hao, Aiming, et al.
Published: (2026)
by: Hao, Aiming, et al.
Published: (2026)
DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO
by: Liu, Henglin, et al.
Published: (2025)
by: Liu, Henglin, et al.
Published: (2025)
AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
by: Yuan, Shihao, et al.
Published: (2025)
by: Yuan, Shihao, et al.
Published: (2025)
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
MotionGRPO: Overcoming Low Intra-Group Diversity in GRPO-Based Egocentric Motion Recovery
by: Yao, Nanjie, et al.
Published: (2026)
by: Yao, Nanjie, et al.
Published: (2026)
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
by: Li, Junzhe, et al.
Published: (2025)
by: Li, Junzhe, et al.
Published: (2025)
$λ$-GRPO: Unifying the GRPO Frameworks with Learnable Token Preferences
by: Wang, Yining, et al.
Published: (2025)
by: Wang, Yining, et al.
Published: (2025)
OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models
by: Zhang, Liyu, et al.
Published: (2026)
by: Zhang, Liyu, et al.
Published: (2026)
BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
by: Li, Yuming, et al.
Published: (2025)
by: Li, Yuming, et al.
Published: (2025)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
by: Yari, Amir Hossein, et al.
Published: (2026)
by: Yari, Amir Hossein, et al.
Published: (2026)
MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting
by: Wei, Kangda, et al.
Published: (2026)
by: Wei, Kangda, et al.
Published: (2026)
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
DEVIS-GRPO: Unleashing GRPO on Dynamic Extreme View Synthesis
by: Zuo, Yi, et al.
Published: (2026)
by: Zuo, Yi, et al.
Published: (2026)
Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation
by: Dai, Yanqi, et al.
Published: (2026)
by: Dai, Yanqi, et al.
Published: (2026)
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
by: Huang, Zhe, et al.
Published: (2025)
by: Huang, Zhe, et al.
Published: (2025)
Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO
by: Wang, Hongcheng, et al.
Published: (2025)
by: Wang, Hongcheng, et al.
Published: (2025)
DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning
by: Chen, Xiwen, et al.
Published: (2025)
by: Chen, Xiwen, et al.
Published: (2025)
Boosting MLLM Reasoning with Text-Debiased Hint-GRPO
by: Huang, Qihan, et al.
Published: (2025)
by: Huang, Qihan, et al.
Published: (2025)
It Takes Two: Your GRPO Is Secretly DPO
by: Wu, Yihong, et al.
Published: (2025)
by: Wu, Yihong, et al.
Published: (2025)
AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
by: Huang, Runhui, et al.
Published: (2026)
by: Huang, Runhui, et al.
Published: (2026)
UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation
by: Liu, Jie, et al.
Published: (2026)
by: Liu, Jie, et al.
Published: (2026)
LithoGRPO: Fast Inverse Lithography via GRPO Reinforced Flow Matching
by: Lai, Yao, et al.
Published: (2026)
by: Lai, Yao, et al.
Published: (2026)
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
by: Chen, Minghan, et al.
Published: (2025)
by: Chen, Minghan, et al.
Published: (2025)
Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization
by: He, Xiaoxuan, et al.
Published: (2026)
by: He, Xiaoxuan, et al.
Published: (2026)
GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning
by: Xu, Yanchen, et al.
Published: (2025)
by: Xu, Yanchen, et al.
Published: (2025)
How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning
by: Tian, Minghao, et al.
Published: (2026)
by: Tian, Minghao, et al.
Published: (2026)
TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
by: Ding, Zheng, et al.
Published: (2025)
by: Ding, Zheng, et al.
Published: (2025)
VMBench: A Benchmark for Perception-Aligned Video Motion Generation
by: Ling, Xinran, et al.
Published: (2025)
by: Ling, Xinran, et al.
Published: (2025)
Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation
by: Mao, Fangyuan, et al.
Published: (2025)
by: Mao, Fangyuan, et al.
Published: (2025)
Improving Visual Representation Alignment Generation with GRPO
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
Advances in GRPO for Generation Models: A Survey
by: Liu, Zexiang, et al.
Published: (2026)
by: Liu, Zexiang, et al.
Published: (2026)
Improving LLM-Generated Code Quality with GRPO
by: Robeyns, Maxime, et al.
Published: (2025)
by: Robeyns, Maxime, et al.
Published: (2025)
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
by: Ren, Yiming, et al.
Published: (2026)
by: Ren, Yiming, et al.
Published: (2026)
What is the Alignment Objective of GRPO?
by: Vojnovic, Milan, et al.
Published: (2025)
by: Vojnovic, Milan, et al.
Published: (2025)
There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
by: Lei, Jiachen, et al.
Published: (2025)
by: Lei, Jiachen, et al.
Published: (2025)
Similar Items
-
STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
by: Ma, Xiaoxiao, et al.
Published: (2025) -
Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation
by: Zhao, Chenxi, et al.
Published: (2026) -
TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
by: He, Xiaoxuan, et al.
Published: (2025) -
DanceGRPO: Unleashing GRPO on Visual Generation
by: Xue, Zeyue, et al.
Published: (2025) -
Ranking-aware Reinforcement Learning for Ordinal Ranking
by: Hao, Aiming, et al.
Published: (2026)