Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yibin, Li, Zhimin, Zang, Yuhang, Zhou, Yujie, Bu, Jiazi, Wang, Chunyu, Lu, Qinglin, Jin, Cheng, Wang, Jiaqi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fine-Grained GRPO for Precise Preference Alignment in Flow Models
by: Zhou, Yujie, et al.
Published: (2025)
by: Zhou, Yujie, et al.
Published: (2025)
UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
Pave-GRPO: Beyond Instantaneous Guidance through Principled Average Velocity Decomposition
by: Ling, Pengyang, et al.
Published: (2026)
by: Ling, Pengyang, et al.
Published: (2026)
Unified Personalized Reward Model for Vision Generation
by: Wang, Yibin, et al.
Published: (2026)
by: Wang, Yibin, et al.
Published: (2026)
From Sparse to Dense: Multi-View GRPO for Flow Models via Augmented Condition Space
by: Bu, Jiazi, et al.
Published: (2026)
by: Bu, Jiazi, et al.
Published: (2026)
DiCache: Let Diffusion Model Determine Its Own Cache
by: Bu, Jiazi, et al.
Published: (2025)
by: Bu, Jiazi, et al.
Published: (2025)
Unified Reward Model for Multimodal Understanding and Generation
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
$λ$-GRPO: Unifying the GRPO Frameworks with Learnable Token Preferences
by: Wang, Yining, et al.
Published: (2025)
by: Wang, Yining, et al.
Published: (2025)
BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
by: Li, Yuming, et al.
Published: (2025)
by: Li, Yuming, et al.
Published: (2025)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
by: Yari, Amir Hossein, et al.
Published: (2026)
by: Yari, Amir Hossein, et al.
Published: (2026)
PrefPO: Pairwise Preference Prompt Optimization
by: Singhal, Rahul, et al.
Published: (2026)
by: Singhal, Rahul, et al.
Published: (2026)
Multi-Reward GRPO for Stable and Prosodic Single-Codebook TTS LLMs at Scale
by: Zhong, Yicheng, et al.
Published: (2025)
by: Zhong, Yicheng, et al.
Published: (2025)
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance
by: Bu, Jiazi, et al.
Published: (2025)
by: Bu, Jiazi, et al.
Published: (2025)
GRPO and Reflection Reward for Mathematical Reasoning in Large Language Models
by: Wang, Zhijie
Published: (2026)
by: Wang, Zhijie
Published: (2026)
LithoGRPO: Fast Inverse Lithography via GRPO Reinforced Flow Matching
by: Lai, Yao, et al.
Published: (2026)
by: Lai, Yao, et al.
Published: (2026)
MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting
by: Wei, Kangda, et al.
Published: (2026)
by: Wei, Kangda, et al.
Published: (2026)
DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO
by: Liu, Henglin, et al.
Published: (2025)
by: Liu, Henglin, et al.
Published: (2025)
Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
by: Luo, Yifu, et al.
Published: (2025)
by: Luo, Yifu, et al.
Published: (2025)
Multi-GRPO: Multi-Group Advantage Estimation for Text-to-Image Generation with Tree-Based Trajectories and Multiple Rewards
by: Lyu, Qiang, et al.
Published: (2025)
by: Lyu, Qiang, et al.
Published: (2025)
ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way
by: Bu, Jiazi, et al.
Published: (2024)
by: Bu, Jiazi, et al.
Published: (2024)
Bridging the Semantic Gap: Contrastive Rewards for Multilingual Text-to-SQL with GRPO
by: Kattamuri, Ashish, et al.
Published: (2025)
by: Kattamuri, Ashish, et al.
Published: (2025)
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
by: Li, Junzhe, et al.
Published: (2025)
by: Li, Junzhe, et al.
Published: (2025)
Improving Generalization in Intent Detection: GRPO with Reward-Based Curriculum Sampling
by: Feng, Zihao, et al.
Published: (2025)
by: Feng, Zihao, et al.
Published: (2025)
MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation
by: Ma, Xiaoxiao, et al.
Published: (2026)
by: Ma, Xiaoxiao, et al.
Published: (2026)
IndoPref: A Multi-Domain Pairwise Preference Dataset for Indonesian
by: Wiyono, Vanessa Rebecca, et al.
Published: (2025)
by: Wiyono, Vanessa Rebecca, et al.
Published: (2025)
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
by: Chen, Minghan, et al.
Published: (2025)
by: Chen, Minghan, et al.
Published: (2025)
AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
by: Yuan, Shihao, et al.
Published: (2025)
by: Yuan, Shihao, et al.
Published: (2025)
UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models
by: Wang, Jiaqi, et al.
Published: (2026)
by: Wang, Jiaqi, et al.
Published: (2026)
GRPO is Secretly a Process Reward Model
by: Sullivan, Michael, et al.
Published: (2025)
by: Sullivan, Michael, et al.
Published: (2025)
STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
by: Ma, Xiaoxiao, et al.
Published: (2025)
by: Ma, Xiaoxiao, et al.
Published: (2025)
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
by: Tan, Hongze, et al.
Published: (2025)
by: Tan, Hongze, et al.
Published: (2025)
DanceGRPO: Unleashing GRPO on Visual Generation
by: Xue, Zeyue, et al.
Published: (2025)
by: Xue, Zeyue, et al.
Published: (2025)
GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning
by: Xu, Yanchen, et al.
Published: (2025)
by: Xu, Yanchen, et al.
Published: (2025)
How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning
by: Tian, Minghao, et al.
Published: (2026)
by: Tian, Minghao, et al.
Published: (2026)
PrefMMT: Modeling Human Preferences in Preference-based Reinforcement Learning with Multimodal Transformers
by: Zhao, Dezhong, et al.
Published: (2024)
by: Zhao, Dezhong, et al.
Published: (2024)
DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning
by: Chen, Xiwen, et al.
Published: (2025)
by: Chen, Xiwen, et al.
Published: (2025)
Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models
by: Wang, Austin, et al.
Published: (2026)
by: Wang, Austin, et al.
Published: (2026)
Similar Items
-
Fine-Grained GRPO for Precise Preference Alignment in Flow Models
by: Zhou, Yujie, et al.
Published: (2025) -
UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
by: Wang, Yibin, et al.
Published: (2025) -
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
by: Wang, Yibin, et al.
Published: (2025) -
Pave-GRPO: Beyond Instantaneous Guidance through Principled Average Velocity Decomposition
by: Ling, Pengyang, et al.
Published: (2026) -
Unified Personalized Reward Model for Vision Generation
by: Wang, Yibin, et al.
Published: (2026)