Multi-GRPO: Multi-Group Advantage Estimation for Text-to-Image Generation with Tree-Based Trajectories and Multiple Rewards
Fuente:
arXiv
Saved in:
| Main Authors: | Lyu, Qiang, Chen, Zicong, Wang, Chongxiao, Shi, Haolin, Gao, Shibo, Piao, Ran, Zeng, Youwei, Si, Jianlou, Ding, Fei, Li, Jing, Lau, Chun Pong, Wang, Weiqiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models
by: Ding, Fei, et al.
Published: (2025)
by: Ding, Fei, et al.
Published: (2025)
Combining Clusters for the Approximate Randomization Test
by: Lau, Chun Pong
Published: (2025)
by: Lau, Chun Pong
Published: (2025)
Sensitivity Analysis for Dynamic Discrete Choice Models
by: Lau, Chun Pong
Published: (2024)
by: Lau, Chun Pong
Published: (2024)
RC-GRPO: Reward-Conditioned Group Relative Policy Optimization for Multi-Turn Tool Calling Agents
by: Zhong, Haitian, et al.
Published: (2026)
by: Zhong, Haitian, et al.
Published: (2026)
Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
MMAD-Purify: A Precision-Optimized Framework for Efficient and Scalable Multi-Modal Attacks
by: Liu, Xinxin, et al.
Published: (2024)
by: Liu, Xinxin, et al.
Published: (2024)
Cluster-robust inference with a single treated cluster using the t-test
by: Lau, Chun Pong, et al.
Published: (2025)
by: Lau, Chun Pong, et al.
Published: (2025)
Multi-Reward GRPO for Stable and Prosodic Single-Codebook TTS LLMs at Scale
by: Zhong, Yicheng, et al.
Published: (2025)
by: Zhong, Yicheng, et al.
Published: (2025)
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
by: Ichihara, Yuki, et al.
Published: (2025)
by: Ichihara, Yuki, et al.
Published: (2025)
ShapE-GRPO: Shapley-Enhanced Reward Allocation for Multi-Candidate LLM Training
by: Ai, Rui, et al.
Published: (2026)
by: Ai, Rui, et al.
Published: (2026)
FedGRPO: Privately Optimizing Foundation Models with Group-Relative Rewards from Domain Client
by: Zhu, Gongxi, et al.
Published: (2026)
by: Zhu, Gongxi, et al.
Published: (2026)
Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO
by: Wang, Hongcheng, et al.
Published: (2025)
by: Wang, Hongcheng, et al.
Published: (2025)
An In-Depth Survey on Virtualization Technologies in 6G Integrated Terrestrial and Non-Terrestrial Networks
by: Ammar, Sahar, et al.
Published: (2023)
by: Ammar, Sahar, et al.
Published: (2023)
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
Trajectory Attention for Fine-grained Video Motion Control
by: Xiao, Zeqi, et al.
Published: (2024)
by: Xiao, Zeqi, et al.
Published: (2024)
PGID: Progressive Guided Inversion and Denoising for Robust Watermark Detection
by: Duong, Minh Quoc, et al.
Published: (2026)
by: Duong, Minh Quoc, et al.
Published: (2026)
Bridging the Semantic Gap: Contrastive Rewards for Multilingual Text-to-SQL with GRPO
by: Kattamuri, Ashish, et al.
Published: (2025)
by: Kattamuri, Ashish, et al.
Published: (2025)
TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
by: Ding, Zheng, et al.
Published: (2025)
by: Ding, Zheng, et al.
Published: (2025)
GRPO and Reflection Reward for Mathematical Reasoning in Large Language Models
by: Wang, Zhijie
Published: (2026)
by: Wang, Zhijie
Published: (2026)
Blockwise Advantage Estimation for Multi-Objective RL with Verifiable Rewards
by: Pavlenko, Kirill, et al.
Published: (2026)
by: Pavlenko, Kirill, et al.
Published: (2026)
Sebica: Lightweight Spatial and Efficient Bidirectional Channel Attention Super Resolution Network
by: Liu, Chongxiao
Published: (2024)
by: Liu, Chongxiao
Published: (2024)
My Face Is Mine, Not Yours: Facial Protection Against Diffusion Model Face Swapping
by: Yam, Hon Ming, et al.
Published: (2025)
by: Yam, Hon Ming, et al.
Published: (2025)
Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting
by: Wang, Cheng, et al.
Published: (2026)
by: Wang, Cheng, et al.
Published: (2026)
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models
by: Zhang, Liyu, et al.
Published: (2026)
by: Zhang, Liyu, et al.
Published: (2026)
Expand and Prune: Maximizing Trajectory Diversity for Effective GRPO in Generative Models
by: Ge, Shiran, et al.
Published: (2025)
by: Ge, Shiran, et al.
Published: (2025)
TIGFlow-GRPO: Trajectory Forecasting via Interaction-Aware Flow Matching and Reward-Guided Optimization
by: Jing, Xuepeng, et al.
Published: (2026)
by: Jing, Xuepeng, et al.
Published: (2026)
MotionGRPO: Overcoming Low Intra-Group Diversity in GRPO-Based Egocentric Motion Recovery
by: Yao, Nanjie, et al.
Published: (2026)
by: Yao, Nanjie, et al.
Published: (2026)
The Dual (Energy) Equivalence Principle, Use Cubic Square and Fiber Curvature to Develop a Next-Generation Quantum Glass Fiber Chip
by: Lie Chun Pong
Published: (2025)
by: Lie Chun Pong
Published: (2025)
An Innovative Chip Circle Road Path (CP) Design
by: Lie Chun Pong
Published: (2025)
by: Lie Chun Pong
Published: (2025)
Improving Generalization in Intent Detection: GRPO with Reward-Based Curriculum Sampling
by: Feng, Zihao, et al.
Published: (2025)
by: Feng, Zihao, et al.
Published: (2025)
Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage
by: Jin, Weiqiang, et al.
Published: (2025)
by: Jin, Weiqiang, et al.
Published: (2025)
MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting
by: Wei, Kangda, et al.
Published: (2026)
by: Wei, Kangda, et al.
Published: (2026)
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
by: Tan, Hongze, et al.
Published: (2025)
by: Tan, Hongze, et al.
Published: (2025)
GRPO is Secretly a Process Reward Model
by: Sullivan, Michael, et al.
Published: (2025)
by: Sullivan, Michael, et al.
Published: (2025)
EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO
by: Guan, Wei, et al.
Published: (2025)
by: Guan, Wei, et al.
Published: (2025)
Style4D-Bench: A Benchmark Suite for 4D Stylization
by: Chen, Beiqi, et al.
Published: (2025)
by: Chen, Beiqi, et al.
Published: (2025)
RealDPO: Real or Not Real, that is the Preference
by: Cheng, Guo, et al.
Published: (2025)
by: Cheng, Guo, et al.
Published: (2025)
Graph-GRPO: Stabilizing Multi-Agent Topology Learning via Group Relative Policy Optimization
by: Cang, Yueyang, et al.
Published: (2026)
by: Cang, Yueyang, et al.
Published: (2026)
Similar Items
-
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models
by: Ding, Fei, et al.
Published: (2025) -
Combining Clusters for the Approximate Randomization Test
by: Lau, Chun Pong
Published: (2025) -
Sensitivity Analysis for Dynamic Discrete Choice Models
by: Lau, Chun Pong
Published: (2024) -
RC-GRPO: Reward-Conditioned Group Relative Policy Optimization for Multi-Turn Tool Calling Agents
by: Zhong, Haitian, et al.
Published: (2026) -
Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
by: Wang, Yibin, et al.
Published: (2025)