Multi-GRPO: Multi-Group Advantage Estimation for Text-to-Image Generation with Tree-Based Trajectories and Multiple Rewards
Fuente:
arXiv
Guardado en:
| Autores principales: | Lyu, Qiang, Chen, Zicong, Wang, Chongxiao, Shi, Haolin, Gao, Shibo, Piao, Ran, Zeng, Youwei, Si, Jianlou, Ding, Fei, Li, Jing, Lau, Chun Pong, Wang, Weiqiang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models
por: Ding, Fei, et al.
Publicado: (2025)
por: Ding, Fei, et al.
Publicado: (2025)
Combining Clusters for the Approximate Randomization Test
por: Lau, Chun Pong
Publicado: (2025)
por: Lau, Chun Pong
Publicado: (2025)
Sensitivity Analysis for Dynamic Discrete Choice Models
por: Lau, Chun Pong
Publicado: (2024)
por: Lau, Chun Pong
Publicado: (2024)
RC-GRPO: Reward-Conditioned Group Relative Policy Optimization for Multi-Turn Tool Calling Agents
por: Zhong, Haitian, et al.
Publicado: (2026)
por: Zhong, Haitian, et al.
Publicado: (2026)
Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
por: Wang, Yibin, et al.
Publicado: (2025)
por: Wang, Yibin, et al.
Publicado: (2025)
MMAD-Purify: A Precision-Optimized Framework for Efficient and Scalable Multi-Modal Attacks
por: Liu, Xinxin, et al.
Publicado: (2024)
por: Liu, Xinxin, et al.
Publicado: (2024)
Cluster-robust inference with a single treated cluster using the t-test
por: Lau, Chun Pong, et al.
Publicado: (2025)
por: Lau, Chun Pong, et al.
Publicado: (2025)
Multi-Reward GRPO for Stable and Prosodic Single-Codebook TTS LLMs at Scale
por: Zhong, Yicheng, et al.
Publicado: (2025)
por: Zhong, Yicheng, et al.
Publicado: (2025)
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
por: Ichihara, Yuki, et al.
Publicado: (2025)
por: Ichihara, Yuki, et al.
Publicado: (2025)
ShapE-GRPO: Shapley-Enhanced Reward Allocation for Multi-Candidate LLM Training
por: Ai, Rui, et al.
Publicado: (2026)
por: Ai, Rui, et al.
Publicado: (2026)
FedGRPO: Privately Optimizing Foundation Models with Group-Relative Rewards from Domain Client
por: Zhu, Gongxi, et al.
Publicado: (2026)
por: Zhu, Gongxi, et al.
Publicado: (2026)
Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO
por: Wang, Hongcheng, et al.
Publicado: (2025)
por: Wang, Hongcheng, et al.
Publicado: (2025)
An In-Depth Survey on Virtualization Technologies in 6G Integrated Terrestrial and Non-Terrestrial Networks
por: Ammar, Sahar, et al.
Publicado: (2023)
por: Ammar, Sahar, et al.
Publicado: (2023)
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
por: Zhang, Xingjian, et al.
Publicado: (2025)
por: Zhang, Xingjian, et al.
Publicado: (2025)
Trajectory Attention for Fine-grained Video Motion Control
por: Xiao, Zeqi, et al.
Publicado: (2024)
por: Xiao, Zeqi, et al.
Publicado: (2024)
PGID: Progressive Guided Inversion and Denoising for Robust Watermark Detection
por: Duong, Minh Quoc, et al.
Publicado: (2026)
por: Duong, Minh Quoc, et al.
Publicado: (2026)
Bridging the Semantic Gap: Contrastive Rewards for Multilingual Text-to-SQL with GRPO
por: Kattamuri, Ashish, et al.
Publicado: (2025)
por: Kattamuri, Ashish, et al.
Publicado: (2025)
TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
por: Ding, Zheng, et al.
Publicado: (2025)
por: Ding, Zheng, et al.
Publicado: (2025)
GRPO and Reflection Reward for Mathematical Reasoning in Large Language Models
por: Wang, Zhijie
Publicado: (2026)
por: Wang, Zhijie
Publicado: (2026)
Blockwise Advantage Estimation for Multi-Objective RL with Verifiable Rewards
por: Pavlenko, Kirill, et al.
Publicado: (2026)
por: Pavlenko, Kirill, et al.
Publicado: (2026)
Sebica: Lightweight Spatial and Efficient Bidirectional Channel Attention Super Resolution Network
por: Liu, Chongxiao
Publicado: (2024)
por: Liu, Chongxiao
Publicado: (2024)
My Face Is Mine, Not Yours: Facial Protection Against Diffusion Model Face Swapping
por: Yam, Hon Ming, et al.
Publicado: (2025)
por: Yam, Hon Ming, et al.
Publicado: (2025)
Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting
por: Wang, Cheng, et al.
Publicado: (2026)
por: Wang, Cheng, et al.
Publicado: (2026)
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
por: Zhao, Yi, et al.
Publicado: (2025)
por: Zhao, Yi, et al.
Publicado: (2025)
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
por: Xu, Yuanda, et al.
Publicado: (2026)
por: Xu, Yuanda, et al.
Publicado: (2026)
OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models
por: Zhang, Liyu, et al.
Publicado: (2026)
por: Zhang, Liyu, et al.
Publicado: (2026)
Expand and Prune: Maximizing Trajectory Diversity for Effective GRPO in Generative Models
por: Ge, Shiran, et al.
Publicado: (2025)
por: Ge, Shiran, et al.
Publicado: (2025)
TIGFlow-GRPO: Trajectory Forecasting via Interaction-Aware Flow Matching and Reward-Guided Optimization
por: Jing, Xuepeng, et al.
Publicado: (2026)
por: Jing, Xuepeng, et al.
Publicado: (2026)
MotionGRPO: Overcoming Low Intra-Group Diversity in GRPO-Based Egocentric Motion Recovery
por: Yao, Nanjie, et al.
Publicado: (2026)
por: Yao, Nanjie, et al.
Publicado: (2026)
The Dual (Energy) Equivalence Principle, Use Cubic Square and Fiber Curvature to Develop a Next-Generation Quantum Glass Fiber Chip
por: Lie Chun Pong
Publicado: (2025)
por: Lie Chun Pong
Publicado: (2025)
An Innovative Chip Circle Road Path (CP) Design
por: Lie Chun Pong
Publicado: (2025)
por: Lie Chun Pong
Publicado: (2025)
Improving Generalization in Intent Detection: GRPO with Reward-Based Curriculum Sampling
por: Feng, Zihao, et al.
Publicado: (2025)
por: Feng, Zihao, et al.
Publicado: (2025)
Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage
por: Jin, Weiqiang, et al.
Publicado: (2025)
por: Jin, Weiqiang, et al.
Publicado: (2025)
MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting
por: Wei, Kangda, et al.
Publicado: (2026)
por: Wei, Kangda, et al.
Publicado: (2026)
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
por: Tan, Hongze, et al.
Publicado: (2025)
por: Tan, Hongze, et al.
Publicado: (2025)
GRPO is Secretly a Process Reward Model
por: Sullivan, Michael, et al.
Publicado: (2025)
por: Sullivan, Michael, et al.
Publicado: (2025)
EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO
por: Guan, Wei, et al.
Publicado: (2025)
por: Guan, Wei, et al.
Publicado: (2025)
Style4D-Bench: A Benchmark Suite for 4D Stylization
por: Chen, Beiqi, et al.
Publicado: (2025)
por: Chen, Beiqi, et al.
Publicado: (2025)
RealDPO: Real or Not Real, that is the Preference
por: Cheng, Guo, et al.
Publicado: (2025)
por: Cheng, Guo, et al.
Publicado: (2025)
Graph-GRPO: Stabilizing Multi-Agent Topology Learning via Group Relative Policy Optimization
por: Cang, Yueyang, et al.
Publicado: (2026)
por: Cang, Yueyang, et al.
Publicado: (2026)
Ejemplares similares
-
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models
por: Ding, Fei, et al.
Publicado: (2025) -
Combining Clusters for the Approximate Randomization Test
por: Lau, Chun Pong
Publicado: (2025) -
Sensitivity Analysis for Dynamic Discrete Choice Models
por: Lau, Chun Pong
Publicado: (2024) -
RC-GRPO: Reward-Conditioned Group Relative Policy Optimization for Multi-Turn Tool Calling Agents
por: Zhong, Haitian, et al.
Publicado: (2026) -
Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
por: Wang, Yibin, et al.
Publicado: (2025)