Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO
Fuente:
arXiv
Saved in:
| Main Authors: | Tong, Yunze, Liu, Mushui, Zhao, Canyu, He, Wanggui, Zhang, Shiyi, Zhang, Hongwei, Zhang, Peng, Liu, Jinlong, Huang, Ju, Wang, Jiamang, Jiang, Hao, Huang, Pipei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PromptEcho: Annotation-Free Reward from Vision-Language Models for Text-to-Image Reinforcement Learning
by: Liu, Jinlong, et al.
Published: (2026)
by: Liu, Jinlong, et al.
Published: (2026)
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
by: Zhang, Peng, et al.
Published: (2026)
by: Zhang, Peng, et al.
Published: (2026)
Boosting MLLM Reasoning with Text-Debiased Hint-GRPO
by: Huang, Qihan, et al.
Published: (2025)
by: Huang, Qihan, et al.
Published: (2025)
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation
by: Yang, Zixuan, et al.
Published: (2026)
by: Yang, Zixuan, et al.
Published: (2026)
MARBLE: Multi-Aspect Reward Balance for Diffusion RL
by: Zhao, Canyu, et al.
Published: (2026)
by: Zhao, Canyu, et al.
Published: (2026)
DenseGRPO: From Sparse to Dense Reward for Flow Matching Model Alignment
by: Deng, Haoyou, et al.
Published: (2026)
by: Deng, Haoyou, et al.
Published: (2026)
E-GRPO: High Entropy Steps Drive Effective Reinforcement Learning for Flow Models
by: Zhang, Shengjun, et al.
Published: (2026)
by: Zhang, Shengjun, et al.
Published: (2026)
VeriGate: Verifier-Gated Step-Level Supervision for GRPO
by: Agrawal, Aakriti, et al.
Published: (2026)
by: Agrawal, Aakriti, et al.
Published: (2026)
MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting
by: Wei, Kangda, et al.
Published: (2026)
by: Wei, Kangda, et al.
Published: (2026)
Capturing Classic Authorial Style in Long-Form Story Generation with GRPO Fine-Tuning
by: Liu, Jinlong, et al.
Published: (2025)
by: Liu, Jinlong, et al.
Published: (2025)
Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture
by: Zhang, Longxiang, et al.
Published: (2026)
by: Zhang, Longxiang, et al.
Published: (2026)
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
Lag-Relative Sparse Attention In Long Context Training
by: Liang, Manlai, et al.
Published: (2025)
by: Liang, Manlai, et al.
Published: (2025)
Towards Enhanced Image Generation Via Multi-modal Chain of Thought in Unified Generative Models
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
LogicReward: Incentivizing LLM Reasoning via Step-Wise Logical Supervision
by: Xu, Jundong, et al.
Published: (2025)
by: Xu, Jundong, et al.
Published: (2025)
Policy Learning for Balancing Short-Term and Long-Term Rewards
by: Wu, Peng, et al.
Published: (2024)
by: Wu, Peng, et al.
Published: (2024)
TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
by: He, Xiaoxuan, et al.
Published: (2025)
by: He, Xiaoxuan, et al.
Published: (2025)
Flow-GRPO: Training Flow Matching Models via Online RL
by: Liu, Jie, et al.
Published: (2025)
by: Liu, Jie, et al.
Published: (2025)
InfoFlow KV: Information-Flow-Aware KV Recomputation for Long Context
by: Teng, Xin, et al.
Published: (2026)
by: Teng, Xin, et al.
Published: (2026)
Step-GRPO: Internalizing Dynamic Early Exit for Efficient Reasoning
by: Chen, Benteng, et al.
Published: (2026)
by: Chen, Benteng, et al.
Published: (2026)
Agentic Reinforcement Learning with Implicit Step Rewards
by: Liu, Xiaoqian, et al.
Published: (2025)
by: Liu, Xiaoqian, et al.
Published: (2025)
LithoGRPO: Fast Inverse Lithography via GRPO Reinforced Flow Matching
by: Lai, Yao, et al.
Published: (2026)
by: Lai, Yao, et al.
Published: (2026)
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
How to Alleviate Catastrophic Forgetting in LLMs Finetuning? Hierarchical Layer-Wise and Element-Wise Regularization
by: Song, Shezheng, et al.
Published: (2025)
by: Song, Shezheng, et al.
Published: (2025)
MotionRAG-Diff: A Retrieval-Augmented Diffusion Framework for Long-Term Music-to-Dance Generation
by: Huang, Mingyang, et al.
Published: (2025)
by: Huang, Mingyang, et al.
Published: (2025)
Diffusion-APO: Trajectory-Aware Direct Preference Alignment for Video Diffusion Transformers
by: Zhu, Jingyuan, et al.
Published: (2026)
by: Zhu, Jingyuan, et al.
Published: (2026)
Smart-GRPO: Smartly Sampling Noise for Efficient RL of Flow-Matching Models
by: Yu, Benjamin, et al.
Published: (2025)
by: Yu, Benjamin, et al.
Published: (2025)
AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
by: Huang, Runhui, et al.
Published: (2026)
by: Huang, Runhui, et al.
Published: (2026)
PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation
by: Huang, Qihan, et al.
Published: (2024)
by: Huang, Qihan, et al.
Published: (2024)
OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models
by: Zhang, Liyu, et al.
Published: (2026)
by: Zhang, Liyu, et al.
Published: (2026)
Improving Generalization in Intent Detection: GRPO with Reward-Based Curriculum Sampling
by: Feng, Zihao, et al.
Published: (2025)
by: Feng, Zihao, et al.
Published: (2025)
ColorFlow: Retrieval-Augmented Image Sequence Colorization
by: Zhuang, Junhao, et al.
Published: (2024)
by: Zhuang, Junhao, et al.
Published: (2024)
CustomVideoX: 3D Reference Attention Driven Dynamic Adaptation for Zero-Shot Customized Video Diffusion Transformers
by: She, D., et al.
Published: (2025)
by: She, D., et al.
Published: (2025)
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
by: Li, Junzhe, et al.
Published: (2025)
by: Li, Junzhe, et al.
Published: (2025)
DanceGRPO: Unleashing GRPO on Visual Generation
by: Xue, Zeyue, et al.
Published: (2025)
by: Xue, Zeyue, et al.
Published: (2025)
DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO
by: Liu, Henglin, et al.
Published: (2025)
by: Liu, Henglin, et al.
Published: (2025)
Improving Long‐Term Glucose Prediction Accuracy with Uncertainty‐Estimated ProbSparse‐Transformer
by: Wei Huang, et al.
Published: (2025)
by: Wei Huang, et al.
Published: (2025)
Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO
by: Zhu, Muzhi, et al.
Published: (2025)
by: Zhu, Muzhi, et al.
Published: (2025)
From Sparse to Dense: Multi-View GRPO for Flow Models via Augmented Condition Space
by: Bu, Jiazi, et al.
Published: (2026)
by: Bu, Jiazi, et al.
Published: (2026)
Similar Items
-
PromptEcho: Annotation-Free Reward from Vision-Language Models for Text-to-Image Reinforcement Learning
by: Liu, Jinlong, et al.
Published: (2026) -
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
by: Zhang, Peng, et al.
Published: (2026) -
Boosting MLLM Reasoning with Text-Debiased Hint-GRPO
by: Huang, Qihan, et al.
Published: (2025) -
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation
by: Yang, Zixuan, et al.
Published: (2026) -
MARBLE: Multi-Aspect Reward Balance for Diffusion RL
by: Zhao, Canyu, et al.
Published: (2026)