Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Hongcheng, Huang, Yinuo, Wang, Sukai, Ren, Guanghui, Dong, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting
by: Wei, Kangda, et al.
Published: (2026)
by: Wei, Kangda, et al.
Published: (2026)
Chasing Progress, Not Perfection: Revisiting Strategies for End-to-End LLM Plan Generation
by: Huang, Sukai, et al.
Published: (2024)
by: Huang, Sukai, et al.
Published: (2024)
Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting
by: Wang, Cheng, et al.
Published: (2026)
by: Wang, Cheng, et al.
Published: (2026)
Evaluating GRPO and DPO for Faithful Chain-of-Thought Reasoning in LLMs
by: Mohammadi, Hadi, et al.
Published: (2025)
by: Mohammadi, Hadi, et al.
Published: (2025)
DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning
by: Chen, Xiwen, et al.
Published: (2025)
by: Chen, Xiwen, et al.
Published: (2025)
Negative Advantages Is a Double-Edged Sword: Calibrating advantages in GRPO for Search Agents
by: Wu, Jiayi, et al.
Published: (2026)
by: Wu, Jiayi, et al.
Published: (2026)
ALIGN: Word Association Learning for Cultural Alignment in Large Language Models
by: Liu, Chunhua, et al.
Published: (2025)
by: Liu, Chunhua, et al.
Published: (2025)
$λ$-GRPO: Unifying the GRPO Frameworks with Learnable Token Preferences
by: Wang, Yining, et al.
Published: (2025)
by: Wang, Yining, et al.
Published: (2025)
Can GRPO Boost Complex Multimodal Table Understanding?
by: Kang, Xiaoqiang, et al.
Published: (2025)
by: Kang, Xiaoqiang, et al.
Published: (2025)
Capturing Classic Authorial Style in Long-Form Story Generation with GRPO Fine-Tuning
by: Liu, Jinlong, et al.
Published: (2025)
by: Liu, Jinlong, et al.
Published: (2025)
Why Can Large Language Models Generate Correct Chain-of-Thoughts?
by: Tutunov, Rasul, et al.
Published: (2023)
by: Tutunov, Rasul, et al.
Published: (2023)
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
Reasoning Topology Matters: Network-of-Thought for Complex Reasoning Tasks
by: Huang, Fan
Published: (2026)
by: Huang, Fan
Published: (2026)
WebCoT: Enhancing Web Agent Reasoning by Reconstructing Chain-of-Thought in Reflection, Branching, and Rollback
by: Hu, Minda, et al.
Published: (2025)
by: Hu, Minda, et al.
Published: (2025)
Why Slop Matters
by: Kommers, Cody, et al.
Published: (2025)
by: Kommers, Cody, et al.
Published: (2025)
Conditional Advantage Estimation for Reinforcement Learning in Large Reasoning Models
by: Chen, Guanxu, et al.
Published: (2025)
by: Chen, Guanxu, et al.
Published: (2025)
Improving Generalization in Intent Detection: GRPO with Reward-Based Curriculum Sampling
by: Feng, Zihao, et al.
Published: (2025)
by: Feng, Zihao, et al.
Published: (2025)
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
by: Tan, Hongze, et al.
Published: (2025)
by: Tan, Hongze, et al.
Published: (2025)
InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment
by: Long, Yuxing, et al.
Published: (2024)
by: Long, Yuxing, et al.
Published: (2024)
M2K-VDG: Model-Adaptive Multimodal Knowledge Anchor Enhanced Video-grounded Dialogue Generation
by: Liu, Hongcheng, et al.
Published: (2024)
by: Liu, Hongcheng, et al.
Published: (2024)
It Takes Two: Your GRPO Is Secretly DPO
by: Wu, Yihong, et al.
Published: (2025)
by: Wu, Yihong, et al.
Published: (2025)
BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
by: Li, Yuming, et al.
Published: (2025)
by: Li, Yuming, et al.
Published: (2025)
RAIDEN-R1: Improving Role-awareness of LLMs via GRPO with Verifiable Reward
by: Wang, Zongsheng, et al.
Published: (2025)
by: Wang, Zongsheng, et al.
Published: (2025)
Not All Correct Answers Are Equal: Why Your Distillation Source Matters
by: Tian, Xiaoyu, et al.
Published: (2025)
by: Tian, Xiaoyu, et al.
Published: (2025)
RATT: A Thought Structure for Coherent and Correct LLM Reasoning
by: Zhang, Jinghan, et al.
Published: (2024)
by: Zhang, Jinghan, et al.
Published: (2024)
Why Chain of Thought Fails in Clinical Text Understanding
by: Wu, Jiageng, et al.
Published: (2025)
by: Wu, Jiageng, et al.
Published: (2025)
Thought Branches: Interpreting LLM Reasoning Requires Resampling
by: Macar, Uzay, et al.
Published: (2025)
by: Macar, Uzay, et al.
Published: (2025)
GIFT: Group-Relative Implicit Fine-Tuning Integrates GRPO with DPO and UNA
by: Wang, Zhichao
Published: (2025)
by: Wang, Zhichao
Published: (2025)
Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
by: Cheng, Zihui, et al.
Published: (2025)
by: Cheng, Zihui, et al.
Published: (2025)
KTAE: A Model-Free Algorithm to Key-Tokens Advantage Estimation in Mathematical Reasoning
by: Sun, Wei, et al.
Published: (2025)
by: Sun, Wei, et al.
Published: (2025)
Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning
by: Li, Ziheng, et al.
Published: (2026)
by: Li, Ziheng, et al.
Published: (2026)
Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervision
by: Zhu, Dawei, et al.
Published: (2025)
by: Zhu, Dawei, et al.
Published: (2025)
How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning
by: Tian, Minghao, et al.
Published: (2026)
by: Tian, Minghao, et al.
Published: (2026)
Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning
by: Wang, Li, et al.
Published: (2026)
by: Wang, Li, et al.
Published: (2026)
Pretraining with Token-Level Adaptive Latent Chain-of-Thought
by: Zeng, Boyi, et al.
Published: (2026)
by: Zeng, Boyi, et al.
Published: (2026)
Structured Style-Rewrite with Chain-of-Thought Planning for Low-Resource Character Dialogue
by: Zhu, Chanhui
Published: (2026)
by: Zhu, Chanhui
Published: (2026)
Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning
by: Paul, Debjit, et al.
Published: (2024)
by: Paul, Debjit, et al.
Published: (2024)
xCoT: Cross-lingual Instruction Tuning for Cross-lingual Chain-of-Thought Reasoning
by: Chai, Linzheng, et al.
Published: (2024)
by: Chai, Linzheng, et al.
Published: (2024)
TL-GRPO: Turn-Level RL for Reasoning-Guided Iterative Optimization
by: Li, Peiji, et al.
Published: (2026)
by: Li, Peiji, et al.
Published: (2026)
ThoughtProbe: Classifier-Guided Thought Space Exploration Leveraging LLM Intrinsic Reasoning
by: Wang, Zijian, et al.
Published: (2025)
by: Wang, Zijian, et al.
Published: (2025)
Similar Items
-
MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting
by: Wei, Kangda, et al.
Published: (2026) -
Chasing Progress, Not Perfection: Revisiting Strategies for End-to-End LLM Plan Generation
by: Huang, Sukai, et al.
Published: (2024) -
Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting
by: Wang, Cheng, et al.
Published: (2026) -
Evaluating GRPO and DPO for Faithful Chain-of-Thought Reasoning in LLMs
by: Mohammadi, Hadi, et al.
Published: (2025) -
DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning
by: Chen, Xiwen, et al.
Published: (2025)