Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Hongcheng, Huang, Yinuo, Wang, Sukai, Ren, Guanghui, Dong, Hao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910012546220032
author Wang, Hongcheng
Huang, Yinuo
Wang, Sukai
Ren, Guanghui
Dong, Hao
author_facet Wang, Hongcheng
Huang, Yinuo
Wang, Sukai
Ren, Guanghui
Dong, Hao
contents Group Relative Policy Optimization (GRPO) trains Chain-of-Thought reasoning with verifiable rewards, but estimating thought-level advantages without value functions often suffers from high variance. Although tree-style branching is used in practice to reduce the variance, it lacks a theoretical explanation of why it works and whether it is important or even potentially necessary. We study thought-level advantage estimation in GRPO from a variance perspective under a minimal tree-style setting where multiple answers are sampled for each thought. Using the multivariate delta method, we reveal an asymmetry in how different sampling dimensions affect variance. Increasing the number of sampled thoughts ($K$) leaves a strictly positive variance floor, whereas increasing the number of answers per thought ($M$) induces a monotonic decrease in variance, asymptotically decreasing it to zero. This implies that accurate thought-level advantage estimation is impossible through scaling thought sampling alone, making branching a potentially necessary mechanism rather than a heuristic. Experiments further provide empirical evidence for both the effectiveness and necessity of answer-level branching, demonstrating improved optimization stability, training efficiency, and final performance not only in math but also across a broad range of vision domains and under different model architectures and sizes.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24494
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO
Wang, Hongcheng
Huang, Yinuo
Wang, Sukai
Ren, Guanghui
Dong, Hao
Computation and Language
Group Relative Policy Optimization (GRPO) trains Chain-of-Thought reasoning with verifiable rewards, but estimating thought-level advantages without value functions often suffers from high variance. Although tree-style branching is used in practice to reduce the variance, it lacks a theoretical explanation of why it works and whether it is important or even potentially necessary. We study thought-level advantage estimation in GRPO from a variance perspective under a minimal tree-style setting where multiple answers are sampled for each thought. Using the multivariate delta method, we reveal an asymmetry in how different sampling dimensions affect variance. Increasing the number of sampled thoughts ($K$) leaves a strictly positive variance floor, whereas increasing the number of answers per thought ($M$) induces a monotonic decrease in variance, asymptotically decreasing it to zero. This implies that accurate thought-level advantage estimation is impossible through scaling thought sampling alone, making branching a potentially necessary mechanism rather than a heuristic. Experiments further provide empirical evidence for both the effectiveness and necessity of answer-level branching, demonstrating improved optimization stability, training efficiency, and final performance not only in math but also across a broad range of vision domains and under different model architectures and sizes.
title Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO
topic Computation and Language
url https://arxiv.org/abs/2509.24494