Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Miyamoto, Sora, Oba, Daisuke, Okazaki, Naoaki
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912894263754752
author Miyamoto, Sora
Oba, Daisuke
Okazaki, Naoaki
author_facet Miyamoto, Sora
Oba, Daisuke
Okazaki, Naoaki
contents Tree-search decoding is an effective form of test-time scaling for large language models (LLMs), but real-world deployment imposes a fixed per-query token budget that varies across settings. Existing tree-search policies are largely budget-agnostic, treating the budget as a termination condition, which can lead to late-stage over-branching or premature termination. We propose {Budget-Guided MCTS} (BG-MCTS), a tree-search decoding algorithm that aligns its search policy with the remaining token budget: it starts with broad exploration, then prioritizes refinement and answer completion as the budget depletes while reducing late-stage branching from shallow nodes. BG-MCTS consistently outperforms budget-agnostic tree-search baselines across different budgets on MATH500 and AIME24/25 with open-weight LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09574
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs
Miyamoto, Sora
Oba, Daisuke
Okazaki, Naoaki
Computation and Language
Artificial Intelligence
Machine Learning
Tree-search decoding is an effective form of test-time scaling for large language models (LLMs), but real-world deployment imposes a fixed per-query token budget that varies across settings. Existing tree-search policies are largely budget-agnostic, treating the budget as a termination condition, which can lead to late-stage over-branching or premature termination. We propose {Budget-Guided MCTS} (BG-MCTS), a tree-search decoding algorithm that aligns its search policy with the remaining token budget: it starts with broad exploration, then prioritizes refinement and answer completion as the budget depletes while reducing late-stage branching from shallow nodes. BG-MCTS consistently outperforms budget-agnostic tree-search baselines across different budgets on MATH500 and AIME24/25 with open-weight LLMs.
title Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2602.09574