Unifying Tree Search Algorithm and Reward Design for LLM Reasoning: A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Jiaqi, Zhang, Xiang, Yang, Yuejin, Huang, Wenxuan, Cao, Juntai, Xu, Sheng, Zhuang, Xiang, Gao, Zhangyang, Abdul-Mageed, Muhammad, Lakshmanan, Laks V. S., You, Chenyu, Ouyang, Wanli, Sun, Siqi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914087611400192
author Wei, Jiaqi
Zhang, Xiang
Yang, Yuejin
Huang, Wenxuan
Cao, Juntai
Xu, Sheng
Zhuang, Xiang
Gao, Zhangyang
Abdul-Mageed, Muhammad
Lakshmanan, Laks V. S.
You, Chenyu
Ouyang, Wanli
Sun, Siqi
author_facet Wei, Jiaqi
Zhang, Xiang
Yang, Yuejin
Huang, Wenxuan
Cao, Juntai
Xu, Sheng
Zhuang, Xiang
Gao, Zhangyang
Abdul-Mageed, Muhammad
Lakshmanan, Laks V. S.
You, Chenyu
Ouyang, Wanli
Sun, Siqi
contents Deliberative tree search is a cornerstone of modern Large Language Model (LLM) research, driving the pivot from brute-force scaling toward algorithmic efficiency. This single paradigm unifies two critical frontiers: \textbf{Test-Time Scaling (TTS)}, which deploys on-demand computation to solve hard problems, and \textbf{Self-Improvement}, which uses search-generated data to durably enhance model parameters. However, this burgeoning field is fragmented and lacks a common formalism, particularly concerning the ambiguous role of the reward signal -- is it a transient heuristic or a durable learning target? This paper resolves this ambiguity by introducing a unified framework that deconstructs search algorithms into three core components: the \emph{Search Mechanism}, \emph{Reward Formulation}, and \emph{Transition Function}. We establish a formal distinction between transient \textbf{Search Guidance} for TTS and durable \textbf{Parametric Reward Modeling} for Self-Improvement. Building on this formalism, we introduce a component-centric taxonomy, synthesize the state-of-the-art, and chart a research roadmap toward more systematic progress in creating autonomous, self-improving agents.
format Preprint
id arxiv_https___arxiv_org_abs_2510_09988
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unifying Tree Search Algorithm and Reward Design for LLM Reasoning: A Survey
Wei, Jiaqi
Zhang, Xiang
Yang, Yuejin
Huang, Wenxuan
Cao, Juntai
Xu, Sheng
Zhuang, Xiang
Gao, Zhangyang
Abdul-Mageed, Muhammad
Lakshmanan, Laks V. S.
You, Chenyu
Ouyang, Wanli
Sun, Siqi
Computation and Language
Deliberative tree search is a cornerstone of modern Large Language Model (LLM) research, driving the pivot from brute-force scaling toward algorithmic efficiency. This single paradigm unifies two critical frontiers: \textbf{Test-Time Scaling (TTS)}, which deploys on-demand computation to solve hard problems, and \textbf{Self-Improvement}, which uses search-generated data to durably enhance model parameters. However, this burgeoning field is fragmented and lacks a common formalism, particularly concerning the ambiguous role of the reward signal -- is it a transient heuristic or a durable learning target? This paper resolves this ambiguity by introducing a unified framework that deconstructs search algorithms into three core components: the \emph{Search Mechanism}, \emph{Reward Formulation}, and \emph{Transition Function}. We establish a formal distinction between transient \textbf{Search Guidance} for TTS and durable \textbf{Parametric Reward Modeling} for Self-Improvement. Building on this formalism, we introduce a component-centric taxonomy, synthesize the state-of-the-art, and chart a research roadmap toward more systematic progress in creating autonomous, self-improving agents.
title Unifying Tree Search Algorithm and Reward Design for LLM Reasoning: A Survey
topic Computation and Language
url https://arxiv.org/abs/2510.09988