Unifying Tree Search Algorithm and Reward Design for LLM Reasoning: A Survey
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914087611400192 |
|---|---|
| author | Wei, Jiaqi Zhang, Xiang Yang, Yuejin Huang, Wenxuan Cao, Juntai Xu, Sheng Zhuang, Xiang Gao, Zhangyang Abdul-Mageed, Muhammad Lakshmanan, Laks V. S. You, Chenyu Ouyang, Wanli Sun, Siqi |
| author_facet | Wei, Jiaqi Zhang, Xiang Yang, Yuejin Huang, Wenxuan Cao, Juntai Xu, Sheng Zhuang, Xiang Gao, Zhangyang Abdul-Mageed, Muhammad Lakshmanan, Laks V. S. You, Chenyu Ouyang, Wanli Sun, Siqi |
| contents | Deliberative tree search is a cornerstone of modern Large Language Model (LLM) research, driving the pivot from brute-force scaling toward algorithmic efficiency. This single paradigm unifies two critical frontiers: \textbf{Test-Time Scaling (TTS)}, which deploys on-demand computation to solve hard problems, and \textbf{Self-Improvement}, which uses search-generated data to durably enhance model parameters. However, this burgeoning field is fragmented and lacks a common formalism, particularly concerning the ambiguous role of the reward signal -- is it a transient heuristic or a durable learning target? This paper resolves this ambiguity by introducing a unified framework that deconstructs search algorithms into three core components: the \emph{Search Mechanism}, \emph{Reward Formulation}, and \emph{Transition Function}. We establish a formal distinction between transient \textbf{Search Guidance} for TTS and durable \textbf{Parametric Reward Modeling} for Self-Improvement. Building on this formalism, we introduce a component-centric taxonomy, synthesize the state-of-the-art, and chart a research roadmap toward more systematic progress in creating autonomous, self-improving agents. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_09988 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Unifying Tree Search Algorithm and Reward Design for LLM Reasoning: A Survey Wei, Jiaqi Zhang, Xiang Yang, Yuejin Huang, Wenxuan Cao, Juntai Xu, Sheng Zhuang, Xiang Gao, Zhangyang Abdul-Mageed, Muhammad Lakshmanan, Laks V. S. You, Chenyu Ouyang, Wanli Sun, Siqi Computation and Language Deliberative tree search is a cornerstone of modern Large Language Model (LLM) research, driving the pivot from brute-force scaling toward algorithmic efficiency. This single paradigm unifies two critical frontiers: \textbf{Test-Time Scaling (TTS)}, which deploys on-demand computation to solve hard problems, and \textbf{Self-Improvement}, which uses search-generated data to durably enhance model parameters. However, this burgeoning field is fragmented and lacks a common formalism, particularly concerning the ambiguous role of the reward signal -- is it a transient heuristic or a durable learning target? This paper resolves this ambiguity by introducing a unified framework that deconstructs search algorithms into three core components: the \emph{Search Mechanism}, \emph{Reward Formulation}, and \emph{Transition Function}. We establish a formal distinction between transient \textbf{Search Guidance} for TTS and durable \textbf{Parametric Reward Modeling} for Self-Improvement. Building on this formalism, we introduce a component-centric taxonomy, synthesize the state-of-the-art, and chart a research roadmap toward more systematic progress in creating autonomous, self-improving agents. |
| title | Unifying Tree Search Algorithm and Reward Design for LLM Reasoning: A Survey |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2510.09988 |