Unifying Tree Search Algorithm and Reward Design for LLM Reasoning: A Survey
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Jiaqi, Zhang, Xiang, Yang, Yuejin, Huang, Wenxuan, Cao, Juntai, Xu, Sheng, Zhuang, Xiang, Gao, Zhangyang, Abdul-Mageed, Muhammad, Lakshmanan, Laks V. S., You, Chenyu, Ouyang, Wanli, Sun, Siqi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reflection Pretraining Enables Token-Level Self-Correction in Biological Sequence Models
by: Zhang, Xiang, et al.
Published: (2025)
by: Zhang, Xiang, et al.
Published: (2025)
Autoregressive + Chain of Thought = Recurrent: Recurrence's Role in Language Models' Computability and a Revisit of Recurrent Transformer
by: Zhang, Xiang, et al.
Published: (2024)
by: Zhang, Xiang, et al.
Published: (2024)
LLM Performance Predictors are good initializers for Architecture Search
by: Jawahar, Ganesh, et al.
Published: (2023)
by: Jawahar, Ganesh, et al.
Published: (2023)
DetoxLLM: A Framework for Detoxification with Explanations
by: Khondaker, Md Tawkat Islam, et al.
Published: (2024)
by: Khondaker, Md Tawkat Islam, et al.
Published: (2024)
When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning
by: Wei, Jiaqi, et al.
Published: (2026)
by: Wei, Jiaqi, et al.
Published: (2026)
Why Prompt Design Matters and Works: A Complexity Analysis of Prompt Search Space in LLMs
by: Zhang, Xiang, et al.
Published: (2025)
by: Zhang, Xiang, et al.
Published: (2025)
Tokenization Constraints in LLMs: A Study of Symbolic and Arithmetic Reasoning Limits
by: Zhang, Xiang, et al.
Published: (2025)
by: Zhang, Xiang, et al.
Published: (2025)
Counting Ability of Large Language Models and Impact of Tokenization
by: Zhang, Xiang, et al.
Published: (2024)
by: Zhang, Xiang, et al.
Published: (2024)
Cross-Modal Consistency in Multimodal Large Language Models
by: Zhang, Xiang, et al.
Published: (2024)
by: Zhang, Xiang, et al.
Published: (2024)
From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery
by: Wei, Jiaqi, et al.
Published: (2025)
by: Wei, Jiaqi, et al.
Published: (2025)
KRAFT: A Knowledge Graph-Based Framework for Automated Map Conflation
by: Hashemi, Farnoosh, et al.
Published: (2025)
by: Hashemi, Farnoosh, et al.
Published: (2025)
A Survey of Densest Subgraph Discovery on Large Graphs
by: Luo, Wensheng, et al.
Published: (2023)
by: Luo, Wensheng, et al.
Published: (2023)
Topology-Aware LLM-Driven Social Simulation: A Unified Framework for Efficient and Realistic Agent Dynamics
by: Xu, Yuwei, et al.
Published: (2026)
by: Xu, Yuwei, et al.
Published: (2026)
Fast Maximum Common Subgraph Search: A Redundancy-Reduced Backtracking Approach
by: Yu, Kaiqiang, et al.
Published: (2025)
by: Yu, Kaiqiang, et al.
Published: (2025)
Retrieval is Not Enough: Enhancing RAG Reasoning through Test-Time Critique and Optimization
by: Wei, Jiaqi, et al.
Published: (2025)
by: Wei, Jiaqi, et al.
Published: (2025)
OCCAM: Towards Cost-Efficient and Accuracy-Aware Classification Inference
by: Ding, Dujian, et al.
Published: (2024)
by: Ding, Dujian, et al.
Published: (2024)
On Barriers to Archival Audio Processing
by: Sullivan, Peter, et al.
Published: (2025)
by: Sullivan, Peter, et al.
Published: (2025)
In-depth Analysis of Densest Subgraph Discovery in a Unified Framework
by: Zhou, Yingli, et al.
Published: (2024)
by: Zhou, Yingli, et al.
Published: (2024)
Multi2: Multi-Agent Test-Time Scalable Framework for Multi-Document Processing
by: Cao, Juntai, et al.
Published: (2025)
by: Cao, Juntai, et al.
Published: (2025)
HarmonyCell: Automating Single-Cell Perturbation Modeling under Semantic and Distribution Shifts
by: Huang, Wenxuan, et al.
Published: (2026)
by: Huang, Wenxuan, et al.
Published: (2026)
Predicting Cascading Failures with a Hyperparametric Diffusion Model
by: Xiang, Bin, et al.
Published: (2024)
by: Xiang, Bin, et al.
Published: (2024)
To Distill or Not to Distill? On the Robustness of Robust Knowledge Distillation
by: Waheed, Abdul, et al.
Published: (2024)
by: Waheed, Abdul, et al.
Published: (2024)
TEMPO: A Realistic Multi-Domain Benchmark for Temporal Reasoning-Intensive Retrieval
by: Abdallah, Abdelrahman, et al.
Published: (2026)
by: Abdallah, Abdelrahman, et al.
Published: (2026)
A Multi-Agent Approach for Claim Verification from Tabular Data Documents
by: Saha, Rudra Ranajee, et al.
Published: (2026)
by: Saha, Rudra Ranajee, et al.
Published: (2026)
A Community-Based Approach for Stance Distribution and Argument Organization
by: Saha, Rudra Ranajee, et al.
Published: (2026)
by: Saha, Rudra Ranajee, et al.
Published: (2026)
Effective Self-Mining of In-Context Examples for Unsupervised Machine Translation with LLMs
by: Mekki, Abdellah El, et al.
Published: (2024)
by: Mekki, Abdellah El, et al.
Published: (2024)
TimeSeriesScientist: A General-Purpose AI Agent for Time Series Analysis
by: Zhao, Haokun, et al.
Published: (2025)
by: Zhao, Haokun, et al.
Published: (2025)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
by: Doan, Khai Duy, et al.
Published: (2024)
by: Doan, Khai Duy, et al.
Published: (2024)
QuantAgent: Price-Driven Multi-Agent LLMs for High-Frequency Trading
by: Xiong, Fei, et al.
Published: (2025)
by: Xiong, Fei, et al.
Published: (2025)
SlideGen: Collaborative Multimodal Agents for Scientific Slide Generation
by: Liang, Xin, et al.
Published: (2025)
by: Liang, Xin, et al.
Published: (2025)
Efficient and Effective Algorithms for A Family of Influence Maximization Problems with A Matroid Constraint
by: Huang, Yiqian, et al.
Published: (2024)
by: Huang, Yiqian, et al.
Published: (2024)
Hyperparametric Robust and Dynamic Influence Maximization
by: Saha, Arkaprava, et al.
Published: (2024)
by: Saha, Arkaprava, et al.
Published: (2024)
Dallah: A Dialect-Aware Multimodal Large Language Model for Arabic
by: Alwajih, Fakhraddin, et al.
Published: (2024)
by: Alwajih, Fakhraddin, et al.
Published: (2024)
uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes
by: Waheed, Abdul, et al.
Published: (2024)
by: Waheed, Abdul, et al.
Published: (2024)
PosterGen: Aesthetic-Aware Multi-Modal Paper-to-Poster Generation via Multi-Agent LLMs
by: Zhang, Zhilin, et al.
Published: (2025)
by: Zhang, Zhilin, et al.
Published: (2025)
Qalam : A Multimodal LLM for Arabic Optical Character and Handwriting Recognition
by: Bhatia, Gagan, et al.
Published: (2024)
by: Bhatia, Gagan, et al.
Published: (2024)
ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification Queries
by: Huang, Keke, et al.
Published: (2025)
by: Huang, Keke, et al.
Published: (2025)
Mixture-of-Supernets: Improving Weight-Sharing Supernet Training with Architecture-Routed Mixture-of-Experts
by: Jawahar, Ganesh, et al.
Published: (2023)
by: Jawahar, Ganesh, et al.
Published: (2023)
On Efficient Approximate Aggregate Nearest Neighbor Queries over Learned Representations
by: Wang, Carrie, et al.
Published: (2025)
by: Wang, Carrie, et al.
Published: (2025)
Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning
by: Hu, Yulan, et al.
Published: (2025)
by: Hu, Yulan, et al.
Published: (2025)
Similar Items
-
Reflection Pretraining Enables Token-Level Self-Correction in Biological Sequence Models
by: Zhang, Xiang, et al.
Published: (2025) -
Autoregressive + Chain of Thought = Recurrent: Recurrence's Role in Language Models' Computability and a Revisit of Recurrent Transformer
by: Zhang, Xiang, et al.
Published: (2024) -
LLM Performance Predictors are good initializers for Architecture Search
by: Jawahar, Ganesh, et al.
Published: (2023) -
DetoxLLM: A Framework for Detoxification with Explanations
by: Khondaker, Md Tawkat Islam, et al.
Published: (2024) -
When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning
by: Wei, Jiaqi, et al.
Published: (2026)