Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Cinquin, Tristan, Pleiss, Geoff, Kristiadi, Agustinus |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Introduction to the Analysis of Probabilistic Decision-Making Algorithms
by: Kristiadi, Agustinus
Published: (2025)
by: Kristiadi, Agustinus
Published: (2025)
A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?
by: Kristiadi, Agustinus, et al.
Published: (2024)
by: Kristiadi, Agustinus, et al.
Published: (2024)
How Useful is Intermittent, Asynchronous Expert Feedback for Bayesian Optimization?
by: Kristiadi, Agustinus, et al.
Published: (2024)
by: Kristiadi, Agustinus, et al.
Published: (2024)
On the Disconnect Between Theory and Practice of Neural Networks: Limits of the NTK Perspective
by: Wenger, Jonathan, et al.
Published: (2023)
by: Wenger, Jonathan, et al.
Published: (2023)
Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs
by: Feng, Zhangying, et al.
Published: (2025)
by: Feng, Zhangying, et al.
Published: (2025)
FSP-Laplace: Function-Space Priors for the Laplace Approximation in Bayesian Deep Learning
by: Cinquin, Tristan, et al.
Published: (2024)
by: Cinquin, Tristan, et al.
Published: (2024)
Uncertainty-Guided Likelihood Tree Search
by: Grosse, Julia, et al.
Published: (2024)
by: Grosse, Julia, et al.
Published: (2024)
DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMs
by: Zhang, Yuanhe, et al.
Published: (2025)
by: Zhang, Yuanhe, et al.
Published: (2025)
rePIRL: Learn PRM with Inverse RL for LLM Reasoning
by: Wu, Xian, et al.
Published: (2026)
by: Wu, Xian, et al.
Published: (2026)
DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
by: Cao, Qi, et al.
Published: (2025)
by: Cao, Qi, et al.
Published: (2025)
Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
Policy Guided Tree Search for Enhanced LLM Reasoning
by: Li, Yang
Published: (2025)
by: Li, Yang
Published: (2025)
Target-Aligned Reinforcement Learning
by: Pleiss, Leonard S., et al.
Published: (2026)
by: Pleiss, Leonard S., et al.
Published: (2026)
Monte Carlo Permutation Search
by: Cazenave, Tristan
Published: (2025)
by: Cazenave, Tristan
Published: (2025)
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
by: Mirzadeh, Iman, et al.
Published: (2024)
by: Mirzadeh, Iman, et al.
Published: (2024)
Regularized KL-Divergence for Well-Defined Function-Space Variational Inference in Bayesian neural networks
by: Cinquin, Tristan, et al.
Published: (2024)
by: Cinquin, Tristan, et al.
Published: (2024)
Quantization Meets Reasoning: Exploring and Mitigating Degradation of Low-Bit LLMs in Mathematical Reasoning
by: Li, Zhen, et al.
Published: (2025)
by: Li, Zhen, et al.
Published: (2025)
Revisiting Tree Search for LLMs: Gumbel and Sequential Halving for Budget-Scalable Reasoning
by: Ugadiarov, Leonid, et al.
Published: (2026)
by: Ugadiarov, Leonid, et al.
Published: (2026)
TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning
by: Zou, Jiaru, et al.
Published: (2025)
by: Zou, Jiaru, et al.
Published: (2025)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
by: Hao, Yuren, et al.
Published: (2025)
by: Hao, Yuren, et al.
Published: (2025)
A Critical Look At Tokenwise Reward-Guided Text Generation
by: Rashid, Ahmad, et al.
Published: (2024)
by: Rashid, Ahmad, et al.
Published: (2024)
Theoretical Limitations of Ensembles in the Age of Overparameterization
by: Dern, Niclas, et al.
Published: (2024)
by: Dern, Niclas, et al.
Published: (2024)
Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
by: Zhang, Zheng
Published: (2025)
by: Zhang, Zheng
Published: (2025)
Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
by: Liang, Zhenwen, et al.
Published: (2025)
by: Liang, Zhenwen, et al.
Published: (2025)
Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical Reasoning
by: Chen, Feng, et al.
Published: (2025)
by: Chen, Feng, et al.
Published: (2025)
A Fragile Number Sense: Probing the Elemental Limits of Numerical Reasoning in LLMs
by: Rahman, Roussel, et al.
Published: (2025)
by: Rahman, Roussel, et al.
Published: (2025)
Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMs
by: Zhou, Yifan, et al.
Published: (2025)
by: Zhou, Yifan, et al.
Published: (2025)
Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
by: Yu, Erxin, et al.
Published: (2025)
by: Yu, Erxin, et al.
Published: (2025)
FlashMD: long-stride, universal prediction of molecular dynamics
by: Bigi, Filippo, et al.
Published: (2025)
by: Bigi, Filippo, et al.
Published: (2025)
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
by: Shao, Zhihong, et al.
Published: (2024)
by: Shao, Zhihong, et al.
Published: (2024)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
by: Lin, Bill Yuchen, et al.
Published: (2025)
by: Lin, Bill Yuchen, et al.
Published: (2025)
Improving Line Search Methods for Large Scale Neural Network Training
by: Kenneweg, Philip, et al.
Published: (2024)
by: Kenneweg, Philip, et al.
Published: (2024)
Value-Guided Search for Efficient Chain-of-Thought Reasoning
by: Wang, Kaiwen, et al.
Published: (2025)
by: Wang, Kaiwen, et al.
Published: (2025)
DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
by: Yue, Murong, et al.
Published: (2024)
by: Yue, Murong, et al.
Published: (2024)
Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning
by: Singh, Joykirat, et al.
Published: (2024)
by: Singh, Joykirat, et al.
Published: (2024)
Towards Cost-Effective Reward Guided Text Generation
by: Rashid, Ahmad, et al.
Published: (2025)
by: Rashid, Ahmad, et al.
Published: (2025)
MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
by: Chen, Jinhao, et al.
Published: (2025)
by: Chen, Jinhao, et al.
Published: (2025)
Discovering Mathematical Formulas from Data via GPT-guided Monte Carlo Tree Search
by: Li, Yanjie, et al.
Published: (2024)
by: Li, Yanjie, et al.
Published: (2024)
Evidence for Limited Metacognition in LLMs
by: Ackerman, Christopher
Published: (2025)
by: Ackerman, Christopher
Published: (2025)
Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search
by: Zhang, Yifei, et al.
Published: (2026)
by: Zhang, Yifei, et al.
Published: (2026)
Similar Items
-
Introduction to the Analysis of Probabilistic Decision-Making Algorithms
by: Kristiadi, Agustinus
Published: (2025) -
A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?
by: Kristiadi, Agustinus, et al.
Published: (2024) -
How Useful is Intermittent, Asynchronous Expert Feedback for Bayesian Optimization?
by: Kristiadi, Agustinus, et al.
Published: (2024) -
On the Disconnect Between Theory and Practice of Neural Networks: Limits of the NTK Perspective
by: Wenger, Jonathan, et al.
Published: (2023) -
Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs
by: Feng, Zhangying, et al.
Published: (2025)