Learning to Reason Efficiently with A* Post-Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Opedal, Andreas, Re, Francesco Ignazio, Saparov, Abulhair, Sachan, Mrinmaya, Schölkopf, Bernhard, Cotterell, Ryan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918519943200768
author Opedal, Andreas
Re, Francesco Ignazio
Saparov, Abulhair
Sachan, Mrinmaya
Schölkopf, Bernhard
Cotterell, Ryan
author_facet Opedal, Andreas
Re, Francesco Ignazio
Saparov, Abulhair
Sachan, Mrinmaya
Schölkopf, Bernhard
Cotterell, Ryan
contents Many applications of large language models (LLMs) require deductive reasoning, yet models frequently produce incorrect or redundant inference steps. We frame natural language inference as a search problem where the final answer is the valid proof itself, requiring a reasoning procedure in which intermediate inferences are correct. Specifically, we investigate whether LLMs can learn to generate correct and efficient proofs with guidance from A* search -- an algorithm that guarantees an optimally efficient path to a goal. We explore two training techniques: supervised fine-tuning on execution traces from A* and reinforcement learning with A*-informed process reward models. Empirically, we find that Llama-3.2 models in the 1B--3B range benefit substantially from A* post training, going from near-zero accuracy to outperforming DeepSeek-V3.2 -- a much larger model. Our analysis uncovers a trade-off: while simple correctness rewards maximize accuracy, A*-informed signals strike a balance between accuracy and efficiency. Furthermore, we find that on larger search spaces, models trained with imperfect heuristics exhibit superior accuracy. Our results demonstrate a promising direction towards reasoning guided by principles derived from classical search algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24597
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning to Reason Efficiently with A* Post-Training
Opedal, Andreas
Re, Francesco Ignazio
Saparov, Abulhair
Sachan, Mrinmaya
Schölkopf, Bernhard
Cotterell, Ryan
Artificial Intelligence
Computation and Language
Machine Learning
Many applications of large language models (LLMs) require deductive reasoning, yet models frequently produce incorrect or redundant inference steps. We frame natural language inference as a search problem where the final answer is the valid proof itself, requiring a reasoning procedure in which intermediate inferences are correct. Specifically, we investigate whether LLMs can learn to generate correct and efficient proofs with guidance from A* search -- an algorithm that guarantees an optimally efficient path to a goal. We explore two training techniques: supervised fine-tuning on execution traces from A* and reinforcement learning with A*-informed process reward models. Empirically, we find that Llama-3.2 models in the 1B--3B range benefit substantially from A* post training, going from near-zero accuracy to outperforming DeepSeek-V3.2 -- a much larger model. Our analysis uncovers a trade-off: while simple correctness rewards maximize accuracy, A*-informed signals strike a balance between accuracy and efficiency. Furthermore, we find that on larger search spaces, models trained with imperfect heuristics exhibit superior accuracy. Our results demonstrate a promising direction towards reasoning guided by principles derived from classical search algorithms.
title Learning to Reason Efficiently with A* Post-Training
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2605.24597