Maximizing Prefix-Confidence at Test-Time Efficiently Improves Mathematical Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Otth, Matthias, Hübotter, Jonas, Hakimi, Ido, Krause, Andreas
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911074433892352
author Otth, Matthias
Hübotter, Jonas
Hakimi, Ido
Krause, Andreas
author_facet Otth, Matthias
Hübotter, Jonas
Hakimi, Ido
Krause, Andreas
contents Recent work has shown that language models can self-improve by maximizing their own confidence in their predictions, without relying on external verifiers or reward signals. In this work, we study the test-time scaling of language models for mathematical reasoning tasks, where the model's own confidence is used to select the most promising attempts. Surprisingly, we find that we can achieve significant performance gains by continuing only the most promising attempt, selected by the model's prefix-confidence. We systematically evaluate prefix-confidence scaling on five mathematical reasoning datasets: the school-level GSM8K and MATH500, and the competition-level AMC23, AIME24, and AIME25. We find that prefix-confidence scaling with prefixes of only 32 tokens achieves a better accuracy-compute trade-off than majority voting. Moreover, prefix-confidence scaling appears less susceptible than BoN to length biases. Finally, we also evaluate test-time training with prefix-confidence and find that, while outperforming the base model, it does not improve over prefix-confidence scaling.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18122
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Maximizing Prefix-Confidence at Test-Time Efficiently Improves Mathematical Reasoning
Otth, Matthias
Hübotter, Jonas
Hakimi, Ido
Krause, Andreas
Machine Learning
Recent work has shown that language models can self-improve by maximizing their own confidence in their predictions, without relying on external verifiers or reward signals. In this work, we study the test-time scaling of language models for mathematical reasoning tasks, where the model's own confidence is used to select the most promising attempts. Surprisingly, we find that we can achieve significant performance gains by continuing only the most promising attempt, selected by the model's prefix-confidence. We systematically evaluate prefix-confidence scaling on five mathematical reasoning datasets: the school-level GSM8K and MATH500, and the competition-level AMC23, AIME24, and AIME25. We find that prefix-confidence scaling with prefixes of only 32 tokens achieves a better accuracy-compute trade-off than majority voting. Moreover, prefix-confidence scaling appears less susceptible than BoN to length biases. Finally, we also evaluate test-time training with prefix-confidence and find that, while outperforming the base model, it does not improve over prefix-confidence scaling.
title Maximizing Prefix-Confidence at Test-Time Efficiently Improves Mathematical Reasoning
topic Machine Learning
url https://arxiv.org/abs/2507.18122