Sample, Scrutinize and Scale: Effective Inference-Time Search by Scaling Verification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Eric, Awasthi, Pranjal, Gollapudi, Sreenivas
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913700232822784
author Zhao, Eric
Awasthi, Pranjal
Gollapudi, Sreenivas
author_facet Zhao, Eric
Awasthi, Pranjal
Gollapudi, Sreenivas
contents Sampling-based search, a simple paradigm for utilizing test-time compute, involves generating multiple candidate responses and selecting the best one -- typically by having models self-verify each response for correctness. In this paper, we study the scaling trends governing sampling-based search. Among our findings is that simply scaling up a minimalist implementation of sampling-based search, using only random sampling and direct self-verification, provides a practical inference method that, for example, elevates the reasoning capabilities of Gemini v1.5 Pro above that of o1-Preview on popular benchmarks. We partially attribute the scalability of sampling-based search to a phenomenon of implicit scaling, where sampling a larger pool of responses in turn improves self-verification accuracy. We further identify two useful principles for improving self-verification capabilities with test-time compute: (1) comparing across responses provides helpful signals about the locations of errors and hallucinations, and (2) different model output styles are useful for different contexts -- chains of thought are useful for reasoning but harder to verify. We also find that, though accurate verification can be elicited, frontier models demonstrate remarkably weak out-of-box verification capabilities and introduce a benchmark to measure progress on these deficiencies.
format Preprint
id arxiv_https___arxiv_org_abs_2502_01839
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sample, Scrutinize and Scale: Effective Inference-Time Search by Scaling Verification
Zhao, Eric
Awasthi, Pranjal
Gollapudi, Sreenivas
Machine Learning
Artificial Intelligence
Sampling-based search, a simple paradigm for utilizing test-time compute, involves generating multiple candidate responses and selecting the best one -- typically by having models self-verify each response for correctness. In this paper, we study the scaling trends governing sampling-based search. Among our findings is that simply scaling up a minimalist implementation of sampling-based search, using only random sampling and direct self-verification, provides a practical inference method that, for example, elevates the reasoning capabilities of Gemini v1.5 Pro above that of o1-Preview on popular benchmarks. We partially attribute the scalability of sampling-based search to a phenomenon of implicit scaling, where sampling a larger pool of responses in turn improves self-verification accuracy. We further identify two useful principles for improving self-verification capabilities with test-time compute: (1) comparing across responses provides helpful signals about the locations of errors and hallucinations, and (2) different model output styles are useful for different contexts -- chains of thought are useful for reasoning but harder to verify. We also find that, though accurate verification can be elicited, frontier models demonstrate remarkably weak out-of-box verification capabilities and introduce a benchmark to measure progress on these deficiencies.
title Sample, Scrutinize and Scale: Effective Inference-Time Search by Scaling Verification
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2502.01839