How Uncertainty Estimation Scales with Sampling in Reasoning Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Del, Maksym, Kängsepp, Markus, Domnich, Marharyta, Tampuu, Ardi, Yankovskaya, Lisa, Kull, Meelis, Fishel, Mark
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918398420582400
author Del, Maksym
Kängsepp, Markus
Domnich, Marharyta
Tampuu, Ardi
Yankovskaya, Lisa
Kull, Meelis
Fishel, Mark
author_facet Del, Maksym
Kängsepp, Markus
Domnich, Marharyta
Tampuu, Ardi
Yankovskaya, Lisa
Kull, Meelis
Fishel, Mark
contents Uncertainty estimation is critical for deploying reasoning language models, yet remains poorly understood under extended chain-of-thought reasoning. We study parallel sampling as a fully black-box approach using verbalized confidence and self-consistency. Across three reasoning models and 17 tasks spanning mathematics, STEM, and humanities, we characterize how these signals scale. Both self-consistency and verbalized confidence scale in reasoning models, but self-consistency exhibits lower initial discrimination and lags behind verbalized confidence under moderate sampling. Most uncertainty gains, however, arise from signal combination: with just two samples, a hybrid estimator improves AUROC by up to $+12$ on average and already outperforms either signal alone even when scaled to much larger budgets, after which returns diminish. These effects are domain-dependent: in mathematics, the native domain of RLVR-style post-training, reasoning models achieve higher uncertainty quality and exhibit both stronger complementarity and faster scaling than in STEM or humanities.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19118
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle How Uncertainty Estimation Scales with Sampling in Reasoning Models
Del, Maksym
Kängsepp, Markus
Domnich, Marharyta
Tampuu, Ardi
Yankovskaya, Lisa
Kull, Meelis
Fishel, Mark
Artificial Intelligence
Computation and Language
Machine Learning
Uncertainty estimation is critical for deploying reasoning language models, yet remains poorly understood under extended chain-of-thought reasoning. We study parallel sampling as a fully black-box approach using verbalized confidence and self-consistency. Across three reasoning models and 17 tasks spanning mathematics, STEM, and humanities, we characterize how these signals scale. Both self-consistency and verbalized confidence scale in reasoning models, but self-consistency exhibits lower initial discrimination and lags behind verbalized confidence under moderate sampling. Most uncertainty gains, however, arise from signal combination: with just two samples, a hybrid estimator improves AUROC by up to $+12$ on average and already outperforms either signal alone even when scaled to much larger budgets, after which returns diminish. These effects are domain-dependent: in mathematics, the native domain of RLVR-style post-training, reasoning models achieve higher uncertainty quality and exhibit both stronger complementarity and faster scaling than in STEM or humanities.
title How Uncertainty Estimation Scales with Sampling in Reasoning Models
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2603.19118