The Consistency Hypothesis in Uncertainty Quantification for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Quan, Bhattacharjya, Debarun, Ganesan, Balaji, Marinescu, Radu, Mirylenka, Katsiaryna, Pham, Nhan H, Glass, Michael, Lee, Junkyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911025963466752
author Xiao, Quan
Bhattacharjya, Debarun
Ganesan, Balaji
Marinescu, Radu
Mirylenka, Katsiaryna
Pham, Nhan H
Glass, Michael
Lee, Junkyu
author_facet Xiao, Quan
Bhattacharjya, Debarun
Ganesan, Balaji
Marinescu, Radu
Mirylenka, Katsiaryna
Pham, Nhan H
Glass, Michael
Lee, Junkyu
contents Estimating the confidence of large language model (LLM) outputs is essential for real-world applications requiring high user trust. Black-box uncertainty quantification (UQ) methods, relying solely on model API access, have gained popularity due to their practical benefits. In this paper, we examine the implicit assumption behind several UQ methods, which use generation consistency as a proxy for confidence, an idea we formalize as the consistency hypothesis. We introduce three mathematical statements with corresponding statistical tests to capture variations of this hypothesis and metrics to evaluate LLM output conformity across tasks. Our empirical investigation, spanning 8 benchmark datasets and 3 tasks (question answering, text summarization, and text-to-SQL), highlights the prevalence of the hypothesis under different settings. Among the statements, we highlight the `Sim-Any' hypothesis as the most actionable, and demonstrate how it can be leveraged by proposing data-free black-box UQ methods that aggregate similarities between generations for confidence estimation. These approaches can outperform the closest baselines, showcasing the practical value of the empirically observed consistency hypothesis.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21849
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Consistency Hypothesis in Uncertainty Quantification for Large Language Models
Xiao, Quan
Bhattacharjya, Debarun
Ganesan, Balaji
Marinescu, Radu
Mirylenka, Katsiaryna
Pham, Nhan H
Glass, Michael
Lee, Junkyu
Computation and Language
Artificial Intelligence
Machine Learning
Estimating the confidence of large language model (LLM) outputs is essential for real-world applications requiring high user trust. Black-box uncertainty quantification (UQ) methods, relying solely on model API access, have gained popularity due to their practical benefits. In this paper, we examine the implicit assumption behind several UQ methods, which use generation consistency as a proxy for confidence, an idea we formalize as the consistency hypothesis. We introduce three mathematical statements with corresponding statistical tests to capture variations of this hypothesis and metrics to evaluate LLM output conformity across tasks. Our empirical investigation, spanning 8 benchmark datasets and 3 tasks (question answering, text summarization, and text-to-SQL), highlights the prevalence of the hypothesis under different settings. Among the statements, we highlight the `Sim-Any' hypothesis as the most actionable, and demonstrate how it can be leveraged by proposing data-free black-box UQ methods that aggregate similarities between generations for confidence estimation. These approaches can outperform the closest baselines, showcasing the practical value of the empirically observed consistency hypothesis.
title The Consistency Hypothesis in Uncertainty Quantification for Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.21849