Why Don't You Know? Evaluating the Impact of Uncertainty Sources on Uncertainty Quantification in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Goloburda, Maiya, Vashurin, Roman, Chernogorskii, Fedor, Laiyk, Nurkhan, Orel, Daniil, Nakov, Preslav, Panov, Maxim
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914615289446400
author Goloburda, Maiya
Vashurin, Roman
Chernogorskii, Fedor
Laiyk, Nurkhan
Orel, Daniil
Nakov, Preslav
Panov, Maxim
author_facet Goloburda, Maiya
Vashurin, Roman
Chernogorskii, Fedor
Laiyk, Nurkhan
Orel, Daniil
Nakov, Preslav
Panov, Maxim
contents As Large Language Models (LLMs) are increasingly deployed in real-world applications, reliable uncertainty quantification (UQ) becomes critical for safe and effective use. Most existing UQ approaches for language models aim to produce a single confidence score -- for example, estimating the probability that a model's answer is correct. However, uncertainty in natural language tasks arises from multiple distinct sources, including model knowledge gaps, output variability, and input ambiguity, which have different implications for system behavior and user interaction. In this work, we study how the source of uncertainty impacts the behavior and effectiveness of existing UQ methods. To enable controlled analysis, we introduce a new dataset that explicitly categorizes uncertainty sources, allowing systematic evaluation of UQ performance under each condition. Our experiments reveal that while many UQ methods perform well when uncertainty stems solely from model knowledge limitations, their performance degrades or becomes misleading when other sources are introduced. These findings highlight the need for uncertainty-aware methods that explicitly account for the source of uncertainty in large language models.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10495
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Why Don't You Know? Evaluating the Impact of Uncertainty Sources on Uncertainty Quantification in LLMs
Goloburda, Maiya
Vashurin, Roman
Chernogorskii, Fedor
Laiyk, Nurkhan
Orel, Daniil
Nakov, Preslav
Panov, Maxim
Computation and Language
As Large Language Models (LLMs) are increasingly deployed in real-world applications, reliable uncertainty quantification (UQ) becomes critical for safe and effective use. Most existing UQ approaches for language models aim to produce a single confidence score -- for example, estimating the probability that a model's answer is correct. However, uncertainty in natural language tasks arises from multiple distinct sources, including model knowledge gaps, output variability, and input ambiguity, which have different implications for system behavior and user interaction. In this work, we study how the source of uncertainty impacts the behavior and effectiveness of existing UQ methods. To enable controlled analysis, we introduce a new dataset that explicitly categorizes uncertainty sources, allowing systematic evaluation of UQ performance under each condition. Our experiments reveal that while many UQ methods perform well when uncertainty stems solely from model knowledge limitations, their performance degrades or becomes misleading when other sources are introduced. These findings highlight the need for uncertainty-aware methods that explicitly account for the source of uncertainty in large language models.
title Why Don't You Know? Evaluating the Impact of Uncertainty Sources on Uncertainty Quantification in LLMs
topic Computation and Language
url https://arxiv.org/abs/2604.10495