Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Xiaoou, Chen, Tiejin, Da, Longchao, Chen, Chacha, Lin, Zhen, Wei, Hua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908392191164416
author Liu, Xiaoou
Chen, Tiejin
Da, Longchao
Chen, Chacha
Lin, Zhen
Wei, Hua
author_facet Liu, Xiaoou
Chen, Tiejin
Da, Longchao
Chen, Chacha
Lin, Zhen
Wei, Hua
contents Large Language Models (LLMs) excel in text generation, reasoning, and decision-making, enabling their adoption in high-stakes domains such as healthcare, law, and transportation. However, their reliability is a major concern, as they often produce plausible but incorrect responses. Uncertainty quantification (UQ) enhances trustworthiness by estimating confidence in outputs, enabling risk mitigation and selective prediction. However, traditional UQ methods struggle with LLMs due to computational constraints and decoding inconsistencies. Moreover, LLMs introduce unique uncertainty sources, such as input ambiguity, reasoning path divergence, and decoding stochasticity, that extend beyond classical aleatoric and epistemic uncertainty. To address this, we introduce a new taxonomy that categorizes UQ methods based on computational efficiency and uncertainty dimensions (input, reasoning, parameter, and prediction uncertainty). We evaluate existing techniques, assess their real-world applicability, and identify open challenges, emphasizing the need for scalable, interpretable, and robust UQ approaches to enhance LLM reliability.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15850
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey
Liu, Xiaoou
Chen, Tiejin
Da, Longchao
Chen, Chacha
Lin, Zhen
Wei, Hua
Computation and Language
Large Language Models (LLMs) excel in text generation, reasoning, and decision-making, enabling their adoption in high-stakes domains such as healthcare, law, and transportation. However, their reliability is a major concern, as they often produce plausible but incorrect responses. Uncertainty quantification (UQ) enhances trustworthiness by estimating confidence in outputs, enabling risk mitigation and selective prediction. However, traditional UQ methods struggle with LLMs due to computational constraints and decoding inconsistencies. Moreover, LLMs introduce unique uncertainty sources, such as input ambiguity, reasoning path divergence, and decoding stochasticity, that extend beyond classical aleatoric and epistemic uncertainty. To address this, we introduce a new taxonomy that categorizes UQ methods based on computational efficiency and uncertainty dimensions (input, reasoning, parameter, and prediction uncertainty). We evaluate existing techniques, assess their real-world applicability, and identify open challenges, emphasizing the need for scalable, interpretable, and robust UQ approaches to enhance LLM reliability.
title Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey
topic Computation and Language
url https://arxiv.org/abs/2503.15850