Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hobelsberger, Christian, Winner, Theresa, Nawroth, Andreas, Mitevski, Oliver, Haensch, Anna-Carolina
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917036566773760
author Hobelsberger, Christian
Winner, Theresa
Nawroth, Andreas
Mitevski, Oliver
Haensch, Anna-Carolina
author_facet Hobelsberger, Christian
Winner, Theresa
Nawroth, Andreas
Mitevski, Oliver
Haensch, Anna-Carolina
contents Large language models (LLMs) produce outputs with varying levels of uncertainty, and, just as often, varying levels of correctness; making their practical reliability far from guaranteed. To quantify this uncertainty, we systematically evaluate four approaches for confidence estimation in LLM outputs: VCE, MSP, Sample Consistency, and CoCoA (Vashurin et al., 2025). For the evaluation of the approaches, we conduct experiments on four question-answering tasks using a state-of-the-art open-source LLM. Our results show that each uncertainty metric captures a different facet of model confidence and that the hybrid CoCoA approach yields the best reliability overall, improving both calibration and discrimination of correct answers. We discuss the trade-offs of each method and provide recommendations for selecting uncertainty measures in LLM applications.
format Preprint
id arxiv_https___arxiv_org_abs_2510_20460
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models
Hobelsberger, Christian
Winner, Theresa
Nawroth, Andreas
Mitevski, Oliver
Haensch, Anna-Carolina
Computation and Language
Applications
Methodology
Large language models (LLMs) produce outputs with varying levels of uncertainty, and, just as often, varying levels of correctness; making their practical reliability far from guaranteed. To quantify this uncertainty, we systematically evaluate four approaches for confidence estimation in LLM outputs: VCE, MSP, Sample Consistency, and CoCoA (Vashurin et al., 2025). For the evaluation of the approaches, we conduct experiments on four question-answering tasks using a state-of-the-art open-source LLM. Our results show that each uncertainty metric captures a different facet of model confidence and that the hybrid CoCoA approach yields the best reliability overall, improving both calibration and discrimination of correct answers. We discuss the trade-offs of each method and provide recommendations for selecting uncertainty measures in LLM applications.
title Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models
topic Computation and Language
Applications
Methodology
url https://arxiv.org/abs/2510.20460