Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917036566773760 |
|---|---|
| author | Hobelsberger, Christian Winner, Theresa Nawroth, Andreas Mitevski, Oliver Haensch, Anna-Carolina |
| author_facet | Hobelsberger, Christian Winner, Theresa Nawroth, Andreas Mitevski, Oliver Haensch, Anna-Carolina |
| contents | Large language models (LLMs) produce outputs with varying levels of uncertainty, and, just as often, varying levels of correctness; making their practical reliability far from guaranteed. To quantify this uncertainty, we systematically evaluate four approaches for confidence estimation in LLM outputs: VCE, MSP, Sample Consistency, and CoCoA (Vashurin et al., 2025). For the evaluation of the approaches, we conduct experiments on four question-answering tasks using a state-of-the-art open-source LLM. Our results show that each uncertainty metric captures a different facet of model confidence and that the hybrid CoCoA approach yields the best reliability overall, improving both calibration and discrimination of correct answers. We discuss the trade-offs of each method and provide recommendations for selecting uncertainty measures in LLM applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_20460 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models Hobelsberger, Christian Winner, Theresa Nawroth, Andreas Mitevski, Oliver Haensch, Anna-Carolina Computation and Language Applications Methodology Large language models (LLMs) produce outputs with varying levels of uncertainty, and, just as often, varying levels of correctness; making their practical reliability far from guaranteed. To quantify this uncertainty, we systematically evaluate four approaches for confidence estimation in LLM outputs: VCE, MSP, Sample Consistency, and CoCoA (Vashurin et al., 2025). For the evaluation of the approaches, we conduct experiments on four question-answering tasks using a state-of-the-art open-source LLM. Our results show that each uncertainty metric captures a different facet of model confidence and that the hybrid CoCoA approach yields the best reliability overall, improving both calibration and discrimination of correct answers. We discuss the trade-offs of each method and provide recommendations for selecting uncertainty measures in LLM applications. |
| title | Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models |
| topic | Computation and Language Applications Methodology |
| url | https://arxiv.org/abs/2510.20460 |