UNCERTAINTY-LINE: Length-Invariant Estimation of Uncertainty for Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Vashurin, Roman, Goloburda, Maiya, Nakov, Preslav, Panov, Maxim
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909805561511936
author Vashurin, Roman
Goloburda, Maiya
Nakov, Preslav
Panov, Maxim
author_facet Vashurin, Roman
Goloburda, Maiya
Nakov, Preslav
Panov, Maxim
contents Large Language Models (LLMs) have become indispensable tools across various applications, making it more important than ever to ensure the quality and the trustworthiness of their outputs. This has led to growing interest in uncertainty quantification (UQ) methods for assessing the reliability of LLM outputs. Many existing UQ techniques rely on token probabilities, which inadvertently introduces a bias with respect to the length of the output. While some methods attempt to account for this, we demonstrate that such biases persist even in length-normalized approaches. To address the problem, here we propose UNCERTAINTY-LINE: (Length-INvariant Estimation), a simple debiasing procedure that regresses uncertainty scores on output length and uses the residuals as corrected, length-invariant estimates. Our method is post-hoc, model-agnostic, and applicable to a range of UQ measures. Through extensive evaluation on machine translation, summarization, and question-answering tasks, we demonstrate that UNCERTAINTY-LINE: consistently improves over even nominally length-normalized UQ methods uncertainty estimates across multiple metrics and models.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19060
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UNCERTAINTY-LINE: Length-Invariant Estimation of Uncertainty for Large Language Models
Vashurin, Roman
Goloburda, Maiya
Nakov, Preslav
Panov, Maxim
Computation and Language
Large Language Models (LLMs) have become indispensable tools across various applications, making it more important than ever to ensure the quality and the trustworthiness of their outputs. This has led to growing interest in uncertainty quantification (UQ) methods for assessing the reliability of LLM outputs. Many existing UQ techniques rely on token probabilities, which inadvertently introduces a bias with respect to the length of the output. While some methods attempt to account for this, we demonstrate that such biases persist even in length-normalized approaches. To address the problem, here we propose UNCERTAINTY-LINE: (Length-INvariant Estimation), a simple debiasing procedure that regresses uncertainty scores on output length and uses the residuals as corrected, length-invariant estimates. Our method is post-hoc, model-agnostic, and applicable to a range of UQ measures. Through extensive evaluation on machine translation, summarization, and question-answering tasks, we demonstrate that UNCERTAINTY-LINE: consistently improves over even nominally length-normalized UQ methods uncertainty estimates across multiple metrics and models.
title UNCERTAINTY-LINE: Length-Invariant Estimation of Uncertainty for Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2505.19060