Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vazhentsev, Artem, Fadeeva, Ekaterina, Xing, Rui, Kuzmin, Gleb, Lazichny, Ivan, Panchenko, Alexander, Nakov, Preslav, Baldwin, Timothy, Panov, Maxim, Shelmanov, Artem
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915961485918208
author Vazhentsev, Artem
Fadeeva, Ekaterina
Xing, Rui
Kuzmin, Gleb
Lazichny, Ivan
Panchenko, Alexander
Nakov, Preslav
Baldwin, Timothy
Panov, Maxim
Shelmanov, Artem
author_facet Vazhentsev, Artem
Fadeeva, Ekaterina
Xing, Rui
Kuzmin, Gleb
Lazichny, Ivan
Panchenko, Alexander
Nakov, Preslav
Baldwin, Timothy
Panov, Maxim
Shelmanov, Artem
contents Uncertainty quantification (UQ) has emerged as a promising approach for detecting hallucinations and low-quality output of Large Language Models (LLMs). However, obtaining proper uncertainty scores is complicated by the conditional dependency between the generation steps of an autoregressive LLM because it is hard to model it explicitly. Here, we propose to learn this dependency from attention-based features. In particular, we train a regression model that leverages LLM attention maps, probabilities on the current generation step, and recurrently computed uncertainty scores from previously generated tokens. To incorporate the recurrent features, we also suggest a two-staged training procedure. Our experimental evaluation on ten datasets and three LLMs shows that the proposed method is highly effective for selective generation, achieving substantial improvements over rivaling unsupervised and supervised approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2408_10692
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models
Vazhentsev, Artem
Fadeeva, Ekaterina
Xing, Rui
Kuzmin, Gleb
Lazichny, Ivan
Panchenko, Alexander
Nakov, Preslav
Baldwin, Timothy
Panov, Maxim
Shelmanov, Artem
Computation and Language
Uncertainty quantification (UQ) has emerged as a promising approach for detecting hallucinations and low-quality output of Large Language Models (LLMs). However, obtaining proper uncertainty scores is complicated by the conditional dependency between the generation steps of an autoregressive LLM because it is hard to model it explicitly. Here, we propose to learn this dependency from attention-based features. In particular, we train a regression model that leverages LLM attention maps, probabilities on the current generation step, and recurrently computed uncertainty scores from previously generated tokens. To incorporate the recurrent features, we also suggest a two-staged training procedure. Our experimental evaluation on ten datasets and three LLMs shows that the proposed method is highly effective for selective generation, achieving substantial improvements over rivaling unsupervised and supervised approaches.
title Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2408.10692