Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915961485918208 |
|---|---|
| author | Vazhentsev, Artem Fadeeva, Ekaterina Xing, Rui Kuzmin, Gleb Lazichny, Ivan Panchenko, Alexander Nakov, Preslav Baldwin, Timothy Panov, Maxim Shelmanov, Artem |
| author_facet | Vazhentsev, Artem Fadeeva, Ekaterina Xing, Rui Kuzmin, Gleb Lazichny, Ivan Panchenko, Alexander Nakov, Preslav Baldwin, Timothy Panov, Maxim Shelmanov, Artem |
| contents | Uncertainty quantification (UQ) has emerged as a promising approach for detecting hallucinations and low-quality output of Large Language Models (LLMs). However, obtaining proper uncertainty scores is complicated by the conditional dependency between the generation steps of an autoregressive LLM because it is hard to model it explicitly. Here, we propose to learn this dependency from attention-based features. In particular, we train a regression model that leverages LLM attention maps, probabilities on the current generation step, and recurrently computed uncertainty scores from previously generated tokens. To incorporate the recurrent features, we also suggest a two-staged training procedure. Our experimental evaluation on ten datasets and three LLMs shows that the proposed method is highly effective for selective generation, achieving substantial improvements over rivaling unsupervised and supervised approaches. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_10692 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models Vazhentsev, Artem Fadeeva, Ekaterina Xing, Rui Kuzmin, Gleb Lazichny, Ivan Panchenko, Alexander Nakov, Preslav Baldwin, Timothy Panov, Maxim Shelmanov, Artem Computation and Language Uncertainty quantification (UQ) has emerged as a promising approach for detecting hallucinations and low-quality output of Large Language Models (LLMs). However, obtaining proper uncertainty scores is complicated by the conditional dependency between the generation steps of an autoregressive LLM because it is hard to model it explicitly. Here, we propose to learn this dependency from attention-based features. In particular, we train a regression model that leverages LLM attention maps, probabilities on the current generation step, and recurrently computed uncertainty scores from previously generated tokens. To incorporate the recurrent features, we also suggest a two-staged training procedure. Our experimental evaluation on ten datasets and three LLMs shows that the proposed method is highly effective for selective generation, achieving substantial improvements over rivaling unsupervised and supervised approaches. |
| title | Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2408.10692 |