Fine-Tuning Large Language Models to Appropriately Abstain with Semantic Entropy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tjandra, Benedict Aaron, Razzak, Muhammed, Kossen, Jannik, Handa, Kunal, Gal, Yarin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914984248737792
author Tjandra, Benedict Aaron
Razzak, Muhammed
Kossen, Jannik
Handa, Kunal
Gal, Yarin
author_facet Tjandra, Benedict Aaron
Razzak, Muhammed
Kossen, Jannik
Handa, Kunal
Gal, Yarin
contents Large Language Models (LLMs) are known to hallucinate, whereby they generate plausible but inaccurate text. This phenomenon poses significant risks in critical applications, such as medicine or law, necessitating robust hallucination mitigation strategies. While recent works have proposed fine-tuning methods to teach LLMs to abstain from answering questions beyond their knowledge or capabilities, these methods rely on the existence of ground-truth labels or are limited to short-form responses. To address these limitations, we propose fine-tuning using semantic entropy, an uncertainty measure derived from introspection into the model which does not require external labels. We demonstrate that our approach matches or outperforms models fine-tuned using prior work and achieves strong performance for both short and long-form generations on a range of datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2410_17234
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fine-Tuning Large Language Models to Appropriately Abstain with Semantic Entropy
Tjandra, Benedict Aaron
Razzak, Muhammed
Kossen, Jannik
Handa, Kunal
Gal, Yarin
Computation and Language
Machine Learning
Large Language Models (LLMs) are known to hallucinate, whereby they generate plausible but inaccurate text. This phenomenon poses significant risks in critical applications, such as medicine or law, necessitating robust hallucination mitigation strategies. While recent works have proposed fine-tuning methods to teach LLMs to abstain from answering questions beyond their knowledge or capabilities, these methods rely on the existence of ground-truth labels or are limited to short-form responses. To address these limitations, we propose fine-tuning using semantic entropy, an uncertainty measure derived from introspection into the model which does not require external labels. We demonstrate that our approach matches or outperforms models fine-tuned using prior work and achieves strong performance for both short and long-form generations on a range of datasets.
title Fine-Tuning Large Language Models to Appropriately Abstain with Semantic Entropy
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2410.17234