I Don't Know: Explicit Modeling of Uncertainty with an [IDK] Token

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cohen, Roi, Dobler, Konstantin, Biran, Eden, de Melo, Gerard
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916514718810112
author Cohen, Roi
Dobler, Konstantin
Biran, Eden
de Melo, Gerard
author_facet Cohen, Roi
Dobler, Konstantin
Biran, Eden
de Melo, Gerard
contents Large Language Models are known to capture real-world knowledge, allowing them to excel in many downstream tasks. Despite recent advances, these models are still prone to what are commonly known as hallucinations, causing them to emit unwanted and factually incorrect text. In this work, we propose a novel calibration method that can be used to combat hallucinations. We add a special [IDK] ("I don't know") token to the model's vocabulary and introduce an objective function that shifts probability mass to the [IDK] token for incorrect predictions. This approach allows the model to express uncertainty in its output explicitly. We evaluate our proposed method across multiple model architectures and factual downstream tasks. We find that models trained with our method are able to express uncertainty in places where they would previously make mistakes while suffering only a small loss of encoded knowledge. We further perform extensive ablation studies of multiple variations of our approach and provide a detailed analysis of the precision-recall tradeoff of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2412_06676
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle I Don't Know: Explicit Modeling of Uncertainty with an [IDK] Token
Cohen, Roi
Dobler, Konstantin
Biran, Eden
de Melo, Gerard
Machine Learning
Computation and Language
Large Language Models are known to capture real-world knowledge, allowing them to excel in many downstream tasks. Despite recent advances, these models are still prone to what are commonly known as hallucinations, causing them to emit unwanted and factually incorrect text. In this work, we propose a novel calibration method that can be used to combat hallucinations. We add a special [IDK] ("I don't know") token to the model's vocabulary and introduce an objective function that shifts probability mass to the [IDK] token for incorrect predictions. This approach allows the model to express uncertainty in its output explicitly. We evaluate our proposed method across multiple model architectures and factual downstream tasks. We find that models trained with our method are able to express uncertainty in places where they would previously make mistakes while suffering only a small loss of encoded knowledge. We further perform extensive ablation studies of multiple variations of our approach and provide a detailed analysis of the precision-recall tradeoff of our method.
title I Don't Know: Explicit Modeling of Uncertainty with an [IDK] Token
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2412.06676