Alleviating Hallucinations of Large Language Models through Induced Hallucinations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yue, Cui, Leyang, Bi, Wei, Shi, Shuming
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916154351550464
author Zhang, Yue
Cui, Leyang
Bi, Wei
Shi, Shuming
author_facet Zhang, Yue
Cui, Leyang
Bi, Wei
Shi, Shuming
contents Despite their impressive capabilities, large language models (LLMs) have been observed to generate responses that include inaccurate or fabricated information, a phenomenon commonly known as ``hallucination''. In this work, we propose a simple \textit{Induce-then-Contrast} Decoding (ICD) strategy to alleviate hallucinations. We first construct a factually weak LLM by inducing hallucinations from the original LLMs. Then, we penalize these induced hallucinations during decoding to enhance the factuality of the generated content. Concretely, we determine the final next-token predictions by amplifying the predictions from the original model and downplaying the induced untruthful predictions via contrastive decoding. Experimental results on both discrimination-based and generation-based hallucination evaluation benchmarks, such as TruthfulQA and \textsc{FActScore}, demonstrate that our proposed ICD methods can effectively enhance the factuality of LLMs across various model sizes and families. For example, when equipped with ICD, Llama2-7B-Chat and Mistral-7B-Instruct achieve performance comparable to ChatGPT and GPT4 on TruthfulQA, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2312_15710
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Alleviating Hallucinations of Large Language Models through Induced Hallucinations
Zhang, Yue
Cui, Leyang
Bi, Wei
Shi, Shuming
Computation and Language
Artificial Intelligence
Despite their impressive capabilities, large language models (LLMs) have been observed to generate responses that include inaccurate or fabricated information, a phenomenon commonly known as ``hallucination''. In this work, we propose a simple \textit{Induce-then-Contrast} Decoding (ICD) strategy to alleviate hallucinations. We first construct a factually weak LLM by inducing hallucinations from the original LLMs. Then, we penalize these induced hallucinations during decoding to enhance the factuality of the generated content. Concretely, we determine the final next-token predictions by amplifying the predictions from the original model and downplaying the induced untruthful predictions via contrastive decoding. Experimental results on both discrimination-based and generation-based hallucination evaluation benchmarks, such as TruthfulQA and \textsc{FActScore}, demonstrate that our proposed ICD methods can effectively enhance the factuality of LLMs across various model sizes and families. For example, when equipped with ICD, Llama2-7B-Chat and Mistral-7B-Instruct achieve performance comparable to ChatGPT and GPT4 on TruthfulQA, respectively.
title Alleviating Hallucinations of Large Language Models through Induced Hallucinations
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2312.15710