On Large Language Models' Hallucination with Regard to Known Facts

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jiang, Che, Qi, Biqing, Hong, Xiangyu, Fu, Dayuan, Cheng, Yang, Meng, Fandong, Yu, Mo, Zhou, Bowen, Zhou, Jie
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912089079021568
author Jiang, Che
Qi, Biqing
Hong, Xiangyu
Fu, Dayuan
Cheng, Yang
Meng, Fandong
Yu, Mo
Zhou, Bowen
Zhou, Jie
author_facet Jiang, Che
Qi, Biqing
Hong, Xiangyu
Fu, Dayuan
Cheng, Yang
Meng, Fandong
Yu, Mo
Zhou, Bowen
Zhou, Jie
contents Large language models are successful in answering factoid questions but are also prone to hallucination. We investigate the phenomenon of LLMs possessing correct answer knowledge yet still hallucinating from the perspective of inference dynamics, an area not previously covered in studies on hallucinations. We are able to conduct this analysis via two key ideas. First, we identify the factual questions that query the same triplet knowledge but result in different answers. The difference between the model behaviors on the correct and incorrect outputs hence suggests the patterns when hallucinations happen. Second, to measure the pattern, we utilize mappings from the residual streams to vocabulary space. We reveal the different dynamics of the output token probabilities along the depths of layers between the correct and hallucinated cases. In hallucinated cases, the output token's information rarely demonstrates abrupt increases and consistent superiority in the later stages of the model. Leveraging the dynamic curve as a feature, we build a classifier capable of accurately detecting hallucinatory predictions with an 88\% success rate. Our study shed light on understanding the reasons for LLMs' hallucinations on their known facts, and more importantly, on accurately predicting when they are hallucinating.
format Preprint
id arxiv_https___arxiv_org_abs_2403_20009
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On Large Language Models' Hallucination with Regard to Known Facts
Jiang, Che
Qi, Biqing
Hong, Xiangyu
Fu, Dayuan
Cheng, Yang
Meng, Fandong
Yu, Mo
Zhou, Bowen
Zhou, Jie
Computation and Language
Machine Learning
Large language models are successful in answering factoid questions but are also prone to hallucination. We investigate the phenomenon of LLMs possessing correct answer knowledge yet still hallucinating from the perspective of inference dynamics, an area not previously covered in studies on hallucinations. We are able to conduct this analysis via two key ideas. First, we identify the factual questions that query the same triplet knowledge but result in different answers. The difference between the model behaviors on the correct and incorrect outputs hence suggests the patterns when hallucinations happen. Second, to measure the pattern, we utilize mappings from the residual streams to vocabulary space. We reveal the different dynamics of the output token probabilities along the depths of layers between the correct and hallucinated cases. In hallucinated cases, the output token's information rarely demonstrates abrupt increases and consistent superiority in the later stages of the model. Leveraging the dynamic curve as a feature, we build a classifier capable of accurately detecting hallucinatory predictions with an 88\% success rate. Our study shed light on understanding the reasons for LLMs' hallucinations on their known facts, and more importantly, on accurately predicting when they are hallucinating.
title On Large Language Models' Hallucination with Regard to Known Facts
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2403.20009