Saved in:
Bibliographic Details
Main Authors: Kim, Hazel, Lamb, Tom A., Bibi, Adel, Torr, Philip, Gal, Yarin
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2412.10246
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915533051396096
author Kim, Hazel
Lamb, Tom A.
Bibi, Adel
Torr, Philip
Gal, Yarin
author_facet Kim, Hazel
Lamb, Tom A.
Bibi, Adel
Torr, Philip
Gal, Yarin
contents Large language models (LLMs) frequently generate confident yet inaccurate responses, introducing significant risks for deployment in safety-critical domains. We present a novel, test-time approach to detecting model hallucination through systematic analysis of information flow across model layers. We target cases when LLMs process inputs with ambiguous or insufficient context. Our investigation reveals that hallucination manifests as usable information deficiencies in inter-layer transmissions. While existing approaches primarily focus on final-layer output analysis, we demonstrate that tracking cross-layer information dynamics ($\mathcal{L}$I) provides robust indicators of model reliability, accounting for both information gain and loss during computation. $\mathcal{L}$I integrates easily with pretrained LLMs without requiring additional training or architectural modifications.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10246
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Detecting LLM Hallucination Through Layer-wise Information Deficiency: Analysis of Ambiguous Prompts and Unanswerable Questions
Kim, Hazel
Lamb, Tom A.
Bibi, Adel
Torr, Philip
Gal, Yarin
Machine Learning
Large language models (LLMs) frequently generate confident yet inaccurate responses, introducing significant risks for deployment in safety-critical domains. We present a novel, test-time approach to detecting model hallucination through systematic analysis of information flow across model layers. We target cases when LLMs process inputs with ambiguous or insufficient context. Our investigation reveals that hallucination manifests as usable information deficiencies in inter-layer transmissions. While existing approaches primarily focus on final-layer output analysis, we demonstrate that tracking cross-layer information dynamics ($\mathcal{L}$I) provides robust indicators of model reliability, accounting for both information gain and loss during computation. $\mathcal{L}$I integrates easily with pretrained LLMs without requiring additional training or architectural modifications.
title Detecting LLM Hallucination Through Layer-wise Information Deficiency: Analysis of Ambiguous Prompts and Unanswerable Questions
topic Machine Learning
url https://arxiv.org/abs/2412.10246