HalluShift: Measuring Distribution Shifts towards Hallucination Detection in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dasgupta, Sharanya, Nath, Sujoy, Basu, Arkaprabha, Shamsolmoali, Pourya, Das, Swagatam
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908317080616960
author Dasgupta, Sharanya
Nath, Sujoy
Basu, Arkaprabha
Shamsolmoali, Pourya
Das, Swagatam
author_facet Dasgupta, Sharanya
Nath, Sujoy
Basu, Arkaprabha
Shamsolmoali, Pourya
Das, Swagatam
contents Large Language Models (LLMs) have recently garnered widespread attention due to their adeptness at generating innovative responses to the given prompts across a multitude of domains. However, LLMs often suffer from the inherent limitation of hallucinations and generate incorrect information while maintaining well-structured and coherent responses. In this work, we hypothesize that hallucinations stem from the internal dynamics of LLMs. Our observations indicate that, during passage generation, LLMs tend to deviate from factual accuracy in subtle parts of responses, eventually shifting toward misinformation. This phenomenon bears a resemblance to human cognition, where individuals may hallucinate while maintaining logical coherence, embedding uncertainty within minor segments of their speech. To investigate this further, we introduce an innovative approach, HalluShift, designed to analyze the distribution shifts in the internal state space and token probabilities of the LLM-generated responses. Our method attains superior performance compared to existing baselines across various benchmark datasets. Our codebase is available at https://github.com/sharanya-dasgupta001/hallushift.
format Preprint
id arxiv_https___arxiv_org_abs_2504_09482
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HalluShift: Measuring Distribution Shifts towards Hallucination Detection in LLMs
Dasgupta, Sharanya
Nath, Sujoy
Basu, Arkaprabha
Shamsolmoali, Pourya
Das, Swagatam
Computation and Language
Artificial Intelligence
Emerging Technologies
Large Language Models (LLMs) have recently garnered widespread attention due to their adeptness at generating innovative responses to the given prompts across a multitude of domains. However, LLMs often suffer from the inherent limitation of hallucinations and generate incorrect information while maintaining well-structured and coherent responses. In this work, we hypothesize that hallucinations stem from the internal dynamics of LLMs. Our observations indicate that, during passage generation, LLMs tend to deviate from factual accuracy in subtle parts of responses, eventually shifting toward misinformation. This phenomenon bears a resemblance to human cognition, where individuals may hallucinate while maintaining logical coherence, embedding uncertainty within minor segments of their speech. To investigate this further, we introduce an innovative approach, HalluShift, designed to analyze the distribution shifts in the internal state space and token probabilities of the LLM-generated responses. Our method attains superior performance compared to existing baselines across various benchmark datasets. Our codebase is available at https://github.com/sharanya-dasgupta001/hallushift.
title HalluShift: Measuring Distribution Shifts towards Hallucination Detection in LLMs
topic Computation and Language
Artificial Intelligence
Emerging Technologies
url https://arxiv.org/abs/2504.09482