Enhancing Hallucination Detection through Noise Injection

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Litian, Pourreza, Reza, Panchal, Sunny, Bhattacharyya, Apratim, Jian, Yubing, Qin, Yao, Memisevic, Roland
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914359539662848
author Liu, Litian
Pourreza, Reza
Panchal, Sunny
Bhattacharyya, Apratim
Jian, Yubing
Qin, Yao
Memisevic, Roland
author_facet Liu, Litian
Pourreza, Reza
Panchal, Sunny
Bhattacharyya, Apratim
Jian, Yubing
Qin, Yao
Memisevic, Roland
contents Large Language Models (LLMs) are prone to generating plausible yet incorrect responses, known as hallucinations. Effectively detecting hallucinations is therefore crucial for the safe deployment of LLMs. Recent research has linked hallucinations to model uncertainty, suggesting that hallucinations can be detected by measuring dispersion over answer distributions obtained from multiple samples drawn from a model. While drawing from the distribution over tokens defined by the model is a natural way to obtain samples, in this work, we argue that it is suboptimal for the purpose of detecting hallucinations. We show that detection can be improved significantly by taking into account model uncertainty in the Bayesian sense. To this end, we propose a very simple, training-free approach based on perturbing an appropriate subset of model parameters, or equivalently hidden unit activations, during sampling. We demonstrate that our approach significantly improves inference-time hallucination detection over standard sampling across diverse datasets, model architectures, and uncertainty metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2502_03799
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Hallucination Detection through Noise Injection
Liu, Litian
Pourreza, Reza
Panchal, Sunny
Bhattacharyya, Apratim
Jian, Yubing
Qin, Yao
Memisevic, Roland
Computation and Language
Systems and Control
Large Language Models (LLMs) are prone to generating plausible yet incorrect responses, known as hallucinations. Effectively detecting hallucinations is therefore crucial for the safe deployment of LLMs. Recent research has linked hallucinations to model uncertainty, suggesting that hallucinations can be detected by measuring dispersion over answer distributions obtained from multiple samples drawn from a model. While drawing from the distribution over tokens defined by the model is a natural way to obtain samples, in this work, we argue that it is suboptimal for the purpose of detecting hallucinations. We show that detection can be improved significantly by taking into account model uncertainty in the Bayesian sense. To this end, we propose a very simple, training-free approach based on perturbing an appropriate subset of model parameters, or equivalently hidden unit activations, during sampling. We demonstrate that our approach significantly improves inference-time hallucination detection over standard sampling across diverse datasets, model architectures, and uncertainty metrics.
title Enhancing Hallucination Detection through Noise Injection
topic Computation and Language
Systems and Control
url https://arxiv.org/abs/2502.03799