Quantifying Genuine Awareness in Hallucination Prediction Beyond Question-Side Shortcuts

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Seo, Yeongbin, Lee, Dongha, Yeo, Jinyoung
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910046823120896
author Seo, Yeongbin
Lee, Dongha
Yeo, Jinyoung
author_facet Seo, Yeongbin
Lee, Dongha
Yeo, Jinyoung
contents Many works have proposed methodologies for language model (LM) hallucination detection and reported seemingly strong performance. However, we argue that the reported performance to date reflects not only a model's genuine awareness of its internal information, but also awareness derived purely from question-side information (e.g., benchmark hacking). While benchmark hacking can be effective for boosting hallucination detection score on existing benchmarks, it does not generalize to out-of-domain settings and practical usage. Nevertheless, disentangling how much of a model's hallucination detection performance arises from question-side awareness is non-trivial. To address this, we propose a methodology for measuring this effect without requiring human labor, Approximate Question-side Effect (AQE). Our analysis using AQE reveals that existing hallucination detection methods rely heavily on benchmark hacking.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15339
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Quantifying Genuine Awareness in Hallucination Prediction Beyond Question-Side Shortcuts
Seo, Yeongbin
Lee, Dongha
Yeo, Jinyoung
Computation and Language
68T50
I.2.7
Many works have proposed methodologies for language model (LM) hallucination detection and reported seemingly strong performance. However, we argue that the reported performance to date reflects not only a model's genuine awareness of its internal information, but also awareness derived purely from question-side information (e.g., benchmark hacking). While benchmark hacking can be effective for boosting hallucination detection score on existing benchmarks, it does not generalize to out-of-domain settings and practical usage. Nevertheless, disentangling how much of a model's hallucination detection performance arises from question-side awareness is non-trivial. To address this, we propose a methodology for measuring this effect without requiring human labor, Approximate Question-side Effect (AQE). Our analysis using AQE reveals that existing hallucination detection methods rely heavily on benchmark hacking.
title Quantifying Genuine Awareness in Hallucination Prediction Beyond Question-Side Shortcuts
topic Computation and Language
68T50
I.2.7
url https://arxiv.org/abs/2509.15339