HaluNet: Learning Hallucination Risk from Internal Signals in LLM Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tong, Chaodong, Zhang, Qi, Jiang, Zhuojun, Jiang, Lei, Liu, Yanbing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918528550961152
author Tong, Chaodong
Zhang, Qi
Jiang, Zhuojun
Jiang, Lei
Liu, Yanbing
author_facet Tong, Chaodong
Zhang, Qi
Jiang, Zhuojun
Jiang, Lei
Liu, Yanbing
contents Large language models (LLMs) achieve strong question answering (QA) performance but can produce fluent answers unsupported by available evidence. Existing hallucination detectors often rely on external verification, repeated sampling, or test-time judge calls, which can be costly for real-time QA. We propose \textbf{HaluNet}, a lightweight hallucination risk estimator that uses internal signals from one model generation. HaluNet jointly models token likelihood, predictive entropy, and hidden-state information, allowing probabilistic, distributional, and semantic evidence to inform an answer-level risk score. It is trained with LLM-as-a-Judge labels as scalable weak supervision and evaluated with independent human and multi-judge assessments. Experiments on SQuAD, TriviaQA, and Natural Questions show that HaluNet improves answer-level risk ranking across in-domain and out-of-domain settings. On a 300-example human evaluation, HaluNet achieves 0.874 AUROC and 0.869 AUPRC; its top 20\% highest-risk answers contain 96.5\% errors, yielding a 2.06$\times$ lift over the base error rate.
format Preprint
id arxiv_https___arxiv_org_abs_2512_24562
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HaluNet: Learning Hallucination Risk from Internal Signals in LLM Question Answering
Tong, Chaodong
Zhang, Qi
Jiang, Zhuojun
Jiang, Lei
Liu, Yanbing
Computation and Language
Large language models (LLMs) achieve strong question answering (QA) performance but can produce fluent answers unsupported by available evidence. Existing hallucination detectors often rely on external verification, repeated sampling, or test-time judge calls, which can be costly for real-time QA. We propose \textbf{HaluNet}, a lightweight hallucination risk estimator that uses internal signals from one model generation. HaluNet jointly models token likelihood, predictive entropy, and hidden-state information, allowing probabilistic, distributional, and semantic evidence to inform an answer-level risk score. It is trained with LLM-as-a-Judge labels as scalable weak supervision and evaluated with independent human and multi-judge assessments. Experiments on SQuAD, TriviaQA, and Natural Questions show that HaluNet improves answer-level risk ranking across in-domain and out-of-domain settings. On a 300-example human evaluation, HaluNet achieves 0.874 AUROC and 0.869 AUPRC; its top 20\% highest-risk answers contain 96.5\% errors, yielding a 2.06$\times$ lift over the base error rate.
title HaluNet: Learning Hallucination Risk from Internal Signals in LLM Question Answering
topic Computation and Language
url https://arxiv.org/abs/2512.24562