From Out-of-Distribution Detection to Hallucination Detection: A Geometric View

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Litian, Pourreza, Reza, Jian, Yubing, Qin, Yao, Memisevic, Roland
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911730388434944
author Liu, Litian
Pourreza, Reza
Jian, Yubing
Qin, Yao
Memisevic, Roland
author_facet Liu, Litian
Pourreza, Reza
Jian, Yubing
Qin, Yao
Memisevic, Roland
contents Detecting hallucinations in large language models is a critical open problem with significant implications for safety and reliability. While existing hallucination detection methods achieve strong performance in question-answering tasks, they remain less effective on tasks requiring reasoning. In this work, we revisit hallucination detection through the lens of out-of-distribution (OOD) detection, a well-studied problem in areas like computer vision. Treating next-token prediction in language models as a classification task allows us to apply OOD techniques, provided appropriate modifications are made to account for the structural differences in large language models. We show that OOD-based approaches yield training-free, single-sample-based detectors, achieving strong accuracy in hallucination detection for reasoning tasks. Overall, our work suggests that reframing hallucination detection as OOD detection provides a promising and scalable pathway toward language model safety.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07253
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Out-of-Distribution Detection to Hallucination Detection: A Geometric View
Liu, Litian
Pourreza, Reza
Jian, Yubing
Qin, Yao
Memisevic, Roland
Artificial Intelligence
Computation and Language
Detecting hallucinations in large language models is a critical open problem with significant implications for safety and reliability. While existing hallucination detection methods achieve strong performance in question-answering tasks, they remain less effective on tasks requiring reasoning. In this work, we revisit hallucination detection through the lens of out-of-distribution (OOD) detection, a well-studied problem in areas like computer vision. Treating next-token prediction in language models as a classification task allows us to apply OOD techniques, provided appropriate modifications are made to account for the structural differences in large language models. We show that OOD-based approaches yield training-free, single-sample-based detectors, achieving strong accuracy in hallucination detection for reasoning tasks. Overall, our work suggests that reframing hallucination detection as OOD detection provides a promising and scalable pathway toward language model safety.
title From Out-of-Distribution Detection to Hallucination Detection: A Geometric View
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2602.07253