From Out-of-Distribution Detection to Hallucination Detection: A Geometric View
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866911730388434944 |
|---|---|
| author | Liu, Litian Pourreza, Reza Jian, Yubing Qin, Yao Memisevic, Roland |
| author_facet | Liu, Litian Pourreza, Reza Jian, Yubing Qin, Yao Memisevic, Roland |
| contents | Detecting hallucinations in large language models is a critical open problem with significant implications for safety and reliability. While existing hallucination detection methods achieve strong performance in question-answering tasks, they remain less effective on tasks requiring reasoning. In this work, we revisit hallucination detection through the lens of out-of-distribution (OOD) detection, a well-studied problem in areas like computer vision. Treating next-token prediction in language models as a classification task allows us to apply OOD techniques, provided appropriate modifications are made to account for the structural differences in large language models. We show that OOD-based approaches yield training-free, single-sample-based detectors, achieving strong accuracy in hallucination detection for reasoning tasks. Overall, our work suggests that reframing hallucination detection as OOD detection provides a promising and scalable pathway toward language model safety. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_07253 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | From Out-of-Distribution Detection to Hallucination Detection: A Geometric View Liu, Litian Pourreza, Reza Jian, Yubing Qin, Yao Memisevic, Roland Artificial Intelligence Computation and Language Detecting hallucinations in large language models is a critical open problem with significant implications for safety and reliability. While existing hallucination detection methods achieve strong performance in question-answering tasks, they remain less effective on tasks requiring reasoning. In this work, we revisit hallucination detection through the lens of out-of-distribution (OOD) detection, a well-studied problem in areas like computer vision. Treating next-token prediction in language models as a classification task allows us to apply OOD techniques, provided appropriate modifications are made to account for the structural differences in large language models. We show that OOD-based approaches yield training-free, single-sample-based detectors, achieving strong accuracy in hallucination detection for reasoning tasks. Overall, our work suggests that reframing hallucination detection as OOD detection provides a promising and scalable pathway toward language model safety. |
| title | From Out-of-Distribution Detection to Hallucination Detection: A Geometric View |
| topic | Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2602.07253 |