TriLens: Per-Layer Logit-Lens Entropy for White-Box Hallucination Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Bohan, Gong, Yijun, Zhang, Zhi, Zhang, Ge, Xing, Wenpeng, Han, Meng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911736388386816
author Yang, Bohan
Gong, Yijun
Zhang, Zhi
Zhang, Ge
Xing, Wenpeng
Han, Meng
author_facet Yang, Bohan
Gong, Yijun
Zhang, Zhi
Zhang, Ge
Xing, Wenpeng
Han, Meng
contents When a language model hallucinates, the final answer is wrong, but the mistake is not necessarily invisible inside the model. Different internal pathways may remain uncertain, disagree in how quickly they sharpen, or commit to competing continuations before the output is produced. We introduce TriLens, a white-box detector that turns this intuition into a compact representation: at every layer, it reads the multi-head self-attention output, the feed-forward output, and the residual stream through the model's own logit lens, then records only the entropy of each readout. The resulting 3L-dimensional trajectory describes how certainty forms across depth and across modules, without storing high-dimensional hidden states or sampling multiple generations. This simple signal yields a strong detector across instruction-tuned LLMs and QA benchmarks, and our analyses show that the three module-wise entropy trajectories provide complementary evidence. TriLens suggests that hallucination detection can benefit from tracking how internal computation settles, not only what the final layer predicts.
format Preprint
id arxiv_https___arxiv_org_abs_2606_01033
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TriLens: Per-Layer Logit-Lens Entropy for White-Box Hallucination Detection
Yang, Bohan
Gong, Yijun
Zhang, Zhi
Zhang, Ge
Xing, Wenpeng
Han, Meng
Artificial Intelligence
When a language model hallucinates, the final answer is wrong, but the mistake is not necessarily invisible inside the model. Different internal pathways may remain uncertain, disagree in how quickly they sharpen, or commit to competing continuations before the output is produced. We introduce TriLens, a white-box detector that turns this intuition into a compact representation: at every layer, it reads the multi-head self-attention output, the feed-forward output, and the residual stream through the model's own logit lens, then records only the entropy of each readout. The resulting 3L-dimensional trajectory describes how certainty forms across depth and across modules, without storing high-dimensional hidden states or sampling multiple generations. This simple signal yields a strong detector across instruction-tuned LLMs and QA benchmarks, and our analyses show that the three module-wise entropy trajectories provide complementary evidence. TriLens suggests that hallucination detection can benefit from tracking how internal computation settles, not only what the final layer predicts.
title TriLens: Per-Layer Logit-Lens Entropy for White-Box Hallucination Detection
topic Artificial Intelligence
url https://arxiv.org/abs/2606.01033