Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bar-Shalom, Guy, Frasca, Fabrizio, Galron, Yaniv, Ziser, Yftah, Maron, Haggai
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908571298430976
author Bar-Shalom, Guy
Frasca, Fabrizio
Galron, Yaniv
Ziser, Yftah
Maron, Haggai
author_facet Bar-Shalom, Guy
Frasca, Fabrizio
Galron, Yaniv
Ziser, Yftah
Maron, Haggai
contents Detecting hallucinations in Large Language Model-generated text is crucial for their safe deployment. While probing classifiers show promise, they operate on isolated layer-token pairs and are LLM-specific, limiting their effectiveness and hindering cross-LLM applications. In this paper, we introduce a novel approach to address these shortcomings. We build on the natural sequential structure of activation data in both axes (layers $\times$ tokens) and advocate treating full activation tensors akin to images. We design ACT-ViT, a Vision Transformer-inspired model that can be effectively and efficiently applied to activation tensors and supports training on data from multiple LLMs simultaneously. Through comprehensive experiments encompassing diverse LLMs and datasets, we demonstrate that ACT-ViT consistently outperforms traditional probing techniques while remaining extremely efficient for deployment. In particular, we show that our architecture benefits substantially from multi-LLM training, achieves strong zero-shot performance on unseen datasets, and can be transferred effectively to new LLMs through fine-tuning. Full code is available at https://github.com/BarSGuy/ACT-ViT.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00296
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT
Bar-Shalom, Guy
Frasca, Fabrizio
Galron, Yaniv
Ziser, Yftah
Maron, Haggai
Machine Learning
Detecting hallucinations in Large Language Model-generated text is crucial for their safe deployment. While probing classifiers show promise, they operate on isolated layer-token pairs and are LLM-specific, limiting their effectiveness and hindering cross-LLM applications. In this paper, we introduce a novel approach to address these shortcomings. We build on the natural sequential structure of activation data in both axes (layers $\times$ tokens) and advocate treating full activation tensors akin to images. We design ACT-ViT, a Vision Transformer-inspired model that can be effectively and efficiently applied to activation tensors and supports training on data from multiple LLMs simultaneously. Through comprehensive experiments encompassing diverse LLMs and datasets, we demonstrate that ACT-ViT consistently outperforms traditional probing techniques while remaining extremely efficient for deployment. In particular, we show that our architecture benefits substantially from multi-LLM training, achieves strong zero-shot performance on unseen datasets, and can be transferred effectively to new LLMs through fine-tuning. Full code is available at https://github.com/BarSGuy/ACT-ViT.
title Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT
topic Machine Learning
url https://arxiv.org/abs/2510.00296