Neural Message-Passing on Attention Graphs for Hallucination Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Frasca, Fabrizio, Bar-Shalom, Guy, Ziser, Yftah, Maron, Haggai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909814680977408
author Frasca, Fabrizio
Bar-Shalom, Guy
Ziser, Yftah
Maron, Haggai
author_facet Frasca, Fabrizio
Bar-Shalom, Guy
Ziser, Yftah
Maron, Haggai
contents Large Language Models (LLMs) often generate incorrect or unsupported content, known as hallucinations. Existing detection methods rely on heuristics or simple models over isolated computational traces such as activations, or attention maps. We unify these signals by representing them as attributed graphs, where tokens are nodes, edges follow attentional flows, and both carry features from attention scores and activations. Our approach, CHARM, casts hallucination detection as a graph learning task and tackles it by applying GNNs over the above attributed graphs. We show that CHARM provably subsumes prior attention-based heuristics and, experimentally, it consistently outperforms other leading approaches across diverse benchmarks. Our results shed light on the relevant role played by the graph structure and on the benefits of combining computational traces, whilst showing CHARM exhibits promising zero-shot performance on cross-dataset transfer.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24770
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Neural Message-Passing on Attention Graphs for Hallucination Detection
Frasca, Fabrizio
Bar-Shalom, Guy
Ziser, Yftah
Maron, Haggai
Machine Learning
Large Language Models (LLMs) often generate incorrect or unsupported content, known as hallucinations. Existing detection methods rely on heuristics or simple models over isolated computational traces such as activations, or attention maps. We unify these signals by representing them as attributed graphs, where tokens are nodes, edges follow attentional flows, and both carry features from attention scores and activations. Our approach, CHARM, casts hallucination detection as a graph learning task and tackles it by applying GNNs over the above attributed graphs. We show that CHARM provably subsumes prior attention-based heuristics and, experimentally, it consistently outperforms other leading approaches across diverse benchmarks. Our results shed light on the relevant role played by the graph structure and on the benefits of combining computational traces, whilst showing CHARM exhibits promising zero-shot performance on cross-dataset transfer.
title Neural Message-Passing on Attention Graphs for Hallucination Detection
topic Machine Learning
url https://arxiv.org/abs/2509.24770