AggTruth: Contextual Hallucination Detection using Aggregated Attention Scores in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Matys, Piotr, Eliasz, Jan, Kiełczyński, Konrad, Langner, Mikołaj, Ferdinan, Teddy, Kocoń, Jan, Kazienko, Przemysław
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916806969524224
author Matys, Piotr
Eliasz, Jan
Kiełczyński, Konrad
Langner, Mikołaj
Ferdinan, Teddy
Kocoń, Jan
Kazienko, Przemysław
author_facet Matys, Piotr
Eliasz, Jan
Kiełczyński, Konrad
Langner, Mikołaj
Ferdinan, Teddy
Kocoń, Jan
Kazienko, Przemysław
contents In real-world applications, Large Language Models (LLMs) often hallucinate, even in Retrieval-Augmented Generation (RAG) settings, which poses a significant challenge to their deployment. In this paper, we introduce AggTruth, a method for online detection of contextual hallucinations by analyzing the distribution of internal attention scores in the provided context (passage). Specifically, we propose four different variants of the method, each varying in the aggregation technique used to calculate attention scores. Across all LLMs examined, AggTruth demonstrated stable performance in both same-task and cross-task setups, outperforming the current SOTA in multiple scenarios. Furthermore, we conducted an in-depth analysis of feature selection techniques and examined how the number of selected attention heads impacts detection performance, demonstrating that careful selection of heads is essential to achieve optimal results.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18628
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AggTruth: Contextual Hallucination Detection using Aggregated Attention Scores in LLMs
Matys, Piotr
Eliasz, Jan
Kiełczyński, Konrad
Langner, Mikołaj
Ferdinan, Teddy
Kocoń, Jan
Kazienko, Przemysław
Artificial Intelligence
Computation and Language
In real-world applications, Large Language Models (LLMs) often hallucinate, even in Retrieval-Augmented Generation (RAG) settings, which poses a significant challenge to their deployment. In this paper, we introduce AggTruth, a method for online detection of contextual hallucinations by analyzing the distribution of internal attention scores in the provided context (passage). Specifically, we propose four different variants of the method, each varying in the aggregation technique used to calculate attention scores. Across all LLMs examined, AggTruth demonstrated stable performance in both same-task and cross-task setups, outperforming the current SOTA in multiple scenarios. Furthermore, we conducted an in-depth analysis of feature selection techniques and examined how the number of selected attention heads impacts detection performance, demonstrating that careful selection of heads is essential to achieve optimal results.
title AggTruth: Contextual Hallucination Detection using Aggregated Attention Scores in LLMs
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.18628