Attention Head Entropy of LLMs Predicts Answer Correctness

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ostmeier, Sophie, Axelrod, Brian, Varma, Maya, Aali, Asad, Zhang, Yabin, Paschali, Magdalini, Koyejo, Sanmi, Langlotz, Curtis, Chaudhari, Akshay
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917274689994752
author Ostmeier, Sophie
Axelrod, Brian
Varma, Maya
Aali, Asad
Zhang, Yabin
Paschali, Magdalini
Koyejo, Sanmi
Langlotz, Curtis
Chaudhari, Akshay
author_facet Ostmeier, Sophie
Axelrod, Brian
Varma, Maya
Aali, Asad
Zhang, Yabin
Paschali, Magdalini
Koyejo, Sanmi
Langlotz, Curtis
Chaudhari, Akshay
contents Large language models (LLMs) often generate plausible yet incorrect answers, posing risks in safety-critical settings such as medicine. Human evaluation is expensive, and LLM-as-judge approaches risk introducing hidden errors. Recent white-box methods detect contextual hallucinations using model internals, focusing on the localization of the attention mass, but two questions remain open: do these approaches extend to predicting answer correctness, and do they generalize out-of-domains? We introduce Head Entropy, a method that predicts answer correctness from attention entropy patterns, specifically measuring the spread of the attention mass. Using sparse logistic regression on per-head 2-Renyi entropies, Head Entropy matches or exceeds baselines in-distribution and generalizes substantially better on out-of-domains, it outperforms the closest baseline on average by +8.5% AUROC. We further show that attention patterns over the question/context alone, before answer generation, already carry predictive signal using Head Entropy with on average +17.7% AUROC over the closest baseline. We evaluate across 5 instruction-tuned LLMs and 3 QA datasets spanning general knowledge, multi-hop reasoning, and medicine.
format Preprint
id arxiv_https___arxiv_org_abs_2602_13699
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Attention Head Entropy of LLMs Predicts Answer Correctness
Ostmeier, Sophie
Axelrod, Brian
Varma, Maya
Aali, Asad
Zhang, Yabin
Paschali, Magdalini
Koyejo, Sanmi
Langlotz, Curtis
Chaudhari, Akshay
Machine Learning
Large language models (LLMs) often generate plausible yet incorrect answers, posing risks in safety-critical settings such as medicine. Human evaluation is expensive, and LLM-as-judge approaches risk introducing hidden errors. Recent white-box methods detect contextual hallucinations using model internals, focusing on the localization of the attention mass, but two questions remain open: do these approaches extend to predicting answer correctness, and do they generalize out-of-domains? We introduce Head Entropy, a method that predicts answer correctness from attention entropy patterns, specifically measuring the spread of the attention mass. Using sparse logistic regression on per-head 2-Renyi entropies, Head Entropy matches or exceeds baselines in-distribution and generalizes substantially better on out-of-domains, it outperforms the closest baseline on average by +8.5% AUROC. We further show that attention patterns over the question/context alone, before answer generation, already carry predictive signal using Head Entropy with on average +17.7% AUROC over the closest baseline. We evaluate across 5 instruction-tuned LLMs and 3 QA datasets spanning general knowledge, multi-hop reasoning, and medicine.
title Attention Head Entropy of LLMs Predicts Answer Correctness
topic Machine Learning
url https://arxiv.org/abs/2602.13699