Attention Head Entropy of LLMs Predicts Answer Correctness
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917274689994752 |
|---|---|
| author | Ostmeier, Sophie Axelrod, Brian Varma, Maya Aali, Asad Zhang, Yabin Paschali, Magdalini Koyejo, Sanmi Langlotz, Curtis Chaudhari, Akshay |
| author_facet | Ostmeier, Sophie Axelrod, Brian Varma, Maya Aali, Asad Zhang, Yabin Paschali, Magdalini Koyejo, Sanmi Langlotz, Curtis Chaudhari, Akshay |
| contents | Large language models (LLMs) often generate plausible yet incorrect answers, posing risks in safety-critical settings such as medicine. Human evaluation is expensive, and LLM-as-judge approaches risk introducing hidden errors. Recent white-box methods detect contextual hallucinations using model internals, focusing on the localization of the attention mass, but two questions remain open: do these approaches extend to predicting answer correctness, and do they generalize out-of-domains? We introduce Head Entropy, a method that predicts answer correctness from attention entropy patterns, specifically measuring the spread of the attention mass. Using sparse logistic regression on per-head 2-Renyi entropies, Head Entropy matches or exceeds baselines in-distribution and generalizes substantially better on out-of-domains, it outperforms the closest baseline on average by +8.5% AUROC. We further show that attention patterns over the question/context alone, before answer generation, already carry predictive signal using Head Entropy with on average +17.7% AUROC over the closest baseline. We evaluate across 5 instruction-tuned LLMs and 3 QA datasets spanning general knowledge, multi-hop reasoning, and medicine. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_13699 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Attention Head Entropy of LLMs Predicts Answer Correctness Ostmeier, Sophie Axelrod, Brian Varma, Maya Aali, Asad Zhang, Yabin Paschali, Magdalini Koyejo, Sanmi Langlotz, Curtis Chaudhari, Akshay Machine Learning Large language models (LLMs) often generate plausible yet incorrect answers, posing risks in safety-critical settings such as medicine. Human evaluation is expensive, and LLM-as-judge approaches risk introducing hidden errors. Recent white-box methods detect contextual hallucinations using model internals, focusing on the localization of the attention mass, but two questions remain open: do these approaches extend to predicting answer correctness, and do they generalize out-of-domains? We introduce Head Entropy, a method that predicts answer correctness from attention entropy patterns, specifically measuring the spread of the attention mass. Using sparse logistic regression on per-head 2-Renyi entropies, Head Entropy matches or exceeds baselines in-distribution and generalizes substantially better on out-of-domains, it outperforms the closest baseline on average by +8.5% AUROC. We further show that attention patterns over the question/context alone, before answer generation, already carry predictive signal using Head Entropy with on average +17.7% AUROC over the closest baseline. We evaluate across 5 instruction-tuned LLMs and 3 QA datasets spanning general knowledge, multi-hop reasoning, and medicine. |
| title | Attention Head Entropy of LLMs Predicts Answer Correctness |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2602.13699 |