INTRYGUE: Induction-Aware Entropy Gating for Reliable RAG Uncertainty Estimation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bazarova, Alexandra, Volodichev, Andrei, Kotova, Daria, Zaytsev, Alexey
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910065090363392
author Bazarova, Alexandra
Volodichev, Andrei
Kotova, Daria
Zaytsev, Alexey
author_facet Bazarova, Alexandra
Volodichev, Andrei
Kotova, Daria
Zaytsev, Alexey
contents While retrieval-augmented generation (RAG) significantly improves the factual reliability of LLMs, it does not eliminate hallucinations, so robust uncertainty quantification (UQ) remains essential. In this paper, we reveal that standard entropy-based UQ methods often fail in RAG settings due to a mechanistic paradox. An internal "tug-of-war" inherent to context utilization appears: while induction heads promote grounded responses by copying the correct answer, they collaterally trigger the previously established "entropy neurons". This interaction inflates predictive entropy, causing the model to signal false uncertainty on accurate outputs. To address this, we propose INTRYGUE (Induction-Aware Entropy Gating for Uncertainty Estimation), a mechanistically grounded method that gates predictive entropy based on the activation patterns of induction heads. Evaluated across four RAG benchmarks and six open-source LLMs (4B to 13B parameters), INTRYGUE consistently matches or outperforms a wide range of UQ baselines. Our findings demonstrate that hallucination detection in RAG benefits from combining predictive uncertainty with interpretable, internal signals of context utilization.
format Preprint
id arxiv_https___arxiv_org_abs_2603_21607
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle INTRYGUE: Induction-Aware Entropy Gating for Reliable RAG Uncertainty Estimation
Bazarova, Alexandra
Volodichev, Andrei
Kotova, Daria
Zaytsev, Alexey
Artificial Intelligence
While retrieval-augmented generation (RAG) significantly improves the factual reliability of LLMs, it does not eliminate hallucinations, so robust uncertainty quantification (UQ) remains essential. In this paper, we reveal that standard entropy-based UQ methods often fail in RAG settings due to a mechanistic paradox. An internal "tug-of-war" inherent to context utilization appears: while induction heads promote grounded responses by copying the correct answer, they collaterally trigger the previously established "entropy neurons". This interaction inflates predictive entropy, causing the model to signal false uncertainty on accurate outputs. To address this, we propose INTRYGUE (Induction-Aware Entropy Gating for Uncertainty Estimation), a mechanistically grounded method that gates predictive entropy based on the activation patterns of induction heads. Evaluated across four RAG benchmarks and six open-source LLMs (4B to 13B parameters), INTRYGUE consistently matches or outperforms a wide range of UQ baselines. Our findings demonstrate that hallucination detection in RAG benefits from combining predictive uncertainty with interpretable, internal signals of context utilization.
title INTRYGUE: Induction-Aware Entropy Gating for Reliable RAG Uncertainty Estimation
topic Artificial Intelligence
url https://arxiv.org/abs/2603.21607