Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Binkowski, Jakub, Adamczewski, Kamil, Kajdanowicz, Tomasz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hallucination Detection in LLMs Using Spectral Features of Attention Maps
von: Binkowski, Jakub, et al.
Veröffentlicht: (2025)
von: Binkowski, Jakub, et al.
Veröffentlicht: (2025)
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
von: Sawczyn, Albert, et al.
Veröffentlicht: (2025)
von: Sawczyn, Albert, et al.
Veröffentlicht: (2025)
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
Empowering Small-Scale Knowledge Graphs: A Strategy of Leveraging General-Purpose Knowledge Graphs for Enriched Embeddings
von: Sawczyn, Albert, et al.
Veröffentlicht: (2024)
von: Sawczyn, Albert, et al.
Veröffentlicht: (2024)
A Geometry-Based View of Mahalanobis OOD Detection
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
von: Peng, Runyu, et al.
Veröffentlicht: (2026)
von: Peng, Runyu, et al.
Veröffentlicht: (2026)
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
von: Son, Seungwoo, et al.
Veröffentlicht: (2024)
von: Son, Seungwoo, et al.
Veröffentlicht: (2024)
Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration
von: Yu, Zhongzhi, et al.
Veröffentlicht: (2024)
von: Yu, Zhongzhi, et al.
Veröffentlicht: (2024)
A Graph Signal Processing Framework for Hallucination Detection in Large Language Models
von: Noël, Valentin
Veröffentlicht: (2025)
von: Noël, Valentin
Veröffentlicht: (2025)
KG-Guard: Graph-Based Hallucination Detection for Knowledge Base Question Answering
von: Sawczyn, Albert, et al.
Veröffentlicht: (2026)
von: Sawczyn, Albert, et al.
Veröffentlicht: (2026)
Scaling Laws for Fine-Grained Mixture of Experts
von: Krajewski, Jakub, et al.
Veröffentlicht: (2024)
von: Krajewski, Jakub, et al.
Veröffentlicht: (2024)
When Attention Sink Emerges in Language Models: An Empirical View
von: Gu, Xiangming, et al.
Veröffentlicht: (2024)
von: Gu, Xiangming, et al.
Veröffentlicht: (2024)
Detecting Hallucinations in Large Language Models via Internal Attention Divergence Signals
von: van Dijk, Gijs
Veröffentlicht: (2026)
von: van Dijk, Gijs
Veröffentlicht: (2026)
Leveraging Graph Structures to Detect Hallucinations in Large Language Models
von: Nonkes, Noa, et al.
Veröffentlicht: (2024)
von: Nonkes, Noa, et al.
Veröffentlicht: (2024)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
von: Fu, Zizhuo, et al.
Veröffentlicht: (2026)
von: Fu, Zizhuo, et al.
Veröffentlicht: (2026)
Unmasking Hallucinations: A Causal Graph-Attention Perspective on Factual Reliability in Large Language Models
von: kurra, Sailesh kiran, et al.
Veröffentlicht: (2026)
von: kurra, Sailesh kiran, et al.
Veröffentlicht: (2026)
Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models
von: Huang, Yanwen, et al.
Veröffentlicht: (2025)
von: Huang, Yanwen, et al.
Veröffentlicht: (2025)
Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2024)
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2024)
TraceDet: Hallucination Detection from the Decoding Trace of Diffusion Large Language Models
von: Chang, Shenxu, et al.
Veröffentlicht: (2025)
von: Chang, Shenxu, et al.
Veröffentlicht: (2025)
Scalable Token-Level Hallucination Detection in Large Language Models
von: Min, Rui, et al.
Veröffentlicht: (2026)
von: Min, Rui, et al.
Veröffentlicht: (2026)
(Im)possibility of Automated Hallucination Detection in Large Language Models
von: Karbasi, Amin, et al.
Veröffentlicht: (2025)
von: Karbasi, Amin, et al.
Veröffentlicht: (2025)
On the Existence and Behavior of Secondary Attention Sinks
von: Wong, Jeffrey T. H., et al.
Veröffentlicht: (2025)
von: Wong, Jeffrey T. H., et al.
Veröffentlicht: (2025)
On Large Language Models' Hallucination with Regard to Known Facts
von: Jiang, Che, et al.
Veröffentlicht: (2024)
von: Jiang, Che, et al.
Veröffentlicht: (2024)
Projected Compression: Trainable Projection for Efficient Transformer Compression
von: Stefaniak, Maciej, et al.
Veröffentlicht: (2025)
von: Stefaniak, Maciej, et al.
Veröffentlicht: (2025)
Manifold-based Sampling for In-Context Hallucination Detection in Large Language Models
von: Vamshi, Bodla Krishna, et al.
Veröffentlicht: (2026)
von: Vamshi, Bodla Krishna, et al.
Veröffentlicht: (2026)
Principled Detection of Hallucinations in Large Language Models via Multiple Testing
von: Li, Jiawei, et al.
Veröffentlicht: (2025)
von: Li, Jiawei, et al.
Veröffentlicht: (2025)
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
von: Mündler, Niels, et al.
Veröffentlicht: (2023)
von: Mündler, Niels, et al.
Veröffentlicht: (2023)
Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
von: Shang, Bingqi, et al.
Veröffentlicht: (2025)
von: Shang, Bingqi, et al.
Veröffentlicht: (2025)
The Geometry of Tokens in Internal Representations of Large Language Models
von: Viswanathan, Karthik, et al.
Veröffentlicht: (2025)
von: Viswanathan, Karthik, et al.
Veröffentlicht: (2025)
Enhanced Structured State Space Models via Grouped FIR Filtering and Attention Sink Mechanisms
von: Meng, Tian, et al.
Veröffentlicht: (2024)
von: Meng, Tian, et al.
Veröffentlicht: (2024)
Sink-Aware Pruning for Diffusion Language Models
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2026)
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2026)
Attention Sinks in Massively Multilingual Neural Machine Translation:Discovery, Analysis, and Mitigation
von: Mutisya, Hillary, et al.
Veröffentlicht: (2026)
von: Mutisya, Hillary, et al.
Veröffentlicht: (2026)
SelfCheck-Eval: A Multi-Module Framework for Zero-Resource Hallucination Detection in Large Language Models
von: Muhammed, Diyana, et al.
Veröffentlicht: (2025)
von: Muhammed, Diyana, et al.
Veröffentlicht: (2025)
The HalluRAG Dataset: Detecting Closed-Domain Hallucinations in RAG Applications Using an LLM's Internal States
von: Ridder, Fabian, et al.
Veröffentlicht: (2024)
von: Ridder, Fabian, et al.
Veröffentlicht: (2024)
Anatomical Heterogeneity in Transformer Language Models
von: Wietrzykowski, Tomasz
Veröffentlicht: (2026)
von: Wietrzykowski, Tomasz
Veröffentlicht: (2026)
Mitigating Hallucinated Translations in Large Language Models with Hallucination-focused Preference Optimization
von: Tang, Zilu, et al.
Veröffentlicht: (2025)
von: Tang, Zilu, et al.
Veröffentlicht: (2025)
Unified Hallucination Detection for Multimodal Large Language Models
von: Chen, Xiang, et al.
Veröffentlicht: (2024)
von: Chen, Xiang, et al.
Veröffentlicht: (2024)
MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
von: Yamashita, Tomoya, et al.
Veröffentlicht: (2025)
von: Yamashita, Tomoya, et al.
Veröffentlicht: (2025)
Hallucinated Span Detection with Multi-View Attention Features
von: Ogasa, Yuya, et al.
Veröffentlicht: (2025)
von: Ogasa, Yuya, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hallucination Detection in LLMs Using Spectral Features of Attention Maps
von: Binkowski, Jakub, et al.
Veröffentlicht: (2025) -
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
von: Sawczyn, Albert, et al.
Veröffentlicht: (2025) -
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
von: Janiak, Denis, et al.
Veröffentlicht: (2025) -
Empowering Small-Scale Knowledge Graphs: A Strategy of Leveraging General-Purpose Knowledge Graphs for Enriched Embeddings
von: Sawczyn, Albert, et al.
Veröffentlicht: (2024) -
A Geometry-Based View of Mahalanobis OOD Detection
von: Janiak, Denis, et al.
Veröffentlicht: (2025)