Why do LLMs attend to the first token?
Fuente:
arXiv
Saved in:
| Main Authors: | Barbero, Federico, Arroyo, Álvaro, Gu, Xiangming, Perivolaropoulos, Christos, Bronstein, Michael, Veličković, Petar, Pascanu, Razvan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Round and Round We Go! What makes Rotary Positional Encodings useful?
by: Barbero, Federico, et al.
Published: (2024)
by: Barbero, Federico, et al.
Published: (2024)
Perplexity Cannot Always Tell Right from Wrong
by: Veličković, Petar, et al.
Published: (2026)
by: Veličković, Petar, et al.
Published: (2026)
Softmax is not Enough (for Sharp Size Generalisation)
by: Veličković, Petar, et al.
Published: (2024)
by: Veličković, Petar, et al.
Published: (2024)
The Illusion of Stochasticity in LLMs
by: Gu, Xiangming, et al.
Published: (2026)
by: Gu, Xiangming, et al.
Published: (2026)
Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models
by: Gu, Xiangming, et al.
Published: (2026)
by: Gu, Xiangming, et al.
Published: (2026)
Filter Equivariant Functions: A symmetric account of length-general extrapolation on lists
by: Lewis, Owen, et al.
Published: (2025)
by: Lewis, Owen, et al.
Published: (2025)
Transformers need glasses! Information over-squashing in language tasks
by: Barbero, Federico, et al.
Published: (2024)
by: Barbero, Federico, et al.
Published: (2024)
How do LLMs Compute Verbal Confidence
by: Kumaran, Dharshan, et al.
Published: (2026)
by: Kumaran, Dharshan, et al.
Published: (2026)
Transformers meet Neural Algorithmic Reasoners
by: Bounsi, Wilfried, et al.
Published: (2024)
by: Bounsi, Wilfried, et al.
Published: (2024)
Latent Space Representations of Neural Algorithmic Reasoners
by: Mirjanić, Vladimir V., et al.
Published: (2023)
by: Mirjanić, Vladimir V., et al.
Published: (2023)
Building on Efficient Foundations: Effectively Training LLMs with Structured Feedforward Layers
by: Wei, Xiuying, et al.
Published: (2024)
by: Wei, Xiuying, et al.
Published: (2024)
Asynchronous Algorithmic Alignment with Cocycles
by: Dudzik, Andrew, et al.
Published: (2023)
by: Dudzik, Andrew, et al.
Published: (2023)
RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling
by: Wei, Xiuying, et al.
Published: (2025)
by: Wei, Xiuying, et al.
Published: (2025)
Investigating Low-Rank Training in Transformer Language Models: Efficiency and Scaling Analysis
by: Wei, Xiuying, et al.
Published: (2024)
by: Wei, Xiuying, et al.
Published: (2024)
How do language models learn facts? Dynamics, curricula and hallucinations
by: Zucchet, Nicolas, et al.
Published: (2025)
by: Zucchet, Nicolas, et al.
Published: (2025)
Mining Generalizable Activation Functions
by: Vitvitskyi, Alex, et al.
Published: (2026)
by: Vitvitskyi, Alex, et al.
Published: (2026)
Interpreting token compositionality in LLMs: A robustness analysis
by: Aljaafari, Nura, et al.
Published: (2024)
by: Aljaafari, Nura, et al.
Published: (2024)
LBPE: Long-token-first Tokenization to Improve Large Language Models
by: Lian, Haoran, et al.
Published: (2024)
by: Lian, Haoran, et al.
Published: (2024)
LLMs can hide text in other text of the same length
by: Norelli, Antonio, et al.
Published: (2025)
by: Norelli, Antonio, et al.
Published: (2025)
Prediction hubs are context-informed frequent tokens in LLMs
by: Nielsen, Beatrix M. G., et al.
Published: (2025)
by: Nielsen, Beatrix M. G., et al.
Published: (2025)
AnomaLLMy -- Detecting anomalous tokens in black-box LLMs through low-confidence single-token predictions
by: Witold, Waligóra
Published: (2024)
by: Witold, Waligóra
Published: (2024)
Jacobian Scopes: token-level causal attributions in LLMs
by: Liu, Toni J. B., et al.
Published: (2026)
by: Liu, Toni J. B., et al.
Published: (2026)
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
by: Singh, Aaditya K., et al.
Published: (2024)
by: Singh, Aaditya K., et al.
Published: (2024)
Finetuning LLMs for EvaCun 2025 token prediction shared task
by: Jon, Josef, et al.
Published: (2025)
by: Jon, Josef, et al.
Published: (2025)
Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations
by: Zhao, Yize, et al.
Published: (2024)
by: Zhao, Yize, et al.
Published: (2024)
An evaluation of LLM code generation capabilities through graded exercises
by: Jiménez, Álvaro Barbero
Published: (2024)
by: Jiménez, Álvaro Barbero
Published: (2024)
Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
by: Shrestha, Adarsha, et al.
Published: (2025)
by: Shrestha, Adarsha, et al.
Published: (2025)
Where is the signal in tokenization space?
by: Geh, Renato Lui, et al.
Published: (2024)
by: Geh, Renato Lui, et al.
Published: (2024)
Comparative analysis of subword tokenization approaches for Indian languages
by: Das, Sudhansu Bala, et al.
Published: (2025)
by: Das, Sudhansu Bala, et al.
Published: (2025)
Contextual morphologically-guided tokenization for Latin encoder models
by: Hudspeth, Marisa, et al.
Published: (2025)
by: Hudspeth, Marisa, et al.
Published: (2025)
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
by: Xu, Yijie, et al.
Published: (2025)
by: Xu, Yijie, et al.
Published: (2025)
Revisiting Dynamic Evaluation: Online Adaptation for Large Language Models
by: Rannen-Triki, Amal, et al.
Published: (2024)
by: Rannen-Triki, Amal, et al.
Published: (2024)
Why are LLMs' abilities emergent?
by: Havlík, Vladimír
Published: (2025)
by: Havlík, Vladimír
Published: (2025)
Looking beyond the next token
by: Thankaraj, Abitha, et al.
Published: (2025)
by: Thankaraj, Abitha, et al.
Published: (2025)
The pitfalls of next-token prediction
by: Bachmann, Gregor, et al.
Published: (2024)
by: Bachmann, Gregor, et al.
Published: (2024)
On multi-token prediction for efficient LLM inference
by: Mehra, Somesh, et al.
Published: (2025)
by: Mehra, Somesh, et al.
Published: (2025)
Amplifying human performance in combinatorial competitive programming
by: Veličković, Petar, et al.
Published: (2024)
by: Veličković, Petar, et al.
Published: (2024)
Better & Faster Large Language Models via Multi-token Prediction
by: Gloeckle, Fabian, et al.
Published: (2024)
by: Gloeckle, Fabian, et al.
Published: (2024)
Collaborative decoding of critical tokens for boosting factuality of large language models
by: Jin, Lifeng, et al.
Published: (2024)
by: Jin, Lifeng, et al.
Published: (2024)
Getting the most out of your tokenizer for pre-training and domain adaptation
by: Dagan, Gautier, et al.
Published: (2024)
by: Dagan, Gautier, et al.
Published: (2024)
Similar Items
-
Round and Round We Go! What makes Rotary Positional Encodings useful?
by: Barbero, Federico, et al.
Published: (2024) -
Perplexity Cannot Always Tell Right from Wrong
by: Veličković, Petar, et al.
Published: (2026) -
Softmax is not Enough (for Sharp Size Generalisation)
by: Veličković, Petar, et al.
Published: (2024) -
The Illusion of Stochasticity in LLMs
by: Gu, Xiangming, et al.
Published: (2026) -
Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models
by: Gu, Xiangming, et al.
Published: (2026)