Transformers need glasses! Information over-squashing in language tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Barbero, Federico, Banino, Andrea, Kapturowski, Steven, Kumaran, Dharshan, Araújo, João G. M., Vitvitskyi, Alex, Pascanu, Razvan, Veličković, Petar |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Round and Round We Go! What makes Rotary Positional Encodings useful?
por: Barbero, Federico, et al.
Publicado: (2024)
por: Barbero, Federico, et al.
Publicado: (2024)
Perplexity Cannot Always Tell Right from Wrong
por: Veličković, Petar, et al.
Publicado: (2026)
por: Veličković, Petar, et al.
Publicado: (2026)
How do LLMs Compute Verbal Confidence
por: Kumaran, Dharshan, et al.
Publicado: (2026)
por: Kumaran, Dharshan, et al.
Publicado: (2026)
Transformers meet Neural Algorithmic Reasoners
por: Bounsi, Wilfried, et al.
Publicado: (2024)
por: Bounsi, Wilfried, et al.
Publicado: (2024)
Softmax is not Enough (for Sharp Size Generalisation)
por: Veličković, Petar, et al.
Publicado: (2024)
por: Veličković, Petar, et al.
Publicado: (2024)
Why do LLMs attend to the first token?
por: Barbero, Federico, et al.
Publicado: (2025)
por: Barbero, Federico, et al.
Publicado: (2025)
Mining Generalizable Activation Functions
por: Vitvitskyi, Alex, et al.
Publicado: (2026)
por: Vitvitskyi, Alex, et al.
Publicado: (2026)
Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models
por: Gu, Xiangming, et al.
Publicado: (2026)
por: Gu, Xiangming, et al.
Publicado: (2026)
How Overconfidence in Initial Choices and Underconfidence Under Criticism Modulate Change of Mind in Large Language Models
por: Kumaran, Dharshan, et al.
Publicado: (2025)
por: Kumaran, Dharshan, et al.
Publicado: (2025)
The Illusion of Stochasticity in LLMs
por: Gu, Xiangming, et al.
Publicado: (2026)
por: Gu, Xiangming, et al.
Publicado: (2026)
Filter Equivariant Functions: A symmetric account of length-general extrapolation on lists
por: Lewis, Owen, et al.
Publicado: (2025)
por: Lewis, Owen, et al.
Publicado: (2025)
What makes a good feedforward computational graph?
por: Vitvitskyi, Alex, et al.
Publicado: (2025)
por: Vitvitskyi, Alex, et al.
Publicado: (2025)
Latent Space Representations of Neural Algorithmic Reasoners
por: Mirjanić, Vladimir V., et al.
Publicado: (2023)
por: Mirjanić, Vladimir V., et al.
Publicado: (2023)
How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals
por: Kumaran, Dharshan, et al.
Publicado: (2026)
por: Kumaran, Dharshan, et al.
Publicado: (2026)
Causal Evidence that Language Models use Confidence to Drive Behavior
por: Kumaran, Dharshan, et al.
Publicado: (2026)
por: Kumaran, Dharshan, et al.
Publicado: (2026)
Amplifying human performance in combinatorial competitive programming
por: Veličković, Petar, et al.
Publicado: (2024)
por: Veličković, Petar, et al.
Publicado: (2024)
Asynchronous Algorithmic Alignment with Cocycles
por: Dudzik, Andrew, et al.
Publicado: (2023)
por: Dudzik, Andrew, et al.
Publicado: (2023)
Investigating Low-Rank Training in Transformer Language Models: Efficiency and Scaling Analysis
por: Wei, Xiuying, et al.
Publicado: (2024)
por: Wei, Xiuying, et al.
Publicado: (2024)
How does over-squashing affect the power of GNNs?
por: Di Giovanni, Francesco, et al.
Publicado: (2023)
por: Di Giovanni, Francesco, et al.
Publicado: (2023)
How do language models learn facts? Dynamics, curricula and hallucinations
por: Zucchet, Nicolas, et al.
Publicado: (2025)
por: Zucchet, Nicolas, et al.
Publicado: (2025)
The CLRS-Text Algorithmic Reasoning Language Benchmark
por: Markeeva, Larisa, et al.
Publicado: (2024)
por: Markeeva, Larisa, et al.
Publicado: (2024)
Building on Efficient Foundations: Effectively Training LLMs with Structured Feedforward Layers
por: Wei, Xiuying, et al.
Publicado: (2024)
por: Wei, Xiuying, et al.
Publicado: (2024)
RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling
por: Wei, Xiuying, et al.
Publicado: (2025)
por: Wei, Xiuying, et al.
Publicado: (2025)
Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings
por: Sharifzadeh, Sahand, et al.
Publicado: (2024)
por: Sharifzadeh, Sahand, et al.
Publicado: (2024)
Language models show human-like content effects on reasoning tasks
por: Dasgupta, Ishita, et al.
Publicado: (2022)
por: Dasgupta, Ishita, et al.
Publicado: (2022)
Spatial ModernBERT: Spatial-Aware Transformer for Table and Key-Value Extraction in Financial Documents at Scale
por: Javis AI Team, et al.
Publicado: (2025)
por: Javis AI Team, et al.
Publicado: (2025)
On the generalization of language models from in-context learning and finetuning: a controlled study
por: Lampinen, Andrew K., et al.
Publicado: (2025)
por: Lampinen, Andrew K., et al.
Publicado: (2025)
Primary school children's conflicted emotions about using their heritage languages in multilingual classroom tasks
por: Koen Van Gorp, et al.
Publicado: (2024)
por: Koen Van Gorp, et al.
Publicado: (2024)
The effect of task authenticity on second language writing product and process: The case of a morphologically complex language – Russian
por: Vita Kogan, et al.
Publicado: (2024)
por: Vita Kogan, et al.
Publicado: (2024)
Compiling the Mimosa programming language to RTOS tasks
por: Huber, Nikolaus, et al.
Publicado: (2025)
por: Huber, Nikolaus, et al.
Publicado: (2025)
Recurrent Aggregators in Neural Algorithmic Reasoning
por: Xu, Kaijia, et al.
Publicado: (2024)
por: Xu, Kaijia, et al.
Publicado: (2024)
Leveraging Classical Algorithms for Graph Neural Networks
por: Wu, Jason, et al.
Publicado: (2025)
por: Wu, Jason, et al.
Publicado: (2025)
Optimizers Qualitatively Alter Solutions And We Should Leverage This
por: Pascanu, Razvan, et al.
Publicado: (2025)
por: Pascanu, Razvan, et al.
Publicado: (2025)
Just-in-time and distributed task representations in language models
por: Li, Yuxuan, et al.
Publicado: (2025)
por: Li, Yuxuan, et al.
Publicado: (2025)
Auxiliary task demands mask the capabilities of smaller language models
por: Hu, Jennifer, et al.
Publicado: (2024)
por: Hu, Jennifer, et al.
Publicado: (2024)
Aviary: training language agents on challenging scientific tasks
por: Narayanan, Siddharth, et al.
Publicado: (2024)
por: Narayanan, Siddharth, et al.
Publicado: (2024)
Revisiting Dynamic Evaluation: Online Adaptation for Large Language Models
por: Rannen-Triki, Amal, et al.
Publicado: (2024)
por: Rannen-Triki, Amal, et al.
Publicado: (2024)
Lattice: Learning to Efficiently Compress the Memory
por: Karami, Mahdi, et al.
Publicado: (2025)
por: Karami, Mahdi, et al.
Publicado: (2025)
Deep Grokking: Would Deep Neural Networks Generalize Better?
por: Fan, Simin, et al.
Publicado: (2024)
por: Fan, Simin, et al.
Publicado: (2024)
Testing AI on language comprehension tasks reveals insensitivity to underlying meaning
por: Dentella, Vittoria, et al.
Publicado: (2023)
por: Dentella, Vittoria, et al.
Publicado: (2023)
Ejemplares similares
-
Round and Round We Go! What makes Rotary Positional Encodings useful?
por: Barbero, Federico, et al.
Publicado: (2024) -
Perplexity Cannot Always Tell Right from Wrong
por: Veličković, Petar, et al.
Publicado: (2026) -
How do LLMs Compute Verbal Confidence
por: Kumaran, Dharshan, et al.
Publicado: (2026) -
Transformers meet Neural Algorithmic Reasoners
por: Bounsi, Wilfried, et al.
Publicado: (2024) -
Softmax is not Enough (for Sharp Size Generalisation)
por: Veličković, Petar, et al.
Publicado: (2024)