Lost in Backpropagation: The LM Head is a Gradient Bottleneck
Fuente:
arXiv
Saved in:
| Main Authors: | Godey, Nathan, Artzi, Yoav |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck
by: Godey, Nathan, et al.
Published: (2024)
by: Godey, Nathan, et al.
Published: (2024)
No Mean Feat: Simple, Strong Baselines for Context Compression
by: Feldman, Yair, et al.
Published: (2025)
by: Feldman, Yair, et al.
Published: (2025)
Talk Less, Interact Better: Evaluating In-context Conversational Adaptation in Multimodal LLMs
by: Hua, Yilun, et al.
Published: (2024)
by: Hua, Yilun, et al.
Published: (2024)
CoGen: Learning from Feedback with Coupled Comprehension and Generation
by: Gul, Mustafa Omer, et al.
Published: (2024)
by: Gul, Mustafa Omer, et al.
Published: (2024)
Post-training for Efficient Communication via Convention Formation
by: Hua, Yilun, et al.
Published: (2025)
by: Hua, Yilun, et al.
Published: (2025)
A Surprising Failure? Multimodal LLMs and the NLVR Challenge
by: Wu, Anne, et al.
Published: (2024)
by: Wu, Anne, et al.
Published: (2024)
A Joint Study of Phrase Grounding and Task Performance in Vision and Language Models
by: Kojima, Noriyuki, et al.
Published: (2023)
by: Kojima, Noriyuki, et al.
Published: (2023)
Success and Cost Elicit Convention Formation for Efficient Communication
by: Vaduguru, Saujas, et al.
Published: (2025)
by: Vaduguru, Saujas, et al.
Published: (2025)
Anisotropy Is Inherent to Self-Attention in Transformers
by: Godey, Nathan, et al.
Published: (2024)
by: Godey, Nathan, et al.
Published: (2024)
Breadcrumbs Reasoning: Memory-Efficient Reasoning with Compression Beacons
by: Monea, Giovanni, et al.
Published: (2025)
by: Monea, Giovanni, et al.
Published: (2025)
On the Scaling Laws of Geographical Representation in Language Models
by: Godey, Nathan, et al.
Published: (2024)
by: Godey, Nathan, et al.
Published: (2024)
Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content
by: Touchent, Rian, et al.
Published: (2025)
by: Touchent, Rian, et al.
Published: (2025)
LLMs Are In-Context Bandit Reinforcement Learners
by: Monea, Giovanni, et al.
Published: (2024)
by: Monea, Giovanni, et al.
Published: (2024)
Ad hoc conventions generalize to new referents
by: Ji, Anya, et al.
Published: (2025)
by: Ji, Anya, et al.
Published: (2025)
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
by: Calderon, Nitay, et al.
Published: (2026)
by: Calderon, Nitay, et al.
Published: (2026)
EvoLM: In Search of Lost Language Model Training Dynamics
by: Qi, Zhenting, et al.
Published: (2025)
by: Qi, Zhenting, et al.
Published: (2025)
IncDSI: Incrementally Updatable Document Retrieval
by: Kishore, Varsha, et al.
Published: (2023)
by: Kishore, Varsha, et al.
Published: (2023)
Lost in Translationese? Reducing Translation Effect Using Abstract Meaning Representation
by: Wein, Shira, et al.
Published: (2023)
by: Wein, Shira, et al.
Published: (2023)
Knot So Simple: A Minimalistic Environment for Spatial Reasoning
by: Chen, Zizhao, et al.
Published: (2025)
by: Chen, Zizhao, et al.
Published: (2025)
Cancer Type, Stage and Prognosis Assessment from Pathology Reports using LLMs
by: Saluja, Rachit, et al.
Published: (2025)
by: Saluja, Rachit, et al.
Published: (2025)
Flash Interpretability: Decoding Specialised Feature Neurons in Large Language Models with the LM-Head
by: Davies, Harry J
Published: (2025)
by: Davies, Harry J
Published: (2025)
Retrospective Learning from Interactions
by: Chen, Zizhao, et al.
Published: (2024)
by: Chen, Zizhao, et al.
Published: (2024)
SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding
by: Plaksin, Anton, et al.
Published: (2026)
by: Plaksin, Anton, et al.
Published: (2026)
The Backpropagation of the Wave Network
by: Zhang, Xin, et al.
Published: (2024)
by: Zhang, Xin, et al.
Published: (2024)
Gaperon: A Peppered English-French Generative Language Model Suite
by: Godey, Nathan, et al.
Published: (2025)
by: Godey, Nathan, et al.
Published: (2025)
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
by: Godey, Nathan, et al.
Published: (2025)
by: Godey, Nathan, et al.
Published: (2025)
BAMBINO-LM: (Bilingual-)Human-Inspired Continual Pretraining of BabyLM
by: Shen, Zhewen, et al.
Published: (2024)
by: Shen, Zhewen, et al.
Published: (2024)
Lost at the Beginning of Reasoning
by: Liao, Baohao, et al.
Published: (2025)
by: Liao, Baohao, et al.
Published: (2025)
Replicating ReLM Results: Validating Large Language Models with ReLM
by: Adamson, Reece, et al.
Published: (2025)
by: Adamson, Reece, et al.
Published: (2025)
BabyLM Turns 3: Call for papers for the 2025 BabyLM workshop
by: Charpentier, Lucas, et al.
Published: (2025)
by: Charpentier, Lucas, et al.
Published: (2025)
Predicting the Emergence of Induction Heads in Language Model Pretraining
by: Aoyama, Tatsuya, et al.
Published: (2025)
by: Aoyama, Tatsuya, et al.
Published: (2025)
Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation
by: Edin, Joakim, et al.
Published: (2025)
by: Edin, Joakim, et al.
Published: (2025)
Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic
by: Alyafeai, Zaid, et al.
Published: (2024)
by: Alyafeai, Zaid, et al.
Published: (2024)
LokiLM: Technical Report
by: Kiefel, Justin, et al.
Published: (2024)
by: Kiefel, Justin, et al.
Published: (2024)
BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop
by: Choshen, Leshem, et al.
Published: (2026)
by: Choshen, Leshem, et al.
Published: (2026)
Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Paired-Focus Constructions
by: Scivetti, Wesley, et al.
Published: (2026)
by: Scivetti, Wesley, et al.
Published: (2026)
SaulLM-54B & SaulLM-141B: Scaling Up Domain Adaptation for the Legal Domain
by: Colombo, Pierre, et al.
Published: (2024)
by: Colombo, Pierre, et al.
Published: (2024)
Memory-Efficient Structured Backpropagation for On-Device LLM Fine-Tuning
by: Park, Juneyoung, et al.
Published: (2026)
by: Park, Juneyoung, et al.
Published: (2026)
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
by: Ruan, Yangjun, et al.
Published: (2023)
by: Ruan, Yangjun, et al.
Published: (2023)
Linear Correlation in LM's Compositional Generalization and Hallucination
by: Peng, Letian, et al.
Published: (2025)
by: Peng, Letian, et al.
Published: (2025)
Similar Items
-
Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck
by: Godey, Nathan, et al.
Published: (2024) -
No Mean Feat: Simple, Strong Baselines for Context Compression
by: Feldman, Yair, et al.
Published: (2025) -
Talk Less, Interact Better: Evaluating In-context Conversational Adaptation in Multimodal LLMs
by: Hua, Yilun, et al.
Published: (2024) -
CoGen: Learning from Feedback with Coupled Comprehension and Generation
by: Gul, Mustafa Omer, et al.
Published: (2024) -
Post-training for Efficient Communication via Convention Formation
by: Hua, Yilun, et al.
Published: (2025)