Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Godey, Nathan, de la Clergerie, Éric, Sagot, Benoît |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Scaling Laws of Geographical Representation in Language Models
von: Godey, Nathan, et al.
Veröffentlicht: (2024)
von: Godey, Nathan, et al.
Veröffentlicht: (2024)
Anisotropy Is Inherent to Self-Attention in Transformers
von: Godey, Nathan, et al.
Veröffentlicht: (2024)
von: Godey, Nathan, et al.
Veröffentlicht: (2024)
Gaperon: A Peppered English-French Generative Language Model Suite
von: Godey, Nathan, et al.
Veröffentlicht: (2025)
von: Godey, Nathan, et al.
Veröffentlicht: (2025)
Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content
von: Touchent, Rian, et al.
Veröffentlicht: (2025)
von: Touchent, Rian, et al.
Veröffentlicht: (2025)
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
von: Godey, Nathan, et al.
Veröffentlicht: (2025)
von: Godey, Nathan, et al.
Veröffentlicht: (2025)
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
von: Antoun, Wissam, et al.
Veröffentlicht: (2024)
von: Antoun, Wissam, et al.
Veröffentlicht: (2024)
Lost in Backpropagation: The LM Head is a Gradient Bottleneck
von: Godey, Nathan, et al.
Veröffentlicht: (2026)
von: Godey, Nathan, et al.
Veröffentlicht: (2026)
Patent Representation Learning via Self-supervision
von: Zuo, You, et al.
Veröffentlicht: (2025)
von: Zuo, You, et al.
Veröffentlicht: (2025)
PatentEval: Understanding Errors in Patent Generation
von: Zuo, You, et al.
Veröffentlicht: (2024)
von: Zuo, You, et al.
Veröffentlicht: (2024)
A Causal Language Modeling Detour Improves Encoder Continued Pretraining
von: Touchent, Rian, et al.
Veröffentlicht: (2026)
von: Touchent, Rian, et al.
Veröffentlicht: (2026)
CamemBERT-bio: Leveraging Continual Pre-training for Cost-Effective Models on French Biomedical Data
von: Touchent, Rian, et al.
Veröffentlicht: (2023)
von: Touchent, Rian, et al.
Veröffentlicht: (2023)
From Text to Source: Results in Detecting Large Language Model-Generated Content
von: Antoun, Wissam, et al.
Veröffentlicht: (2023)
von: Antoun, Wissam, et al.
Veröffentlicht: (2023)
Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?
von: Riabi, Arij, et al.
Veröffentlicht: (2021)
von: Riabi, Arij, et al.
Veröffentlicht: (2021)
Disentangling meaning from language in LLM-based machine translation
von: Lasnier, Théo, et al.
Veröffentlicht: (2026)
von: Lasnier, Théo, et al.
Veröffentlicht: (2026)
In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation
von: Zebaze, Armel, et al.
Veröffentlicht: (2024)
von: Zebaze, Armel, et al.
Veröffentlicht: (2024)
How Should We Model the Probability of a Language?
von: Dent, Rasul, et al.
Veröffentlicht: (2026)
von: Dent, Rasul, et al.
Veröffentlicht: (2026)
ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
von: Antoun, Wissam, et al.
Veröffentlicht: (2025)
von: Antoun, Wissam, et al.
Veröffentlicht: (2025)
Language-Switching Triggers Take a Latent Detour Through Language Models
von: Kulumba, Francis, et al.
Veröffentlicht: (2026)
von: Kulumba, Francis, et al.
Veröffentlicht: (2026)
KréyoLID From Language Identification Towards Language Mining
von: Dent, Rasul, et al.
Veröffentlicht: (2025)
von: Dent, Rasul, et al.
Veröffentlicht: (2025)
Tree of Problems: Improving structured problem solving with compositionality
von: Zebaze, Armel, et al.
Veröffentlicht: (2024)
von: Zebaze, Armel, et al.
Veröffentlicht: (2024)
Making Sentence Embeddings Robust to User-Generated Content
von: Nishimwe, Lydia, et al.
Veröffentlicht: (2024)
von: Nishimwe, Lydia, et al.
Veröffentlicht: (2024)
A French Version of the OLDI Seed Corpus
von: Marmonier, Malik, et al.
Veröffentlicht: (2025)
von: Marmonier, Malik, et al.
Veröffentlicht: (2025)
Testing the Deliteralization Hypothesis in Human and Machine Translation
von: Marmonier, Malik, et al.
Veröffentlicht: (2026)
von: Marmonier, Malik, et al.
Veröffentlicht: (2026)
LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
von: Zebaze, Armel, et al.
Veröffentlicht: (2025)
von: Zebaze, Armel, et al.
Veröffentlicht: (2025)
Explicit Learning and the LLM in Machine Translation
von: Marmonier, Malik, et al.
Veröffentlicht: (2025)
von: Marmonier, Malik, et al.
Veröffentlicht: (2025)
Hindsight Quality Prediction Experiments in Multi-Candidate Human-Post-Edited Machine Translation
von: Marmonier, Malik, et al.
Veröffentlicht: (2026)
von: Marmonier, Malik, et al.
Veröffentlicht: (2026)
TopXGen: Topic-Diverse Parallel Data Generation for Low-Resource Machine Translation
von: Zebaze, Armel, et al.
Veröffentlicht: (2025)
von: Zebaze, Armel, et al.
Veröffentlicht: (2025)
Compositional Translation: A Novel LLM-based Approach for Low-resource Machine Translation
von: Zebaze, Armel, et al.
Veröffentlicht: (2025)
von: Zebaze, Armel, et al.
Veröffentlicht: (2025)
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
von: Karamolegkou, Antonia, et al.
Veröffentlicht: (2026)
von: Karamolegkou, Antonia, et al.
Veröffentlicht: (2026)
Why Softmax Attention Outperforms Linear Attention
von: Deng, Yichuan, et al.
Veröffentlicht: (2023)
von: Deng, Yichuan, et al.
Veröffentlicht: (2023)
Why do language models perform worse for morphologically complex languages?
von: Arnett, Catherine, et al.
Veröffentlicht: (2024)
von: Arnett, Catherine, et al.
Veröffentlicht: (2024)
Molyé: A Corpus-based Approach to Language Contact in Colonial France
von: Dent, Rasul, et al.
Veröffentlicht: (2024)
von: Dent, Rasul, et al.
Veröffentlicht: (2024)
Towards Zero-Shot Multimodal Machine Translation
von: Futeral, Matthieu, et al.
Veröffentlicht: (2024)
von: Futeral, Matthieu, et al.
Veröffentlicht: (2024)
InkubaLM: A small language model for low-resource African languages
von: Tonja, Atnafu Lambebo, et al.
Veröffentlicht: (2024)
von: Tonja, Atnafu Lambebo, et al.
Veröffentlicht: (2024)
When your Cousin has the Right Connections: Unsupervised Bilingual Lexicon Induction for Related Data-Imbalanced Languages
von: Bafna, Niyati, et al.
Veröffentlicht: (2023)
von: Bafna, Niyati, et al.
Veröffentlicht: (2023)
Why transformers are obviously good models of language
von: Hill, Felix
Veröffentlicht: (2024)
von: Hill, Felix
Veröffentlicht: (2024)
BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?
von: Chambon, Pierre, et al.
Veröffentlicht: (2025)
von: Chambon, Pierre, et al.
Veröffentlicht: (2025)
Self-Adjust Softmax
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2025)
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2025)
BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding
von: Yuan, Jiayi, et al.
Veröffentlicht: (2025)
von: Yuan, Jiayi, et al.
Veröffentlicht: (2025)
To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
von: Gonsior, Julius, et al.
Veröffentlicht: (2022)
von: Gonsior, Julius, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
On the Scaling Laws of Geographical Representation in Language Models
von: Godey, Nathan, et al.
Veröffentlicht: (2024) -
Anisotropy Is Inherent to Self-Attention in Transformers
von: Godey, Nathan, et al.
Veröffentlicht: (2024) -
Gaperon: A Peppered English-French Generative Language Model Suite
von: Godey, Nathan, et al.
Veröffentlicht: (2025) -
Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content
von: Touchent, Rian, et al.
Veröffentlicht: (2025) -
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
von: Godey, Nathan, et al.
Veröffentlicht: (2025)