Need a Small Specialized Language Model? Plan Early!
Fuente:
arXiv
Saved in:
| Main Authors: | Grangier, David, Katharopoulos, Angelos, Ablin, Pierre, Hannun, Awni |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
by: Ablin, Pierre, et al.
Published: (2025)
by: Ablin, Pierre, et al.
Published: (2025)
Compute-Optimal Quantization-Aware Training
by: Dremov, Aleksandr, et al.
Published: (2025)
by: Dremov, Aleksandr, et al.
Published: (2025)
Optimal Splitting of Language Models from Mixtures to Specialized Domains
by: Seto, Skyler, et al.
Published: (2026)
by: Seto, Skyler, et al.
Published: (2026)
No Need to Talk: Asynchronous Mixture of Language Models
by: Filippova, Anastasiia, et al.
Published: (2024)
by: Filippova, Anastasiia, et al.
Published: (2024)
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
by: Grangier, David, et al.
Published: (2024)
by: Grangier, David, et al.
Published: (2024)
Dynamic Gradient Alignment for Online Data Mixing
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
by: Bethune, Louis, et al.
Published: (2025)
by: Bethune, Louis, et al.
Published: (2025)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
by: Saada, Thiziri Nait, et al.
Published: (2025)
by: Saada, Thiziri Nait, et al.
Published: (2025)
Partial Parameter Updates for Efficient Distributed Training
by: Filippova, Anastasiia, et al.
Published: (2025)
by: Filippova, Anastasiia, et al.
Published: (2025)
The AdEMAMix Optimizer: Better, Faster, Older
by: Pagliardini, Matteo, et al.
Published: (2024)
by: Pagliardini, Matteo, et al.
Published: (2024)
Nectar: Neural Estimation of Cached-Token Attention via Regression
by: Monteiro, João, et al.
Published: (2026)
by: Monteiro, João, et al.
Published: (2026)
Scaling Laws for Mixture Pretraining Under Data Constraints
by: Sedova, Anastasiia, et al.
Published: (2026)
by: Sedova, Anastasiia, et al.
Published: (2026)
Training Bilingual LMs with Data Constraints in the Targeted Language
by: Seto, Skyler, et al.
Published: (2024)
by: Seto, Skyler, et al.
Published: (2024)
ADaPT: As-Needed Decomposition and Planning with Language Models
by: Prasad, Archiki, et al.
Published: (2023)
by: Prasad, Archiki, et al.
Published: (2023)
Pretraining with hierarchical memories: separating long-tail and common knowledge
by: Pouransari, Hadi, et al.
Published: (2025)
by: Pouransari, Hadi, et al.
Published: (2025)
Projectable Models: One-Shot Generation of Small Specialized Transformers from Large Ones
by: Zhmoginov, Andrey, et al.
Published: (2025)
by: Zhmoginov, Andrey, et al.
Published: (2025)
Small or Large? Zero-Shot or Finetuned? Guiding Language Model Choice for Specialized Applications in Healthcare
by: Gondara, Lovedeep, et al.
Published: (2025)
by: Gondara, Lovedeep, et al.
Published: (2025)
NurseLLM: The First Specialized Language Model for Nursing
by: Khondaker, Md Tawkat Islam, et al.
Published: (2025)
by: Khondaker, Md Tawkat Islam, et al.
Published: (2025)
All Language Models Large and Small
by: Chen, Zhixun, et al.
Published: (2024)
by: Chen, Zhixun, et al.
Published: (2024)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
by: Tyukin, Georgy, et al.
Published: (2024)
by: Tyukin, Georgy, et al.
Published: (2024)
Language Models Need Inductive Biases to Count Inductively
by: Chang, Yingshan, et al.
Published: (2024)
by: Chang, Yingshan, et al.
Published: (2024)
Self-MoE: Towards Compositional Large Language Models with Self-Specialized Experts
by: Kang, Junmo, et al.
Published: (2024)
by: Kang, Junmo, et al.
Published: (2024)
PlaSma: Making Small Language Models Better Procedural Knowledge Models for (Counterfactual) Planning
by: Brahman, Faeze, et al.
Published: (2023)
by: Brahman, Faeze, et al.
Published: (2023)
Secure multiparty computations in floating-point arithmetic
by: Guo, Chuan, et al.
Published: (2020)
by: Guo, Chuan, et al.
Published: (2020)
Detecting and Characterizing Planning in Language Models
by: Nainani, Jatin, et al.
Published: (2025)
by: Nainani, Jatin, et al.
Published: (2025)
The Geometries of Truth Are Orthogonal Across Tasks
by: Azizian, Waiss, et al.
Published: (2025)
by: Azizian, Waiss, et al.
Published: (2025)
H-Probes: Extracting Hierarchical Structures From Latent Representations of Language Models
by: Dawes, Cutter, et al.
Published: (2026)
by: Dawes, Cutter, et al.
Published: (2026)
Small Language Models Improve Giants by Rewriting Their Outputs
by: Vernikos, Giorgos, et al.
Published: (2023)
by: Vernikos, Giorgos, et al.
Published: (2023)
Self-Taught Self-Correction for Small Language Models
by: Moskvoretskii, Viktor, et al.
Published: (2025)
by: Moskvoretskii, Viktor, et al.
Published: (2025)
An Analysis for Reasoning Bias of Language Models with Small Initialization
by: Yao, Junjie, et al.
Published: (2025)
by: Yao, Junjie, et al.
Published: (2025)
Diffusion Language Models Generation Can Be Halted Early
by: Vaina, Sofia Maria Lo Cicero, et al.
Published: (2023)
by: Vaina, Sofia Maria Lo Cicero, et al.
Published: (2023)
Task-Specific Efficiency Analysis: When Small Language Models Outperform Large Language Models
by: Cao, Jinghan, et al.
Published: (2026)
by: Cao, Jinghan, et al.
Published: (2026)
HierRouter: Coordinated Routing of Specialized Large Language Models via Reinforcement Learning
by: Gupta, Nikunj, et al.
Published: (2025)
by: Gupta, Nikunj, et al.
Published: (2025)
From Small to Large Language Models: Revisiting the Federalist Papers
by: Jeong, So Won, et al.
Published: (2025)
by: Jeong, So Won, et al.
Published: (2025)
Domain-Adaptive Continued Pre-Training of Small Language Models
by: Faroz, Salman
Published: (2025)
by: Faroz, Salman
Published: (2025)
Beyond Early-Token Bias: Model-Specific and Language-Specific Position Effects in Multilingual LLMs
by: Menschikov, Mikhail, et al.
Published: (2025)
by: Menschikov, Mikhail, et al.
Published: (2025)
What Happens When Small Is Made Smaller? Exploring the Impact of Compression on Small Data Pretrained Language Models
by: Awobade, Busayo, et al.
Published: (2024)
by: Awobade, Busayo, et al.
Published: (2024)
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
by: Li, Pengyi, et al.
Published: (2025)
by: Li, Pengyi, et al.
Published: (2025)
Can Small Language Models Learn, Unlearn, and Retain Noise Patterns?
by: Scaria, Nicy, et al.
Published: (2024)
by: Scaria, Nicy, et al.
Published: (2024)
Domain-Adaptive Small Language Models for Structured Tax Code Prediction
by: Nath, Souvik, et al.
Published: (2025)
by: Nath, Souvik, et al.
Published: (2025)
Similar Items
-
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
by: Ablin, Pierre, et al.
Published: (2025) -
Compute-Optimal Quantization-Aware Training
by: Dremov, Aleksandr, et al.
Published: (2025) -
Optimal Splitting of Language Models from Mixtures to Specialized Domains
by: Seto, Skyler, et al.
Published: (2026) -
No Need to Talk: Asynchronous Mixture of Language Models
by: Filippova, Anastasiia, et al.
Published: (2024) -
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
by: Grangier, David, et al.
Published: (2024)