Compact Language Models via Pruning and Knowledge Distillation
Fuente:
arXiv
Salvato in:
| Autori principali: | Muralidharan, Saurav, Sreenivas, Sharath Turuvekere, Joshi, Raviraj, Chochowski, Marcin, Patwary, Mostofa, Shoeybi, Mohammad, Catanzaro, Bryan, Kautz, Jan, Molchanov, Pavlo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Minitron-SSM: Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2025)
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2025)
LLM Pruning and Distillation in Practice: The Minitron Approach
di: Sreenivas, Sharath Turuvekere, et al.
Pubblicazione: (2024)
di: Sreenivas, Sharath Turuvekere, et al.
Pubblicazione: (2024)
Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2025)
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2025)
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
di: Sreenivas, Sharath Turuvekere, et al.
Pubblicazione: (2026)
di: Sreenivas, Sharath Turuvekere, et al.
Pubblicazione: (2026)
Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models
di: Parmar, Jupinder, et al.
Pubblicazione: (2024)
di: Parmar, Jupinder, et al.
Pubblicazione: (2024)
RLP: Reinforcement as a Pretraining Objective
di: Hatamizadeh, Ali, et al.
Pubblicazione: (2025)
di: Hatamizadeh, Ali, et al.
Pubblicazione: (2025)
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
di: Mahabadi, Rabeeh Karimi, et al.
Pubblicazione: (2025)
di: Mahabadi, Rabeeh Karimi, et al.
Pubblicazione: (2025)
Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining
di: Feng, Steven, et al.
Pubblicazione: (2024)
di: Feng, Steven, et al.
Pubblicazione: (2024)
Flextron: Many-in-One Flexible Large Language Model
di: Cai, Ruisi, et al.
Pubblicazione: (2024)
di: Cai, Ruisi, et al.
Pubblicazione: (2024)
Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
di: Akter, Syeda Nahida, et al.
Pubblicazione: (2025)
di: Akter, Syeda Nahida, et al.
Pubblicazione: (2025)
FusionFactory: Fusing LLM Capabilities with Multi-LLM Log Data
di: Feng, Tao, et al.
Pubblicazione: (2025)
di: Feng, Tao, et al.
Pubblicazione: (2025)
FeatSharp: Your Vision Model Features, Sharper
di: Ranzinger, Mike, et al.
Pubblicazione: (2025)
di: Ranzinger, Mike, et al.
Pubblicazione: (2025)
MaskLLM: Learnable Semi-Structured Sparsity for Large Language Models
di: Fang, Gongfan, et al.
Pubblicazione: (2024)
di: Fang, Gongfan, et al.
Pubblicazione: (2024)
MIND: Math Informed syNthetic Dialogues for Pretraining LLMs
di: Akter, Syeda Nahida, et al.
Pubblicazione: (2024)
di: Akter, Syeda Nahida, et al.
Pubblicazione: (2024)
Minifinetuning: Low-Data Generation Domain Adaptation through Corrective Self-Distillation
di: Belcak, Peter, et al.
Pubblicazione: (2025)
di: Belcak, Peter, et al.
Pubblicazione: (2025)
Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
di: Su, Dan, et al.
Pubblicazione: (2024)
di: Su, Dan, et al.
Pubblicazione: (2024)
Data, Data Everywhere: A Guide for Pretraining Dataset Construction
di: Parmar, Jupinder, et al.
Pubblicazione: (2024)
di: Parmar, Jupinder, et al.
Pubblicazione: (2024)
PHI-S: Distribution Balancing for Label-Free Multi-Teacher Distillation
di: Ranzinger, Mike, et al.
Pubblicazione: (2024)
di: Ranzinger, Mike, et al.
Pubblicazione: (2024)
Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget Control
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2026)
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2026)
C-RADIOv4 (Tech Report)
di: Ranzinger, Mike, et al.
Pubblicazione: (2026)
di: Ranzinger, Mike, et al.
Pubblicazione: (2026)
RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models
di: Heinrich, Greg, et al.
Pubblicazione: (2024)
di: Heinrich, Greg, et al.
Pubblicazione: (2024)
VILA: On Pre-training for Visual Language Models
di: Lin, Ji, et al.
Pubblicazione: (2023)
di: Lin, Ji, et al.
Pubblicazione: (2023)
Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning
di: Lu, Ximing, et al.
Pubblicazione: (2025)
di: Lu, Ximing, et al.
Pubblicazione: (2025)
AM-RADIO: Agglomerative Vision Foundation Model -- Reduce All Domains Into One
di: Ranzinger, Mike, et al.
Pubblicazione: (2023)
di: Ranzinger, Mike, et al.
Pubblicazione: (2023)
Nemotron-CrossThink: Scaling Self-Learning beyond Math Reasoning
di: Akter, Syeda Nahida, et al.
Pubblicazione: (2025)
di: Akter, Syeda Nahida, et al.
Pubblicazione: (2025)
Small Language Models are the Future of Agentic AI
di: Belcak, Peter, et al.
Pubblicazione: (2025)
di: Belcak, Peter, et al.
Pubblicazione: (2025)
Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training
di: Diao, Shizhe, et al.
Pubblicazione: (2025)
di: Diao, Shizhe, et al.
Pubblicazione: (2025)
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
di: Liu, Zihan, et al.
Pubblicazione: (2024)
di: Liu, Zihan, et al.
Pubblicazione: (2024)
On Importance of Pruning and Distillation for Efficient Low Resource NLP
di: Mirashi, Aishwarya, et al.
Pubblicazione: (2024)
di: Mirashi, Aishwarya, et al.
Pubblicazione: (2024)
LITA: Language Instructed Temporal-Localization Assistant
di: Huang, De-An, et al.
Pubblicazione: (2024)
di: Huang, De-An, et al.
Pubblicazione: (2024)
Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning
di: Jung, Jaehun, et al.
Pubblicazione: (2025)
di: Jung, Jaehun, et al.
Pubblicazione: (2025)
Towards Building Efficient Sentence BERT Models using Layer Pruning
di: Shelke, Anushka, et al.
Pubblicazione: (2024)
di: Shelke, Anushka, et al.
Pubblicazione: (2024)
RAVEN: In-Context Learning with Retrieval-Augmented Encoder-Decoder Language Models
di: Huang, Jie, et al.
Pubblicazione: (2023)
di: Huang, Jie, et al.
Pubblicazione: (2023)
On Data Engineering for Scaling LLM Terminal Capabilities
di: Pi, Renjie, et al.
Pubblicazione: (2026)
di: Pi, Renjie, et al.
Pubblicazione: (2026)
Adaptive Sharpness-Aware Pruning for Robust Sparse Networks
di: Bair, Anna, et al.
Pubblicazione: (2023)
di: Bair, Anna, et al.
Pubblicazione: (2023)
On Importance of Layer Pruning for Smaller BERT Models and Low Resource Languages
di: Shirke, Mayur, et al.
Pubblicazione: (2025)
di: Shirke, Mayur, et al.
Pubblicazione: (2025)
Universal Deep Research: Bring Your Own Model and Strategy
di: Belcak, Peter, et al.
Pubblicazione: (2025)
di: Belcak, Peter, et al.
Pubblicazione: (2025)
MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
di: Lin, Sheng-Chieh, et al.
Pubblicazione: (2024)
di: Lin, Sheng-Chieh, et al.
Pubblicazione: (2024)
Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights
di: Krajewski, Jakub, et al.
Pubblicazione: (2025)
di: Krajewski, Jakub, et al.
Pubblicazione: (2025)
From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models
di: Xu, Chejian, et al.
Pubblicazione: (2025)
di: Xu, Chejian, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Minitron-SSM: Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2025) -
LLM Pruning and Distillation in Practice: The Minitron Approach
di: Sreenivas, Sharath Turuvekere, et al.
Pubblicazione: (2024) -
Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2025) -
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
di: Sreenivas, Sharath Turuvekere, et al.
Pubblicazione: (2026) -
Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models
di: Parmar, Jupinder, et al.
Pubblicazione: (2024)