Minitron-SSM: Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Taghibakhshi, Ali, Sreenivas, Sharath Turuvekere, Muralidharan, Saurav, Chochowski, Marcin, Karnati, Yashaswi, Joshi, Raviraj, Mahabaleshwarkar, Ameya Sunil, Chen, Zijia, Suhara, Yoshi, Olabiyi, Oluwatobi, Korzekwa, Daniel, Patwary, Mostofa, Shoeybi, Mohammad, Kautz, Jan, Catanzaro, Bryan, Aithal, Ashwath, Tajbakhsh, Nima, Molchanov, Pavlo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs
von: Taghibakhshi, Ali, et al.
Veröffentlicht: (2025)
von: Taghibakhshi, Ali, et al.
Veröffentlicht: (2025)
LLM Pruning and Distillation in Practice: The Minitron Approach
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2024)
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2024)
Compact Language Models via Pruning and Knowledge Distillation
von: Muralidharan, Saurav, et al.
Veröffentlicht: (2024)
von: Muralidharan, Saurav, et al.
Veröffentlicht: (2024)
Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget Control
von: Taghibakhshi, Ali, et al.
Veröffentlicht: (2026)
von: Taghibakhshi, Ali, et al.
Veröffentlicht: (2026)
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2026)
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2026)
Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
When2Call: When (not) to Call Tools
von: Ross, Hayley, et al.
Veröffentlicht: (2025)
von: Ross, Hayley, et al.
Veröffentlicht: (2025)
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
von: Mahabadi, Rabeeh Karimi, et al.
Veröffentlicht: (2025)
von: Mahabadi, Rabeeh Karimi, et al.
Veröffentlicht: (2025)
Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining
von: Feng, Steven, et al.
Veröffentlicht: (2024)
von: Feng, Steven, et al.
Veröffentlicht: (2024)
Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2025)
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2025)
Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights
von: Krajewski, Jakub, et al.
Veröffentlicht: (2025)
von: Krajewski, Jakub, et al.
Veröffentlicht: (2025)
Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery
von: Xin, Meng, et al.
Veröffentlicht: (2026)
von: Xin, Meng, et al.
Veröffentlicht: (2026)
FusionFactory: Fusing LLM Capabilities with Multi-LLM Log Data
von: Feng, Tao, et al.
Veröffentlicht: (2025)
von: Feng, Tao, et al.
Veröffentlicht: (2025)
SSM: Population Health
Veröffentlicht: (2016)
Veröffentlicht: (2016)
SSM - Mental Health
Veröffentlicht: (2021)
Veröffentlicht: (2021)
SSM - Health Systems
Veröffentlicht: (2025)
Veröffentlicht: (2025)
Source Identification in Abstractive Summarization
von: Suhara, Yoshi, et al.
Veröffentlicht: (2024)
von: Suhara, Yoshi, et al.
Veröffentlicht: (2024)
Heterogeneous Parallelism for Multimodal Large Language Model Training
von: Karnati, Yashaswi, et al.
Veröffentlicht: (2026)
von: Karnati, Yashaswi, et al.
Veröffentlicht: (2026)
MIND: Math Informed syNthetic Dialogues for Pretraining LLMs
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2024)
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2024)
RLP: Reinforcement as a Pretraining Objective
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2025)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2025)
SSM: Qualitative Research in Health
Veröffentlicht: (2021)
Veröffentlicht: (2021)
The SSM with Suppressed SUSY Charge
von: Dixon, John
Veröffentlicht: (2016)
von: Dixon, John
Veröffentlicht: (2016)
Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
von: Su, Dan, et al.
Veröffentlicht: (2024)
von: Su, Dan, et al.
Veröffentlicht: (2024)
Data, Data Everywhere: A Guide for Pretraining Dataset Construction
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
ASecond-Order SpikingSSM for Wearables
von: Agrawal, Kartikay, et al.
Veröffentlicht: (2025)
von: Agrawal, Kartikay, et al.
Veröffentlicht: (2025)
Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training
von: Diao, Shizhe, et al.
Veröffentlicht: (2025)
von: Diao, Shizhe, et al.
Veröffentlicht: (2025)
Upcycling Large Language Models into Mixture of Experts
von: He, Ethan, et al.
Veröffentlicht: (2024)
von: He, Ethan, et al.
Veröffentlicht: (2024)
Llama 3 Meets MoE: Efficient Upcycling
von: Vavre, Aditya, et al.
Veröffentlicht: (2024)
von: Vavre, Aditya, et al.
Veröffentlicht: (2024)
Noisy Pairing and Partial Supervision for Stylized Opinion Summarization
von: Iso, Hayate, et al.
Veröffentlicht: (2022)
von: Iso, Hayate, et al.
Veröffentlicht: (2022)
Large Language Models are Inconsistent and Biased Evaluators
von: Stureborg, Rickard, et al.
Veröffentlicht: (2024)
von: Stureborg, Rickard, et al.
Veröffentlicht: (2024)
Retrieval-Aware Distillation for Transformer-SSM Hybrids
von: Bick, Aviv, et al.
Veröffentlicht: (2026)
von: Bick, Aviv, et al.
Veröffentlicht: (2026)
Bottenstock SSM_502_L_2 från Riddarholmsskeppet
von: Museer och konst, Medeltidsmuseet, et al.
Veröffentlicht: (2026)
von: Museer och konst, Medeltidsmuseet, et al.
Veröffentlicht: (2026)
Topptimmer SSM_502_G_1 från Riddarholmsskeppet.
von: Museer och konst, Stadsmuseet, et al.
Veröffentlicht: (2026)
von: Museer och konst, Stadsmuseet, et al.
Veröffentlicht: (2026)
Bottenstock SSM_502_K_3 från Riddarholmsskeppet
von: Museer och konst, Stadsmuseet, et al.
Veröffentlicht: (2026)
von: Museer och konst, Stadsmuseet, et al.
Veröffentlicht: (2026)
Topptimmer SSM_502_H_1 från Riddarholmsskeppet
von: Museer och konst, Stadsmuseet, et al.
Veröffentlicht: (2026)
von: Museer och konst, Stadsmuseet, et al.
Veröffentlicht: (2026)
Spant SSM_502_B_1 från Riddarholmsskeppet
von: Museer och Konst, Medeltidsmuseet, et al.
Veröffentlicht: (2026)
von: Museer och Konst, Medeltidsmuseet, et al.
Veröffentlicht: (2026)
Spant SSM_502_AF_1 från Riddarholmsskeppet.
von: Museer och konst, Medeltidsmuseet, et al.
Veröffentlicht: (2026)
von: Museer och konst, Medeltidsmuseet, et al.
Veröffentlicht: (2026)
A SSM is Polymerized from Multivariate Time Series
von: Wu, Haixiang
Veröffentlicht: (2024)
von: Wu, Haixiang
Veröffentlicht: (2024)
ECGMamba: Towards Efficient ECG Classification with BiSSM
von: Qiang, Yupeng, et al.
Veröffentlicht: (2024)
von: Qiang, Yupeng, et al.
Veröffentlicht: (2024)
Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning
von: Lu, Ximing, et al.
Veröffentlicht: (2025)
von: Lu, Ximing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs
von: Taghibakhshi, Ali, et al.
Veröffentlicht: (2025) -
LLM Pruning and Distillation in Practice: The Minitron Approach
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2024) -
Compact Language Models via Pruning and Knowledge Distillation
von: Muralidharan, Saurav, et al.
Veröffentlicht: (2024) -
Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget Control
von: Taghibakhshi, Ali, et al.
Veröffentlicht: (2026) -
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2026)