Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Taghibakhshi, Ali, Sreenivas, Sharath Turuvekere, Muralidharan, Saurav, Cai, Ruisi, Chochowski, Marcin, Mahabaleshwarkar, Ameya Sunil, Suhara, Yoshi, Olabiyi, Oluwatobi, Korzekwa, Daniel, Patwary, Mostofa, Shoeybi, Mohammad, Kautz, Jan, Catanzaro, Bryan, Aithal, Ashwath, Tajbakhsh, Nima, Molchanov, Pavlo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Minitron-SSM: Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning
von: Taghibakhshi, Ali, et al.
Veröffentlicht: (2025)
von: Taghibakhshi, Ali, et al.
Veröffentlicht: (2025)
Compact Language Models via Pruning and Knowledge Distillation
von: Muralidharan, Saurav, et al.
Veröffentlicht: (2024)
von: Muralidharan, Saurav, et al.
Veröffentlicht: (2024)
Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget Control
von: Taghibakhshi, Ali, et al.
Veröffentlicht: (2026)
von: Taghibakhshi, Ali, et al.
Veröffentlicht: (2026)
LLM Pruning and Distillation in Practice: The Minitron Approach
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2024)
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2024)
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2026)
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2026)
When2Call: When (not) to Call Tools
von: Ross, Hayley, et al.
Veröffentlicht: (2025)
von: Ross, Hayley, et al.
Veröffentlicht: (2025)
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
von: Mahabadi, Rabeeh Karimi, et al.
Veröffentlicht: (2025)
von: Mahabadi, Rabeeh Karimi, et al.
Veröffentlicht: (2025)
Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training
von: Diao, Shizhe, et al.
Veröffentlicht: (2025)
von: Diao, Shizhe, et al.
Veröffentlicht: (2025)
Nemotron-4 15B Technical Report
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
von: Su, Dan, et al.
Veröffentlicht: (2024)
von: Su, Dan, et al.
Veröffentlicht: (2024)
Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
Nemotron-CrossThink: Scaling Self-Learning beyond Math Reasoning
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2025)
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2025)
Flextron: Many-in-One Flexible Large Language Model
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)
Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery
von: Xin, Meng, et al.
Veröffentlicht: (2026)
von: Xin, Meng, et al.
Veröffentlicht: (2026)
Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining
von: Feng, Steven, et al.
Veröffentlicht: (2024)
von: Feng, Steven, et al.
Veröffentlicht: (2024)
Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2025)
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2025)
Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights
von: Krajewski, Jakub, et al.
Veröffentlicht: (2025)
von: Krajewski, Jakub, et al.
Veröffentlicht: (2025)
FusionFactory: Fusing LLM Capabilities with Multi-LLM Log Data
von: Feng, Tao, et al.
Veröffentlicht: (2025)
von: Feng, Tao, et al.
Veröffentlicht: (2025)
Source Identification in Abstractive Summarization
von: Suhara, Yoshi, et al.
Veröffentlicht: (2024)
von: Suhara, Yoshi, et al.
Veröffentlicht: (2024)
Hymba: A Hybrid-head Architecture for Small Language Models
von: Dong, Xin, et al.
Veröffentlicht: (2024)
von: Dong, Xin, et al.
Veröffentlicht: (2024)
MIND: Math Informed syNthetic Dialogues for Pretraining LLMs
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2024)
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2024)
RLP: Reinforcement as a Pretraining Objective
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2025)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2025)
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
von: Chen, Yang, et al.
Veröffentlicht: (2025)
von: Chen, Yang, et al.
Veröffentlicht: (2025)
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
Data, Data Everywhere: A Guide for Pretraining Dataset Construction
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
NVIDIA Nemotron Nano V2 VL
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
Upcycling Large Language Models into Mixture of Experts
von: He, Ethan, et al.
Veröffentlicht: (2024)
von: He, Ethan, et al.
Veröffentlicht: (2024)
Llama 3 Meets MoE: Efficient Upcycling
von: Vavre, Aditya, et al.
Veröffentlicht: (2024)
von: Vavre, Aditya, et al.
Veröffentlicht: (2024)
Noisy Pairing and Partial Supervision for Stylized Opinion Summarization
von: Iso, Hayate, et al.
Veröffentlicht: (2022)
von: Iso, Hayate, et al.
Veröffentlicht: (2022)
Large Language Models are Inconsistent and Biased Evaluators
von: Stureborg, Rickard, et al.
Veröffentlicht: (2024)
von: Stureborg, Rickard, et al.
Veröffentlicht: (2024)
Nemotron-4 340B Technical Report
von: Nvidia, et al.
Veröffentlicht: (2024)
von: Nvidia, et al.
Veröffentlicht: (2024)
MaskLLM: Learnable Semi-Structured Sparsity for Large Language Models
von: Fang, Gongfan, et al.
Veröffentlicht: (2024)
von: Fang, Gongfan, et al.
Veröffentlicht: (2024)
Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning
von: Lu, Ximing, et al.
Veröffentlicht: (2025)
von: Lu, Ximing, et al.
Veröffentlicht: (2025)
Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
von: Wang, Boxin, et al.
Veröffentlicht: (2025)
von: Wang, Boxin, et al.
Veröffentlicht: (2025)
Small Language Models are the Future of Agentic AI
von: Belcak, Peter, et al.
Veröffentlicht: (2025)
von: Belcak, Peter, et al.
Veröffentlicht: (2025)
Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
Llama-Nemotron: Efficient Reasoning Models
von: Bercovich, Akhiad, et al.
Veröffentlicht: (2025)
von: Bercovich, Akhiad, et al.
Veröffentlicht: (2025)
Universal Deep Research: Bring Your Own Model and Strategy
von: Belcak, Peter, et al.
Veröffentlicht: (2025)
von: Belcak, Peter, et al.
Veröffentlicht: (2025)
PHI-S: Distribution Balancing for Label-Free Multi-Teacher Distillation
von: Ranzinger, Mike, et al.
Veröffentlicht: (2024)
von: Ranzinger, Mike, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Minitron-SSM: Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning
von: Taghibakhshi, Ali, et al.
Veröffentlicht: (2025) -
Compact Language Models via Pruning and Knowledge Distillation
von: Muralidharan, Saurav, et al.
Veröffentlicht: (2024) -
Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget Control
von: Taghibakhshi, Ali, et al.
Veröffentlicht: (2026) -
LLM Pruning and Distillation in Practice: The Minitron Approach
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2024) -
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2026)