Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget Control
Fuente:
arXiv
Salvato in:
| Autori principali: | Taghibakhshi, Ali, Cai, Ruisi, Muralidharan, Saurav, Sreenivas, Sharath Turuvekere, Vavre, Aditya, Mahabaleshwarkar, Ameya Sunil, Kartal, Bilal, Liang, Sheldon, Chochowski, Marcin, Chen, Zijia, Bercovich, Akhiad, Zilberstein, Ran, El-Yaniv, Ran, Geifman, Yonatan, Korzekwa, Daniel, Suhara, Yoshi, Olabiyi, Oluwatobi, Aithal, Ashwath, Tajbakhsh, Nima, Molchanov, Pavlo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2025)
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2025)
Minitron-SSM: Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2025)
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2025)
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
di: Sreenivas, Sharath Turuvekere, et al.
Pubblicazione: (2026)
di: Sreenivas, Sharath Turuvekere, et al.
Pubblicazione: (2026)
LLM Pruning and Distillation in Practice: The Minitron Approach
di: Sreenivas, Sharath Turuvekere, et al.
Pubblicazione: (2024)
di: Sreenivas, Sharath Turuvekere, et al.
Pubblicazione: (2024)
Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery
di: Xin, Meng, et al.
Pubblicazione: (2026)
di: Xin, Meng, et al.
Pubblicazione: (2026)
Compact Language Models via Pruning and Knowledge Distillation
di: Muralidharan, Saurav, et al.
Pubblicazione: (2024)
di: Muralidharan, Saurav, et al.
Pubblicazione: (2024)
When2Call: When (not) to Call Tools
di: Ross, Hayley, et al.
Pubblicazione: (2025)
di: Ross, Hayley, et al.
Pubblicazione: (2025)
CRoCoDiL: Continuous and Robust Conditioned Diffusion for Language
di: Uziel, Roy, et al.
Pubblicazione: (2026)
di: Uziel, Roy, et al.
Pubblicazione: (2026)
Llama 3 Meets MoE: Efficient Upcycling
di: Vavre, Aditya, et al.
Pubblicazione: (2024)
di: Vavre, Aditya, et al.
Pubblicazione: (2024)
Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights
di: Krajewski, Jakub, et al.
Pubblicazione: (2025)
di: Krajewski, Jakub, et al.
Pubblicazione: (2025)
FFN Fusion: Rethinking Sequential Computation in Large Language Models
di: Bercovich, Akhiad, et al.
Pubblicazione: (2025)
di: Bercovich, Akhiad, et al.
Pubblicazione: (2025)
Source Identification in Abstractive Summarization
di: Suhara, Yoshi, et al.
Pubblicazione: (2024)
di: Suhara, Yoshi, et al.
Pubblicazione: (2024)
Extending Puzzle for Mixture-of-Experts Reasoning Models with Application to GPT-OSS Acceleration
di: Bercovich, Akhiad, et al.
Pubblicazione: (2026)
di: Bercovich, Akhiad, et al.
Pubblicazione: (2026)
Puzzle: Distillation-Based NAS for Inference-Optimized LLMs
di: Bercovich, Akhiad, et al.
Pubblicazione: (2024)
di: Bercovich, Akhiad, et al.
Pubblicazione: (2024)
Changing Base Without Losing Pace: A GPU-Efficient Alternative to MatMul in DNNs
di: Ailon, Nir, et al.
Pubblicazione: (2025)
di: Ailon, Nir, et al.
Pubblicazione: (2025)
Noisy Pairing and Partial Supervision for Stylized Opinion Summarization
di: Iso, Hayate, et al.
Pubblicazione: (2022)
di: Iso, Hayate, et al.
Pubblicazione: (2022)
Large Language Models are Inconsistent and Biased Evaluators
di: Stureborg, Rickard, et al.
Pubblicazione: (2024)
di: Stureborg, Rickard, et al.
Pubblicazione: (2024)
Hymba: A Hybrid-head Architecture for Small Language Models
di: Dong, Xin, et al.
Pubblicazione: (2024)
di: Dong, Xin, et al.
Pubblicazione: (2024)
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
di: Abramovich, Talor, et al.
Pubblicazione: (2026)
di: Abramovich, Talor, et al.
Pubblicazione: (2026)
Flextron: Many-in-One Flexible Large Language Model
di: Cai, Ruisi, et al.
Pubblicazione: (2024)
di: Cai, Ruisi, et al.
Pubblicazione: (2024)
Your Large Language Models Are Leaving Fingerprints
di: McGovern, Hope, et al.
Pubblicazione: (2024)
di: McGovern, Hope, et al.
Pubblicazione: (2024)
Elucidating Optimal Reward-Diversity Tradeoffs in Text-to-Image Diffusion Models
di: Jena, Rohit, et al.
Pubblicazione: (2024)
di: Jena, Rohit, et al.
Pubblicazione: (2024)
Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training
di: Diao, Shizhe, et al.
Pubblicazione: (2025)
di: Diao, Shizhe, et al.
Pubblicazione: (2025)
Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
di: Park, Dongmin, et al.
Pubblicazione: (2025)
di: Park, Dongmin, et al.
Pubblicazione: (2025)
Optimization of Wind Turbine Placement for Maximum Energy Output in Bangladesh
di: Winner, Olabiyi
Pubblicazione: (2026)
di: Winner, Olabiyi
Pubblicazione: (2026)
Policy and Investment Strategies for Expanding Wind Energy Projects in Bangladesh
di: Winner, Olabiyi
Pubblicazione: (2026)
di: Winner, Olabiyi
Pubblicazione: (2026)
Potential Blind Spots and Hot Spots for Nutrient Loss
di: Kaine Korzekwa
Pubblicazione: (2024)
di: Kaine Korzekwa
Pubblicazione: (2024)
Evaluating Overhead Irrigation Sprinkler Packages in Utah
di: Kaine Korzekwa
Pubblicazione: (2024)
di: Kaine Korzekwa
Pubblicazione: (2024)
What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design
di: Bercovich, Ivan
Pubblicazione: (2026)
di: Bercovich, Ivan
Pubblicazione: (2026)
Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference
di: Bagrov, Natan, et al.
Pubblicazione: (2025)
di: Bagrov, Natan, et al.
Pubblicazione: (2025)
NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment
di: Shen, Gerald, et al.
Pubblicazione: (2024)
di: Shen, Gerald, et al.
Pubblicazione: (2024)
Multiplication Operator Semigroups on Banach lattice valued continuous function spaces
di: Olabiyi, Tobi David
Pubblicazione: (2025)
di: Olabiyi, Tobi David
Pubblicazione: (2025)
Regioselective Ring Opening of Epoxides with Amines Using Silica-bonded S-sulfonic Acid under Solvent-free Conditions
di: Mahmood Tajbakhsh
Pubblicazione: (2012)
di: Mahmood Tajbakhsh
Pubblicazione: (2012)
Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
di: Iso, Hayate, et al.
Pubblicazione: (2026)
di: Iso, Hayate, et al.
Pubblicazione: (2026)
The Herskovits Legacy In African Narrative Analysis And Beyond
di: Olabiyi Babalola J. Gai
Pubblicazione: (2002)
di: Olabiyi Babalola J. Gai
Pubblicazione: (2002)
AI-driven GPT Prompts for Industry Analysis Scholarly Article
di: Aithal, Sreeramana
Pubblicazione: (2025)
di: Aithal, Sreeramana
Pubblicazione: (2025)
Ethical Leadership in the Light of the Bhagavad Gita: Exploring Dharma-Centered Management Practices for the 21st Century
di: Aithal, Sreeramana
Pubblicazione: (2025)
di: Aithal, Sreeramana
Pubblicazione: (2025)
Heterogeneous Parallelism for Multimodal Large Language Model Training
di: Karnati, Yashaswi, et al.
Pubblicazione: (2026)
di: Karnati, Yashaswi, et al.
Pubblicazione: (2026)
Outcome Logic: A Unified Approach to the Metatheory of Program Logics with Branching Effects
di: Zilberstein, Noam
Pubblicazione: (2024)
di: Zilberstein, Noam
Pubblicazione: (2024)
AN ARTIFICIAL INTELLIGENCE BASED DROUGHT PREDICTIONS IN PART OF THE TROPICS
di: Aiyelokun Oluwatobi
Pubblicazione: (2017)
di: Aiyelokun Oluwatobi
Pubblicazione: (2017)
Documenti analoghi
-
Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2025) -
Minitron-SSM: Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning
di: Taghibakhshi, Ali, et al.
Pubblicazione: (2025) -
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
di: Sreenivas, Sharath Turuvekere, et al.
Pubblicazione: (2026) -
LLM Pruning and Distillation in Practice: The Minitron Approach
di: Sreenivas, Sharath Turuvekere, et al.
Pubblicazione: (2024) -
Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery
di: Xin, Meng, et al.
Pubblicazione: (2026)