Advancing Model Refinement: Muon-Optimized Distillation and Quantization for LLM Deployment
Fuente:
arXiv
Saved in:
| Main Authors: | Sander, Jacob, Jalaian, Brian, Dasari, Venkat R. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On Accelerating Edge AI: Optimizing Resource-Constrained Environments
by: Sander, Jacob, et al.
Published: (2025)
by: Sander, Jacob, et al.
Published: (2025)
Constrained Edge AI Deployment: Fine-Tuning vs Distillation for LLM Compression
by: Sander, Jacob, et al.
Published: (2025)
by: Sander, Jacob, et al.
Published: (2025)
Towards Interpretable Adversarial Examples via Sparse Adversarial Attack
by: Lin, Fudong, et al.
Published: (2025)
by: Lin, Fudong, et al.
Published: (2025)
Q-realign: Piggybacking Realignment on Quantization for Safe and Efficient LLM Deployment
by: Tan, Qitao, et al.
Published: (2026)
by: Tan, Qitao, et al.
Published: (2026)
SLMQuant:Benchmarking Small Language Model Quantization for Practical Deployment
by: Wang, Jiacheng, et al.
Published: (2025)
by: Wang, Jiacheng, et al.
Published: (2025)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
by: Lee, Deokjae, et al.
Published: (2025)
by: Lee, Deokjae, et al.
Published: (2025)
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
by: Huang, Feihu, et al.
Published: (2026)
by: Huang, Feihu, et al.
Published: (2026)
The Newton-Muon Optimizer
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Muon is Scalable for LLM Training
by: Liu, Jingyuan, et al.
Published: (2025)
by: Liu, Jingyuan, et al.
Published: (2025)
Post-Training Quantization of OpenPangu Models for Efficient Deployment on Atlas A2
by: Luo, Yilun, et al.
Published: (2025)
by: Luo, Yilun, et al.
Published: (2025)
Intrinsic Muon: Spectral Optimization on Riemannian Matrix Manifolds
by: Li, Yibang, et al.
Published: (2026)
by: Li, Yibang, et al.
Published: (2026)
Muon Optimizer Accelerates Grokking
by: Tveit, Amund, et al.
Published: (2025)
by: Tveit, Amund, et al.
Published: (2025)
Activation Outliers in Transformer Quantization: Reproduction, Statistical Analysis, and Deployment Tradeoffs
by: Kaliaperumal, Pranav Kumar
Published: (2026)
by: Kaliaperumal, Pranav Kumar
Published: (2026)
Critic-Driven Voronoi-Quantization for Distilling Deep RL Policies to Explainable Models
by: Deproost, Senne, et al.
Published: (2026)
by: Deproost, Senne, et al.
Published: (2026)
Refinement Provenance Inference: Detecting LLM-Refined Training Prompts from Model Behavior
by: Yin, Bo, et al.
Published: (2026)
by: Yin, Bo, et al.
Published: (2026)
A Reliable Knowledge Processing Framework for Combustion Science using Foundation Models
by: Sharma, Vansh, et al.
Published: (2023)
by: Sharma, Vansh, et al.
Published: (2023)
DynMuon: A Dynamic Spectral Shaping View of Muon
by: Wu, Fangzhou, et al.
Published: (2026)
by: Wu, Fangzhou, et al.
Published: (2026)
TrasMuon: Trust-Region Adaptive Scaling for Orthogonalized Momentum Optimizers
by: Cheng, Peng, et al.
Published: (2026)
by: Cheng, Peng, et al.
Published: (2026)
Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning
by: Liu, Ziyue, et al.
Published: (2026)
by: Liu, Ziyue, et al.
Published: (2026)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
by: Zhang, Tuo, et al.
Published: (2025)
by: Zhang, Tuo, et al.
Published: (2025)
OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
by: Hu, Xing, et al.
Published: (2025)
by: Hu, Xing, et al.
Published: (2025)
ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation
by: Singh, Aasheesh, et al.
Published: (2025)
by: Singh, Aasheesh, et al.
Published: (2025)
AmoebaLLM: Constructing Any-Shape Large Language Models for Efficient and Instant Deployment
by: Fu, Yonggan, et al.
Published: (2024)
by: Fu, Yonggan, et al.
Published: (2024)
Distillation Traps and Guards: A Calibration Knob for LLM Distillability
by: Zhan, Weixiao, et al.
Published: (2026)
by: Zhan, Weixiao, et al.
Published: (2026)
RAFT: Data Refinement and Adaptive Distillation for Domain Fine-Tuning with Alleviated Forgetting
by: Li, Yuduo, et al.
Published: (2026)
by: Li, Yuduo, et al.
Published: (2026)
Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision
by: Ye, Yaowen, et al.
Published: (2025)
by: Ye, Yaowen, et al.
Published: (2025)
EdgeRazor: A Lightweight Framework for Large Language Models via Mixed-Precision Quantization-Aware Distillation
by: Zhang, Shu-Hao, et al.
Published: (2026)
by: Zhang, Shu-Hao, et al.
Published: (2026)
MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale Deployment
by: Huang, Hanxian, et al.
Published: (2026)
by: Huang, Hanxian, et al.
Published: (2026)
CosmicFish-HRM: Adaptive Reasoning via Hierarchical Recurrent Mechanisms in Compact Language Models
by: Lakkapragada, Venkat Akhil
Published: (2026)
by: Lakkapragada, Venkat Akhil
Published: (2026)
SDQ: Sparse Decomposed Quantization for LLM Inference
by: Jeong, Geonhwa, et al.
Published: (2024)
by: Jeong, Geonhwa, et al.
Published: (2024)
Membership and Memorization in LLM Knowledge Distillation
by: Zhang, Ziqi, et al.
Published: (2025)
by: Zhang, Ziqi, et al.
Published: (2025)
Grasp-HGN: Grasping the Unexpected
by: Zandigohar, Mehrshad, et al.
Published: (2025)
by: Zandigohar, Mehrshad, et al.
Published: (2025)
LLM and GNN are Complementary: Distilling LLM for Multimodal Graph Learning
by: Xu, Junjie, et al.
Published: (2024)
by: Xu, Junjie, et al.
Published: (2024)
Optimizer-Induced Mode Connectivity: From AdamW to Muon
by: Zhang, Fangzhao, et al.
Published: (2026)
by: Zhang, Fangzhao, et al.
Published: (2026)
Exploiting LLM Quantization
by: Egashira, Kazuki, et al.
Published: (2024)
by: Egashira, Kazuki, et al.
Published: (2024)
HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
by: Ding, Ken
Published: (2026)
by: Ding, Ken
Published: (2026)
Refining Adaptive Zeroth-Order Optimization at Ease
by: Shu, Yao, et al.
Published: (2025)
by: Shu, Yao, et al.
Published: (2025)
Reinforcement-aware Knowledge Distillation for LLM Reasoning
by: Zhang, Zhaoyang, et al.
Published: (2026)
by: Zhang, Zhaoyang, et al.
Published: (2026)
A Cost-Benefit Analysis of On-Premise Large Language Model Deployment: Breaking Even with Commercial LLM Services
by: Pan, Guanzhong, et al.
Published: (2025)
by: Pan, Guanzhong, et al.
Published: (2025)
Neurosymbolic Artificial Intelligence for Robust Network Intrusion Detection: From Scratch to Transfer Learning
by: Tran, Huynh T. T., et al.
Published: (2025)
by: Tran, Huynh T. T., et al.
Published: (2025)
Similar Items
-
On Accelerating Edge AI: Optimizing Resource-Constrained Environments
by: Sander, Jacob, et al.
Published: (2025) -
Constrained Edge AI Deployment: Fine-Tuning vs Distillation for LLM Compression
by: Sander, Jacob, et al.
Published: (2025) -
Towards Interpretable Adversarial Examples via Sparse Adversarial Attack
by: Lin, Fudong, et al.
Published: (2025) -
Q-realign: Piggybacking Realignment on Quantization for Safe and Efficient LLM Deployment
by: Tan, Qitao, et al.
Published: (2026) -
SLMQuant:Benchmarking Small Language Model Quantization for Practical Deployment
by: Wang, Jiacheng, et al.
Published: (2025)