MUON+: Towards More Effective Muon via One Additional Normalization Step for LLM Pre-training
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Ruijie, Zhao, Yequan, Liu, Ziyue, Wang, Zhengyang, Su, Yupeng, Tan, Liyan, Zhang, Zheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization
by: Su, Yupeng, et al.
Published: (2026)
by: Su, Yupeng, et al.
Published: (2026)
Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning
by: Liu, Ziyue, et al.
Published: (2026)
by: Liu, Ziyue, et al.
Published: (2026)
TEON: Tensorized Orthonormalization Beyond Layer-Wise Muon for Large Language Model Pre-Training
by: Zhang, Ruijie, et al.
Published: (2026)
by: Zhang, Ruijie, et al.
Published: (2026)
LaX: Boosting Low-Rank Training of Foundation Models via Latent Crossing
by: Zhang, Ruijie, et al.
Published: (2025)
by: Zhang, Ruijie, et al.
Published: (2025)
FuRA: Full-Rank Parameter-Efficient Fine-Tuning with Spectral Preconditioning
by: Zhao, Yequan, et al.
Published: (2026)
by: Zhao, Yequan, et al.
Published: (2026)
ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload
by: Liu, Ziyue, et al.
Published: (2026)
by: Liu, Ziyue, et al.
Published: (2026)
Masked Structural Growth for 2x Faster Language Model Pre-training
by: Yao, Yiqun, et al.
Published: (2023)
by: Yao, Yiqun, et al.
Published: (2023)
CoLA: Compute-Efficient Pre-Training of LLMs via Low-Rank Activation
by: Liu, Ziyue, et al.
Published: (2025)
by: Liu, Ziyue, et al.
Published: (2025)
Revisiting Diffusion Models: From Generative Pre-training to One-Step Generation
by: Zheng, Bowen, et al.
Published: (2025)
by: Zheng, Bowen, et al.
Published: (2025)
SUDA-Muon: Structural Design Principles and Boundaries for Fully Decentralized Muon
by: Zhang, Hengrui, et al.
Published: (2026)
by: Zhang, Hengrui, et al.
Published: (2026)
BiPFT: Binary Pre-trained Foundation Transformer with Low-rank Estimation of Binarization Residual Polynomials
by: Xing, Xingrun, et al.
Published: (2023)
by: Xing, Xingrun, et al.
Published: (2023)
The Missing Piece in Pre-trained Model Evaluation: Reward-Guided Decoding Unlocks Task-Oriented Behavior Without Parameter Updates
by: Wang, Shaobo, et al.
Published: (2026)
by: Wang, Shaobo, et al.
Published: (2026)
Distance-Aware Muon: Adaptive Step Scaling for Normalized Optimization
by: Demidovich, Yury, et al.
Published: (2026)
by: Demidovich, Yury, et al.
Published: (2026)
BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
by: Wang, Zhengyang, et al.
Published: (2025)
by: Wang, Zhengyang, et al.
Published: (2025)
One-Step Diffusion Distillation via Deep Equilibrium Models
by: Geng, Zhengyang, et al.
Published: (2023)
by: Geng, Zhengyang, et al.
Published: (2023)
Generic Knowledge Boosted Pre-training For Remote Sensing Images
by: Huang, Ziyue, et al.
Published: (2024)
by: Huang, Ziyue, et al.
Published: (2024)
Large EEG-U-Transformer for Time-Step Level Detection Without Pre-Training
by: Wu, Kerui, et al.
Published: (2025)
by: Wu, Kerui, et al.
Published: (2025)
Muon is Scalable for LLM Training
by: Liu, Jingyuan, et al.
Published: (2025)
by: Liu, Jingyuan, et al.
Published: (2025)
nanoLM: an Affordable LLM Pre-training Benchmark via Accurate Loss Prediction across Scales
by: Yao, Yiqun, et al.
Published: (2023)
by: Yao, Yiqun, et al.
Published: (2023)
Real-Time FJ/MAC PDE Solvers via Tensorized, Back-Propagation-Free Optical PINN Training
by: Zhao, Yequan, et al.
Published: (2023)
by: Zhao, Yequan, et al.
Published: (2023)
David and Goliath: Small One-step Model Beats Large Diffusion with Score Post-training
by: Luo, Weijian, et al.
Published: (2024)
by: Luo, Weijian, et al.
Published: (2024)
Towards Pre-trained Graph Condensation via Optimal Transport
by: Yan, Yeyu, et al.
Published: (2025)
by: Yan, Yeyu, et al.
Published: (2025)
Separable Operator Networks
by: Yu, Xinling, et al.
Published: (2024)
by: Yu, Xinling, et al.
Published: (2024)
Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
by: Zhang, Xiangdong, et al.
Published: (2025)
by: Zhang, Xiangdong, et al.
Published: (2025)
Towards a General Framework for Continual Learning with Pre-training
by: Wang, Liyuan, et al.
Published: (2023)
by: Wang, Liyuan, et al.
Published: (2023)
MUON TELESCOPE (MUTE): A FIRST STUDY USING GEANT4
by: H. Asorey
Published: (2017)
by: H. Asorey
Published: (2017)
GPR: Towards a Generative Pre-trained One-Model Paradigm for Large-Scale Advertising Recommendation
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
SparseLLM: Towards Global Pruning for Pre-trained Language Models
by: Bai, Guangji, et al.
Published: (2024)
by: Bai, Guangji, et al.
Published: (2024)
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
by: Tian, Changxin, et al.
Published: (2025)
by: Tian, Changxin, et al.
Published: (2025)
MuonEq: Balancing Before Orthogonalization with Lightweight Equilibration
by: Chang, Da, et al.
Published: (2026)
by: Chang, Da, et al.
Published: (2026)
Towards Effective and Efficient Continual Pre-training of Large Language Models
by: Chen, Jie, et al.
Published: (2024)
by: Chen, Jie, et al.
Published: (2024)
Delayed Bottlenecking: Alleviating Forgetting in Pre-trained Graph Neural Networks
by: Zhao, Zhe, et al.
Published: (2024)
by: Zhao, Zhe, et al.
Published: (2024)
Towards Pre-training an Effective Respiratory Audio Foundation Model
by: Niizumi, Daisuke, et al.
Published: (2025)
by: Niizumi, Daisuke, et al.
Published: (2025)
MutDet: Mutually Optimizing Pre-training for Remote Sensing Object Detection
by: Huang, Ziyue, et al.
Published: (2024)
by: Huang, Ziyue, et al.
Published: (2024)
ASTROPARTICLE TECHNIQUES: COLOMBIA ACTIVE VOLCANO CANDIDATES FOR MUON TELESCOPE OBSERVATION SITES
by: H. Asorey
Published: (2017)
by: H. Asorey
Published: (2017)
Towards Faster Graph Partitioning via Pre-training and Inductive Inference
by: Qin, Meng, et al.
Published: (2024)
by: Qin, Meng, et al.
Published: (2024)
Effectiveness of Pre-training for Few-shot Intent Classification
by: Zhang, Haode, et al.
Published: (2021)
by: Zhang, Haode, et al.
Published: (2021)
HonestFace: Towards Honest Face Restoration with One-Step Diffusion Model
by: Wang, Jingkai, et al.
Published: (2025)
by: Wang, Jingkai, et al.
Published: (2025)
Poor Man's Training on MCUs: A Memory-Efficient Quantized Back-Propagation-Free Approach
by: Zhao, Yequan, et al.
Published: (2024)
by: Zhao, Yequan, et al.
Published: (2024)
CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
by: Zhao, Shuai, et al.
Published: (2023)
by: Zhao, Shuai, et al.
Published: (2023)
Similar Items
-
MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization
by: Su, Yupeng, et al.
Published: (2026) -
Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning
by: Liu, Ziyue, et al.
Published: (2026) -
TEON: Tensorized Orthonormalization Beyond Layer-Wise Muon for Large Language Model Pre-Training
by: Zhang, Ruijie, et al.
Published: (2026) -
LaX: Boosting Low-Rank Training of Foundation Models via Latent Crossing
by: Zhang, Ruijie, et al.
Published: (2025) -
FuRA: Full-Rank Parameter-Efficient Fine-Tuning with Spectral Preconditioning
by: Zhao, Yequan, et al.
Published: (2026)