MuonEq: Balancing Before Orthogonalization with Lightweight Equilibration
Fuente:
arXiv
Saved in:
| Main Authors: | Chang, Da, Shi, Qiankun, Zhang, Lvgang, Li, Yu, Zhang, Ruijie, Lu, Yao, Liu, Yongxiang, Yuan, Ganzhao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Convergence of Muon and Beyond
by: Chang, Da, et al.
Published: (2025)
by: Chang, Da, et al.
Published: (2025)
AlphaAdam:Asynchronous Masked Optimization with Dynamic Alpha for Selective Updates
by: Chang, Da, et al.
Published: (2025)
by: Chang, Da, et al.
Published: (2025)
A Block Coordinate Descent Method for Nonsmooth Composite Optimization under Orthogonality Constraints
by: Yuan, Ganzhao
Published: (2023)
by: Yuan, Ganzhao
Published: (2023)
AMO: Adaptive Muon Orthogonalization
by: Zhuang, Xinlin, et al.
Published: (2026)
by: Zhuang, Xinlin, et al.
Published: (2026)
Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning
by: Lu, Binghang, et al.
Published: (2026)
by: Lu, Binghang, et al.
Published: (2026)
OptEMA: Adaptive Exponential Moving Average for Stochastic Optimization with Zero-Noise Optimality
by: Yuan, Ganzhao
Published: (2026)
by: Yuan, Ganzhao
Published: (2026)
ADMM for Structured Fractional Minimization
by: Yuan, Ganzhao
Published: (2024)
by: Yuan, Ganzhao
Published: (2024)
Adaptive Lipschitz-Free Conditional Gradient Methods for Stochastic Composite Nonconvex Optimization
by: Yuan, Ganzhao
Published: (2026)
by: Yuan, Ganzhao
Published: (2026)
MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization
by: Su, Yupeng, et al.
Published: (2026)
by: Su, Yupeng, et al.
Published: (2026)
Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning
by: Liu, Ziyue, et al.
Published: (2026)
by: Liu, Ziyue, et al.
Published: (2026)
MuonBP: Faster Muon via Block-Periodic Orthogonalization
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
by: Zhang, Minxin, et al.
Published: (2025)
by: Zhang, Minxin, et al.
Published: (2025)
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
by: He, Di, et al.
Published: (2025)
by: He, Di, et al.
Published: (2025)
Preconditioning Benefits of Spectral Orthogonalization in Muon
by: Ma, Jianhao, et al.
Published: (2026)
by: Ma, Jianhao, et al.
Published: (2026)
FedMuon: Accelerating Federated Learning with Matrix Orthogonalization
by: Liu, Junkang, et al.
Published: (2025)
by: Liu, Junkang, et al.
Published: (2025)
Calibrating and Rotating: A Unified Framework for Weight Conditioning in PEFT
by: Chang, Da, et al.
Published: (2025)
by: Chang, Da, et al.
Published: (2025)
TrasMuon: Trust-Region Adaptive Scaling for Orthogonalized Momentum Optimizers
by: Cheng, Peng, et al.
Published: (2026)
by: Cheng, Peng, et al.
Published: (2026)
Effect Decomposition of Functional-Output Computer Experiments via Orthogonal Additive Gaussian Processes
by: Tan, Yu, et al.
Published: (2025)
by: Tan, Yu, et al.
Published: (2025)
MUON+: Towards More Effective Muon via One Additional Normalization Step for LLM Pre-training
by: Zhang, Ruijie, et al.
Published: (2026)
by: Zhang, Ruijie, et al.
Published: (2026)
ADMM for Nonsmooth Composite Optimization under Orthogonality Constraints
by: Yuan, Ganzhao
Published: (2024)
by: Yuan, Ganzhao
Published: (2024)
TEON: Tensorized Orthonormalization Beyond Layer-Wise Muon for Large Language Model Pre-Training
by: Zhang, Ruijie, et al.
Published: (2026)
by: Zhang, Ruijie, et al.
Published: (2026)
A Single-Loop Bilevel Deep Learning Method for Optimal Control of Obstacle Problems
by: Song, Yongcun, et al.
Published: (2026)
by: Song, Yongcun, et al.
Published: (2026)
Orthogonal Subspace Clustering: Enhancing High-Dimensional Data Analysis through Adaptive Dimensionality Reduction and Efficient Clustering
by: Wen, Qing-Yuan, et al.
Published: (2026)
by: Wen, Qing-Yuan, et al.
Published: (2026)
DP-Muon: Differentially Private Optimization via Matrix-Orthogonalized Momentum
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
AdaMuon: Adaptive Muon Optimizer
by: Si, Chongjie, et al.
Published: (2025)
by: Si, Chongjie, et al.
Published: (2025)
Lightweight Transformer for EEG Classification via Balanced Signed Graph Algorithm Unrolling
by: Yao, Junyi, et al.
Published: (2025)
by: Yao, Junyi, et al.
Published: (2025)
Sparse Orthogonal Parameters Tuning for Continual Learning
by: Ning, Kun-Peng, et al.
Published: (2024)
by: Ning, Kun-Peng, et al.
Published: (2024)
AURORA: Contextual Orthogonalization for Geometric Representation Learning in Healthcare Foundation Models
by: Zhang, Yuanyun, et al.
Published: (2026)
by: Zhang, Yuanyun, et al.
Published: (2026)
Adam Improves Muon: Adaptive Moment Estimation with Orthogonalized Momentum
by: Zhang, Minxin, et al.
Published: (2026)
by: Zhang, Minxin, et al.
Published: (2026)
Muon with Spectral Guidance: Efficient Optimization for Scientific Machine Learning
by: Lu, Binghang, et al.
Published: (2026)
by: Lu, Binghang, et al.
Published: (2026)
When Does Value-Aware KV Eviction Help? A Fixed-Contract Diagnostic for Non-Monotone Cache Compression
by: Zhang, Ruijie, et al.
Published: (2026)
by: Zhang, Ruijie, et al.
Published: (2026)
RED-DiffEq: Regularization by denoising diffusion models for solving inverse PDE problems with application to full waveform inversion
by: Shan, Siming, et al.
Published: (2025)
by: Shan, Siming, et al.
Published: (2025)
Recent Advances of NeuroDiffEq -- An Open-Source Library for Physics-Informed Neural Networks
by: Liu, Shuheng, et al.
Published: (2025)
by: Liu, Shuheng, et al.
Published: (2025)
BI-EqNO: Generalized Approximate Bayesian Inference with an Equivariant Neural Operator Framework
by: Zhou, Xu-Hui, et al.
Published: (2024)
by: Zhou, Xu-Hui, et al.
Published: (2024)
EqDeepRx: Learning a Scalable MIMO Receiver
by: Honkala, Mikko, et al.
Published: (2026)
by: Honkala, Mikko, et al.
Published: (2026)
G2LoRA: Gradient Orthogonal Low-Rank Adaptation Framework for Graph Continual Learning on Text-Attributed Graphs
by: Wang, Yuhan, et al.
Published: (2026)
by: Wang, Yuhan, et al.
Published: (2026)
Orthogonal Model Merging
by: Yang, Sihan, et al.
Published: (2026)
by: Yang, Sihan, et al.
Published: (2026)
Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence
by: Nguyen, Tien-Phat, et al.
Published: (2026)
by: Nguyen, Tien-Phat, et al.
Published: (2026)
NorMuon: Making Muon more efficient and scalable
by: Li, Zichong, et al.
Published: (2025)
by: Li, Zichong, et al.
Published: (2025)
On Provable Benefits of Muon in Federated Learning
by: Zhang, Xinwen, et al.
Published: (2025)
by: Zhang, Xinwen, et al.
Published: (2025)
Similar Items
-
On the Convergence of Muon and Beyond
by: Chang, Da, et al.
Published: (2025) -
AlphaAdam:Asynchronous Masked Optimization with Dynamic Alpha for Selective Updates
by: Chang, Da, et al.
Published: (2025) -
A Block Coordinate Descent Method for Nonsmooth Composite Optimization under Orthogonality Constraints
by: Yuan, Ganzhao
Published: (2023) -
AMO: Adaptive Muon Orthogonalization
by: Zhuang, Xinlin, et al.
Published: (2026) -
Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning
by: Lu, Binghang, et al.
Published: (2026)