HTMuon: Improving Muon via Heavy-Tailed Spectral Correction
Fuente:
arXiv
Saved in:
| Main Authors: | Pang, Tianyu, Fang, Yujie, Liu, Zihang, Deng, Shenyang, Hsiung, Lei, Yu, Shuhua, Yang, Yaoqing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Spikes to Heavy Tails: Unveiling the Spectral Evolution of Neural Networks
by: Kothapalli, Vignesh, et al.
Published: (2024)
by: Kothapalli, Vignesh, et al.
Published: (2024)
RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization
by: Deng, Shenyang, et al.
Published: (2026)
by: Deng, Shenyang, et al.
Published: (2026)
Learning to Discover Iterative Spectral Algorithms
by: Liu, Zihang, et al.
Published: (2026)
by: Liu, Zihang, et al.
Published: (2026)
Depth, Not Data: An Analysis of Hessian Spectral Bifurcation
by: Deng, Shenyang, et al.
Published: (2026)
by: Deng, Shenyang, et al.
Published: (2026)
LIFT the Veil for the Truth: Principal Weights Emerge after Rank Reduction for Reasoning-Focused Supervised Fine-Tuning
by: Liu, Zihang, et al.
Published: (2025)
by: Liu, Zihang, et al.
Published: (2025)
Muon Outperforms Adam in Tail-End Associative Memory Learning
by: Wang, Shuche, et al.
Published: (2025)
by: Wang, Shuche, et al.
Published: (2025)
PRISM: Structured Optimization via Anisotropic Spectral Shaping
by: Yang, Yujie
Published: (2026)
by: Yang, Yujie
Published: (2026)
DynMuon: A Dynamic Spectral Shaping View of Muon
by: Wu, Fangzhou, et al.
Published: (2026)
by: Wu, Fangzhou, et al.
Published: (2026)
Iterative Refinement Neural Operators are Learned Fixed-Point Solvers: A Principled Approach to Spectral Bias Mitigation
by: Liu, Xiaotian, et al.
Published: (2026)
by: Liu, Xiaotian, et al.
Published: (2026)
Preconditioning Benefits of Spectral Orthogonalization in Muon
by: Ma, Jianhao, et al.
Published: (2026)
by: Ma, Jianhao, et al.
Published: (2026)
Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning
by: Liu, Ziyue, et al.
Published: (2026)
by: Liu, Ziyue, et al.
Published: (2026)
Suspicious Alignment of SGD: A Fine-Grained Step Size Condition Analysis
by: Deng, Shenyang, et al.
Published: (2026)
by: Deng, Shenyang, et al.
Published: (2026)
Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks
by: Pang, Tianyu, et al.
Published: (2026)
by: Pang, Tianyu, et al.
Published: (2026)
Adaptive Heavy-Tailed Stochastic Gradient Descent
by: Gong, Bodu, et al.
Published: (2025)
by: Gong, Bodu, et al.
Published: (2025)
Intrinsic Muon: Spectral Optimization on Riemannian Matrix Manifolds
by: Li, Yibang, et al.
Published: (2026)
by: Li, Yibang, et al.
Published: (2026)
TailedTS: Benchmark Dataset for Heavy-Tailed Time Series Prediction and Periodicity Quantification
by: Chen, Xinyu, et al.
Published: (2026)
by: Chen, Xinyu, et al.
Published: (2026)
Robust Offline Reinforcement learning with Heavy-Tailed Rewards
by: Zhu, Jin, et al.
Published: (2023)
by: Zhu, Jin, et al.
Published: (2023)
REFORMER: A ChatGPT-Driven Data Synthesis Framework Elevating Text-to-SQL Models
by: Liu, Shenyang, et al.
Published: (2025)
by: Liu, Shenyang, et al.
Published: (2025)
Muon Dynamics as a Spectral Wasserstein Flow
by: Peyré, Gabriel
Published: (2026)
by: Peyré, Gabriel
Published: (2026)
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
by: Huang, Feihu, et al.
Published: (2026)
by: Huang, Feihu, et al.
Published: (2026)
$(ε, u)$-Adaptive Regret Minimization in Heavy-Tailed Bandits
by: Genalti, Gianmarco, et al.
Published: (2023)
by: Genalti, Gianmarco, et al.
Published: (2023)
A Mathematics Framework of Artificial Shifted Population Risk and Its Further Understanding Related to Consistency Regularization
by: Yang, Xiliang, et al.
Published: (2025)
by: Yang, Xiliang, et al.
Published: (2025)
Eigenspectrum Analysis of Neural Networks without Aspect Ratio Bias
by: Hu, Yuanzhe, et al.
Published: (2025)
by: Hu, Yuanzhe, et al.
Published: (2025)
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
by: He, Di, et al.
Published: (2026)
by: He, Di, et al.
Published: (2026)
Robust Federated Learning Over the Air: Combating Heavy-Tailed Noise with Median Anchored Clipping
by: Li, Jiaxing, et al.
Published: (2024)
by: Li, Jiaxing, et al.
Published: (2024)
Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence
by: Nguyen, Tien-Phat, et al.
Published: (2026)
by: Nguyen, Tien-Phat, et al.
Published: (2026)
Phase-Type Variational Autoencoders for Heavy-Tailed Data
by: Ziani, Abdelhakim, et al.
Published: (2026)
by: Ziani, Abdelhakim, et al.
Published: (2026)
Markov Chain Decoders Overcome the Heavy-Tail Limitations of Lipschitz Generative Models
by: Ziani, Abdelhakim, et al.
Published: (2026)
by: Ziani, Abdelhakim, et al.
Published: (2026)
Heavy-Tailed Class-Conditional Priors for Long-Tailed Generative Modeling
by: Bouayed, Aymene Mohammed, et al.
Published: (2025)
by: Bouayed, Aymene Mohammed, et al.
Published: (2025)
StyleRec: A Benchmark Dataset for Prompt Recovery in Writing Style Transformation
by: Liu, Shenyang, et al.
Published: (2025)
by: Liu, Shenyang, et al.
Published: (2025)
Tackling Heavy-Tailed Rewards in Reinforcement Learning with Function Approximation: Minimax Optimal and Instance-Dependent Regret Bounds
by: Huang, Jiayi, et al.
Published: (2023)
by: Huang, Jiayi, et al.
Published: (2023)
Spectral Insights into Data-Oblivious Critical Layers in Large Language Models
by: Liu, Xuyuan, et al.
Published: (2025)
by: Liu, Xuyuan, et al.
Published: (2025)
PI-Mamba: Linear-Time Protein Backbone Generation via Spectrally Initialized Flow Matching
by: Wu, Tianyu, et al.
Published: (2026)
by: Wu, Tianyu, et al.
Published: (2026)
The Malignant Tail: Spectral Segregation of Label Noise in Over-Parameterized Networks
by: Wang, Zice
Published: (2026)
by: Wang, Zice
Published: (2026)
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
by: He, Di, et al.
Published: (2025)
by: He, Di, et al.
Published: (2025)
Unveiling Multi-regime Patterns in SciML: Distinct Failure Modes and Regime-specific Optimization
by: Wang, Yuxin, et al.
Published: (2026)
by: Wang, Yuxin, et al.
Published: (2026)
KCES: Training-Free Defense for Robust Graph Neural Networks via Kernel Complexity
by: Jia, Yaning, et al.
Published: (2025)
by: Jia, Yaning, et al.
Published: (2025)
Provable Benefit of Sign Descent: A Minimal Model Under Heavy-Tailed Class Imbalance
by: Yadav, Robin, et al.
Published: (2025)
by: Yadav, Robin, et al.
Published: (2025)
Rectify and Align GPS Points to Parking Spots via Rank-1 Constraint
by: Deng, Jiaxing, et al.
Published: (2025)
by: Deng, Jiaxing, et al.
Published: (2025)
FedMuon: Accelerating Federated Learning with Matrix Orthogonalization
by: Liu, Junkang, et al.
Published: (2025)
by: Liu, Junkang, et al.
Published: (2025)
Similar Items
-
From Spikes to Heavy Tails: Unveiling the Spectral Evolution of Neural Networks
by: Kothapalli, Vignesh, et al.
Published: (2024) -
RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization
by: Deng, Shenyang, et al.
Published: (2026) -
Learning to Discover Iterative Spectral Algorithms
by: Liu, Zihang, et al.
Published: (2026) -
Depth, Not Data: An Analysis of Hessian Spectral Bifurcation
by: Deng, Shenyang, et al.
Published: (2026) -
LIFT the Veil for the Truth: Principal Weights Emerge after Rank Reduction for Reasoning-Focused Supervised Fine-Tuning
by: Liu, Zihang, et al.
Published: (2025)