Saved in:
| Main Author: | Yi, Albert |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.26459 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NuMuon: Nuclear-Norm-Constrained Muon for Compressible LLM Training
by: Dolatabadi, Hadi Mohaghegh, et al.
Published: (2026)
by: Dolatabadi, Hadi Mohaghegh, et al.
Published: (2026)
Muon is Scalable for LLM Training
by: Liu, Jingyuan, et al.
Published: (2025)
by: Liu, Jingyuan, et al.
Published: (2025)
MuLoCo: Muon is a practical inner optimizer for DiLoCo
by: Thérien, Benjamin, et al.
Published: (2025)
by: Thérien, Benjamin, et al.
Published: (2025)
SoftAdaClip: A Smooth Clipping Strategy for Fair and Private Model Training
by: Soleymani, Dorsa, et al.
Published: (2025)
by: Soleymani, Dorsa, et al.
Published: (2025)
Drop-Muon: Update Less, Converge Faster
by: Gruntkowska, Kaja, et al.
Published: (2025)
by: Gruntkowska, Kaja, et al.
Published: (2025)
Beyond the Ideal: Analyzing the Inexact Muon Update
by: Shulgin, Egor, et al.
Published: (2025)
by: Shulgin, Egor, et al.
Published: (2025)
Gradient Shaping Beyond Clipping: A Functional Perspective on Update Magnitude Control
by: You, Haochen, et al.
Published: (2025)
by: You, Haochen, et al.
Published: (2025)
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
by: Marshall, Noah, et al.
Published: (2024)
by: Marshall, Noah, et al.
Published: (2024)
AGGC: Adaptive Group Gradient Clipping for Stabilizing Large Language Model Training
by: Li, Zhiyuan, et al.
Published: (2026)
by: Li, Zhiyuan, et al.
Published: (2026)
AdaGrad Meets Muon: Adaptive Stepsizes for Orthogonal Updates
by: Zhang, Minxin, et al.
Published: (2025)
by: Zhang, Minxin, et al.
Published: (2025)
Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning
by: Lu, Binghang, et al.
Published: (2026)
by: Lu, Binghang, et al.
Published: (2026)
Robust and Fast Training via Per-Sample Clipping
by: Nobile, Davide, et al.
Published: (2026)
by: Nobile, Davide, et al.
Published: (2026)
AdaMuon: Adaptive Muon Optimizer
by: Si, Chongjie, et al.
Published: (2025)
by: Si, Chongjie, et al.
Published: (2025)
LionMuon: Alternating Spectral and Sign Descent for Efficient Training
by: Bolatov, Arman, et al.
Published: (2026)
by: Bolatov, Arman, et al.
Published: (2026)
Muon: Training and Trade-offs with Latent Attention and MoE
by: Mehta, Sushant, et al.
Published: (2025)
by: Mehta, Sushant, et al.
Published: (2025)
Regularized Gradient Clipping Provably Trains Wide and Deep Neural Networks
by: Tucat, Matteo, et al.
Published: (2024)
by: Tucat, Matteo, et al.
Published: (2024)
ProAct: Progressive Training for Hybrid Clipped Activation Function to Enhance Resilience of DNNs
by: Mousavi, Seyedhamidreza, et al.
Published: (2024)
by: Mousavi, Seyedhamidreza, et al.
Published: (2024)
Logits Replay + MoClip: Stabilized, Low-Cost Post-Training with Minimal Forgetting
by: Qiu, Suming, et al.
Published: (2025)
by: Qiu, Suming, et al.
Published: (2025)
PPO-Clip Attains Global Optimality: Towards Deeper Understandings of Clipping
by: Huang, Nai-Chieh, et al.
Published: (2023)
by: Huang, Nai-Chieh, et al.
Published: (2023)
ConfClip: Confidence-Weighted and Clipped Reward for Reinforcement Learning in LLMs
by: Zhang, Bonan, et al.
Published: (2025)
by: Zhang, Bonan, et al.
Published: (2025)
AdaDPIGU: Differentially Private SGD with Adaptive Clipping and Importance-Based Gradient Updates for Deep Neural Networks
by: Zhang, Huiqi, et al.
Published: (2025)
by: Zhang, Huiqi, et al.
Published: (2025)
MuViT: Multi-Resolution Vision Transformers for Learning Across Scales in Microscopy
by: Mantes, Albert Dominguez, et al.
Published: (2026)
by: Mantes, Albert Dominguez, et al.
Published: (2026)
Muon in Associative Memory Learning: Training Dynamics and Scaling Laws
by: Li, Binghui, et al.
Published: (2026)
by: Li, Binghui, et al.
Published: (2026)
Differentially Private Clipped-SGD: High-Probability Convergence with Arbitrary Clipping Level
by: Khah, Saleh Vatan, et al.
Published: (2025)
by: Khah, Saleh Vatan, et al.
Published: (2025)
GeoClip: Geometry-Aware Clipping for Differentially Private SGD
by: Gilani, Atefeh, et al.
Published: (2025)
by: Gilani, Atefeh, et al.
Published: (2025)
When Muon Optimizer Meets Adversarial Training: A Theoretical and Empirical Study
by: Yan, Jun, et al.
Published: (2026)
by: Yan, Jun, et al.
Published: (2026)
Clip-Low Increases Entropy and Clip-High Decreases Entropy in Reinforcement Learning of Large Language Models
by: Park, Jaesung R., et al.
Published: (2025)
by: Park, Jaesung R., et al.
Published: (2025)
HELENE: Hessian Layer-wise Clipping and Gradient Annealing for Accelerating Fine-tuning LLM with Zeroth-order Optimization
by: Zhao, Huaqin, et al.
Published: (2024)
by: Zhao, Huaqin, et al.
Published: (2024)
ClipFormer: Key-Value Clipping of Transformers on Memristive Crossbars for Write Noise Mitigation
by: Bhattacharjee, Abhiroop, et al.
Published: (2024)
by: Bhattacharjee, Abhiroop, et al.
Published: (2024)
Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
by: Li, Chenliang, et al.
Published: (2025)
by: Li, Chenliang, et al.
Published: (2025)
MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training
by: Li, Jiacheng, et al.
Published: (2026)
by: Li, Jiacheng, et al.
Published: (2026)
Phases of Muon: When Muon Eclipses SignSGD
by: Paquette, Elliot, et al.
Published: (2026)
by: Paquette, Elliot, et al.
Published: (2026)
NorMuon: Making Muon more efficient and scalable
by: Li, Zichong, et al.
Published: (2025)
by: Li, Zichong, et al.
Published: (2025)
On the Convergence of Muon and Beyond
by: Chang, Da, et al.
Published: (2025)
by: Chang, Da, et al.
Published: (2025)
It's Not You, It's Clipping: A Soft Trust-Region via Probability Smoothing for LLM RL
by: Dwyer, Madeleine, et al.
Published: (2025)
by: Dwyer, Madeleine, et al.
Published: (2025)
Adaptive-Boundary-Clipping GRPO: Ensuring Bounded Ratios for Stable and Generalizable Training
by: Liu, Chi, et al.
Published: (2026)
by: Liu, Chi, et al.
Published: (2026)
How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning
by: Tian, Minghao, et al.
Published: (2026)
by: Tian, Minghao, et al.
Published: (2026)
Advancing Model Refinement: Muon-Optimized Distillation and Quantization for LLM Deployment
by: Sander, Jacob, et al.
Published: (2026)
by: Sander, Jacob, et al.
Published: (2026)
Clipped SGD Algorithms for Performative Prediction: Tight Bounds for Clipping Bias and Remedies
by: Li, Qiang, et al.
Published: (2024)
by: Li, Qiang, et al.
Published: (2024)
Variance-Adaptive Muon: Accelerating LLM Pretraining with NSR-Modulated and Variance-Scaled Momentum
by: Li, Jingru, et al.
Published: (2026)
by: Li, Jingru, et al.
Published: (2026)
Similar Items
-
NuMuon: Nuclear-Norm-Constrained Muon for Compressible LLM Training
by: Dolatabadi, Hadi Mohaghegh, et al.
Published: (2026) -
Muon is Scalable for LLM Training
by: Liu, Jingyuan, et al.
Published: (2025) -
MuLoCo: Muon is a practical inner optimizer for DiLoCo
by: Thérien, Benjamin, et al.
Published: (2025) -
SoftAdaClip: A Smooth Clipping Strategy for Fair and Private Model Training
by: Soleymani, Dorsa, et al.
Published: (2025) -
Drop-Muon: Update Less, Converge Faster
by: Gruntkowska, Kaja, et al.
Published: (2025)