Convergence Bound and Critical Batch Size of Muon Optimizer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sato, Naoki, Naganuma, Hiroki, Iiduka, Hideaki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum
von: Kamo, Keisuke, et al.
Veröffentlicht: (2025)
von: Kamo, Keisuke, et al.
Veröffentlicht: (2025)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
von: Nagashima, Shuntaro, et al.
Veröffentlicht: (2026)
von: Nagashima, Shuntaro, et al.
Veröffentlicht: (2026)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
von: Oowada, Kanata, et al.
Veröffentlicht: (2025)
von: Oowada, Kanata, et al.
Veröffentlicht: (2025)
Explicit and Implicit Graduated Optimization in Deep Neural Networks
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
von: Iiduka, Hideaki
Veröffentlicht: (2026)
von: Iiduka, Hideaki
Veröffentlicht: (2026)
Both Asymptotic and Non-Asymptotic Convergence of Quasi-Hyperbolic Momentum using Increasing Batch Size
von: Imaizumi, Kento, et al.
Veröffentlicht: (2025)
von: Imaizumi, Kento, et al.
Veröffentlicht: (2025)
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
von: Harada, Hinata, et al.
Veröffentlicht: (2024)
von: Harada, Hinata, et al.
Veröffentlicht: (2024)
Using Stochastic Gradient Descent to Smooth Nonconvex Functions: Analysis of Implicit Graduated Optimization
von: Sato, Naoki, et al.
Veröffentlicht: (2023)
von: Sato, Naoki, et al.
Veröffentlicht: (2023)
Lipschitz Multiscale Deep Equilibrium Models: A Theoretically Guaranteed and Accelerated Approach
von: Sato, Naoki, et al.
Veröffentlicht: (2026)
von: Sato, Naoki, et al.
Veröffentlicht: (2026)
Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
Scaled Conjugate Gradient Method for Nonconvex Optimization in Deep Neural Networks
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
Accelerating SGDM via Learning Rate and Batch Size Schedules: A Lyapunov-Based Analysis
von: Kondo, Yuichi, et al.
Veröffentlicht: (2025)
von: Kondo, Yuichi, et al.
Veröffentlicht: (2025)
Relationship between Batch Size and Number of Steps Needed for Nonconvex Optimization of Stochastic Gradient Descent using Armijo Line Search
von: Tsukada, Yuki, et al.
Veröffentlicht: (2023)
von: Tsukada, Yuki, et al.
Veröffentlicht: (2023)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
von: Umeda, Hikaru, et al.
Veröffentlicht: (2024)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2024)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
von: Umeda, Hikaru, et al.
Veröffentlicht: (2025)
Convergence Analysis of SGD under Expected Smoothness
von: Kawamoto, Yuta, et al.
Veröffentlicht: (2025)
von: Kawamoto, Yuta, et al.
Veröffentlicht: (2025)
Iteration and Stochastic First-order Oracle Complexities of Stochastic Gradient Descent using Constant and Decaying Learning Rates
von: Imaizumi, Kento, et al.
Veröffentlicht: (2024)
von: Imaizumi, Kento, et al.
Veröffentlicht: (2024)
Mini-Batch Stochastic Halpern Algorithm for Nonexpansive Fixed point Problems
von: Iiduka, Hideaki
Veröffentlicht: (2026)
von: Iiduka, Hideaki
Veröffentlicht: (2026)
Mini-Batch Stochastic Krasnosel'ski\uı-Mann Algorithm for Nonexpansive Fixed Point Problems
von: Iiduka, Hideaki
Veröffentlicht: (2026)
von: Iiduka, Hideaki
Veröffentlicht: (2026)
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
von: Naganuma, Hiroki, et al.
Veröffentlicht: (2026)
von: Naganuma, Hiroki, et al.
Veröffentlicht: (2026)
On the Convergence of Muon and Beyond
von: Chang, Da, et al.
Veröffentlicht: (2025)
von: Chang, Da, et al.
Veröffentlicht: (2025)
Towards Understanding Variants of Invariant Risk Minimization through the Lens of Calibration
von: Yoshida, Kotaro, et al.
Veröffentlicht: (2024)
von: Yoshida, Kotaro, et al.
Veröffentlicht: (2024)
Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training
von: Merrill, William, et al.
Veröffentlicht: (2025)
von: Merrill, William, et al.
Veröffentlicht: (2025)
Geometric Insights into Focal Loss: Reducing Curvature for Enhanced Model Calibration
von: Kimura, Masanari, et al.
Veröffentlicht: (2024)
von: Kimura, Masanari, et al.
Veröffentlicht: (2024)
AdaMuon: Adaptive Muon Optimizer
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
Convergence of Muon with Newton-Schulz
von: Kim, Gyu Yeol, et al.
Veröffentlicht: (2026)
von: Kim, Gyu Yeol, et al.
Veröffentlicht: (2026)
On the Convergence Analysis of Muon
von: Shen, Wei, et al.
Veröffentlicht: (2025)
von: Shen, Wei, et al.
Veröffentlicht: (2025)
A Deep State-Space Model Compression Method using Upper Bound on Output Error
von: Sakamoto, Hiroki, et al.
Veröffentlicht: (2025)
von: Sakamoto, Hiroki, et al.
Veröffentlicht: (2025)
How Does Critical Batch Size Scale in Pre-training?
von: Zhang, Hanlin, et al.
Veröffentlicht: (2024)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2024)
Collaborative Batch Size Optimization for Federated Learning
von: Geimer, Arno, et al.
Veröffentlicht: (2025)
von: Geimer, Arno, et al.
Veröffentlicht: (2025)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
von: Ostroukhov, Petr, et al.
Veröffentlicht: (2024)
von: Ostroukhov, Petr, et al.
Veröffentlicht: (2024)
Takeuchi's Information Criteria as Generalization Measures for DNNs Close to NTK Regime
von: Naganuma, Hiroki, et al.
Veröffentlicht: (2026)
von: Naganuma, Hiroki, et al.
Veröffentlicht: (2026)
Drop-Muon: Update Less, Converge Faster
von: Gruntkowska, Kaja, et al.
Veröffentlicht: (2025)
von: Gruntkowska, Kaja, et al.
Veröffentlicht: (2025)
Muon Does Not Converge on Convex Lipschitz Functions
von: Parshakova, Tetiana, et al.
Veröffentlicht: (2026)
von: Parshakova, Tetiana, et al.
Veröffentlicht: (2026)
Full-Graph vs. Mini-Batch Training: Comprehensive Analysis from a Batch Size and Fan-Out Size Perspective
von: Liu, Mengfan, et al.
Veröffentlicht: (2026)
von: Liu, Mengfan, et al.
Veröffentlicht: (2026)
Batch-Size Independent Regret Bounds for Combinatorial Semi-Bandits with Probabilistically Triggered Arms or Independent Arms
von: Liu, Xutong, et al.
Veröffentlicht: (2022)
von: Liu, Xutong, et al.
Veröffentlicht: (2022)
LiMuon: Light and Fast Muon Optimizer for Large Models
von: Huang, Feihu, et al.
Veröffentlicht: (2025)
von: Huang, Feihu, et al.
Veröffentlicht: (2025)
Effective Quantization of Muon Optimizer States
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization
von: Su, Yupeng, et al.
Veröffentlicht: (2026)
von: Su, Yupeng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum
von: Kamo, Keisuke, et al.
Veröffentlicht: (2025) -
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
von: Nagashima, Shuntaro, et al.
Veröffentlicht: (2026) -
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
von: Oowada, Kanata, et al.
Veröffentlicht: (2025) -
Explicit and Implicit Graduated Optimization in Deep Neural Networks
von: Sato, Naoki, et al.
Veröffentlicht: (2024) -
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
von: Iiduka, Hideaki
Veröffentlicht: (2026)