Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs)
Fuente:
arXiv
Saved in:
| Main Authors: | Riabinin, Artem, Shulgin, Egor, Gruntkowska, Kaja, Richtárik, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Local LMO: Constrained Gradient Optimization via a Local Linear Minimization Oracle
by: Richtárik, Peter, et al.
Published: (2026)
by: Richtárik, Peter, et al.
Published: (2026)
Non-Euclidean Broximal Point Method: A Blueprint for Geometry-Aware Optimization
by: Gruntkowska, Kaja, et al.
Published: (2025)
by: Gruntkowska, Kaja, et al.
Published: (2025)
Error Feedback for Muon and Friends
by: Gruntkowska, Kaja, et al.
Published: (2025)
by: Gruntkowska, Kaja, et al.
Published: (2025)
Drop-Muon: Update Less, Converge Faster
by: Gruntkowska, Kaja, et al.
Published: (2025)
by: Gruntkowska, Kaja, et al.
Published: (2025)
Improving the Worst-Case Bidirectional Communication Complexity for Nonconvex Distributed Optimization under Function Similarity
by: Gruntkowska, Kaja, et al.
Published: (2024)
by: Gruntkowska, Kaja, et al.
Published: (2024)
Freya PAGE: First Optimal Time Complexity for Large-Scale Nonconvex Finite-Sum Optimization with Heterogeneous Asynchronous Computations
by: Tyurin, Alexander, et al.
Published: (2024)
by: Tyurin, Alexander, et al.
Published: (2024)
Beyond the Ideal: Analyzing the Inexact Muon Update
by: Shulgin, Egor, et al.
Published: (2025)
by: Shulgin, Egor, et al.
Published: (2025)
Tighter Performance Theory of FedExProx
by: Anyszka, Wojciech, et al.
Published: (2024)
by: Anyszka, Wojciech, et al.
Published: (2024)
The Ball-Proximal (="Broximal") Point Method: a New Algorithm, Convergence Theory, and Applications
by: Gruntkowska, Kaja, et al.
Published: (2025)
by: Gruntkowska, Kaja, et al.
Published: (2025)
On the Convergence of DP-SGD with Adaptive Clipping
by: Shulgin, Egor, et al.
Published: (2024)
by: Shulgin, Egor, et al.
Published: (2024)
A Novel Unified Parametric Assumption for Nonconvex Optimization
by: Riabinin, Artem, et al.
Published: (2025)
by: Riabinin, Artem, et al.
Published: (2025)
First Provable Guarantees for Practical Private FL: Beyond Restrictive Assumptions
by: Shulgin, Egor, et al.
Published: (2025)
by: Shulgin, Egor, et al.
Published: (2025)
Towards a Better Theoretical Understanding of Independent Subnetwork Training
by: Shulgin, Egor, et al.
Published: (2023)
by: Shulgin, Egor, et al.
Published: (2023)
Smoothed Normalization for Efficient Distributed Private Optimization
by: Shulgin, Egor, et al.
Published: (2025)
by: Shulgin, Egor, et al.
Published: (2025)
Better LMO-based Momentum Methods with Second-Order Information
by: Khirirat, Sarit, et al.
Published: (2025)
by: Khirirat, Sarit, et al.
Published: (2025)
FedMuon: Federated Learning with Bias-corrected LMO-based Optimization
by: Takezawa, Yuki, et al.
Published: (2025)
by: Takezawa, Yuki, et al.
Published: (2025)
SPAM: Stochastic Proximal Point Method with Momentum Variance Reduction for Non-convex Cross-Device Federated Learning
by: Karagulyan, Avetik, et al.
Published: (2024)
by: Karagulyan, Avetik, et al.
Published: (2024)
Communication Compression for Byzantine Robust Learning: New Efficient Algorithms and Improved Rates
by: Rammal, Ahmad, et al.
Published: (2023)
by: Rammal, Ahmad, et al.
Published: (2023)
Broximal Alignment for Global Non-Convex Optimization
by: Gruntkowska, Kaja, et al.
Published: (2026)
by: Gruntkowska, Kaja, et al.
Published: (2026)
MAST: Model-Agnostic Sparsified Training
by: Demidovich, Yury, et al.
Published: (2023)
by: Demidovich, Yury, et al.
Published: (2023)
Stabilized Proximal Point Method via Trust Region Control
by: Li, Hanmin, et al.
Published: (2026)
by: Li, Hanmin, et al.
Published: (2026)
Ringmaster LMO: Asynchronous Linear Minimization Oracle Momentum Method
by: Sadiev, Abdurakhmon, et al.
Published: (2026)
by: Sadiev, Abdurakhmon, et al.
Published: (2026)
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
by: Riabinin, Artem, et al.
Published: (2026)
by: Riabinin, Artem, et al.
Published: (2026)
Muon is Provably Faster with Momentum Variance Reduction
by: Qian, Xun, et al.
Published: (2025)
by: Qian, Xun, et al.
Published: (2025)
MARINA-P: Superior Performance in Non-smooth Federated Optimization with Adaptive Stepsizes
by: Sokolov, Igor, et al.
Published: (2024)
by: Sokolov, Igor, et al.
Published: (2024)
Convergence Analysis of the PAGE Stochastic Algorithm for Weakly Convex Finite-Sum Optimization
by: Condat, Laurent, et al.
Published: (2025)
by: Condat, Laurent, et al.
Published: (2025)
A Unified Theory of Stochastic Proximal Point Methods without Smoothness
by: Richtárik, Peter, et al.
Published: (2024)
by: Richtárik, Peter, et al.
Published: (2024)
BiCoLoR: Communication-Efficient Optimization with Bidirectional Compression and Local Training
by: Condat, Laurent, et al.
Published: (2026)
by: Condat, Laurent, et al.
Published: (2026)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
by: Petrov, Egor, et al.
Published: (2025)
by: Petrov, Egor, et al.
Published: (2025)
Second-order Optimization under Heavy-Tailed Noise: Hessian Clipping and Sample Complexity Limits
by: Sadiev, Abdurakhmon, et al.
Published: (2025)
by: Sadiev, Abdurakhmon, et al.
Published: (2025)
A Computation and Communication Efficient Method for Distributed Nonconvex Problems in the Partial Participation Setting
by: Tyurin, Alexander, et al.
Published: (2022)
by: Tyurin, Alexander, et al.
Published: (2022)
TAMUNA: Doubly Accelerated Distributed Optimization with Local Training, Compression, and Partial Participation
by: Condat, Laurent, et al.
Published: (2023)
by: Condat, Laurent, et al.
Published: (2023)
LiMuon: Light and Fast Muon Optimizer for Large Models
by: Huang, Feihu, et al.
Published: (2025)
by: Huang, Feihu, et al.
Published: (2025)
Bridging Theory and Practice: A Stochastic Learning-Optimization Model for Resilient Automotive Supply Chains
by: Shahnawaz, Muhammad, et al.
Published: (2025)
by: Shahnawaz, Muhammad, et al.
Published: (2025)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
by: Nagashima, Shuntaro, et al.
Published: (2026)
by: Nagashima, Shuntaro, et al.
Published: (2026)
Muon Optimizes Under Spectral Norm Constraints
by: Chen, Lizhang, et al.
Published: (2025)
by: Chen, Lizhang, et al.
Published: (2025)
Lions and Muons: Optimization via Stochastic Frank-Wolfe
by: Sfyraki, Maria-Eleni, et al.
Published: (2025)
by: Sfyraki, Maria-Eleni, et al.
Published: (2025)
Sparse-ProxSkip: Accelerated Sparse-to-Sparse Training in Federated Learning
by: Meinhardt, Georg, et al.
Published: (2024)
by: Meinhardt, Georg, et al.
Published: (2024)
Shadowheart SGD: Distributed Asynchronous SGD with Optimal Time Complexity Under Arbitrary Computation and Communication Heterogeneity
by: Tyurin, Alexander, et al.
Published: (2024)
by: Tyurin, Alexander, et al.
Published: (2024)
The Newton-Muon Optimizer
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Similar Items
-
Local LMO: Constrained Gradient Optimization via a Local Linear Minimization Oracle
by: Richtárik, Peter, et al.
Published: (2026) -
Non-Euclidean Broximal Point Method: A Blueprint for Geometry-Aware Optimization
by: Gruntkowska, Kaja, et al.
Published: (2025) -
Error Feedback for Muon and Friends
by: Gruntkowska, Kaja, et al.
Published: (2025) -
Drop-Muon: Update Less, Converge Faster
by: Gruntkowska, Kaja, et al.
Published: (2025) -
Improving the Worst-Case Bidirectional Communication Complexity for Nonconvex Distributed Optimization under Function Similarity
by: Gruntkowska, Kaja, et al.
Published: (2024)