LionMuon: Alternating Spectral and Sign Descent for Efficient Training
Fuente:
arXiv
Saved in:
| Main Authors: | Bolatov, Arman, Riabinin, Artem, Kornilov, Nikita, Veprikov, Andrey, Horváth, Samuel, Takáč, Martin, Beznosikov, Aleksandr |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
by: Riabinin, Artem, et al.
Published: (2026)
by: Riabinin, Artem, et al.
Published: (2026)
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
by: Veprikov, Andrey, et al.
Published: (2025)
by: Veprikov, Andrey, et al.
Published: (2025)
Faster Than SVD, Smarter Than SGD: The OPLoRA Alternating Update
by: Almansoori, Abdulla Jasem, et al.
Published: (2025)
by: Almansoori, Abdulla Jasem, et al.
Published: (2025)
FRUGAL: Memory-Efficient Optimization by Reducing State Overhead for Scalable Training
by: Zmushko, Philip, et al.
Published: (2024)
by: Zmushko, Philip, et al.
Published: (2024)
Beyond SGD, Without SVD: Proximal Subspace Iteration LoRA with Diagonal Fractional K-FAC
by: Almansoori, Abdulla Jasem, et al.
Published: (2026)
by: Almansoori, Abdulla Jasem, et al.
Published: (2026)
Byzantine-Robust Optimization under $(L_0, L_1)$-Smoothness
by: Bolatov, Arman, et al.
Published: (2026)
by: Bolatov, Arman, et al.
Published: (2026)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
by: Petrov, Egor, et al.
Published: (2025)
by: Petrov, Egor, et al.
Published: (2025)
Softsign: Smooth Sign in Your Optimizer For Better Parameter Heterogeneity Handling
by: Feoktistov, Dmitrii, et al.
Published: (2026)
by: Feoktistov, Dmitrii, et al.
Published: (2026)
Random-reshuffled SARAH does not need a full gradient computations
by: Beznosikov, Aleksandr, et al.
Published: (2021)
by: Beznosikov, Aleksandr, et al.
Published: (2021)
WeightLoRA: Keep Only Necessary Adapters
by: Veprikov, Andrey, et al.
Published: (2025)
by: Veprikov, Andrey, et al.
Published: (2025)
Who to Trust? Aggregating Client Predictions in Federated Distillation
by: Kovalchuk, Viktor, et al.
Published: (2025)
by: Kovalchuk, Viktor, et al.
Published: (2025)
Similarity, Compression and Local Steps: Three Pillars of Efficient Communications for Distributed Variational Inequalities
by: Beznosikov, Aleksandr, et al.
Published: (2023)
by: Beznosikov, Aleksandr, et al.
Published: (2023)
Aligning Distributionally Robust Optimization with Practical Deep Learning Needs
by: Feoktistov, Dmitrii, et al.
Published: (2025)
by: Feoktistov, Dmitrii, et al.
Published: (2025)
Zero-Order Optimization for LLM Fine-Tuning via Learnable Direction Sampling
by: Parfenov, Valery, et al.
Published: (2026)
by: Parfenov, Valery, et al.
Published: (2026)
Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs)
by: Riabinin, Artem, et al.
Published: (2025)
by: Riabinin, Artem, et al.
Published: (2025)
Collaborative and Efficient Personalization with Mixtures of Adaptors
by: Almansoori, Abdulla Jasem, et al.
Published: (2024)
by: Almansoori, Abdulla Jasem, et al.
Published: (2024)
Sign Operator for Coping with Heavy-Tailed Noise in Non-Convex Optimization: High Probability Bounds Under $(L_0, L_1)$-Smoothness
by: Kornilov, Nikita, et al.
Published: (2025)
by: Kornilov, Nikita, et al.
Published: (2025)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
by: Chezhegov, Savelii, et al.
Published: (2024)
by: Chezhegov, Savelii, et al.
Published: (2024)
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
by: Chezhegov, Savelii, et al.
Published: (2025)
by: Chezhegov, Savelii, et al.
Published: (2025)
Gradient Descent Fails to Learn High-frequency Functions and Modular Arithmetic
by: Takhanov, Rustem, et al.
Published: (2023)
by: Takhanov, Rustem, et al.
Published: (2023)
On Biased Compression for Distributed Learning
by: Beznosikov, Aleksandr, et al.
Published: (2020)
by: Beznosikov, Aleksandr, et al.
Published: (2020)
Efficient Conformal Prediction under Data Heterogeneity
by: Plassier, Vincent, et al.
Published: (2023)
by: Plassier, Vincent, et al.
Published: (2023)
Simplex Deep Linear Discriminant Analysis
by: Tezekbayev, Maxat, et al.
Published: (2026)
by: Tezekbayev, Maxat, et al.
Published: (2026)
Generalising Battery Control in Net-Zero Buildings via Personalised Federated RL
by: Avila, Nicolas M Cuadrado, et al.
Published: (2024)
by: Avila, Nicolas M Cuadrado, et al.
Published: (2024)
Thinking like a CHEMIST: Combined Heterogeneous Embedding Model Integrating Structure and Tokens
by: Rekut, Nikolai, et al.
Published: (2025)
by: Rekut, Nikolai, et al.
Published: (2025)
Simple Stepsize for Quasi-Newton Methods with Global Convergence Guarantees
by: Agafonov, Artem, et al.
Published: (2025)
by: Agafonov, Artem, et al.
Published: (2025)
On the Equivalence of Optimal Transport Problem and Action Matching with Optimal Vector Fields
by: Kornilov, Nikita, et al.
Published: (2025)
by: Kornilov, Nikita, et al.
Published: (2025)
New Aspects of Black Box Conditional Gradient: Variance Reduction and One Point Feedback
by: Veprikov, Andrey, et al.
Published: (2024)
by: Veprikov, Andrey, et al.
Published: (2024)
PaDPaF: Partial Disentanglement with Partially-Federated GANs
by: Almansoori, Abdulla Jasem, et al.
Published: (2022)
by: Almansoori, Abdulla Jasem, et al.
Published: (2022)
A Novel Unified Parametric Assumption for Nonconvex Optimization
by: Riabinin, Artem, et al.
Published: (2025)
by: Riabinin, Artem, et al.
Published: (2025)
Stochastic Gradient Methods with Preconditioned Updates
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
by: Sadiev, Abdurakhmon, et al.
Published: (2022)
Federated Learning Can Find Friends That Are Advantageous
by: Tupitsa, Nazarii, et al.
Published: (2024)
by: Tupitsa, Nazarii, et al.
Published: (2024)
Generalized Policy Learning for Smart Grids: FL TRPO Approach
by: Li, Yunxiang, et al.
Published: (2024)
by: Li, Yunxiang, et al.
Published: (2024)
DyKAF: Dynamical Kronecker Approximation of the Fisher Information Matrix for Gradient Preconditioning
by: Yudin, Nikolay, et al.
Published: (2025)
by: Yudin, Nikolay, et al.
Published: (2025)
Can Muon Fine-tune Adam-Pretrained Models?
by: Qu, Xingyu, et al.
Published: (2026)
by: Qu, Xingyu, et al.
Published: (2026)
FedPeWS: Personalized Warmup via Subnetworks for Enhanced Heterogeneous Federated Learning
by: Tastan, Nurbek, et al.
Published: (2024)
by: Tastan, Nurbek, et al.
Published: (2024)
Deep Linear Discriminant Analysis Revisited
by: Tezekbayev, Maxat, et al.
Published: (2026)
by: Tezekbayev, Maxat, et al.
Published: (2026)
Ito Diffusion Approximation of Universal Ito Chains for Sampling, Optimization and Boosting
by: Ustimenko, Aleksei, et al.
Published: (2023)
by: Ustimenko, Aleksei, et al.
Published: (2023)
SignMuon: Communication-Efficient Distributed Muon Optimization
by: Mishra, Neel, et al.
Published: (2026)
by: Mishra, Neel, et al.
Published: (2026)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)
by: Fan, Chen, et al.
Published: (2025)
Similar Items
-
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
by: Riabinin, Artem, et al.
Published: (2026) -
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
by: Veprikov, Andrey, et al.
Published: (2025) -
Faster Than SVD, Smarter Than SGD: The OPLoRA Alternating Update
by: Almansoori, Abdulla Jasem, et al.
Published: (2025) -
FRUGAL: Memory-Efficient Optimization by Reducing State Overhead for Scalable Training
by: Zmushko, Philip, et al.
Published: (2024) -
Beyond SGD, Without SVD: Proximal Subspace Iteration LoRA with Diagonal Fractional K-FAC
by: Almansoori, Abdulla Jasem, et al.
Published: (2026)