LionMuon: Alternating Spectral and Sign Descent for Efficient Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bolatov, Arman, Riabinin, Artem, Kornilov, Nikita, Veprikov, Andrey, Horváth, Samuel, Takáč, Martin, Beznosikov, Aleksandr
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911698457198592
author Bolatov, Arman
Riabinin, Artem
Kornilov, Nikita
Veprikov, Andrey
Horváth, Samuel
Takáč, Martin
Beznosikov, Aleksandr
author_facet Bolatov, Arman
Riabinin, Artem
Kornilov, Nikita
Veprikov, Andrey
Horváth, Samuel
Takáč, Martin
Beznosikov, Aleksandr
contents In large-scale optimization, the cheapness and effectiveness of update steps are the most crucial factors for a successful optimizer. Sign-based optimizers like Lion or Signum produce cheap per-step updates, whereas Muon's spectral matrix-sign update gives a much stronger direction at a substantially higher per-step cost. In this work, we propose LionMuon, which retains the effectiveness of Muon steps while considerably cutting the averaged iteration cost, similar to sign-based methods. It alternates between Lion's and Muon's updates on a fixed period P, sharing a single dual-EMA momentum buffer between them. The optimizer state memory therefore matches Lion and is exactly half of AdamW's. A simpler single-EMA variant, SignMuon, by itself already outperforms pure Muon. At P = 2, LionMuon Pareto-dominates Muon, Lion, Signum, and AdamW on every dataset and architecture we tested at 124M model size, reaching lower validation loss at lower compute, and the same advantage persists at 355M and 720M scale. On the theory side, we prove sharp complexity bounds under heavy-tailed noise which are governed by period-averaged smoothness and noise that interpolate between Muon's and Lion's constants. These bounds predict the compute-optimal period and the conditions under which LionMuon outruns Muon and Lion. Code: https://github.com/brain-lab-research/lion-muon
format Preprint
id arxiv_https___arxiv_org_abs_2605_19811
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LionMuon: Alternating Spectral and Sign Descent for Efficient Training
Bolatov, Arman
Riabinin, Artem
Kornilov, Nikita
Veprikov, Andrey
Horváth, Samuel
Takáč, Martin
Beznosikov, Aleksandr
Machine Learning
In large-scale optimization, the cheapness and effectiveness of update steps are the most crucial factors for a successful optimizer. Sign-based optimizers like Lion or Signum produce cheap per-step updates, whereas Muon's spectral matrix-sign update gives a much stronger direction at a substantially higher per-step cost. In this work, we propose LionMuon, which retains the effectiveness of Muon steps while considerably cutting the averaged iteration cost, similar to sign-based methods. It alternates between Lion's and Muon's updates on a fixed period P, sharing a single dual-EMA momentum buffer between them. The optimizer state memory therefore matches Lion and is exactly half of AdamW's. A simpler single-EMA variant, SignMuon, by itself already outperforms pure Muon. At P = 2, LionMuon Pareto-dominates Muon, Lion, Signum, and AdamW on every dataset and architecture we tested at 124M model size, reaching lower validation loss at lower compute, and the same advantage persists at 355M and 720M scale. On the theory side, we prove sharp complexity bounds under heavy-tailed noise which are governed by period-averaged smoothness and noise that interpolate between Muon's and Lion's constants. These bounds predict the compute-optimal period and the conditions under which LionMuon outruns Muon and Lion. Code: https://github.com/brain-lab-research/lion-muon
title LionMuon: Alternating Spectral and Sign Descent for Efficient Training
topic Machine Learning
url https://arxiv.org/abs/2605.19811