Dion2: A Simple Method to Shrink Matrix in Muon

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ahn, Kwangjun, Amsel, Noah, Langford, John
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912775711752192
author Ahn, Kwangjun
Amsel, Noah
Langford, John
author_facet Ahn, Kwangjun
Amsel, Noah
Langford, John
contents The Muon optimizer enjoys strong empirical performance and theoretical grounding. However, the super-linear cost of its orthonormalization step introduces increasing overhead with scale. To alleviate this cost, several works have attempted to reduce the size of the matrix entering the orthonormalization step. We introduce Dion2, a much simpler method for shrinking the matrix involved in Muon's computation compared to prior approaches. At a high level, Dion2 selects a fraction of rows or columns at each iteration and orthonormalizes only those. This sampling procedure makes the update sparse, reducing both computation and communication costs which in turn improves the scalability of Muon.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16928
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dion2: A Simple Method to Shrink Matrix in Muon
Ahn, Kwangjun
Amsel, Noah
Langford, John
Machine Learning
Distributed, Parallel, and Cluster Computing
The Muon optimizer enjoys strong empirical performance and theoretical grounding. However, the super-linear cost of its orthonormalization step introduces increasing overhead with scale. To alleviate this cost, several works have attempted to reduce the size of the matrix entering the orthonormalization step. We introduce Dion2, a much simpler method for shrinking the matrix involved in Muon's computation compared to prior approaches. At a high level, Dion2 selects a fraction of rows or columns at each iteration and orthonormalizes only those. This sampling procedure makes the update sparse, reducing both computation and communication costs which in turn improves the scalability of Muon.
title Dion2: A Simple Method to Shrink Matrix in Muon
topic Machine Learning
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2512.16928