When and Why Grouping Attention Heads Accelerates Muon Optimization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Hongtao, Zhou, Wenjie, Chen, Wei, Cheng, Xueqi
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915997613555712
author Zhang, Hongtao
Zhou, Wenjie
Chen, Wei
Cheng, Xueqi
author_facet Zhang, Hongtao
Zhou, Wenjie
Chen, Wei
Cheng, Xueqi
contents Muon orthogonalizes matrix updates, but multi-head attention naturally operates at the level of heads. This granularity mismatch raises the question of whether Muon should be applied to the full attention projection, to individual heads, or to intermediate head groups. We study this question through a one-step descent comparison between full-matrix Muon and group-wise Muon. Our analysis reveals a trade-off between the \textbf{group-wise whitening gain} from group-wise updates and the \textbf{grouping-induced norm cost}, an additional update-norm cost caused by replacing full-matrix whitening with group-wise whitening. Motivated by this trade-off, we propose \textbf{Group Muon}, which treats head group size and grouping rule as optimizer hyperparameters. On GPT-2 Small trained on FineWeb, appropriate grouping improves validation loss over both full-QKV Muon and fully head-wise MuonSplit.
format Preprint
id arxiv_https___arxiv_org_abs_2605_08933
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When and Why Grouping Attention Heads Accelerates Muon Optimization
Zhang, Hongtao
Zhou, Wenjie
Chen, Wei
Cheng, Xueqi
Machine Learning
Muon orthogonalizes matrix updates, but multi-head attention naturally operates at the level of heads. This granularity mismatch raises the question of whether Muon should be applied to the full attention projection, to individual heads, or to intermediate head groups. We study this question through a one-step descent comparison between full-matrix Muon and group-wise Muon. Our analysis reveals a trade-off between the \textbf{group-wise whitening gain} from group-wise updates and the \textbf{grouping-induced norm cost}, an additional update-norm cost caused by replacing full-matrix whitening with group-wise whitening. Motivated by this trade-off, we propose \textbf{Group Muon}, which treats head group size and grouping rule as optimizer hyperparameters. On GPT-2 Small trained on FineWeb, appropriate grouping improves validation loss over both full-QKV Muon and fully head-wise MuonSplit.
title When and Why Grouping Attention Heads Accelerates Muon Optimization
topic Machine Learning
url https://arxiv.org/abs/2605.08933