Optimistic Dual Averaging Unifies Modern Optimizers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pethick, Thomas, Xie, Wanyun, Machacek, Roman, Cevher, Volkan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910209867251712
author Pethick, Thomas
Xie, Wanyun
Machacek, Roman
Cevher, Volkan
author_facet Pethick, Thomas
Xie, Wanyun
Machacek, Roman
Cevher, Volkan
contents We introduce SODA, a generalization of Optimistic Dual Averaging, which provides a common perspective on state-of-the-art optimizers like Muon, Lion, AdEMAMix and NAdam, showing that they can all be viewed as optimistic instances of this framework. Based on this framing, we propose a practical SODA wrapper for any base optimizer that eliminates weight decay tuning through a theoretically-grounded $1/k$ decay schedule. Empirical results across various scales and training horizons show that SODA consistently improves performance without any additional hyperparameter tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11172
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Optimistic Dual Averaging Unifies Modern Optimizers
Pethick, Thomas
Xie, Wanyun
Machacek, Roman
Cevher, Volkan
Machine Learning
We introduce SODA, a generalization of Optimistic Dual Averaging, which provides a common perspective on state-of-the-art optimizers like Muon, Lion, AdEMAMix and NAdam, showing that they can all be viewed as optimistic instances of this framework. Based on this framing, we propose a practical SODA wrapper for any base optimizer that eliminates weight decay tuning through a theoretically-grounded $1/k$ decay schedule. Empirical results across various scales and training horizons show that SODA consistently improves performance without any additional hyperparameter tuning.
title Optimistic Dual Averaging Unifies Modern Optimizers
topic Machine Learning
url https://arxiv.org/abs/2605.11172