The Implicit Bias of Adam on Separable Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Chenyang, Zou, Difan, Cao, Yuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911919176155136
author Zhang, Chenyang
Zou, Difan
Cao, Yuan
author_facet Zhang, Chenyang
Zou, Difan
Cao, Yuan
contents Adam has become one of the most favored optimizers in deep learning problems. Despite its success in practice, numerous mysteries persist regarding its theoretical understanding. In this paper, we study the implicit bias of Adam in linear logistic regression. Specifically, we show that when the training data are linearly separable, Adam converges towards a linear classifier that achieves the maximum $\ell_\infty$-margin. Notably, for a general class of diminishing learning rates, this convergence occurs within polynomial time. Our result shed light on the difference between Adam and (stochastic) gradient descent from a theoretical perspective.
format Preprint
id arxiv_https___arxiv_org_abs_2406_10650
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Implicit Bias of Adam on Separable Data
Zhang, Chenyang
Zou, Difan
Cao, Yuan
Machine Learning
Adam has become one of the most favored optimizers in deep learning problems. Despite its success in practice, numerous mysteries persist regarding its theoretical understanding. In this paper, we study the implicit bias of Adam in linear logistic regression. Specifically, we show that when the training data are linearly separable, Adam converges towards a linear classifier that achieves the maximum $\ell_\infty$-margin. Notably, for a general class of diminishing learning rates, this convergence occurs within polynomial time. Our result shed light on the difference between Adam and (stochastic) gradient descent from a theoretical perspective.
title The Implicit Bias of Adam on Separable Data
topic Machine Learning
url https://arxiv.org/abs/2406.10650