A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean Descent

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xie, Shuo, Wang, Tianhao, Wu, Beining, Li, Zhiyuan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918217997352960
author Xie, Shuo
Wang, Tianhao
Wu, Beining
Li, Zhiyuan
author_facet Xie, Shuo
Wang, Tianhao
Wu, Beining
Li, Zhiyuan
contents Adaptive optimizers can reduce to normalized steepest descent (NSD) when only adapting to the current gradient, suggesting a close connection between the two algorithmic families. A key distinction between their analyses, however, lies in the geometries, e.g., smoothness notions, they rely on. In the convex setting, adaptive optimizers are governed by a stronger adaptive smoothness condition, while NSD relies on the standard notion of smoothness. We extend the theory of adaptive smoothness to the nonconvex setting and show that it precisely characterizes the convergence of adaptive optimizers. Moreover, we establish that adaptive smoothness enables acceleration of adaptive optimizers with Nesterov momentum in the convex setting, a guarantee unattainable under standard smoothness for certain non-Euclidean geometry. We further develop an analogous comparison for stochastic optimization by introducing adaptive gradient variance, which parallels adaptive smoothness and leads to dimension-free convergence guarantees that cannot be achieved under standard gradient variance for certain non-Euclidean geometry.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20584
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean Descent
Xie, Shuo
Wang, Tianhao
Wu, Beining
Li, Zhiyuan
Machine Learning
Adaptive optimizers can reduce to normalized steepest descent (NSD) when only adapting to the current gradient, suggesting a close connection between the two algorithmic families. A key distinction between their analyses, however, lies in the geometries, e.g., smoothness notions, they rely on. In the convex setting, adaptive optimizers are governed by a stronger adaptive smoothness condition, while NSD relies on the standard notion of smoothness. We extend the theory of adaptive smoothness to the nonconvex setting and show that it precisely characterizes the convergence of adaptive optimizers. Moreover, we establish that adaptive smoothness enables acceleration of adaptive optimizers with Nesterov momentum in the convex setting, a guarantee unattainable under standard smoothness for certain non-Euclidean geometry. We further develop an analogous comparison for stochastic optimization by introducing adaptive gradient variance, which parallels adaptive smoothness and leads to dimension-free convergence guarantees that cannot be achieved under standard gradient variance for certain non-Euclidean geometry.
title A Tale of Two Geometries: Adaptive Optimizers and Non-Euclidean Descent
topic Machine Learning
url https://arxiv.org/abs/2511.20584