PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Yao, Fan, Dengdong, Zhang, Shixun, Tian, Yonghong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911671176396800
author Lu, Yao
Fan, Dengdong
Zhang, Shixun
Tian, Yonghong
author_facet Lu, Yao
Fan, Dengdong
Zhang, Shixun
Tian, Yonghong
contents Adaptive optimizers, most notably Adam, have become the default standard for training large-scale neural networks such as Transformers. These methods maintain running estimates of gradient first and second moments, incurring substantial memory overhead. We introduce PowerStep, a memory-efficient optimizer that achieves coordinate-wise adaptivity without storing second-moment statistics. Motivated by steepest descent under an $\ell_p$-norm geometry, we show that applying a nonlinear transform directly to a momentum buffer yields coordinate-wise adaptivity. We prove that PowerStep converges at the optimal $O(1/\sqrt{T})$ rate for non-convex stochastic optimization. Extensive experiments on Transformer models ranging from 124M to 235B parameters demonstrate that PowerStep matches Adam's convergence speed while halving optimizer memory. Furthermore, when combined with aggressive \texttt{int8} quantization, PowerStep remains numerically stable and reduces optimizer memory by $\sim\!8\times$ compared to full-precision Adam. PowerStep thus provides a principled, scalable and resource-efficient alternative for large-scale training. Code is available at https://github.com/yaolubrain/PowerStep.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10335
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent
Lu, Yao
Fan, Dengdong
Zhang, Shixun
Tian, Yonghong
Machine Learning
Artificial Intelligence
Computation and Language
Numerical Analysis
Optimization and Control
Adaptive optimizers, most notably Adam, have become the default standard for training large-scale neural networks such as Transformers. These methods maintain running estimates of gradient first and second moments, incurring substantial memory overhead. We introduce PowerStep, a memory-efficient optimizer that achieves coordinate-wise adaptivity without storing second-moment statistics. Motivated by steepest descent under an $\ell_p$-norm geometry, we show that applying a nonlinear transform directly to a momentum buffer yields coordinate-wise adaptivity. We prove that PowerStep converges at the optimal $O(1/\sqrt{T})$ rate for non-convex stochastic optimization. Extensive experiments on Transformer models ranging from 124M to 235B parameters demonstrate that PowerStep matches Adam's convergence speed while halving optimizer memory. Furthermore, when combined with aggressive \texttt{int8} quantization, PowerStep remains numerically stable and reduces optimizer memory by $\sim\!8\times$ compared to full-precision Adam. PowerStep thus provides a principled, scalable and resource-efficient alternative for large-scale training. Code is available at https://github.com/yaolubrain/PowerStep.
title PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent
topic Machine Learning
Artificial Intelligence
Computation and Language
Numerical Analysis
Optimization and Control
url https://arxiv.org/abs/2605.10335