Generalized Gradient Norm Clipping & Non-Euclidean $(L_0,L_1)$-Smoothness

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pethick, Thomas, Xie, Wanyun, Erdogan, Mete, Antonakopoulos, Kimon, Silveti-Falls, Antonio, Cevher, Volkan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914305215037440
author Pethick, Thomas
Xie, Wanyun
Erdogan, Mete
Antonakopoulos, Kimon
Silveti-Falls, Antonio
Cevher, Volkan
author_facet Pethick, Thomas
Xie, Wanyun
Erdogan, Mete
Antonakopoulos, Kimon
Silveti-Falls, Antonio
Cevher, Volkan
contents This work introduces a hybrid non-Euclidean optimization method which generalizes gradient norm clipping by combining steepest descent and conditional gradient approaches. The method achieves the best of both worlds by establishing a descent property under a generalized notion of ($L_0$,$L_1$)-smoothness. Weight decay is incorporated in a principled manner by identifying a connection to the Frank-Wolfe short step. In the stochastic case, we show an order optimal $O(n^{-1/4})$ convergence rate by leveraging a momentum based gradient estimator. We discuss how to instantiate the algorithms for deep learning, which we dub Clipped Scion, and demonstrate their properties on image classification and language modeling. The code is available at https://github.com/LIONS-EPFL/ClippedScion.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01913
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generalized Gradient Norm Clipping & Non-Euclidean $(L_0,L_1)$-Smoothness
Pethick, Thomas
Xie, Wanyun
Erdogan, Mete
Antonakopoulos, Kimon
Silveti-Falls, Antonio
Cevher, Volkan
Machine Learning
This work introduces a hybrid non-Euclidean optimization method which generalizes gradient norm clipping by combining steepest descent and conditional gradient approaches. The method achieves the best of both worlds by establishing a descent property under a generalized notion of ($L_0$,$L_1$)-smoothness. Weight decay is incorporated in a principled manner by identifying a connection to the Frank-Wolfe short step. In the stochastic case, we show an order optimal $O(n^{-1/4})$ convergence rate by leveraging a momentum based gradient estimator. We discuss how to instantiate the algorithms for deep learning, which we dub Clipped Scion, and demonstrate their properties on image classification and language modeling. The code is available at https://github.com/LIONS-EPFL/ClippedScion.
title Generalized Gradient Norm Clipping & Non-Euclidean $(L_0,L_1)$-Smoothness
topic Machine Learning
url https://arxiv.org/abs/2506.01913