A second-order-like optimizer with adaptive gradient scaling for deep learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Bolte, Jérôme, Boustany, Ryan, Pauwels, Edouard, Purica, Andrei
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912152841879552
author Bolte, Jérôme
Boustany, Ryan
Pauwels, Edouard
Purica, Andrei
author_facet Bolte, Jérôme
Boustany, Ryan
Pauwels, Edouard
Purica, Andrei
contents In this empirical article, we introduce INNAprop, an optimization algorithm that combines the INNA method with the RMSprop adaptive gradient scaling. It leverages second-order information and rescaling while keeping the memory requirements of standard DL methods as AdamW or SGD with momentum. After giving geometrical insights, we evaluate INNAprop on CIFAR-10, Food101, and ImageNet with ResNets, VGG, DenseNet, and ViT, and on GPT-2 (OpenWebText) train from scratch and with LoRA fine-tuning (E2E). INNAprop consistently matches or outperforms AdamW both in training speed and accuracy, with minimal hyperparameter tuning in large-scale settings. Our code is publicly available at \url{https://github.com/innaprop/innaprop}.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05871
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A second-order-like optimizer with adaptive gradient scaling for deep learning
Bolte, Jérôme
Boustany, Ryan
Pauwels, Edouard
Purica, Andrei
Machine Learning
Artificial Intelligence
Optimization and Control
In this empirical article, we introduce INNAprop, an optimization algorithm that combines the INNA method with the RMSprop adaptive gradient scaling. It leverages second-order information and rescaling while keeping the memory requirements of standard DL methods as AdamW or SGD with momentum. After giving geometrical insights, we evaluate INNAprop on CIFAR-10, Food101, and ImageNet with ResNets, VGG, DenseNet, and ViT, and on GPT-2 (OpenWebText) train from scratch and with LoRA fine-tuning (E2E). INNAprop consistently matches or outperforms AdamW both in training speed and accuracy, with minimal hyperparameter tuning in large-scale settings. Our code is publicly available at \url{https://github.com/innaprop/innaprop}.
title A second-order-like optimizer with adaptive gradient scaling for deep learning
topic Machine Learning
Artificial Intelligence
Optimization and Control
url https://arxiv.org/abs/2410.05871