Torque-Aware Momentum

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Malviya, Pranshu, Mordido, Goncalo, Baratin, Aristide, Harikandeh, Reza Babanezhad, Dziugaite, Gintare Karolina, Pascanu, Razvan, Chandar, Sarath
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915080738701312
author Malviya, Pranshu
Mordido, Goncalo
Baratin, Aristide
Harikandeh, Reza Babanezhad
Dziugaite, Gintare Karolina
Pascanu, Razvan
Chandar, Sarath
author_facet Malviya, Pranshu
Mordido, Goncalo
Baratin, Aristide
Harikandeh, Reza Babanezhad
Dziugaite, Gintare Karolina
Pascanu, Razvan
Chandar, Sarath
contents Efficiently exploring complex loss landscapes is key to the performance of deep neural networks. While momentum-based optimizers are widely used in state-of-the-art setups, classical momentum can still struggle with large, misaligned gradients, leading to oscillations. To address this, we propose Torque-Aware Momentum (TAM), which introduces a damping factor based on the angle between the new gradients and previous momentum, stabilizing the update direction during training. Empirical results show that TAM, which can be combined with both SGD and Adam, enhances exploration, handles distribution shifts more effectively, and improves generalization performance across various tasks, including image classification and large language model fine-tuning, when compared to classical momentum-based optimizers.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18790
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Torque-Aware Momentum
Malviya, Pranshu
Mordido, Goncalo
Baratin, Aristide
Harikandeh, Reza Babanezhad
Dziugaite, Gintare Karolina
Pascanu, Razvan
Chandar, Sarath
Machine Learning
Artificial Intelligence
Efficiently exploring complex loss landscapes is key to the performance of deep neural networks. While momentum-based optimizers are widely used in state-of-the-art setups, classical momentum can still struggle with large, misaligned gradients, leading to oscillations. To address this, we propose Torque-Aware Momentum (TAM), which introduces a damping factor based on the angle between the new gradients and previous momentum, stabilizing the update direction during training. Empirical results show that TAM, which can be combined with both SGD and Adam, enhances exploration, handles distribution shifts more effectively, and improves generalization performance across various tasks, including image classification and large language model fine-tuning, when compared to classical momentum-based optimizers.
title Torque-Aware Momentum
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2412.18790