Torque-Aware Momentum
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915080738701312 |
|---|---|
| author | Malviya, Pranshu Mordido, Goncalo Baratin, Aristide Harikandeh, Reza Babanezhad Dziugaite, Gintare Karolina Pascanu, Razvan Chandar, Sarath |
| author_facet | Malviya, Pranshu Mordido, Goncalo Baratin, Aristide Harikandeh, Reza Babanezhad Dziugaite, Gintare Karolina Pascanu, Razvan Chandar, Sarath |
| contents | Efficiently exploring complex loss landscapes is key to the performance of deep neural networks. While momentum-based optimizers are widely used in state-of-the-art setups, classical momentum can still struggle with large, misaligned gradients, leading to oscillations. To address this, we propose Torque-Aware Momentum (TAM), which introduces a damping factor based on the angle between the new gradients and previous momentum, stabilizing the update direction during training. Empirical results show that TAM, which can be combined with both SGD and Adam, enhances exploration, handles distribution shifts more effectively, and improves generalization performance across various tasks, including image classification and large language model fine-tuning, when compared to classical momentum-based optimizers. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_18790 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Torque-Aware Momentum Malviya, Pranshu Mordido, Goncalo Baratin, Aristide Harikandeh, Reza Babanezhad Dziugaite, Gintare Karolina Pascanu, Razvan Chandar, Sarath Machine Learning Artificial Intelligence Efficiently exploring complex loss landscapes is key to the performance of deep neural networks. While momentum-based optimizers are widely used in state-of-the-art setups, classical momentum can still struggle with large, misaligned gradients, leading to oscillations. To address this, we propose Torque-Aware Momentum (TAM), which introduces a damping factor based on the angle between the new gradients and previous momentum, stabilizing the update direction during training. Empirical results show that TAM, which can be combined with both SGD and Adam, enhances exploration, handles distribution shifts more effectively, and improves generalization performance across various tasks, including image classification and large language model fine-tuning, when compared to classical momentum-based optimizers. |
| title | Torque-Aware Momentum |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2412.18790 |