Pay Attention to Small Weights

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Chao, Jacobs, Tom, Gadhikar, Advait, Burkholz, Rebekka
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911226875871232
author Zhou, Chao
Jacobs, Tom
Gadhikar, Advait
Burkholz, Rebekka
author_facet Zhou, Chao
Jacobs, Tom
Gadhikar, Advait
Burkholz, Rebekka
contents Finetuning large pretrained neural networks is known to be resource-intensive, both in terms of memory and computational cost. To mitigate this, a common approach is to restrict training to a subset of the model parameters. By analyzing the relationship between gradients and weights during finetuning, we observe a notable pattern: large gradients are often associated with small-magnitude weights. This correlation is more pronounced in finetuning settings than in training from scratch. Motivated by this observation, we propose NANOADAM, which dynamically updates only the small-magnitude weights during finetuning and offers several practical advantages: first, this criterion is gradient-free -- the parameter subset can be determined without gradient computation; second, it preserves large-magnitude weights, which are likely to encode critical features learned during pretraining, thereby reducing the risk of catastrophic forgetting; thirdly, it permits the use of larger learning rates and consistently leads to better generalization performance in experiments. We demonstrate this for both NLP and vision tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21374
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pay Attention to Small Weights
Zhou, Chao
Jacobs, Tom
Gadhikar, Advait
Burkholz, Rebekka
Machine Learning
Artificial Intelligence
Finetuning large pretrained neural networks is known to be resource-intensive, both in terms of memory and computational cost. To mitigate this, a common approach is to restrict training to a subset of the model parameters. By analyzing the relationship between gradients and weights during finetuning, we observe a notable pattern: large gradients are often associated with small-magnitude weights. This correlation is more pronounced in finetuning settings than in training from scratch. Motivated by this observation, we propose NANOADAM, which dynamically updates only the small-magnitude weights during finetuning and offers several practical advantages: first, this criterion is gradient-free -- the parameter subset can be determined without gradient computation; second, it preserves large-magnitude weights, which are likely to encode critical features learned during pretraining, thereby reducing the risk of catastrophic forgetting; thirdly, it permits the use of larger learning rates and consistently leads to better generalization performance in experiments. We demonstrate this for both NLP and vision tasks.
title Pay Attention to Small Weights
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.21374