DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shin, Baekrok, Oh, Junsoo, Cho, Hanseul, Yun, Chulhee
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915001454821376
author Shin, Baekrok
Oh, Junsoo
Cho, Hanseul
Yun, Chulhee
author_facet Shin, Baekrok
Oh, Junsoo
Cho, Hanseul
Yun, Chulhee
contents Warm-starting neural network training by initializing networks with previously learned weights is appealing, as practical neural networks are often deployed under a continuous influx of new data. However, it often leads to loss of plasticity, where the network loses its ability to learn new information, resulting in worse generalization than training from scratch. This occurs even under stationary data distributions, and its underlying mechanism is poorly understood. We develop a framework emulating real-world neural network training and identify noise memorization as the primary cause of plasticity loss when warm-starting on stationary data. Motivated by this, we propose Direction-Aware SHrinking (DASH), a method aiming to mitigate plasticity loss by selectively forgetting memorized noise while preserving learned features. We validate our approach on vision tasks, demonstrating improvements in test accuracy and training efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2410_23495
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity
Shin, Baekrok
Oh, Junsoo
Cho, Hanseul
Yun, Chulhee
Machine Learning
Artificial Intelligence
Warm-starting neural network training by initializing networks with previously learned weights is appealing, as practical neural networks are often deployed under a continuous influx of new data. However, it often leads to loss of plasticity, where the network loses its ability to learn new information, resulting in worse generalization than training from scratch. This occurs even under stationary data distributions, and its underlying mechanism is poorly understood. We develop a framework emulating real-world neural network training and identify noise memorization as the primary cause of plasticity loss when warm-starting on stationary data. Motivated by this, we propose Direction-Aware SHrinking (DASH), a method aiming to mitigate plasticity loss by selectively forgetting memorized noise while preserving learned features. We validate our approach on vision tasks, demonstrating improvements in test accuracy and training efficiency.
title DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.23495