Revisiting Softmax Masking: Stop Gradient for Enhancing Stability in Replay-based Continual Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kim, Hoyong, Kwon, Minchan, Kim, Kangil
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910306334146560
author Kim, Hoyong
Kwon, Minchan
Kim, Kangil
author_facet Kim, Hoyong
Kwon, Minchan
Kim, Kangil
contents In replay-based methods for continual learning, replaying input samples in episodic memory has shown its effectiveness in alleviating catastrophic forgetting. However, the potential key factor of cross-entropy loss with softmax in causing catastrophic forgetting has been underexplored. In this paper, we analyze the effect of softmax and revisit softmax masking with negative infinity to shed light on its ability to mitigate catastrophic forgetting. Based on the analyses, it is found that negative infinity masked softmax is not always compatible with dark knowledge. To improve the compatibility, we propose a general masked softmax that controls the stability by adjusting the gradient scale to old and new classes. We demonstrate that utilizing our method on other replay-based methods results in better performance, primarily by enhancing model stability in continual learning benchmarks, even when the buffer size is set to an extremely small value.
format Preprint
id arxiv_https___arxiv_org_abs_2309_14808
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Revisiting Softmax Masking: Stop Gradient for Enhancing Stability in Replay-based Continual Learning
Kim, Hoyong
Kwon, Minchan
Kim, Kangil
Machine Learning
Artificial Intelligence
In replay-based methods for continual learning, replaying input samples in episodic memory has shown its effectiveness in alleviating catastrophic forgetting. However, the potential key factor of cross-entropy loss with softmax in causing catastrophic forgetting has been underexplored. In this paper, we analyze the effect of softmax and revisit softmax masking with negative infinity to shed light on its ability to mitigate catastrophic forgetting. Based on the analyses, it is found that negative infinity masked softmax is not always compatible with dark knowledge. To improve the compatibility, we propose a general masked softmax that controls the stability by adjusting the gradient scale to old and new classes. We demonstrate that utilizing our method on other replay-based methods results in better performance, primarily by enhancing model stability in continual learning benchmarks, even when the buffer size is set to an extremely small value.
title Revisiting Softmax Masking: Stop Gradient for Enhancing Stability in Replay-based Continual Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2309.14808