Does RoBERTa Perform Better than BERT in Continual Learning: An Attention Sink Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bai, Xueying, Sun, Yifan, Balasubramanian, Niranjan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913536022675456
author Bai, Xueying
Sun, Yifan
Balasubramanian, Niranjan
author_facet Bai, Xueying
Sun, Yifan
Balasubramanian, Niranjan
contents Continual learning (CL) aims to train models that can sequentially learn new tasks without forgetting previous tasks' knowledge. Although previous works observed that pre-training can benefit CL, it remains unclear whether a pre-trained model with higher downstream capacity also performs better in CL. In this paper, we observe that pre-trained models may allocate high attention scores to some 'sink' tokens, such as [SEP] tokens, which are ubiquitous across various tasks. Such attention sinks may lead to models' over-smoothing in single-task learning and interference in sequential tasks' learning, which may compromise the models' CL performance despite their high pre-trained capabilities. To reduce these effects, we propose a pre-scaling mechanism that encourages attention diversity across all tokens. Specifically, it first scales the task's attention to the non-sink tokens in a probing stage, and then fine-tunes the model with scaling. Experiments show that pre-scaling yields substantial improvements in CL without experience replay, or progressively storing parameters from previous tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05648
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Does RoBERTa Perform Better than BERT in Continual Learning: An Attention Sink Perspective
Bai, Xueying
Sun, Yifan
Balasubramanian, Niranjan
Machine Learning
Computation and Language
Continual learning (CL) aims to train models that can sequentially learn new tasks without forgetting previous tasks' knowledge. Although previous works observed that pre-training can benefit CL, it remains unclear whether a pre-trained model with higher downstream capacity also performs better in CL. In this paper, we observe that pre-trained models may allocate high attention scores to some 'sink' tokens, such as [SEP] tokens, which are ubiquitous across various tasks. Such attention sinks may lead to models' over-smoothing in single-task learning and interference in sequential tasks' learning, which may compromise the models' CL performance despite their high pre-trained capabilities. To reduce these effects, we propose a pre-scaling mechanism that encourages attention diversity across all tokens. Specifically, it first scales the task's attention to the non-sink tokens in a probing stage, and then fine-tunes the model with scaling. Experiments show that pre-scaling yields substantial improvements in CL without experience replay, or progressively storing parameters from previous tasks.
title Does RoBERTa Perform Better than BERT in Continual Learning: An Attention Sink Perspective
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2410.05648