ConPET: Continual Parameter-Efficient Tuning for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Chenyang, Han, Xu, Zeng, Zheni, Li, Kuai, Chen, Chen, Liu, Zhiyuan, Sun, Maosong, Yang, Tao
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909365797126144
author Song, Chenyang
Han, Xu
Zeng, Zheni
Li, Kuai
Chen, Chen
Liu, Zhiyuan
Sun, Maosong
Yang, Tao
author_facet Song, Chenyang
Han, Xu
Zeng, Zheni
Li, Kuai
Chen, Chen
Liu, Zhiyuan
Sun, Maosong
Yang, Tao
contents Continual learning necessitates the continual adaptation of models to newly emerging tasks while minimizing the catastrophic forgetting of old ones. This is extremely challenging for large language models (LLMs) with vanilla full-parameter tuning due to high computation costs, memory consumption, and forgetting issue. Inspired by the success of parameter-efficient tuning (PET), we propose Continual Parameter-Efficient Tuning (ConPET), a generalizable paradigm for continual task adaptation of LLMs with task-number-independent training complexity. ConPET includes two versions with different application scenarios. First, Static ConPET can adapt former continual learning methods originally designed for relatively smaller models to LLMs through PET and a dynamic replay strategy, which largely reduces the tuning costs and alleviates the over-fitting and forgetting issue. Furthermore, to maintain scalability, Dynamic ConPET adopts separate PET modules for different tasks and a PET module selector for dynamic optimal selection. In our extensive experiments, the adaptation of Static ConPET helps multiple former methods reduce the scale of tunable parameters by over 3,000 times and surpass the PET-only baseline by at least 5 points on five smaller benchmarks, while Dynamic ConPET gains its advantage on the largest dataset. The codes and datasets are available at https://github.com/Raincleared-Song/ConPET.
format Preprint
id arxiv_https___arxiv_org_abs_2309_14763
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle ConPET: Continual Parameter-Efficient Tuning for Large Language Models
Song, Chenyang
Han, Xu
Zeng, Zheni
Li, Kuai
Chen, Chen
Liu, Zhiyuan
Sun, Maosong
Yang, Tao
Computation and Language
I.2.7
Continual learning necessitates the continual adaptation of models to newly emerging tasks while minimizing the catastrophic forgetting of old ones. This is extremely challenging for large language models (LLMs) with vanilla full-parameter tuning due to high computation costs, memory consumption, and forgetting issue. Inspired by the success of parameter-efficient tuning (PET), we propose Continual Parameter-Efficient Tuning (ConPET), a generalizable paradigm for continual task adaptation of LLMs with task-number-independent training complexity. ConPET includes two versions with different application scenarios. First, Static ConPET can adapt former continual learning methods originally designed for relatively smaller models to LLMs through PET and a dynamic replay strategy, which largely reduces the tuning costs and alleviates the over-fitting and forgetting issue. Furthermore, to maintain scalability, Dynamic ConPET adopts separate PET modules for different tasks and a PET module selector for dynamic optimal selection. In our extensive experiments, the adaptation of Static ConPET helps multiple former methods reduce the scale of tunable parameters by over 3,000 times and surpass the PET-only baseline by at least 5 points on five smaller benchmarks, while Dynamic ConPET gains its advantage on the largest dataset. The codes and datasets are available at https://github.com/Raincleared-Song/ConPET.
title ConPET: Continual Parameter-Efficient Tuning for Large Language Models
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2309.14763