Adapter Incremental Continual Learning of Efficient Audio Spectrogram Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Selvaraj, Nithish Muthuchamy, Guo, Xiaobao, Kong, Adams, Shen, Bingquan, Kot, Alex
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929196024987648
author Selvaraj, Nithish Muthuchamy
Guo, Xiaobao
Kong, Adams
Shen, Bingquan
Kot, Alex
author_facet Selvaraj, Nithish Muthuchamy
Guo, Xiaobao
Kong, Adams
Shen, Bingquan
Kot, Alex
contents Continual learning involves training neural networks incrementally for new tasks while retaining the knowledge of previous tasks. However, efficiently fine-tuning the model for sequential tasks with minimal computational resources remains a challenge. In this paper, we propose Task Incremental Continual Learning (TI-CL) of audio classifiers with both parameter-efficient and compute-efficient Audio Spectrogram Transformers (AST). To reduce the trainable parameters without performance degradation for TI-CL, we compare several Parameter Efficient Transfer (PET) methods and propose AST with Convolutional Adapters for TI-CL, which has less than 5% of trainable parameters of the fully fine-tuned counterparts. To reduce the computational complexity, we introduce a novel Frequency-Time factorized Attention (FTA) method that replaces the traditional self-attention in transformers for audio spectrograms. FTA achieves competitive performance with only a factor of the computations required by Global Self-Attention (GSA). Finally, we formulate our method for TI-CL, called Adapter Incremental Continual Learning (AI-CL), as a combination of the "parameter-efficient" Convolutional Adapter and the "compute-efficient" FTA. Experiments on ESC-50, SpeechCommandsV2 (SCv2), and Audio-Visual Event (AVE) benchmarks show that our proposed method prevents catastrophic forgetting in TI-CL while maintaining a lower computational budget.
format Preprint
id arxiv_https___arxiv_org_abs_2302_14314
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Adapter Incremental Continual Learning of Efficient Audio Spectrogram Transformers
Selvaraj, Nithish Muthuchamy
Guo, Xiaobao
Kong, Adams
Shen, Bingquan
Kot, Alex
Sound
Audio and Speech Processing
Continual learning involves training neural networks incrementally for new tasks while retaining the knowledge of previous tasks. However, efficiently fine-tuning the model for sequential tasks with minimal computational resources remains a challenge. In this paper, we propose Task Incremental Continual Learning (TI-CL) of audio classifiers with both parameter-efficient and compute-efficient Audio Spectrogram Transformers (AST). To reduce the trainable parameters without performance degradation for TI-CL, we compare several Parameter Efficient Transfer (PET) methods and propose AST with Convolutional Adapters for TI-CL, which has less than 5% of trainable parameters of the fully fine-tuned counterparts. To reduce the computational complexity, we introduce a novel Frequency-Time factorized Attention (FTA) method that replaces the traditional self-attention in transformers for audio spectrograms. FTA achieves competitive performance with only a factor of the computations required by Global Self-Attention (GSA). Finally, we formulate our method for TI-CL, called Adapter Incremental Continual Learning (AI-CL), as a combination of the "parameter-efficient" Convolutional Adapter and the "compute-efficient" FTA. Experiments on ESC-50, SpeechCommandsV2 (SCv2), and Audio-Visual Event (AVE) benchmarks show that our proposed method prevents catastrophic forgetting in TI-CL while maintaining a lower computational budget.
title Adapter Incremental Continual Learning of Efficient Audio Spectrogram Transformers
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2302.14314