ASiT: Local-Global Audio Spectrogram vIsion Transformer for Event Classification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Atito, Sara, Awais, Muhammad, Wang, Wenwu, Plumbley, Mark D, Kittler, Josef
Natura: Preprint
Pubblicazione: 2022
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914912379338752
author Atito, Sara
Awais, Muhammad
Wang, Wenwu
Plumbley, Mark D
Kittler, Josef
author_facet Atito, Sara
Awais, Muhammad
Wang, Wenwu
Plumbley, Mark D
Kittler, Josef
contents Transformers, which were originally developed for natural language processing, have recently generated significant interest in the computer vision and audio communities due to their flexibility in learning long-range relationships. Constrained by the data hungry nature of transformers and the limited amount of labelled data, most transformer-based models for audio tasks are finetuned from ImageNet pretrained models, despite the huge gap between the domain of natural images and audio. This has motivated the research in self-supervised pretraining of audio transformers, which reduces the dependency on large amounts of labeled data and focuses on extracting concise representations of audio spectrograms. In this paper, we propose \textbf{L}ocal-\textbf{G}lobal \textbf{A}udio \textbf{S}pectrogram v\textbf{I}sion \textbf{T}ransformer, namely ASiT, a novel self-supervised learning framework that captures local and global contextual information by employing group masked model learning and self-distillation. We evaluate our pretrained models on both audio and speech classification tasks, including audio event classification, keyword spotting, and speaker identification. We further conduct comprehensive ablation studies, including evaluations of different pretraining strategies. The proposed ASiT framework significantly boosts the performance on all tasks and sets a new state-of-the-art performance in five audio and speech classification tasks, outperforming recent methods, including the approaches that use additional datasets for pretraining.
format Preprint
id arxiv_https___arxiv_org_abs_2211_13189
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle ASiT: Local-Global Audio Spectrogram vIsion Transformer for Event Classification
Atito, Sara
Awais, Muhammad
Wang, Wenwu
Plumbley, Mark D
Kittler, Josef
Sound
Computer Vision and Pattern Recognition
Audio and Speech Processing
Transformers, which were originally developed for natural language processing, have recently generated significant interest in the computer vision and audio communities due to their flexibility in learning long-range relationships. Constrained by the data hungry nature of transformers and the limited amount of labelled data, most transformer-based models for audio tasks are finetuned from ImageNet pretrained models, despite the huge gap between the domain of natural images and audio. This has motivated the research in self-supervised pretraining of audio transformers, which reduces the dependency on large amounts of labeled data and focuses on extracting concise representations of audio spectrograms. In this paper, we propose \textbf{L}ocal-\textbf{G}lobal \textbf{A}udio \textbf{S}pectrogram v\textbf{I}sion \textbf{T}ransformer, namely ASiT, a novel self-supervised learning framework that captures local and global contextual information by employing group masked model learning and self-distillation. We evaluate our pretrained models on both audio and speech classification tasks, including audio event classification, keyword spotting, and speaker identification. We further conduct comprehensive ablation studies, including evaluations of different pretraining strategies. The proposed ASiT framework significantly boosts the performance on all tasks and sets a new state-of-the-art performance in five audio and speech classification tasks, outperforming recent methods, including the approaches that use additional datasets for pretraining.
title ASiT: Local-Global Audio Spectrogram vIsion Transformer for Event Classification
topic Sound
Computer Vision and Pattern Recognition
Audio and Speech Processing
url https://arxiv.org/abs/2211.13189