A-JEPA: Joint-Embedding Predictive Architecture Can Listen

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Fei, Zhengcong, Fan, Mingyuan, Huang, Junshi
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909069359448064
author Fei, Zhengcong
Fan, Mingyuan
Huang, Junshi
author_facet Fei, Zhengcong
Fan, Mingyuan
Huang, Junshi
contents This paper presents that the masked-modeling principle driving the success of large foundational vision models can be effectively applied to audio by making predictions in a latent space. We introduce Audio-based Joint-Embedding Predictive Architecture (A-JEPA), a simple extension method for self-supervised learning from the audio spectrum. Following the design of I-JEPA, our A-JEPA encodes visible audio spectrogram patches with a curriculum masking strategy via context encoder, and predicts the representations of regions sampled at well-designed locations. The target representations of those regions are extracted by the exponential moving average of context encoder, \emph{i.e.}, target encoder, on the whole spectrogram. We find it beneficial to transfer random block masking into time-frequency aware masking in a curriculum manner, considering the complexity of highly correlated in local time and frequency in audio spectrograms. To enhance contextual semantic understanding and robustness, we fine-tune the encoder with a regularized masking on target datasets, instead of input dropping or zero. Empirically, when built with Vision Transformers structure, we find A-JEPA to be highly scalable and sets new state-of-the-art performance on multiple audio and speech classification tasks, outperforming other recent models that use externally supervised pre-training.
format Preprint
id arxiv_https___arxiv_org_abs_2311_15830
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A-JEPA: Joint-Embedding Predictive Architecture Can Listen
Fei, Zhengcong
Fan, Mingyuan
Huang, Junshi
Sound
Computer Vision and Pattern Recognition
Audio and Speech Processing
This paper presents that the masked-modeling principle driving the success of large foundational vision models can be effectively applied to audio by making predictions in a latent space. We introduce Audio-based Joint-Embedding Predictive Architecture (A-JEPA), a simple extension method for self-supervised learning from the audio spectrum. Following the design of I-JEPA, our A-JEPA encodes visible audio spectrogram patches with a curriculum masking strategy via context encoder, and predicts the representations of regions sampled at well-designed locations. The target representations of those regions are extracted by the exponential moving average of context encoder, \emph{i.e.}, target encoder, on the whole spectrogram. We find it beneficial to transfer random block masking into time-frequency aware masking in a curriculum manner, considering the complexity of highly correlated in local time and frequency in audio spectrograms. To enhance contextual semantic understanding and robustness, we fine-tune the encoder with a regularized masking on target datasets, instead of input dropping or zero. Empirically, when built with Vision Transformers structure, we find A-JEPA to be highly scalable and sets new state-of-the-art performance on multiple audio and speech classification tasks, outperforming other recent models that use externally supervised pre-training.
title A-JEPA: Joint-Embedding Predictive Architecture Can Listen
topic Sound
Computer Vision and Pattern Recognition
Audio and Speech Processing
url https://arxiv.org/abs/2311.15830