Masked Audio Modeling with CLAP and Multi-Objective Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xin, Yifei, Peng, Xiulian, Lu, Yan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911766407020544
author Xin, Yifei
Peng, Xiulian
Lu, Yan
author_facet Xin, Yifei
Peng, Xiulian
Lu, Yan
contents Most existing masked audio modeling (MAM) methods learn audio representations by masking and reconstructing local spectrogram patches. However, the reconstruction loss mainly accounts for the signal-level quality of the reconstructed spectrogram and is still limited in extracting high-level audio semantics. In this paper, we propose to enhance the semantic modeling of MAM by distilling cross-modality knowledge from contrastive language-audio pretraining (CLAP) representations for both masked and unmasked regions (MAM-CLAP) and leveraging a multi-objective learning strategy with a supervised classification branch (SupMAM), thereby providing more semantic knowledge for MAM and enabling it to effectively learn global features from labels. Experiments show that our methods significantly improve the performance on multiple downstream tasks. Furthermore, by combining our MAM-CLAP with SupMAM, we can achieve new state-of-the-art results on various audio and speech classification tasks, exceeding previous self-supervised learning and supervised pretraining methods.
format Preprint
id arxiv_https___arxiv_org_abs_2401_15953
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Masked Audio Modeling with CLAP and Multi-Objective Learning
Xin, Yifei
Peng, Xiulian
Lu, Yan
Sound
Audio and Speech Processing
Most existing masked audio modeling (MAM) methods learn audio representations by masking and reconstructing local spectrogram patches. However, the reconstruction loss mainly accounts for the signal-level quality of the reconstructed spectrogram and is still limited in extracting high-level audio semantics. In this paper, we propose to enhance the semantic modeling of MAM by distilling cross-modality knowledge from contrastive language-audio pretraining (CLAP) representations for both masked and unmasked regions (MAM-CLAP) and leveraging a multi-objective learning strategy with a supervised classification branch (SupMAM), thereby providing more semantic knowledge for MAM and enabling it to effectively learn global features from labels. Experiments show that our methods significantly improve the performance on multiple downstream tasks. Furthermore, by combining our MAM-CLAP with SupMAM, we can achieve new state-of-the-art results on various audio and speech classification tasks, exceeding previous self-supervised learning and supervised pretraining methods.
title Masked Audio Modeling with CLAP and Multi-Objective Learning
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2401.15953