EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kim, Jongsuk, Lee, Hyeongkeun, Rho, Kyeongha, Kim, Junmo, Chung, Joon Son
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916294994952192
author Kim, Jongsuk
Lee, Hyeongkeun
Rho, Kyeongha
Kim, Junmo
Chung, Joon Son
author_facet Kim, Jongsuk
Lee, Hyeongkeun
Rho, Kyeongha
Kim, Junmo
Chung, Joon Son
contents Recent advancements in self-supervised audio-visual representation learning have demonstrated its potential to capture rich and comprehensive representations. However, despite the advantages of data augmentation verified in many learning methods, audio-visual learning has struggled to fully harness these benefits, as augmentations can easily disrupt the correspondence between input pairs. To address this limitation, we introduce EquiAV, a novel framework that leverages equivariance for audio-visual contrastive learning. Our approach begins with extending equivariance to audio-visual learning, facilitated by a shared attention-based transformation predictor. It enables the aggregation of features from diverse augmentations into a representative embedding, providing robust supervision. Notably, this is achieved with minimal computational overhead. Extensive ablation studies and qualitative results verify the effectiveness of our method. EquiAV outperforms previous works across various audio-visual benchmarks. The code is available on https://github.com/JongSuk1/EquiAV.
format Preprint
id arxiv_https___arxiv_org_abs_2403_09502
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning
Kim, Jongsuk
Lee, Hyeongkeun
Rho, Kyeongha
Kim, Junmo
Chung, Joon Son
Machine Learning
Artificial Intelligence
Multimedia
Recent advancements in self-supervised audio-visual representation learning have demonstrated its potential to capture rich and comprehensive representations. However, despite the advantages of data augmentation verified in many learning methods, audio-visual learning has struggled to fully harness these benefits, as augmentations can easily disrupt the correspondence between input pairs. To address this limitation, we introduce EquiAV, a novel framework that leverages equivariance for audio-visual contrastive learning. Our approach begins with extending equivariance to audio-visual learning, facilitated by a shared attention-based transformation predictor. It enables the aggregation of features from diverse augmentations into a representative embedding, providing robust supervision. Notably, this is achieved with minimal computational overhead. Extensive ablation studies and qualitative results verify the effectiveness of our method. EquiAV outperforms previous works across various audio-visual benchmarks. The code is available on https://github.com/JongSuk1/EquiAV.
title EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning
topic Machine Learning
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2403.09502