Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ishikawa, Yuchi, Nakada, Shota, Munakata, Hokuto, Saito, Kazuhiro, Komatsu, Tatsuya, Aoki, Yoshimitsu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908451893936128
author Ishikawa, Yuchi
Nakada, Shota
Munakata, Hokuto
Saito, Kazuhiro
Komatsu, Tatsuya
Aoki, Yoshimitsu
author_facet Ishikawa, Yuchi
Nakada, Shota
Munakata, Hokuto
Saito, Kazuhiro
Komatsu, Tatsuya
Aoki, Yoshimitsu
contents In this paper, we propose Language-Guided Contrastive Audio-Visual Masked Autoencoders (LG-CAV-MAE) to improve audio-visual representation learning. LG-CAV-MAE integrates a pretrained text encoder into contrastive audio-visual masked autoencoders, enabling the model to learn across audio, visual and text modalities. To train LG-CAV-MAE, we introduce an automatic method to generate audio-visual-text triplets from unlabeled videos. We first generate frame-level captions using an image captioning model and then apply CLAP-based filtering to ensure strong alignment between audio and captions. This approach yields high-quality audio-visual-text triplets without requiring manual annotations. We evaluate LG-CAV-MAE on audio-visual retrieval tasks, as well as an audio-visual classification task. Our method significantly outperforms existing approaches, achieving up to a 5.6% improvement in recall@10 for retrieval tasks and a 3.2% improvement for the classification task.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11967
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
Ishikawa, Yuchi
Nakada, Shota
Munakata, Hokuto
Saito, Kazuhiro
Komatsu, Tatsuya
Aoki, Yoshimitsu
Computer Vision and Pattern Recognition
Audio and Speech Processing
Image and Video Processing
In this paper, we propose Language-Guided Contrastive Audio-Visual Masked Autoencoders (LG-CAV-MAE) to improve audio-visual representation learning. LG-CAV-MAE integrates a pretrained text encoder into contrastive audio-visual masked autoencoders, enabling the model to learn across audio, visual and text modalities. To train LG-CAV-MAE, we introduce an automatic method to generate audio-visual-text triplets from unlabeled videos. We first generate frame-level captions using an image captioning model and then apply CLAP-based filtering to ensure strong alignment between audio and captions. This approach yields high-quality audio-visual-text triplets without requiring manual annotations. We evaluate LG-CAV-MAE on audio-visual retrieval tasks, as well as an audio-visual classification task. Our method significantly outperforms existing approaches, achieving up to a 5.6% improvement in recall@10 for retrieval tasks and a 3.2% improvement for the classification task.
title Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
topic Computer Vision and Pattern Recognition
Audio and Speech Processing
Image and Video Processing
url https://arxiv.org/abs/2507.11967