Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sudarsanam, Parthasaarathy, Martín-Morató, Irene, Virtanen, Tuomas
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915295172493312
author Sudarsanam, Parthasaarathy
Martín-Morató, Irene
Virtanen, Tuomas
author_facet Sudarsanam, Parthasaarathy
Martín-Morató, Irene
Virtanen, Tuomas
contents This paper proposes a single-stage training approach that semantically aligns three modalities - audio, visual, and text using a contrastive learning framework. Contrastive training has gained prominence for multimodal alignment, utilizing large-scale unlabeled data to learn shared representations. Existing deep learning approach for trimodal alignment involves two-stages, that separately align visual-text and audio-text modalities. This approach suffers from mismatched data distributions, resulting in suboptimal alignment. Leveraging the AVCaps dataset, which provides audio, visual and audio-visual captions for video clips, our method jointly optimizes the representation of all the modalities using contrastive training. Our results demonstrate that the single-stage approach outperforms the two-stage method, achieving a two-fold improvement in audio based visual retrieval, highlighting the advantages of unified multimodal representation learning.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14562
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
Sudarsanam, Parthasaarathy
Martín-Morató, Irene
Virtanen, Tuomas
Sound
Multimedia
Audio and Speech Processing
This paper proposes a single-stage training approach that semantically aligns three modalities - audio, visual, and text using a contrastive learning framework. Contrastive training has gained prominence for multimodal alignment, utilizing large-scale unlabeled data to learn shared representations. Existing deep learning approach for trimodal alignment involves two-stages, that separately align visual-text and audio-text modalities. This approach suffers from mismatched data distributions, resulting in suboptimal alignment. Leveraging the AVCaps dataset, which provides audio, visual and audio-visual captions for video clips, our method jointly optimizes the representation of all the modalities using contrastive training. Our results demonstrate that the single-stage approach outperforms the two-stage method, achieving a two-fold improvement in audio based visual retrieval, highlighting the advantages of unified multimodal representation learning.
title Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
topic Sound
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2505.14562