Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915295172493312 |
|---|---|
| author | Sudarsanam, Parthasaarathy Martín-Morató, Irene Virtanen, Tuomas |
| author_facet | Sudarsanam, Parthasaarathy Martín-Morató, Irene Virtanen, Tuomas |
| contents | This paper proposes a single-stage training approach that semantically aligns three modalities - audio, visual, and text using a contrastive learning framework. Contrastive training has gained prominence for multimodal alignment, utilizing large-scale unlabeled data to learn shared representations. Existing deep learning approach for trimodal alignment involves two-stages, that separately align visual-text and audio-text modalities. This approach suffers from mismatched data distributions, resulting in suboptimal alignment. Leveraging the AVCaps dataset, which provides audio, visual and audio-visual captions for video clips, our method jointly optimizes the representation of all the modalities using contrastive training. Our results demonstrate that the single-stage approach outperforms the two-stage method, achieving a two-fold improvement in audio based visual retrieval, highlighting the advantages of unified multimodal representation learning. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_14562 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Sudarsanam, Parthasaarathy Martín-Morató, Irene Virtanen, Tuomas Sound Multimedia Audio and Speech Processing This paper proposes a single-stage training approach that semantically aligns three modalities - audio, visual, and text using a contrastive learning framework. Contrastive training has gained prominence for multimodal alignment, utilizing large-scale unlabeled data to learn shared representations. Existing deep learning approach for trimodal alignment involves two-stages, that separately align visual-text and audio-text modalities. This approach suffers from mismatched data distributions, resulting in suboptimal alignment. Leveraging the AVCaps dataset, which provides audio, visual and audio-visual captions for video clips, our method jointly optimizes the representation of all the modalities using contrastive training. Our results demonstrate that the single-stage approach outperforms the two-stage method, achieving a two-fold improvement in audio based visual retrieval, highlighting the advantages of unified multimodal representation learning. |
| title | Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities |
| topic | Sound Multimedia Audio and Speech Processing |
| url | https://arxiv.org/abs/2505.14562 |