Do Audio-Language Models Understand Linguistic Variations?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Selvakumar, Ramaneswaran, Kumar, Sonal, Giri, Hemant Kumar, Anand, Nishit, Seth, Ashish, Ghosh, Sreyan, Manocha, Dinesh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912238506344448
author Selvakumar, Ramaneswaran
Kumar, Sonal
Giri, Hemant Kumar
Anand, Nishit
Seth, Ashish
Ghosh, Sreyan
Manocha, Dinesh
author_facet Selvakumar, Ramaneswaran
Kumar, Sonal
Giri, Hemant Kumar
Anand, Nishit
Seth, Ashish
Ghosh, Sreyan
Manocha, Dinesh
contents Open-vocabulary audio language models (ALMs), like Contrastive Language Audio Pretraining (CLAP), represent a promising new paradigm for audio-text retrieval using natural language queries. In this paper, for the first time, we perform controlled experiments on various benchmarks to show that existing ALMs struggle to generalize to linguistic variations in textual queries. To address this issue, we propose RobustCLAP, a novel and compute-efficient technique to learn audio-language representations agnostic to linguistic variations. Specifically, we reformulate the contrastive loss used in CLAP architectures by introducing a multi-view contrastive learning objective, where paraphrases are treated as different views of the same audio scene and use this for training. Our proposed approach improves the text-to-audio retrieval performance of CLAP by 0.8%-13% across benchmarks and enhances robustness to linguistic variation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16505
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Do Audio-Language Models Understand Linguistic Variations?
Selvakumar, Ramaneswaran
Kumar, Sonal
Giri, Hemant Kumar
Anand, Nishit
Seth, Ashish
Ghosh, Sreyan
Manocha, Dinesh
Sound
Machine Learning
Audio and Speech Processing
Open-vocabulary audio language models (ALMs), like Contrastive Language Audio Pretraining (CLAP), represent a promising new paradigm for audio-text retrieval using natural language queries. In this paper, for the first time, we perform controlled experiments on various benchmarks to show that existing ALMs struggle to generalize to linguistic variations in textual queries. To address this issue, we propose RobustCLAP, a novel and compute-efficient technique to learn audio-language representations agnostic to linguistic variations. Specifically, we reformulate the contrastive loss used in CLAP architectures by introducing a multi-view contrastive learning objective, where paraphrases are treated as different views of the same audio scene and use this for training. Our proposed approach improves the text-to-audio retrieval performance of CLAP by 0.8%-13% across benchmarks and enhances robustness to linguistic variation.
title Do Audio-Language Models Understand Linguistic Variations?
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2410.16505