SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Mei, Xinhao, Lan, Gael Le, Liu, Haohe, Ni, Zhaoheng, Nagaraja, Varun, Liu, Yang, Shi, Yangyang, Chandra, Vikas
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914263230054400
author Mei, Xinhao
Lan, Gael Le
Liu, Haohe
Ni, Zhaoheng
Nagaraja, Varun
Liu, Yang
Shi, Yangyang
Chandra, Vikas
author_facet Mei, Xinhao
Lan, Gael Le
Liu, Haohe
Ni, Zhaoheng
Nagaraja, Varun
Liu, Yang
Shi, Yangyang
Chandra, Vikas
contents Contrastive language-audio pretraining (CLAP) has achieved notable success in learning semantically rich audio representations and is widely adopted for various audio-related tasks. However, current CLAP models face several key limitations. First, they are typically trained on relatively small datasets, often comprising a few million audio samples. Second, existing CLAP models are restricted to short and fixed duration, which constrains their usage in real-world scenarios with variable-duration audio. Third, the standard contrastive training objective operates on global representations, which may hinder the learning of dense, fine-grained audio features. To address these challenges, we introduce Scalable Language-Audio Pretraining (SLAP), which scales language-audio pretraining to 109 million audio-text pairs with variable audio durations and incorporates multiple training objectives. SLAP unifies contrastive loss with additional self-supervised and captioning losses in a single-stage training, facilitating the learning of richer dense audio representations. The proposed SLAP model achieves new state-of-the-art performance on audio-text retrieval and zero-shot audio classification tasks, demonstrating its effectiveness across diverse benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2601_12594
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training
Mei, Xinhao
Lan, Gael Le
Liu, Haohe
Ni, Zhaoheng
Nagaraja, Varun
Liu, Yang
Shi, Yangyang
Chandra, Vikas
Audio and Speech Processing
Artificial Intelligence
Sound
Contrastive language-audio pretraining (CLAP) has achieved notable success in learning semantically rich audio representations and is widely adopted for various audio-related tasks. However, current CLAP models face several key limitations. First, they are typically trained on relatively small datasets, often comprising a few million audio samples. Second, existing CLAP models are restricted to short and fixed duration, which constrains their usage in real-world scenarios with variable-duration audio. Third, the standard contrastive training objective operates on global representations, which may hinder the learning of dense, fine-grained audio features. To address these challenges, we introduce Scalable Language-Audio Pretraining (SLAP), which scales language-audio pretraining to 109 million audio-text pairs with variable audio durations and incorporates multiple training objectives. SLAP unifies contrastive loss with additional self-supervised and captioning losses in a single-stage training, facilitating the learning of richer dense audio representations. The proposed SLAP model achieves new state-of-the-art performance on audio-text retrieval and zero-shot audio classification tasks, demonstrating its effectiveness across diverse benchmarks.
title SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2601.12594