GLAP: General contrastive audio-text pretraining across domains and languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dinkel, Heinrich, Yan, Zhiyong, Wang, Tianzi, Wang, Yongqing, Sun, Xingwei, Niu, Yadong, Liu, Jizhong, Li, Gang, Zhang, Junbo, Luan, Jian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914262011609088
author Dinkel, Heinrich
Yan, Zhiyong
Wang, Tianzi
Wang, Yongqing
Sun, Xingwei
Niu, Yadong
Liu, Jizhong
Li, Gang
Zhang, Junbo
Luan, Jian
author_facet Dinkel, Heinrich
Yan, Zhiyong
Wang, Tianzi
Wang, Yongqing
Sun, Xingwei
Niu, Yadong
Liu, Jizhong
Li, Gang
Zhang, Junbo
Luan, Jian
contents Contrastive Language Audio Pretraining (CLAP) is a widely-used method to bridge the gap between audio and text domains. Current CLAP methods enable sound and music retrieval in English, ignoring multilingual spoken content. To address this, we introduce general language audio pretraining (GLAP), which expands CLAP with multilingual and multi-domain abilities. GLAP demonstrates its versatility by achieving competitive performance on standard audio-text retrieval benchmarks like Clotho and AudioCaps, while significantly surpassing existing methods in speech retrieval and classification tasks. Additionally, GLAP achieves strong results on widely used sound-event zero-shot benchmarks, while simultaneously outperforming previous methods on speech content benchmarks. Further keyword spotting evaluations across 50 languages emphasize GLAP's advanced multilingual capabilities. Finally, multilingual sound and music understanding is evaluated across four languages. Checkpoints and Source: https://github.com/xiaomi-research/dasheng-glap.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11350
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GLAP: General contrastive audio-text pretraining across domains and languages
Dinkel, Heinrich
Yan, Zhiyong
Wang, Tianzi
Wang, Yongqing
Sun, Xingwei
Niu, Yadong
Liu, Jizhong
Li, Gang
Zhang, Junbo
Luan, Jian
Sound
Computation and Language
Audio and Speech Processing
Contrastive Language Audio Pretraining (CLAP) is a widely-used method to bridge the gap between audio and text domains. Current CLAP methods enable sound and music retrieval in English, ignoring multilingual spoken content. To address this, we introduce general language audio pretraining (GLAP), which expands CLAP with multilingual and multi-domain abilities. GLAP demonstrates its versatility by achieving competitive performance on standard audio-text retrieval benchmarks like Clotho and AudioCaps, while significantly surpassing existing methods in speech retrieval and classification tasks. Additionally, GLAP achieves strong results on widely used sound-event zero-shot benchmarks, while simultaneously outperforming previous methods on speech content benchmarks. Further keyword spotting evaluations across 50 languages emphasize GLAP's advanced multilingual capabilities. Finally, multilingual sound and music understanding is evaluated across four languages. Checkpoints and Source: https://github.com/xiaomi-research/dasheng-glap.
title GLAP: General contrastive audio-text pretraining across domains and languages
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2506.11350