CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abootorabi, Mohammad Mahdi, Asgari, Ehsaneddin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912288552779776
author Abootorabi, Mohammad Mahdi
Asgari, Ehsaneddin
author_facet Abootorabi, Mohammad Mahdi
Asgari, Ehsaneddin
contents This study introduces CLASP (Contrastive Language-Speech Pretraining), a multilingual, multimodal representation tailored for audio-text information retrieval. CLASP leverages the synergy between spoken content and textual data. During training, we utilize our newly introduced speech-text dataset, which encompasses 15 diverse categories ranging from fiction to religion. CLASP's audio component integrates audio spectrograms with a pre-trained self-supervised speech model, while its language encoding counterpart employs a sentence encoder pre-trained on over 100 languages. This unified lightweight model bridges the gap between various modalities and languages, enhancing its effectiveness in handling and retrieving multilingual and multimodal data. Our evaluations across multiple languages demonstrate that CLASP establishes new benchmarks in HITS@1, MRR, and meanR metrics, outperforming traditional ASR-based retrieval methods that rely on transcribing speech into text for subsequent text retrieval, especially in specific scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13071
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
Abootorabi, Mohammad Mahdi
Asgari, Ehsaneddin
Computation and Language
Information Retrieval
Sound
Audio and Speech Processing
This study introduces CLASP (Contrastive Language-Speech Pretraining), a multilingual, multimodal representation tailored for audio-text information retrieval. CLASP leverages the synergy between spoken content and textual data. During training, we utilize our newly introduced speech-text dataset, which encompasses 15 diverse categories ranging from fiction to religion. CLASP's audio component integrates audio spectrograms with a pre-trained self-supervised speech model, while its language encoding counterpart employs a sentence encoder pre-trained on over 100 languages. This unified lightweight model bridges the gap between various modalities and languages, enhancing its effectiveness in handling and retrieving multilingual and multimodal data. Our evaluations across multiple languages demonstrate that CLASP establishes new benchmarks in HITS@1, MRR, and meanR metrics, outperforming traditional ASR-based retrieval methods that rely on transcribing speech into text for subsequent text retrieval, especially in specific scenarios.
title CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
topic Computation and Language
Information Retrieval
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2412.13071