CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Abootorabi, Mohammad Mahdi, Asgari, Ehsaneddin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
by: Lin, Chyi-Jiunn, et al.
Published: (2024)
by: Lin, Chyi-Jiunn, et al.
Published: (2024)
Multi-Modal Retrieval For Large Language Model Based Speech Recognition
by: Kolehmainen, Jari, et al.
Published: (2024)
by: Kolehmainen, Jari, et al.
Published: (2024)
Transforming LLMs into Cross-modal and Cross-lingual Retrieval Systems
by: Gomez, Frank Palma, et al.
Published: (2024)
by: Gomez, Frank Palma, et al.
Published: (2024)
Language-based Audio Retrieval with Co-Attention Networks
by: Sun, Haoran, et al.
Published: (2024)
by: Sun, Haoran, et al.
Published: (2024)
Analyzing Byte-Pair Encoding on Monophonic and Polyphonic Symbolic Music: A Focus on Musical Phrase Segmentation
by: Le, Dinh-Viet-Toan, et al.
Published: (2024)
by: Le, Dinh-Viet-Toan, et al.
Published: (2024)
Pretrained Conformers for Audio Fingerprinting and Retrieval
by: Altwlkany, Kemal, et al.
Published: (2025)
by: Altwlkany, Kemal, et al.
Published: (2025)
TALKPLAY: Multimodal Music Recommendation with Large Language Models
by: Doh, Seungheon, et al.
Published: (2025)
by: Doh, Seungheon, et al.
Published: (2025)
Navigating Speech Recording Collections with AI-Generated Illustrations
by: Håland, Sirina, et al.
Published: (2025)
by: Håland, Sirina, et al.
Published: (2025)
DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval
by: Xin, Yifei, et al.
Published: (2024)
by: Xin, Yifei, et al.
Published: (2024)
Multiscale Matching Driven by Cross-Modal Similarity Consistency for Audio-Text Retrieval
by: Wang, Qian, et al.
Published: (2024)
by: Wang, Qian, et al.
Published: (2024)
Application of Audio Fingerprinting Techniques for Real-Time Scalable Speech Retrieval and Speech Clusterization
by: Altwlkany, Kemal, et al.
Published: (2024)
by: Altwlkany, Kemal, et al.
Published: (2024)
Expressivity-aware Music Performance Retrieval using Mid-level Perceptual Features and Emotion Word Embeddings
by: Chowdhury, Shreyan, et al.
Published: (2024)
by: Chowdhury, Shreyan, et al.
Published: (2024)
JEPOO: Highly Accurate Joint Estimation of Pitch, Onset and Offset for Music Information Retrieval
by: Wei, Haojie, et al.
Published: (2023)
by: Wei, Haojie, et al.
Published: (2023)
Natural Language Processing Methods for Symbolic Music Generation and Information Retrieval: a Survey
by: Le, Dinh-Viet-Toan, et al.
Published: (2024)
by: Le, Dinh-Viet-Toan, et al.
Published: (2024)
Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling
by: Yeo, Eunjung, et al.
Published: (2026)
by: Yeo, Eunjung, et al.
Published: (2026)
LARP: Language Audio Relational Pre-training for Cold-Start Playlist Continuation
by: Salganik, Rebecca, et al.
Published: (2024)
by: Salganik, Rebecca, et al.
Published: (2024)
Music Discovery Dialogue Generation Using Human Intent Analysis and Large Language Models
by: Doh, SeungHeon, et al.
Published: (2024)
by: Doh, SeungHeon, et al.
Published: (2024)
A SOUND APPROACH: Using Large Language Models to generate audio descriptions for egocentric text-audio retrieval
by: Oncescu, Andreea-Maria, et al.
Published: (2024)
by: Oncescu, Andreea-Maria, et al.
Published: (2024)
Music Era Recognition Using Supervised Contrastive Learning and Artist Information
by: He, Qiqi, et al.
Published: (2024)
by: He, Qiqi, et al.
Published: (2024)
Enriching Music Descriptions with a Finetuned-LLM and Metadata for Text-to-Music Retrieval
by: Doh, SeungHeon, et al.
Published: (2024)
by: Doh, SeungHeon, et al.
Published: (2024)
I can listen but cannot read: An evaluation of two-tower multimodal systems for instrument recognition
by: Vasilakis, Yannis, et al.
Published: (2024)
by: Vasilakis, Yannis, et al.
Published: (2024)
A GEN AI Framework for Medical Note Generation
by: Leong, Hui Yi, et al.
Published: (2024)
by: Leong, Hui Yi, et al.
Published: (2024)
More than words: Advancements and challenges in speech recognition for singing
by: Kruspe, Anna
Published: (2024)
by: Kruspe, Anna
Published: (2024)
WikiMuTe: A web-sourced dataset of semantic descriptions for music audio
by: Weck, Benno, et al.
Published: (2023)
by: Weck, Benno, et al.
Published: (2023)
Beyond Musical Descriptors: Extracting Preference-Bearing Intent in Music Queries
by: Baranes, Marion, et al.
Published: (2026)
by: Baranes, Marion, et al.
Published: (2026)
Technical Report on classification of literature related to children speech disorder
by: Wang, Ziang, et al.
Published: (2025)
by: Wang, Ziang, et al.
Published: (2025)
Diff4Steer: Steerable Diffusion Prior for Generative Music Retrieval with Semantic Guidance
by: Bao, Xuchan, et al.
Published: (2024)
by: Bao, Xuchan, et al.
Published: (2024)
SpeechTaxi: On Multilingual Semantic Speech Classification
by: Keller, Lennart, et al.
Published: (2024)
by: Keller, Lennart, et al.
Published: (2024)
Teaching a Multilingual Large Language Model to Understand Multilingual Speech via Multi-Instructional Training
by: Denisov, Pavel, et al.
Published: (2024)
by: Denisov, Pavel, et al.
Published: (2024)
wav2graph: A Framework for Supervised Learning Knowledge Graph from Speech
by: Le-Duc, Khai, et al.
Published: (2024)
by: Le-Duc, Khai, et al.
Published: (2024)
SECP: A Speech Enhancement-Based Curation Pipeline For Scalable Acquisition Of Clean Speech
by: Sabra, Adam, et al.
Published: (2024)
by: Sabra, Adam, et al.
Published: (2024)
Analyzing and reducing the synthetic-to-real transfer gap in Music Information Retrieval: the task of automatic drum transcription
by: Zehren, Mickaël, et al.
Published: (2024)
by: Zehren, Mickaël, et al.
Published: (2024)
Exploring Diverse Sounds: Identifying Outliers in a Music Corpus
by: Cai, Le, et al.
Published: (2024)
by: Cai, Le, et al.
Published: (2024)
Track Role Prediction of Single-Instrumental Sequences
by: Han, Changheon, et al.
Published: (2024)
by: Han, Changheon, et al.
Published: (2024)
Personalized Dynamic Music Emotion Recognition with Dual-Scale Attention-Based Meta-Learning
by: Zhang, Dengming, et al.
Published: (2024)
by: Zhang, Dengming, et al.
Published: (2024)
Multi-Sample Dynamic Time Warping for Few-Shot Keyword Spotting
by: Wilkinghoff, Kevin, et al.
Published: (2024)
by: Wilkinghoff, Kevin, et al.
Published: (2024)
Do Captioning Metrics Reflect Music Semantic Alignment?
by: Lee, Jinwoo, et al.
Published: (2024)
by: Lee, Jinwoo, et al.
Published: (2024)
Automatic Estimation of Singing Voice Musical Dynamics
by: Narang, Jyoti, et al.
Published: (2024)
by: Narang, Jyoti, et al.
Published: (2024)
Engraving Oriented Joint Estimation of Pitch Spelling and Local and Global Keys
by: Bouquillard, Augustin, et al.
Published: (2024)
by: Bouquillard, Augustin, et al.
Published: (2024)
Towards Computational Analysis of Pansori Singing
by: Park, Sangheon, et al.
Published: (2024)
by: Park, Sangheon, et al.
Published: (2024)
Similar Items
-
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
by: Lin, Chyi-Jiunn, et al.
Published: (2024) -
Multi-Modal Retrieval For Large Language Model Based Speech Recognition
by: Kolehmainen, Jari, et al.
Published: (2024) -
Transforming LLMs into Cross-modal and Cross-lingual Retrieval Systems
by: Gomez, Frank Palma, et al.
Published: (2024) -
Language-based Audio Retrieval with Co-Attention Networks
by: Sun, Haoran, et al.
Published: (2024) -
Analyzing Byte-Pair Encoding on Monophonic and Polyphonic Symbolic Music: A Focus on Musical Phrase Segmentation
by: Le, Dinh-Viet-Toan, et al.
Published: (2024)