Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language
Fuente:
arXiv
Saved in:
| Main Authors: | Hamilton, Mark, Zisserman, Andrew, Hershey, John R., Freeman, William T. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2024)
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2024)
A SOUND APPROACH: Using Large Language Models to generate audio descriptions for egocentric text-audio retrieval
by: Oncescu, Andreea-Maria, et al.
Published: (2024)
by: Oncescu, Andreea-Maria, et al.
Published: (2024)
Transforming LLMs into Cross-modal and Cross-lingual Retrieval Systems
by: Gomez, Frank Palma, et al.
Published: (2024)
by: Gomez, Frank Palma, et al.
Published: (2024)
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
by: Lin, Chyi-Jiunn, et al.
Published: (2024)
by: Lin, Chyi-Jiunn, et al.
Published: (2024)
Analyzing Byte-Pair Encoding on Monophonic and Polyphonic Symbolic Music: A Focus on Musical Phrase Segmentation
by: Le, Dinh-Viet-Toan, et al.
Published: (2024)
by: Le, Dinh-Viet-Toan, et al.
Published: (2024)
Multi-Modal Retrieval For Large Language Model Based Speech Recognition
by: Kolehmainen, Jari, et al.
Published: (2024)
by: Kolehmainen, Jari, et al.
Published: (2024)
Exploring Diverse Sounds: Identifying Outliers in a Music Corpus
by: Cai, Le, et al.
Published: (2024)
by: Cai, Le, et al.
Published: (2024)
WikiMuTe: A web-sourced dataset of semantic descriptions for music audio
by: Weck, Benno, et al.
Published: (2023)
by: Weck, Benno, et al.
Published: (2023)
Beyond Musical Descriptors: Extracting Preference-Bearing Intent in Music Queries
by: Baranes, Marion, et al.
Published: (2026)
by: Baranes, Marion, et al.
Published: (2026)
Technical Report on classification of literature related to children speech disorder
by: Wang, Ziang, et al.
Published: (2025)
by: Wang, Ziang, et al.
Published: (2025)
I can listen but cannot read: An evaluation of two-tower multimodal systems for instrument recognition
by: Vasilakis, Yannis, et al.
Published: (2024)
by: Vasilakis, Yannis, et al.
Published: (2024)
A GEN AI Framework for Medical Note Generation
by: Leong, Hui Yi, et al.
Published: (2024)
by: Leong, Hui Yi, et al.
Published: (2024)
More than words: Advancements and challenges in speech recognition for singing
by: Kruspe, Anna
Published: (2024)
by: Kruspe, Anna
Published: (2024)
Navigating Speech Recording Collections with AI-Generated Illustrations
by: Håland, Sirina, et al.
Published: (2025)
by: Håland, Sirina, et al.
Published: (2025)
Implicit Self-supervised Language Representation for Spoken Language Diarization
by: Mishra, Jagabandhu, et al.
Published: (2023)
by: Mishra, Jagabandhu, et al.
Published: (2023)
Language-based Audio Retrieval with Co-Attention Networks
by: Sun, Haoran, et al.
Published: (2024)
by: Sun, Haoran, et al.
Published: (2024)
TALKPLAY: Multimodal Music Recommendation with Large Language Models
by: Doh, Seungheon, et al.
Published: (2025)
by: Doh, Seungheon, et al.
Published: (2025)
Can Impressions of Music be Extracted from Thumbnail Images?
by: Harada, Takashi, et al.
Published: (2025)
by: Harada, Takashi, et al.
Published: (2025)
LARP: Language Audio Relational Pre-training for Cold-Start Playlist Continuation
by: Salganik, Rebecca, et al.
Published: (2024)
by: Salganik, Rebecca, et al.
Published: (2024)
Music Discovery Dialogue Generation Using Human Intent Analysis and Large Language Models
by: Doh, SeungHeon, et al.
Published: (2024)
by: Doh, SeungHeon, et al.
Published: (2024)
AV2Wav: Diffusion-Based Re-synthesis from Continuous Self-supervised Features for Audio-Visual Speech Enhancement
by: Chou, Ju-Chieh, et al.
Published: (2023)
by: Chou, Ju-Chieh, et al.
Published: (2023)
Multimodal Transformer Distillation for Audio-Visual Synchronization
by: Chen, Xuanjun, et al.
Published: (2022)
by: Chen, Xuanjun, et al.
Published: (2022)
SCORE: Self-supervised Correspondence Fine-tuning for Improved Content Representations
by: Meghanani, Amit, et al.
Published: (2024)
by: Meghanani, Amit, et al.
Published: (2024)
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
by: Fan, Ruchao, et al.
Published: (2024)
by: Fan, Ruchao, et al.
Published: (2024)
Separate This, and All of these Things Around It: Music Source Separation via Hyperellipsoidal Queries
by: Watcharasupat, Karn N., et al.
Published: (2025)
by: Watcharasupat, Karn N., et al.
Published: (2025)
Self-supervised Speech Representations Still Struggle with African American Vernacular English
by: Chang, Kalvin, et al.
Published: (2024)
by: Chang, Kalvin, et al.
Published: (2024)
SSHR: Leveraging Self-supervised Hierarchical Representations for Multilingual Automatic Speech Recognition
by: Xue, Hongfei, et al.
Published: (2023)
by: Xue, Hongfei, et al.
Published: (2023)
LASER: Learning by Aligning Self-supervised Representations of Speech for Improving Content-related Tasks
by: Meghanani, Amit, et al.
Published: (2024)
by: Meghanani, Amit, et al.
Published: (2024)
Revisiting Self-supervised Learning of Speech Representation from a Mutual Information Perspective
by: Liu, Alexander H., et al.
Published: (2024)
by: Liu, Alexander H., et al.
Published: (2024)
Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
by: Meghanani, Amit, et al.
Published: (2026)
by: Meghanani, Amit, et al.
Published: (2026)
Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
by: Meghanani, Amit, et al.
Published: (2024)
by: Meghanani, Amit, et al.
Published: (2024)
The Kolmogorov Complexity of Irish traditional dance music
by: McGettrick, Michael, et al.
Published: (2024)
by: McGettrick, Michael, et al.
Published: (2024)
Distance Sampling-based Paraphraser Leveraging ChatGPT for Text Data Manipulation
by: Oh, Yoori, et al.
Published: (2024)
by: Oh, Yoori, et al.
Published: (2024)
Emergent musical properties of a transformer under contrastive self-supervised learning
by: Kong, Yuexuan, et al.
Published: (2025)
by: Kong, Yuexuan, et al.
Published: (2025)
VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering
by: Rackauckas, Zackary, et al.
Published: (2025)
by: Rackauckas, Zackary, et al.
Published: (2025)
Track Role Prediction of Single-Instrumental Sequences
by: Han, Changheon, et al.
Published: (2024)
by: Han, Changheon, et al.
Published: (2024)
Personalized Dynamic Music Emotion Recognition with Dual-Scale Attention-Based Meta-Learning
by: Zhang, Dengming, et al.
Published: (2024)
by: Zhang, Dengming, et al.
Published: (2024)
Expressivity-aware Music Performance Retrieval using Mid-level Perceptual Features and Emotion Word Embeddings
by: Chowdhury, Shreyan, et al.
Published: (2024)
by: Chowdhury, Shreyan, et al.
Published: (2024)
DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval
by: Xin, Yifei, et al.
Published: (2024)
by: Xin, Yifei, et al.
Published: (2024)
Exploring GPT's Ability as a Judge in Music Understanding
by: Fang, Kun, et al.
Published: (2025)
by: Fang, Kun, et al.
Published: (2025)
Similar Items
-
CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2024) -
A SOUND APPROACH: Using Large Language Models to generate audio descriptions for egocentric text-audio retrieval
by: Oncescu, Andreea-Maria, et al.
Published: (2024) -
Transforming LLMs into Cross-modal and Cross-lingual Retrieval Systems
by: Gomez, Frank Palma, et al.
Published: (2024) -
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
by: Lin, Chyi-Jiunn, et al.
Published: (2024) -
Analyzing Byte-Pair Encoding on Monophonic and Polyphonic Symbolic Music: A Focus on Musical Phrase Segmentation
by: Le, Dinh-Viet-Toan, et al.
Published: (2024)