Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ryu, Hyeonggon, Kim, Seongyu, Chung, Joon Son, Senocak, Arda
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908281384992768
author Ryu, Hyeonggon
Kim, Seongyu
Chung, Joon Son
Senocak, Arda
author_facet Ryu, Hyeonggon
Kim, Seongyu
Chung, Joon Son
Senocak, Arda
contents We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited to handling either speech or non-speech sounds independently, or at best, together but sequentially without mixing. This limitation prevents them from capturing the complexity of real-world audio sources that are often mixed. Our approach introduces a 'mix-and-separate' framework with audio-visual alignment objectives that jointly learn correspondence and disentanglement using mixed audio. Through these objectives, our model learns to produce distinct embeddings for each audio type, enabling effective disentanglement and grounding across mixed audio sources. Additionally, we created a new dataset to evaluate simultaneous grounding of mixed audio sources, demonstrating that our model outperforms prior methods. Our approach also achieves comparable or better performance in standard segmentation and cross-modal retrieval tasks, highlighting the benefits of our mix-and-separate approach.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18880
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes
Ryu, Hyeonggon
Kim, Seongyu
Chung, Joon Son
Senocak, Arda
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited to handling either speech or non-speech sounds independently, or at best, together but sequentially without mixing. This limitation prevents them from capturing the complexity of real-world audio sources that are often mixed. Our approach introduces a 'mix-and-separate' framework with audio-visual alignment objectives that jointly learn correspondence and disentanglement using mixed audio. Through these objectives, our model learns to produce distinct embeddings for each audio type, enabling effective disentanglement and grounding across mixed audio sources. Additionally, we created a new dataset to evaluate simultaneous grounding of mixed audio sources, demonstrating that our model outperforms prior methods. Our approach also achieves comparable or better performance in standard segmentation and cross-modal retrieval tasks, highlighting the benefits of our mix-and-separate approach.
title Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes
topic Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2503.18880